EDBT 2026 Demo / reviewers in the wild / expert
Lei Jin 0003
dblp:59/3349-3
· DBLP profile ↗
40ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0003-4855-2464ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 19 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Security and privacy · 6 · 4 first-author · 1 since 2021Computer networks · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DrawMotion: Generating 3D Human Motions by Freehand Drawing
Tao Wang 0011, Lei Jin 0003, Qiaozhi He, Jiaming Chu, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Shuicheng Yan, Li Wang 0039 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | CE-CLIP: Cloud-edge collaborative fine-tuning for multimodal adaptation
Kejun Ren, Yuntao Du 0001, Lianming Xu, Yunxiang Yao, Lei Jin 0003, Li Wang 0039 |
Pattern Recognit. | 6 |
| 2026 | SynSP++: General Pose Sequences Refinement via Synergy of Smoothness and Precision
Lei Jin 0003, Tao Wang 0011, Junliang Xing, Jian Zhao 0006, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | CCIN: Compositional Conflict Identification and Neutralization for Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a multi-modal task that seeks to retrieve target images by harmonizing a reference image with a modified instruction. A key challenge in CIR lies in compositional conflicts between the reference image (e.g., blue, long sleeve) and the modified instruction (e.g., grey, short sleeve). Previous works attempt to mitigate such conflicts through feature-level manipulation, commonly employing learnable masks to obscure conflicting features within the reference image. However, the inherent complexity of feature spaces poses significant challenges in precise conflict neutralization, thereby leading to uncontrollable results. To this end, this paper proposes the Compositional Conflict Identification and Neutralization (CCIN) framework, which sequentially identifies and neutralizes compositional conflicts for effective CIR. Specifically, CCIN comprises two core modules: 1) Compositional Conflict Identification module, which utilizes LLM-based analysis to identify specific conflicting attributes, and 2) Compositional Conflict Neutralization module, which first generates a kept instruction to preserve non-conflicting attributes, then neutralizes conflicts under collaborative guidance of both the kept and modified instructions. Extensive experiments demonstrate the superiority of CCIN over the state-of-the-arts. Code repository: https://github.com/LikaiTian/CCIN. Likai Tian, Jian Zhao 0006, Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Lei Jin 0003, Zheng Wang 0007, Xuelong Li 0001 |
CVPR | 6 |
| 2025 | StickMotion: Generating 3D Human Motions by Drawing a StickmanabstractText-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text inputs. This paper introduces StickMotion, an efficient diffusion-based network designed for multi-condition scenarios, which generates desired motions based on traditional text and our proposed stickman conditions for global and local control of these motions, respectively. We address the challenges introduced by the user-friendly stickman from three perspectives: 1) Data generation. We develop an algorithm to generate hand-drawn stickmen automatically across different dataset formats. 2) Multi-condition fusion. We propose a multi-condition module that integrates into the diffusion process and obtains outputs of all possible condition combinations, reducing computational complexity and enhancing StickMotion’s performance compared to conventional approaches with the self-attention module. 3) Dynamic supervision. We empower StickMotion to make minor adjustments to the stickman’s position within the output sequences, generating more natural movements through our proposed dynamic supervision strategy. Through quantitative experiments and user studies, sketching stickmen saves users about 51.5% of their time generating motions consistent with their imagination. Our codes, demos, and relevant data will be released in https:// github.com/InvertedForest/StickMotion. Tao Wang 0011, Qiaozhi He, Jiaming Chu, Ling Qian, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Lei Jin 0003 |
CVPR | 9 |
| 2025 | JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV SystemsabstractUnmanned Aerial Vehicles (UAVs) are widely adopted across various fields, yet they raise significant privacy and safety concerns, demanding robust monitoring solutions. Existing anti-UAV methods primarily focus on position tracking but fail to capture UAV behavior and intent. To address this, we introduce a novel task—UAV Tracking and Intent Understanding (UTIU)—which aims to track UAVs while inferring and describing their motion states and intent for a more comprehensive monitoring approach. To tackle the task, we propose JTD-UAV, the first joint tracking, and intent description framework based on large language models. Our dual-branch architecture integrates UAV tracking with Visual Question Answering (VQA), allowing simultaneous localization and behavior description. To benchmark this task, we introduce the TDUAV dataset, the largest dataset for joint UAV tracking and intent understanding, featuring 1,328 challenging video sequences, over 163K annotated thermal frames, and 3K VQA pairs. Our benchmark demonstrates the effectiveness of JTD-UAV. Jian Zhao 0006, Zhaoxin Fan, Xin Zhang 0093, Yudian Zhang, Lei Jin 0003, Gang Wang 0031, Mengxi Jia, Xuelong Li 0001 |
CVPR | 7 |
| 2025 | Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
Xiaojian Lin, Wenxin Zhang 0005, Yuchu Jiang, Wangyu Wu, Kangxu Wang, Zongzheng Zhang, Guijin Wang, Lei Jin 0003, Hao Zhao 0002 |
ACM Multimedia | 9 |
| 2025 | MT-Agent: Constructing a GUI Agent via Modality Enhancement and Text-Guided FusionabstractGraphical User Interfaces (GUIs) play a crucial role in facilitating user-computer interactions, making them an essential focus of research. However, current automated GUI agents face significant challenges in effectively associating task implementations with specific visual elements, and the resolution constraints of Vision-Language Models (VLMs) also limit the richness of visual information. To this end, we propose a novel multi-modal agent named MT-Agent, which enhances both textual and visual input modalities to enable the model to perceive visual elements in GUIs more effectively. Specifically, Textual Modality Enhancement improves the semantic richness of input text by capturing task-specific details via an external VLM, while Visual Modality Enhancement incorporates fine-grained visual details to better represent critical GUI elements. In addition, we introduce an innovative text-guided directional feature fusion mechanism, which leverages enriched text features to guide the integration with visual information. In experiments, MT-Agent demonstrated exceptional performance on AITZ dataset, achieving an action type prediction accuracy of 84.80% and a step prediction accuracy of 58.07%, surpassing previous state-of-the-art models. Furthermore, on the GUI Odyssey benchmark, MT-Agent achieves performance comparable to previous state-of-the-art models while using only about 1/20 of their trainable parameters. Our codes, demos, and relevant data will be released to facilitate further research and validation within the scientific community. Jinhan Dong, Lei Jin 0003, Zhihong Zhang 0006, Runqing Zhang, Liqiang Xu, Junliang Xing |
IEEE Internet Things J. | 2 |
| 2025 | Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing ImagesabstractCloud segmentation is a critical challenge in remote sensing image interpretation, as its accuracy directly impacts the effectiveness of subsequent data processing and analysis. Recently, vision foundation models (VFM) have demonstrated powerful generalization capabilities across various visual tasks. In this paper, we present a parameter-efficient adaptive approach, termed Cloud-Adapter, designed to enhance the accuracy and robustness of cloud segmentation. Our method leverages a VFM pretrained on general domain data, which remains frozen, eliminating the need for additional training. Cloud-Adapter incorporates a lightweight spatial perception module that initially utilizes a convolutional neural network (ConvNet) to extract dense spatial representations. These multi-scale features are then aggregated and serve as contextual inputs to an adapting module, which modulates the frozen transformer layers within the VFM. Experimental results demonstrate that the Cloud-Adapter approach, utilizing only 0.6% of the trainable parameters of the frozen backbone, achieves substantial performance gains. Cloud-Adapter consistently achieves state-of-the-art performance across various cloud segmentation datasets from multiple satellite sources, sensor series, data processing levels, land cover scenarios, and annotation granularities. Code and model checkpoints are available at https://xavierjiezou.github.io/Cloud-Adapter/. Xuechao Zou, Kai Li 0023, Junliang Xing, Lei Jin 0003, Congyan Lang, Pin Tao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | CMoA: Contrastive Mixture of Adapters for Generalized Few-Shot Continual LearningabstractThe goal of Few-Shot Continual Learning (FSCL) is to incrementally learn novel tasks with limited labeled samples and preserve previous capabilities simultaneously. However, current FSCL works lack research on domain increment and domain generalization ability, which cannot cope with changes in the visual perception environment. In this paper, we set up a Generalized FSCL (GFSCL) protocol involving both class- and domain-incremental scenarios together with domain generalization assessment. Firstly, two benchmark datasets and protocols are newly arranged, and detailed baselines are provided for this unexplored configuration. Furthermore, we find that common continual learning methods have poor generalization ability on unseen domains and cannot better tackle catastrophic forgetting issue in cross-incremental tasks. Hence, we propose a rehearsal-free framework based on Vision Transformer (ViT) named Contrastive Mixture of Adapters (CMoA). It contains two non-conflicting parts: (1) By applying the fast-adaptation characteristic of adapter-embedded ViT, the mixture of Adapters (MoA) module is incorporated into ViT. For stability purpose, cosine similarity regularization and dynamic weighting are designed to make each adapter learn specific knowledge and concentrate on particular classes. (2) To further enhance domain generalization ability, we alleviate the intra-class variation by prototype-calibrated contrastive learning to improve domain-invariant representation learning. Finally, six evaluation indicators showing the overall performance and forgetting are compared by comprehensive experiments on two benchmark datasets to validate the efficacy of CMoA, and the results illustrate that CMoA can achieve comparative performance with rehearsal-based continual learning methods. Yawen Cui, Jian Zhao 0006, Zitong Yu, Rizhao Cai, Lei Jin 0003, Alex Chichung Kot, Li Liu 0002, Xuelong Li 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | SynSP: Synergy of Smoothness and Precision in Pose Sequences RefinementabstractPredicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily concentrated on striking a balance between the two objectives, i.e., smoothness and precision, while optimizing the predicted pose sequences. However, it has come to our attention that the tension between these two objectives can provide additional quality cues about the predicted pose sequences. These cues, in turn, are able to aid the network in optimizing lower-quality poses. To leverage this quality information, we propose a motion refinement network, termed SynSP, to achieve a Synergy of Smoothness and Precision in the sequence refinement tasks. Moreover, SynSP can also address multi-view poses of one person simultaneously, fixing inaccuracies in predicted poses through heightened attention to similar poses from other views, thereby amplifying the resultant quality cues and overall performance. Compared with previous methods, SynSP benefits from both pose quality and multi-view information with a much shorter input sequence length, achieving state-of-the-art results among four challenging datasets involving 2D, 3D, and SMPL pose representations in both single-view and multi-view scenes. Github code: https://github.com/InvertedForest/SynSP. Tao Wang 0011, Lei Jin 0003, Zheng Wang 0007, Jianshu Li, Liang Li 0003, Fang Zhao 0006, Yu Cheng 0009, Li Yuan 0007, Junliang Xing, Jian Zhao 0006 |
CVPR | 2 |
| 2024 | DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingabstractVision-centric autonomous driving has recently raised wide attention due to its lower cost. Pretraining is essential for extracting a universal representation. However, current vision-centric pretraining typically relies on either 2D or 3D pre-text tasks, overlooking the temporal characteristics of autonomous driving as a 4D scene understanding task. In this paper, we address this challenge by introducing a world model-based autonomous driving 4D representation learning framework, dubbed DriveWorld, which is capable of pretraining from multi-camera driving videos in a spatiotemporal fashion. Specifically, we propose a Memory State-Space Model for spatiotemporal modelling, which consists of a Dynamic Memory Bank module for learning temporal-aware latent dynamics to predict future changes and a Static Scene Propagation module for learning spatial-aware latent statics to offer comprehensive scene contexts. We additionally introduce a Task Prompt to decouple task-aware features for various downstream tasks. The experiments demonstrate that DriveWorld delivers promising results on various autonomous driving tasks. When pretrained with the OpenScene dataset, DriveWorld achieves a 7.5% increase in mAP for 3D object detection, a 3.0% increase in IoU for online mapping, a 5.0% increase in AMOTA for multi-object tracking, a 0.1m decrease in minADE for motionforecasting, a 3.0% increase in IoU for occupancy prediction, and a 0.34m reduction in average L2 error for planning. Dawei Zhao 0003, Liang Xiao 0007, Jian Zhao 0006, Xinli Xu, Lei Jin 0003, Jianshu Li, Yulan Guo, Junliang Xing, Liping Jing, Yiming Nie, Bin Dai 0001 |
CVPR | 7 |
| 2024 | ASQuery: A Query-based Model for Action SegmentationabstractFor the task of temporal action segmentation, existing works commonly treat it as a frame-wise classification problem. In this paper, we propose a straight but effective model namely ASQuery by learning central representation of each action category, which transforms the classification problem to the similarity calculation between category-specific queries and frame features. These central representations are dynamically generated through our Transformer decoder module, endowing them more flexible and comprehensive perception of the whole video. Moreover, we first introduce the boundary query for refining segmentation results, aiding to alleviating the troublesome over-segmentation problem. ASQuery demonstrates superior performance compared to state-of-the-art models, achieving improvements of 0.9% and 4.1% in the mean metrics on two public action segmentation datasets, i.e., Breakfast and Assembly101, respectively. The source codes are available at https://github.com/zlngan/ASQuery. Ziliang Gan, Lei Jin 0003, Zheng Wang 0007, Liang Li 0003, Zhecan Wang, Jianshu Li, Junliang Xing, Jian Zhao 0006 |
ICME | 2 |
| 2024 | Unified Single-Stage Transformer Network for Efficient RGB-T Tracking
Jianqiang Xia, Dian-xi Shi, Linna Song, Songchang Jin, Chenran Zhao, Yu Cheng 0009, Lei Jin 0003, Jianan Li 0001, Gang Wang 0031, Junliang Xing, Jian Zhao 0006 |
IJCAI | 9 |
| 2024 | GOP: A Group Object Perception Framework for Optical Remote Sensing
Lei Jin 0003, Xuechao Zou, Jian Zhao 0006, Junliang Xing |
PRCV (12) | 2 |
| 2024 | SkatingVerse: A large-scale benchmark for comprehensive evaluation on human action understandingabstractAbstract Human action understanding (HAU) is a broad topic that involves specific tasks, such as action localisation, recognition, and assessment. However, most popular HAU datasets are bound to one task based on particular actions. Combining different but relevant HAU tasks to establish a unified action understanding system is challenging due to the disparate actions across datasets. A large‐scale and comprehensive benchmark, namely SkatingVerse is constructed for action recognition, segmentation, proposal, and assessment. SkatingVerse focus on fine‐grained sport action, hence figure skating is chosen as the task object, which eliminates the biases of the object, scene, and space that exist in most previous datasets. In addition, skating actions have inherent complexity and similarity, which is an enormous challenge for current algorithms. A total of 1687 official figure skating competition videos was collected with a total of 184.4 h, exceeding four times over other datasets with a similar topic. SkatingVerse enables to formulate a unified task to output fine‐grained human action classification and assessment results from a raw figure skating competition video. In addition, SkatingVerse can facilitate the study of HAU foundation model due to its large scale and abundant categories. Moreover, image modality is incorporated for human pose estimation task into SkatingVerse . Extensive experimental results show that (1) SkatingVerse significantly helps the training and evaluation of HAU methods, (2) the performance of existing HAU methods has much room to improve, and SkatingVerse helps to reduce such gaps, and (3) unifying relevant tasks in HAU through a uniform dataset can facilitate more practical applications. SkatingVerse will be publicly available to facilitate further studies on relevant problems. Ziliang Gan, Lei Jin 0003, Yu Cheng 0009, Yinglei Teng, Zun Li 0001, Yawen Li 0001, Wenhan Yang, Junliang Xing, Jian Zhao 0006 |
IET Comput. Vis. | 2 |
| 2024 | DiffCR: A Fast Conditional Diffusion Framework for Cloud Removal From Optical Satellite ImagesabstractOptical satellite images are a critical data source; however, cloud cover often compromises their quality, hindering image applications and analysis. Consequently, effectively removing clouds from optical satellite images has emerged as a prominent research direction. Recent advances in deep learning-based cloud removal methods have been significant, but image generation quality still needs improvement. Diffusion models have demonstrated remarkable success in diverse image-generation tasks, showcasing their potential in addressing this challenge. This paper presents a novel framework called DiffCR, which leverages conditional guided diffusion with deep convolutional networks for high-performance cloud removal for optical satellite imagery. Specifically, we introduce a decoupled encoder for conditional image feature extraction, providing a robust color representation to ensure the close similarity of appearance information between the conditional input and the synthesized output. Moreover, we propose a novel and efficient time and condition fusion block within the cloud removal model to accurately simulate the correspondence between the appearance in the conditional image and the target image at a low computational cost. Extensive experimental evaluations on three commonly used benchmark datasets demonstrate that DiffCR consistently achieves state-of-the-art performance on all metrics, with parameter and computational complexities amounting to only 5.1% and 5.4%, respectively, of those previous best methods. The source code, pre-trained models, and all the experimental results will be publicly available at https://github.com/XavierJiezou/DiffCR upon the paper’s acceptance of this work. Xuechao Zou, Kai Li 0023, Junliang Xing, Yu Zhang 0165, Lei Jin 0003, Pin Tao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | UniParser: Multi-Human Parsing With Unified Correlation Representation LearningabstractMulti-human parsing is an image segmentation task necessitating both instance-level and fine-grained category-level information. However, prior research has typically processed these two types of information through distinct branch types and output formats, leading to inefficient and redundant frameworks. This paper introduces UniParser, which integrates instance-level and category-level representations in three key aspects: 1) we propose a unified correlation representation learning approach, allowing our network to learn instance and category features within the cosine space; 2) we unify the form of outputs of each modules as pixel-level results while supervising instance and category features using a homogeneous label accompanied by an auxiliary loss; and 3) we design a joint optimization procedure to fuse instance and category representations. By unifying instance-level and category-level output, UniParser circumvents manually designed post-processing techniques and surpasses state-of-the-art methods, achieving 49.3% AP on MHPv2.0 and 60.4% AP on CIHP. We have released our source code, pretrained models, and demos to facilitate future studies on https://github.com/cjm-sfw/Uniparser. Jiaming Chu, Lei Jin 0003, Yinglei Teng, Jianshu Li, Yunchao Wei, Zheng Wang 0007, Junliang Xing, Shuicheng Yan, Jian Zhao 0006 |
IEEE Trans. Image Process. | 2 |
| 2024 | Rethinking the Person Localization for Single-Stage Multi-Person Pose EstimationabstractSingle-stage models for multi-person pose estimation have garnered significant attention due to their streamlined approach in generating person position localization and body structure perception in a single pass. These two parts, however, are processed individually by existing methods, leading to suboptimal results, e.g., candidates with high confidences for person localization while poor structure estimations. To this end, we propose a simple yet effective approach, namely Structure-guided Person Localization (SPL), jointly leveraging the advantages of the two aspects to solve the multi-person pose estimation problem, with two complementary novelties. First, we propose to incorporate body structure perception to guide person position localization, consequently, we introduce the Structure-guided Center Learning (SCL) to unify the quality of the body structure perception in the displacement map with the confidence of the person existence in the center map, thus achieving more accurate keypoint position localization results even with extreme poses. Second, to facilitate the end-to-end training of SPL, we propose the efficient Agency-based Scale-adaptive Learning (ASL). Specifically, we predict an agency map of the same size as the center map, which focuses on the foreground area and can adaptively adjust the scale size for each central area with the body structure perception confidence. Comprehensive experiments on challenging benchmarks including COCO and CrowdPose clearly verify the superiority of our framework, which achieves new state-of-the-art single-stage multi-person pose estimation results. Specifically, SPL obtains 72.1 AP scores and 69.5 AP scores in COCO test-dev2017 and CrowdPose test set, respectively. Lei Jin 0003, Xuecheng Nie, Wendong Wang 0003, Yandong Guo, Shuicheng Yan, Jian Zhao 0006 |
IEEE Trans. Multim. | 1 |
| 2024 | Review and Analysis of RGBT Single Object Tracking Methods: A Fusion PerspectiveabstractVisual tracking is a fundamental task in computer vision with significant practical applications in various domains, including surveillance, security, robotics, and human-computer interaction. However, it may face limitations in visible light data, such as low-light environments, occlusion, and camouflage, which can significantly reduce its accuracy. To cope with these challenges, researchers have explored the potential of combining the visible and infrared modalities to improve tracking performance. By leveraging the complementary strengths of visible and infrared data, RGB-infrared fusion tracking has emerged as a promising approach to address these limitations and improve tracking accuracy in challenging scenarios. In this article, we present a review on RGB-infrared fusion tracking. Specifically, we categorize existing RGBT tracking methods into four categories based on their underlying architectures, feature representations, and fusion strategies, namely feature decoupling based method, feature selecting based method, collaborative graph tracking method, and traditional fusion method. Furthermore, we provide a critical analysis of their strengths, limitations, representative methods, and future research directions. To further demonstrate the advantages and disadvantages of these methods, we present a review of publicly available RGBT tracking datasets and analyze the main results on public datasets. Moreover, we discuss some limitations in RGBT tracking at present and provide some opportunities and future directions for RGBT visual tracking, such as dataset diversity, unsupervised and weakly supervised applications. In conclusion, our survey aims to serve as a useful resource for researchers and practitioners interested in the emerging field of RGBT tracking, and to promote further progress and innovation in this area. Jun Wang 0041, Shengjie Li 0003, Lei Jin 0003, Hao Wu 0098, Jian Zhao 0006, Bo Zhang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Modality Meets Long-Term Tracker: A Siamese Dual Fusion Framework for Tracking UAVabstractTracking an Unmanned Aerial Vehicle (UAV) to obtain its locations and trajectory is a crucial task to avoid the unlawful use of UAVs. However, most existing UAV tracking methods fail when facing cluster environments, out-of-view, and occlusions because of their insufficient representation of global context information capacity. To mitigate these issues, we propose a new tracker, namely SiamFusion, to innovate a dual fusion procedure that leverages the advantages in both the feature and decision levels. In particular, we propose a novel feature fusion module named Modality-Fusion to utilize multi-modal information, enhancing the perception of the target. From the decision level, we further develop a local-global converter based on a multi-modal fusion decision-making mechanism to reduce the accumulation during tracking, which significantly increases the robustness of the tracking process. Extensive experiments demonstrate the superiority of the proposed SiamFusion, which achieves the best performance on Anti-UAV in terms of accuracy and speed. In particular, we exceed the state-of-the-art tracking algorithm in the tracking accuracy by 4.2% at a similar frame rate. Our source codes, pre-trained models, and online demos will be released upon acceptance. Lei Jin 0003, Shengjie Li 0003, Jianqiang Xia, Jun Wang 0041, Zun Li 0001, Wenhan Yang, Pengfei Zhang 0016, Jian Zhao 0006, Bo Zhang 0007 |
ICIP | 2 |
| 2023 | Single-Stage Multi-human Parsing via Point Sets and Center-based OffsetsabstractThis work studies the multi-human parsing problem. Existing methods, either following top-down or bottom-up two-stage paradigms, usually involve expensive computational costs. We instead present a high-performance Single-stage Multi-human Parsing (SMP) deep architecture that decouples the multi-human parsing problem into two fine-grained sub-problems,i.e., locating the human body and parts. SMP leverages the point features in the barycenter positions to obtain their segmentation and then generates a series of offsets from the barycenter of the human body to the barycenters of parts, thus performing human body and parts matching without the grouping process. Within the SMP architecture, we propose a Refined Feature Retain module to extract the global feature of instances through generated mask attention and a Mask of Interest Reclassify module as a trainable plug-in module to refine the classification results with the predicted segmentation. Extensive experiments on the MHPv2.0 dataset demonstrate the best effectiveness and efficiency of the proposed method, surpassing the state-of-the-art method by 2.1% in AP50p, 1.0% in APvolpsup>, and 1.2% in PCP50. Moreover, SMP also achieves superior performance in DensePose-COCO, verifying generalization of the model. In particular, the proposed method requires fewer training epochs and a less complex model architecture. Our codes are released in https://github.com/cjm-sfw/SMP. Jiaming Chu, Lei Jin 0003, Xiaojin Fan, Yinglei Teng, Yunchao Wei, Yuqiang Fang, Junliang Xing, Jian Zhao 0006 |
ACM Multimedia | 2 |
| 2023 | DecenterNet: Bottom-Up Human Pose Estimation Via Decentralized Pose RepresentationabstractMulti-person pose estimation in crowded scenes remains a very challenging task. This paper finds that most previous methods fail to estimate or group visible keypoints in crowded scenes rather than reasoning invisible keypoints. We thus categorize the crowded scenes into entanglement and occlusion based on the visibility of human parts and observe that entanglement is a significant problem in crowded scenes. With this observation, we propose DecenterNet, an end-to-end deep architecture to perform robust and efficient pose estimation in crowded scenes. Within DecenterNet, we introduce a decentralized pose representation that uses all visible keypoints as the root points to represent human poses, which is more robust in the entanglement area. We also propose a decoupled pose assessment mechanism, which introduces a location map to adaptively select optimal poses in the offset map. In addition, we have constructed a new dataset named SkatingPose, containing more entangled scenes. The proposed DecenterNet surpasses the best method on SkatingPose by 1.8 AP. Furthermore, DecenterNet obtains 71.2 AP and 71.4 AP on the COCO and CrowdPose datasets, respectively, demonstrating the superiority of our method. We will release our source code, trained models, and dataset to facilitate further studies in this research direction. Our code and dataset are available in https://github.com/InvertedForest/DecenterNet. Tao Wang 0011, Lei Jin 0003, Xiaojin Fan, Yu Cheng 0009, Yinglei Teng, Junliang Xing, Jian Zhao 0006 |
ACM Multimedia | 2 |
| 2023 | Joint coupled representation and homogeneous reconstruction for multi-resolution small sample face recognition
Xiaojin Fan, Mengmeng Liao, Jingfeng Xue, Hao Wu 0098, Lei Jin 0003, Jian Zhao 0006, Liehuang Zhu |
Neurocomputing | 5 |
| 2023 | Pruning-and-distillation: One-stage joint compression framework for CNNs via clustering
Tao Niu, Yinglei Teng, Lei Jin 0003, Panpan Zou |
Image Vis. Comput. | 3 |
| 2023 | Grouping by Center: Predicting Centripetal Offsets for the Bottom-up Human Pose EstimationabstractWe introduce Grouping by Center, a novel grouping approach for the bottom-up human pose estimation, which detects human joint first and then does grouping. The grouping strategy is the critical factor for the bottom-up pose estimation. To increase the conciseness and accuracy, we propose to use the center of the body as a grouping clue. More concretely, we predict the offsets from the keypoints to the body centers. Keypoints with aligned shifted results will be grouped as one person. However, the multi-scale variance of people can affect the prediction of the grouping clue, which has been neglected in previous research. To resolve the scale variance of the offset, we put forward a Multi-scale Translation Layer and an iterative refinement. Furthermore, we scheme a greedy grouping strategy with a dynamic threshold due to the various scales of instances. Through a comprehensive comparison, our framework is validated to be effective and practical. We also lay out the state-of-the-art performance revolving the bottom-up multi-person pose estimation on the MS-COCO dataset and the CrowdPose dataset. Lei Jin 0003, Xuecheng Nie, Luoqi Liu, Yandong Guo, Jian Zhao 0006 |
IEEE Trans. Multim. | 1 |
| 2022 | Single-Stage is Enough: Multi-Person Absolute 3D Pose EstimationabstractThe existing multi-person absolute 3D pose estimation methods are mainly based on two-stage paradigm, i.e., top-down or bottom-up, leading to redundant pipelines with high computation cost. We argue that it is more desirable to simplify such two-stage paradigm to a single-stage one to promote both efficiency and performance. To this end, we present an efficient single-stage solution, Decoupled Regression Model (DRM), with three distinct novelties. First, DRM introduces a new decoupled representation for 3D pose, which expresses the 2D pose in image plane and depth information of each 3D human instance via 2D center point (center of visible keypoints) and root point (denoted as pelvis), respectively. Second, to learn better feature representation for the human depth regression, DRM introduces a 2D Pose-guided Depth Query Module (PDQM) to extract the features in 2D pose regression branch, enabling the depth regression branch to perceive the scale information of instances. Third, DRM leverages a Decoupled Absolute Pose Loss (DAPL) to facilitate the absolute root depth and root-relative depth estimation, thus improving the accuracy of absolute 3D pose. Comprehensive experiments on challenging benchmarks including MuPoTS-3D and Panoptic clearly verify the superiority of our framework, which outperforms the state-of-the-art bottom-up absolute 3D pose estimation methods. Lei Jin 0003, Yabo Xiao, Yandong Guo, Xuecheng Nie, Jian Zhao 0006 |
CVPR | 1 |
| 2022 | QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level QueryabstractWe propose a sparse end-to-end multi-person pose regression framework, termed QueryPose, which can directly predict multi-person keypoint sequences from the input image. The existing end-to-end methods rely on dense representations to preserve the spatial detail and structure for precise keypoint localization. However, the dense paradigm introduces complex and redundant post-processes during inference. In our framework, each human instance is encoded by several learnable spatial-aware part-level queries associated with an instance-level query. First, we propose the Spatial Part Embedding Generation Module (SPEGM) that considers the local spatial attention mechanism to generate several spatial-sensitive part embeddings, which contain spatial details and structural information for enhancing the part-level queries. Second, we introduce the Selective Iteration Module (SIM) to adaptively update the sparse part-level queries via the generated spatial-sensitive part embeddings stage-by-stage. Based on the two proposed modules, the part-level queries are able to fully encode the spatial details and structural information for precise keypoint regression. With the bipartite matching, QueryPose avoids the hand-designed post-processes. Without bells and whistles, QueryPose surpasses the existing dense end-to-end methods with 73.6 AP on MS COCO mini-val set and 72.7 AP on CrowdPose test set. Code is available at https://github.com/buptxyb666/QueryPose. Yabo Xiao, Dongdong Yu, Lei Jin 0003, Mingshu He, Zehuan Yuan |
NeurIPS | 5 |
| 2021 | Deep-Feature-Based Autoencoder Network for Few-Shot Malicious Traffic DetectionabstractWith the increase of Internet visits and connections, it is becoming essential and arduous to protect the networks and different devices of the Internet of Things (IoT) from malicious attacks. The intrusion detection systems (IDSs) based on supervised machine learning (ML) methods require a large number of labeled samples. However, the number of abnormal behaviors is far less than that of normal behaviors, let alone that the shots of malicious behavior samples which can be intercepted as training dataset are actually limited. Consequently, it is a key research topic to conduct the anomaly detection for the small number of abnormal behavior samples. This paper proposes an anomaly detection model with a few abnormal samples to solve the problem in few-shot detection based on convolutional neural networks (CNN) and autoencoder (AE). This model mainly consists of the CNN-based supervised pretraining module and the AE-based data reconstruction module. Only a few abnormal samples are utilized to the pretrain module to build the structure of extracting deep features. The data reconstruction module simply chooses the deep features of normal samples as training data. There also exist some effective attention mechanisms in the pretraining module. Through the pretraining of small samples, the accuracy of abnormal detection is improved compared with merely training normal samples with AE. The simulation results prove that this solution can solve the above problems occurring in network behavior anomaly detection. In comparison to the original AE model and other clustering methods, the proposed model advances the detection results in a visible way. Mingshu He, Junhua Zhou, Yuanyuan Xi, Lei Jin 0003 |
Secur. Commun. Networks | 5 |
| 2018 | k-Trustee: Location injection attack-resilient anonymization for location privacy
Lei Jin 0003, Chao Li 0023, Balaji Palanisamy, James B. D. Joshi |
Comput. Secur. | 1 |
| 2016 | Characterizing users' check-in activities using their scores in a location-based social network
Lei Jin 0003, Xuelian Long, Ke Zhang 0013, Yu-Ru Lin, James B. D. Joshi |
Multim. Syst. | 1 |
| 2016 | Towards understanding the gamification upon users' scores in a location-based social network
Lei Jin 0003, Ke Zhang 0013, Jianfeng Lu 0002, Yu-Ru Lin |
Multim. Tools Appl. | 1 |
| 2015 | Towards complexity analysis of User Authorization Query problem in RBAC
Jianfeng Lu 0002, James B. D. Joshi, Lei Jin 0003 |
Comput. Secur. | 3 |
| 2014 | POSTER: Compromising Cloaking-based Location Privacy Preserving Mechanisms with Location Injection AttacksabstractCloaking-based location privacy preserving mechanisms have been widely adopted to protect users' location privacy while traveling on road networks. However, a fundamental limitation of such mechanisms is that users in the system are inherently trusted and assumed to always report their true locations. Such vulnerability can lead to a new class of attacks called location injection attacks which can successfully break users' anonymity among a set of users through the injection of fake user accounts and incorrect location updates. In this paper, we characterize location injection attacks, demonstrate their effectiveness through experiments on real-world geographic maps and discuss possible defense mechanisms to protect against such attacks. Lei Jin 0003, Balaji Palanisamy, James B. D. Joshi |
CCS | 1 |
| 2014 | On the complexity of role updating feasibility problem in RBAC
Jianfeng Lu 0002, Dewu Xu, Lei Jin 0003, Jianmin Han, Hao Peng 0002 |
Inf. Process. Lett. | 3 |
| 2013 | Understanding venue popularity in FoursquareabstractRecently, social media has become an increasingly important part of business and marketing. More and more businesses use social media as part of their marketing platforms. Moreover, the fast development of the 4th generation mobile network and the ubiquity of the advanced mobile devices in which GP Xuelian Long, Lei Jin 0003, James B. D. Joshi |
CollaborateCom | 2 |
| 2013 | Towards understanding traveler behavior in Location-based Social NetworksabstractUnderstanding users' behavior in Location-based Social Networks (LBSNs) is becoming an interesting research topic. In LBSNs, users can explore the places of interest around their current locations, check in at these locations and share such check-ins with their friends or the public. Therefore, the check-ins are valuable information for studying user behavior. Many services would benefit from the research of user behavior. For example, it can help the urban design and development based on the user mobility patterns and it could also improve the location recommendations to help users to find their places of interests. Intrinsically, traveler's activities in LBSN are distinctive, especially when compared with the a local user's activities. Therefore, a study of travelers' activities in LBSNs can help understand traveler behavior and then help LBSNs provider to improve their services, e.g. location recommendation service to new visitors. The location recommendation is especially important for a new visitor to a city. However, in the literature, there is little work specially focusing on the research of travelers' behavior in LBSNs. In this paper, we take the first step towards understanding such user behavior in LBSNs. Our research is based on the travelers' check-in information created in the greater Pittsburgh area in Foursquare. At first, we empirically study the venues and the check-ins created on such venues based on venue category information. After that, we investigate the temporal features of travelers' check-ins, and examine the evolution of check-ins created at the venues related to four categories using spatio-temporal information. Besides the empirical study, we employ the notion of user entropy to investigate the diversity of the travelers' check-ins. Through the research of the user entropy as a function of the user's check-ins, we find that the majority travelers usually exhibit higher diversity in their activities. Moreover, we also use the Latent Dirichlet Allocation (LDA) to generate travelers' mobility patterns. These human centric latent topics cannot only help to cluster the venues but also address the hot spots in a city based on the crowd level. Xuelian Long, Lei Jin 0003, James B. D. Joshi |
GLOBECOM | 2 |
| 2013 | Mutual-friend based attacks in social network systems
Lei Jin 0003, James B. D. Joshi, Mohd Anwar |
Comput. Secur. | 1 |
| 2012 | Exploring trajectory-driven local geographic topics in foursquareabstractThe location based social networking services (LBSNSs) are becoming very popular today. In LBSNSs, such as Foursquare, users can explore their places of interests around their current locations, check in at these places to share their locations with their friends, etc. These check-ins contain rich information and imply human mobility patterns; thus, they can greatly facilitate mining and analysis of local geographic topics driven by users' trajectories. The local geographic topics indicate the potential and intrinsic relations among the locations in accordance with users' trajectories. These relations are useful for users in both location and friend recommendations. In this paper, we focus on exploring the local geographic topics through check-ins in Pittsburgh area in Foursquare. We use the Latent Dirichlet Allocation (LDA) model to discover the local geographic topics from the checkins. We also compare the local geographic topics on weekdays with those at weekends. Our results show that LDA works well in finding the related places of interests. Xuelian Long, Lei Jin 0003, James B. D. Joshi |
UbiComp | 2 |
| 2011 | Towards active detection of identity clone attacks on online social networksabstractOnline social networks (OSNs) are becoming increasingly popular and Identity Clone Attacks (ICAs) that aim at creating fake identities for malicious purposes on OSNs are becoming a significantly growing concern. Such attacks severely affect the trust relationships a victim has built with other users if no active protection is applied. In this paper, we first analyze and characterize the behaviors of ICAs. Then we propose a detection framework that is focused on discovering suspicious identities and then validating them. Towards detecting suspicious identities, we propose two approaches based on attribute similarity and similarity of friend networks. The first approach addresses a simpler scenario where mutual friends in friend networks are considered; and the second one captures the scenario where similar friend identities are involved. We also present experimental results to demonstrate flexibility and effectiveness of the proposed approaches. Finally, we discuss some feasible solutions to validate suspicious identities. Lei Jin 0003, Hassan Takabi, James B. D. Joshi |
CODASPY | 1 |