EDBT 2026 Demo / reviewers in the wild / expert
Jian Xue 0002
dblp:21/628-2
· DBLP profile ↗
66ranked-venue papers
2as first author
47since 2021 · last 2026
0000-0002-9460-802XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 32 since 2021Artificial intelligence and machine learning · 12 · 11 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mambafusion: State-space model-driven object-scene fusion for multi-modal 3D object detection
Tong Ning, Ke Lu 0002, Jian Xue 0002 |
Pattern Recognit. | 4 |
| 2026 | CLSS: A Plug-and-Play Closed-Loop Semantic Supervisor for Actor-Critic Reinforcement Learning
Jian Xue 0002, Ke Lu 0002 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2026 | DreamAssemble: Complex Multi-Object Text-to-3D Generation via Multi-Density Neural Fields
Bin Huang 0016, Jinbao Wang 0001, Dongmei Jiang, Hongjuan Pei, Qiulu Li, Jian Xue 0002, Ke Lu 0002 |
IEEE Trans. Image Process. | 6 |
| 2026 | Improving Unsupervised Ultrasonic Image Anomaly Detection via Frequency-Spatial Feature Filtering and Gaussian Mixture ModelingabstractUltrasonic image anomaly detection faces significant challenges due to limited labeled data, strong structural and random noise, and highly diverse defect manifestations. To overcome these obstacles, we introduce UltraChip, a new large-scale C-scan benchmark containing about 8,000 real-world images from various chip packaging types, each meticulously annotated with pixel-level masks for cracks, holes, and layers. Building on this resource, we present FSGM-Net, a fully unsupervised framework tailored for anomaly detection. FSGM-Net leverages an adaptive Frequency-Spatial feature filtering mechanism: a learnable FFT-Spatial patch filter first suppresses noise and dynamically assigns normality weights to Vision Transformer (ViT) patch features. Subsequently, an Adaptive Gaussian Mixture Model (Ada-GMM) captures the distribution of normal features and guides a deep-shallow multi-scale interaction decoder for accurate, pixel-level anomaly inference. In addition, we propose a filter loss that enforces encoder-filter consistency and entropy-based sparse gating, together with a distributional loss that encourages both feature reconstruction and confident Gaussian mixture modeling. Extensive experiments demonstrate that FSGM-Net not only achieves state-of-the-art results on UltraChip but also exhibits superior cross-domain generalization to MVTec-AD and VisA, while supporting real-time inference on a single GPU. Together, the dataset and framework advance robust, annotation-free ultrasonic NDT in practical applications. The UltraChip dataset can be obtained via https://iiplab.net/ultrachip/. Ke Lu 0002, Jinbao Wang 0001, Can Gao, Jian Xue 0002 |
IEEE Trans. Image Process. | 6 |
| 2026 | MotionFlow: Efficient Motion Generation With Latent Flow MatchingabstractIn the field of human centric multimedia, text-driven human motion generation is a significant pursuit with wide-ranging applications across diverse scenarios. Despite substantial advancements, existing methods often suffer from a trade-off between inference latency and high-quality generation. To overcome this gap, we propose the Motion Latent Flow Matching model (MotionFlow), a novel and powerful framework for motion generation. It introduces flow matching algorithm in the latent space, which can achieve superior performance with just one-step inference. In addition to the text-driven task, we further extend our method to controllable motion generation. Specifically, we integrate a control encoder into the latent space and further decode the predicted latent code into motion space to support explicit supervision, ensuring the synthesized motion can tightly align with the input signals. Extensive experiments demonstrate that our MotionFlow not only outperforms current leading approaches for the text-driven task, but also delivers remarkable capabilities in controllable motion generation. Kun Dong 0001, Jian Xue 0002, Xing Lan, Qingyuan Liu 0001, Ke Lu 0002 |
IEEE Trans. Multim. | 2 |
| 2025 | Learning Hierarchical Continuous Dynamics for Facial Action Unit Intensity EstimationabstractDynamic facial action recognition is key to understanding human emotions and behaviors, yet estimating facial action units (AUs) intensities in videos is difficult due to subtle muscle motions and complex spatial-temporal dependencies. Existing methods often use fixed or coarse graphs, limiting the ability to capture intricate AU relations and long-range dynamics. This paper presents a hierarchical framework CDAU to effectively capture Continuous Dynamics for AU intensity estimation task. Our approach dynamically constructs multiscale graphs for fine-grained spatiotemporal AU interactions and adaptively fusing information across levels. A bidirectional state-space module further captures long-range temporal dependencies. Extensive experiments on FEAFA and DISFA show that CD-AU outperforms existing methods in both ICC and MAE metrics, validating its generalization capability and stability across subjects and expressions. Ke Lu 0002, Yan Li 0121, Menghao Hu, Guohong Hu, Dongmei Jiang, Jian Xue 0002 |
BIBM | 7 |
| 2025 | Text-to-Any-Skeleton Motion Generation Without Retargeting
Qingyuan Liu 0001, Ke Lu 0002, Kun Dong 0001, Jian Xue 0002, Zehai Niu, Jinbao Wang 0001 |
ICCV | 4 |
| 2025 | Continuous Action Unit Intensity Modeling for Micro-Expression RecognitionabstractMicro-Expression Recognition (MER) remains challenging due to the subtle and transient nature of facial muscle movements. While recent methods leverage Action Unit (AU) labels for MER, they often tend to ignore continuous AU intensity variations, which are critical for capturing nuanced facial expressions. To address these limitations, we propose a novel framework integrating continuous AU intensity with hierarchical motion modeling. Our approach begins with a lightweight model that regresses in-frame AU intensity values. These AU intensities are fed into our proposed Continuous AU Transformer (CAUT), which employs a temporal Transformer and a spatial Transformer to model AU evolution across frames and inter-AU dependencies. Simultaneously, a two-stage Transformer architecture extracts hierarchical optical flow features, fused with AU semantics via a multi-scale region-based fusion strategy for enhancing facial motion features. Extensive experiments demonstrate the proposed method’s state-of-the-art performance, validating the effectiveness of continuous AU intensity modeling and hierarchical feature integration for MER. Hanyu Jiang 0004, Jiayi Lyu, Xing Lan, Jian Xue 0002 |
ICIP | 4 |
| 2025 | One General Plug-In for Facial Heatmap-based Keypoint DetectionabstractIn this paper, we systematically investigate the error distribution in predicted heatmaps for face alignment, and point out that previous works are unreliable in following the rule that decodes coordinates by locating the maximum-score pixel. Our research reveals that the majority of ground-truth positions do not match that pixel but rather lie within a range of a few pixels. Building on this phenomenon, we transform the model’s objective from predicting inaccurate landmarks to identifying precise proposals with that range. We propose a simple but effective module, termed the Response Aware Module (RAM), leveraging response scores in the proposal to regress the proposal offset, which can be used as a plug-and-play layer integrated into public models. Furthermore, we present a novel Heatmap RCNN framework to exploit the distribution of multi-scale heatmaps. Extensive experiments have demonstrated that the trained RAM can be integrated seamlessly as a ready-to-use plugin with the model, yielding impressive improvements. Meanwhile, Heatmap RCNN performs far superior to SOTA results, with 3.82 NME on WFLW, 3.09 on COFW, and 2.90 on 300W. Hanyu Jiang 0004, Jian Xue 0002, Xing Lan, Ke Lu 0002 |
ICME | 2 |
| 2025 | DC-BEV: Depth-Completed Bird's Eye View Representation for Multi-Modal 3D Object DetectionabstractThe bird’s eye view (BEV) representation is essential for accurate 3D perception tasks (e.g., 3D object detection) in autonomous driving for its precise localization, scale consistency, and modality independence. However, traditional depth-prediction-based methods, which rely on predicted depth distributions derived from semantic image features for BEV transformation, face challenges such as ambiguous depth-prior information and dependence on predefined depth distribution types. To address these limitations, we propose a novel multimodal 3D object detection method, Depth Completed BEV (DC-BEV), which leverages ground-truth sparse LiDAR depth to guide BEV transformations, significantly enhancing depth estimation accuracy. Specifically, we introduce a Multimodal-Depth Completion (MDC) mechanism, which enriches sparse LiDAR depth into dense depth maps by integrating semantic and geometric cues from images. Additionally, to mitigate gradient instability caused by inadequate implicit supervision, we present an Explicit Depth Supervision (EDS) mechanism that directly supervises depth predictions using a dedicated depth loss. Comprehensive experiments conducted on the nuScenes dataset demonstrate that DC-BEV achieves superior performance, notably improving detection accuracy through enhanced depth estimation quality and robust BEV representation. Tong Ning, Ke Lu 0002, Jian Xue 0002 |
SMC | 4 |
| 2025 | Learning Multi-Scale Spatial Features Representation in Frequency Domain for Gait RecognitionabstractGait recognition is a highly promising biometric technology because of its robust performance in long-distance scenarios. However, existing methods typically rely on geometric approaches to extract local and global features, which may lack the accuracy required for truly discriminative representations. To overcome this problem, we propose a frequency-based channel attention mechanism that captures both global and local information in the frequency domain, where the signal preserves a rich set of latent details that can boost the recognition accuracy. Furthermore, assembling large-scale gait datasets remains prohibitively expensive, constraining research progress. To alleviate this issue, we introduce a simple yet effective data augmentation strategy, where each raw gait image is horizontally split into two parts, which are then randomly recombined to create new synthetic identities. Extensive experiments on the FVG and CASIA-B datasets demonstrate that our method achieves competitive performance. Tong Ning, Jian Xue 0002 |
SMC | 3 |
| 2025 | Adaptive Multi-Layer Prioritized Fictitious Self-play in Multi-Agent Reinforcement LearningabstractIn the field of Multi-Agent Reinforcement Learning (MARL), strategy selection and optimization are key challenges in improving agent performance. The Policy Space Response Oracles (PSRO) algorithm is widely used in MARL, and the meta-solver is one of its cores. However, existing meta-solvers may exhibit limitations in complex MARL environments, such as significant consumption of computing resources, poor convergence, instability, and so on. Therefore, an adaptive multi-layer Prioritized Fictitious Self-Play (PFSP) is proposed in this paper as a meta-solver method for the PSRO algorithm, which further improves the effectiveness of strategy optimization by utilizing the game results of the meta-game in a more efficient and reasonable way. Adaptive multi-layer PFSP can flexibly make strategy choices based on payoff at different layers, thus overcoming the shortcomings of traditional meta-solvers in dealing with complex strategy spaces. The experimental results show that the proposed method significantly improves the convergence speed and performance of strategies in complex MARL environments, especially in the training and testing of the Google Research Football environment, demonstrating its potential and advantages in practical applications. Jian Xue 0002, Ke Lu 0002 |
SMC | 2 |
| 2025 | Multimodal Emotional Talking Face Generation Based on Action UnitsabstractTalking face generation focuses on creating natural facial animations that align with the provided text or audio input. Current methods in this field primarily rely on facial landmarks to convey emotional changes. However, spatial key-points are valuable, yet limited in capturing the intricate dynamics and subtle nuances of emotional expressions due to their restricted spatial coverage. Consequently, this reliance on sparse landmarks can result in decreased accuracy and visual quality, especially when representing complex emotional states. To address this issue, we propose a novel method called Emotional Talking with Action Unit (ETAU), which seamlessly integrates facial Action Units (AUs) into the generation process. Unlike previous works that solely rely on facial landmarks, ETAU employs both Action Units and landmarks to comprehensively represent facial expressions through interpretable representations. Our method provides a detailed and dynamic representation of emotions by capturing the complex interactions among facial muscle movements. Moreover, ETAU adopts a multi-modal strategy by seamlessly integrating emotion prompts, driving videos, and target images, and by leveraging various input data effectively, it generates highly realistic and emotional talking-face videos. Through extensive evaluations across multiple datasets, including MEAD, LRW, GRID and HDTF, ETAU outperforms previous methods, showcasing its superior ability to generate high-quality, expressive talking faces with improved visual fidelity and synchronization. Moreover, ETAU exhibits a significant improvement on the emotion accuracy of the generated results, reaching an impressive average accuracy of 84% on the MEAD dataset. Jiayi Lyu, Xing Lan, Guohong Hu, Hanyu Jiang 0004, Jinbao Wang 0001, Jian Xue 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | DinoQuery: Promoting Small 3D Object Detection With Textual PromptabstractQuery-based 3D object detection has gained significant success in the application of autonomous driving due to its ability to achieve good performance while maintaining low computational cost. However, it still struggles with the reliable detection of small objects such as bicycles and pedestrians. To address this challenge, this paper introduces a novel sparse query-based approach, termed DinoQuery. This approach utilizes Grounding-DINO with textual prompts to select small-sized objects and generate 2D category-aware queries. These 2D category-aware queries combined with 2D global queries are then lifted to 3D queries by associating each sampled query with its respective 3D position, orientation, and size. The validity of these 3D queries, along with the 2D queries, is verified by the Comprehensive Contrastive Learning (CCL) mechanism. This is achieved by aligning all 2D and 3D queries with their respective 2D and 3D ground truth labels, and computing similarity to select true positive and false positive queries. Then a contrastive loss is introduced to enhance true positive queries and weaken false positive ones based on geometric and semantic similarity. The DinoQuery was tested on the nuScenes dataset and demonstrated excellent performance. Notably, the largest increase of our method is 3.2% on NDS and 3.1% on mAP. Tong Ning, Ke Lu 0002, Hongjuan Pei, Jian Xue 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | FoodSAM: Any Food SegmentationabstractIn this paper, we explore the zero-shot capability of the Segment Anything Model (SAM) for food image segmentation. To address the lack of class-specific information in SAM-generated masks, we propose a novel framework, calledFoodSAM. This innovative approach integrates the coarse semantic mask with SAM-generated masks to enhance semantic segmentation quality. Besides, we recognize that the ingredients in food can be supposed as independent individuals, which motivated us to perform instance segmentation on food images. Furthermore, FoodSAM extends its zero-shot capability to encompass panoptic segmentation by incorporating an object detector, which renders FoodSAM to effectively capture non-food object information. Drawing inspiration from the recent success of promptable segmentation, we also extend FoodSAM to promptable segmentation, supporting various prompt variants. Consequently, FoodSAM emerges as an all-encompassing solution capable of segmenting food items at multiple levels of granularity. Remarkably, this pioneering framework stands as the first-ever work to achieve instance, panoptic, and promptable segmentation on food images. Extensive experiments demonstrate the feasibility and impressing performance of FoodSAM, validating SAM's potential as a prominent and influential tool within the domain of food image segmentation. Xing Lan, Jiayi Lyu, Hanyu Jiang 0004, Kun Dong 0001, Zehai Niu, Yi Zhang 0162, Jian Xue 0002 |
IEEE Trans. Multim. | 7 |
| 2025 | ExpLLM: Towards Chain of Thought for Facial Expression RecognitionabstractFacial expression recognition (FER) is a critical task in multimedia with significant implications across various domains. However, analyzing the causes of facial expressions is essential for accurately recognizing them. Current approaches, such as those based on facial action units (AUs), typically provide AU names and intensities but lack insight into the interactions and relationships between AUs and the overall expression. In this paper, we propose a novel method called ExpLLM, which leverages large language models to generate an accurate chain of thought (CoT) for facial expression recognition. Specifically, we have designed the CoT mechanism from three key perspectives: key observations, overall emotional interpretation, and conclusion. The key observations describe the AU's name, intensity, and associated emotions. The overall emotional interpretation provides an analysis based on multiple AUs and their interactions, identifying the dominant emotions and their relationships. Finally, the conclusion presents the final expression label derived from the preceding analysis. Furthermore, we also introduce the Exp-CoT Engine, designed to construct this expression CoT and generate instruction-description data for training our ExpLLM. Extensive experiments on the RAF-DB and AffectNet datasets demonstrate that ExpLLM outperforms current state-of-the-art FER methods. ExpLLM also surpasses the latest GPT-4o in expression CoT generation, particularly in recognizing micro-expressions where GPT-4o frequently fails. Xing Lan, Jian Xue 0002, Ji Qi 0003, Dongmei Jiang, Ke Lu 0002, Tat-Seng Chua |
IEEE Trans. Multim. | 2 |
| 2025 | Beyond 3D: Generic IoU for 3D Object DetectionabstractObject detection from point clouds is a fundamental task for 3D scene understanding and has a wide range of applications in the field of multimedia data processing and analysis, such as autonomous driving and virtual interaction. The IoU evaluates the overlap between the two bounding boxes to ensure consistency across network optimization and testing, becoming a recognized regression loss in the field of 3D object detection. However, there is a kind of error coupling between the IoU and the angle, i.e., the IoU does not decrease as the angle error increases and vice versa. This problem leads to sub-optimal solutions for the neural network model, which severely hampers the improvement of 3D object detection accuracy. In this paper, a novel 4DIoU method is introduced for detecting 3D objects from point clouds, which provides a comprehensive rethinking of IoU computation by integrating angular information as an additional dimension. 4DIoU not only solves the problem of error coupling between IoU and angular but also facilitates neural network optimization using angle information. Furthermore, to solve the different impacts of various object shapes on IoU variations, a special 4DIoU called TV4DIoU is proposed to fuse shape information based on three orthogonal projection views, which can adaptively learn the information of objects with different shapes. In addition, to enhance the generalization of the 4DIoU method, a high-flexibility anchor encoding method and a cyclic consistent computation formula for angular errors are designed to make 4DIoU a plug-and-play module for both anchor-based and anchor-free frameworks. Extensive evaluations conducted on the nuScenes, Waymo, and KITTI datasets have confirmed the effectiveness of the proposed method. Hengsheng Lun, Ke Lu 0002, Liping Hou, Jian Xue 0002 |
IEEE Trans. Multim. | 5 |
| 2025 | DA-Net: Density-Aware 3D Object Detection Network for Point Cloudsabstract3D object detection is an important but demanding task, which has become an active research topic in the field of multimedia. Much recent research has been devoted to exploiting end-to-end trainable object detection networks with point clouds. However, most state-of-the-art methods have bottlenecks in detecting occluded objects and small objects, because the sparseness of point clouds is exacerbated on these objects. In this paper, a Density-Aware 3D object detection network (DA-Net) is proposed to improve the perception performance for detecting occluded and small objects, which contains four components: a backbone module with an inverse density scoring module (IDM) and a point-wise attention module (PAM), a 3D intersection over union Estimation Module (3DEM), a Consistent Label Assignment (CLA) method and an Adaptive-Soft-NMS method. The proposed backbone module makes the network concentrate on low-density points of occluded objects, and suppresses outliers and background points. Then, the 3DEM is introduced to evaluate the localization quality of the prediction boxes. Furthermore, the proposed CLA method can more accurately select positive and negative samples for small objects. Finally, Adaptive-Soft-NMS is proposed in our method to reduce the number of false detections during inference and thereby improve detection performance substantially. Extensive experiments demonstrated that the proposed method achieves state-of-the-art performance on two large-scale datasets, SUN RGB-D (62.1% in terms of [email protected]) and ScanNetV2 (67.1% in terms of [email protected]), and in particular, the detection accuracy of small objects and occluded objects are extremely improved. Ke Lu 0002, Jian Xue 0002, Yang Zhao 0028 |
IEEE Trans. Multim. | 3 |
| 2024 | SRA-YOLO: Spatial Resolution Adaptive YOLO for Semi-supervised Cross-Domain Aerial Object Detection
Jian Xue 0002, Yuqiu Li, Ke Lu 0002 |
ICANN (2) | 2 |
| 2024 | 3Dlaneformer: Rethinking Learning Views for 3D Lane DetectionabstractAccurate 3D lane detection from monocular images is crucial for autonomous driving. Recent advances leverage either front-view (FV) or bird’s-eye-view (BEV) features for prediction, inevitably limiting their ability to perceive driving environments precisely and resulting in suboptimal performance. To overcome the limitations of using features from a single view, we design a novel dual-view cross-attention mechanism, which leverages features from FV and BEV simultaneously. Based on this mechanism, we propose 3DLaneFormer, a powerful framework for 3D lane detection. It outperforms the latest BEV-based or FV-based approaches through extensive experiments on challenging benchmarks and thus verifies the necessity and benefits of utilizing features in both views. Kun Dong 0001, Jian Xue 0002, Xing Lan, Ke Lu 0002 |
ICIP | 2 |
| 2024 | From 3D to 4D: Fixing the Erroneous Coupling between IoU and Angle for Optimizing 3D Object DetectionabstractThe IoU metric directly measures the overlap between two boxes, maintaining consistency in model optimization and testing stages. It has emerged as a highly regarded regression loss in the field of 3D object detection. However, the optimization of IoU often leads to an increased angular error. This erroneous coupling phenomenon renders the model susceptible to settling into sub-optimal solutions, which have not been extensively analyzed and addressed, significantly impeding further advancements in the accuracy of 3D object detection. In this paper, a novel concept "4DIoU" is introduced for 3D object detection, where the angle information is integrated as an additional dimension in the IoU calculation, and a new formula for measuring angle correlation is proposed. The 4DIoU not only resolves the erroneous coupling between IoU and angles but also capitalizes on angle information to enhance network optimization. Furthermore, a new encoding and decoding paradigm is proposed, which is more compatible with 4DIoU for object detection in point clouds. Extensive experiments on nuScenes, Waymo and KITTI datasets demonstrate the effectiveness of our method. The plug-and-play design of our approach proves to be highly versatile. Hengsheng Lun, Ke Lu 0002, Liping Hou, Jian Xue 0002 |
ICME | 5 |
| 2024 | ETAU: Towards Emotional Talking Head Generation Via Facial Action UnitabstractCreating expressive talking heads is crucial for multimedia applications involving virtual human. Existing approaches predominantly rely on facial landmarks to convey emotional changes. However, these spatial keypoints struggle to capture subtle emotional intricacies due to their limited spatial coverage, consequently decreasing accuracy and visual quality, particularly in emotion representation. To address this issue, we introduce a novel method called Emotional Talking with Action Unit (ETAU), which introduces the additional facial Action Units (AUs) to generate talking head video that accurately portray the target emotions. Unlike previous works, ETAU comprehensively quantify facial expressions through Action Units, which provides a detailed and dynamic representation of emotion. To the best of our knowledge, this work pioneers the integration of Action Units for emotional talking head generation. Extensive evaluations on the MEAD dataset showcase ETAU’s state-of-the-art performance with 21.89 PSNR and 0.68 SSIM. Critically, ETAU achieves significant improvement in emotion accuracy of the generated results, reaching 84%, confirming its feasibility in representing emotional expressions. Jiayi Lyu, Xing Lan, Guohong Hu, Hanyu Jiang 0004, Jian Xue 0002 |
ICME | 6 |
| 2024 | VS3D: A Vote-Based Semi-Supervised 3D Object Detection Framework for Point CloudsabstractIn recent years, the 3D object detection method has undergone rapid evolution, heavily relying on substantial amounts of high-quality labeled data. However, the process of annotating 3D data is both time-consuming and costly. In response to this challenge, we propose a vote-based semi-supervised 3D object detection framework called VS3D. First, a data augmentation technique named Random Grid Deleting (RGD) is proposed to detect occluded objects and small objects more robustly. Then, an auxiliary branch with Voting Consistency Learning (VCL) is added to predict object centers more accurately. Additionally, a Teacher-Student Matching (TSM) module with stricter consistency constraints is designed to accelerate network convergence and improve detection performance. Our method can integrate any vote-based fully supervised network seamlessly. Extensive experiments on SUN RGB-D and ScanNet V2 datasets demonstrate that the proposed method outperforms the state-of-the-art fully supervised model when using only 70% labeled data. Ke Lu 0002, Yang Zhao 0028, Hengsheng Lun, Zehai Niu, Jian Xue 0002 |
ICME | 6 |
| 2024 | MISTA: A Large-Scale Dataset for Multi-Modal Instruction Tuning on Aerial ImagesabstractThis paper introduces MISTA, a novel dataset for visual instruction tuning on aerial imagery, designed to enhance large multi-modal model applications in remote sensing. Originating from the renowned DOTA-v2.0 aerial object detection benchmark, MISTA uniformly processes high-resolution images into 2048×2048 pixels, creating a detailed and complex dataset tailored for remote sensing analysis. To craft this dataset, we design an automated annotation pipeline, employing advanced language models such as GPT-4 and LLaVA-1.5, to generate diverse and specialized instruction-following data. The annotations include various instruction types like multi-turn conversation, detailed description, and complex reasoning, each reflecting the intricacies inherent in remote sensing tasks. The innovative approach of subdividing aerial images into individually annotated sub-patches significantly enhances the richness of the dataset and allows for a more granular analysis of visual content. As a robust foundation for multi-modal model development in remote sensing, MISTA represents a significant advancement, setting the stage for future research and further applications in the field. Ke Lu 0002, Yuqiu Li, Jian Xue 0002 |
ICME | 5 |
| 2024 | MA-Mamba: Multi-Agent Reinforcement Learning with State Space Model
Jian Xue 0002, Ke Lu 0002 |
ICONIP (3) | 2 |
| 2024 | Real-time Integration of Fine-tuned Large Language Model for Improved Decision-Making in Reinforcement LearningabstractIn this paper, we investigate a novel and efficacious methodology for training reinforcement learning agents. We also further reveal the potential of the Large Language Model (LLM) in intricate decision-making environments. Reinforcement learning, one main approach in training decision-making capabilities for intelligent agents, suffers from the challenges of sparse rewards and inefficient exploration. Considering the current phenomenal performance of pre-trained LLMs, certain studies have adopted the LLMs to shepherd the actions of intelligent agents. However, there are some limitations in the application of LLMs, such as the lack of domain knowledge of LLMs for specific fields, and LLMs generally access a slower response time but consume a higher economic cost. Embarking from these constraints, this paper takes on the formidable challenge of the Unmanned Air Vehicle (UAV) air combat simulations environment, where decision-making is notably circumscribed by temporal limitations. We initially fine-tuned the LLMs to a domain-specific expertise while concurrently constructing a knowledge base. Hence, upon gaining profound insight into the field of air combat, it evolves into an adept model capable of effectively guiding UAV decisions. Further addressing the difficulties of real-time accessing the LLM, we propose a novel approach named Reward Shaping with Large Language Model (LLM-RS) to augment the autonomous decision-making competency of UAVs within the context of air combat simulations. We compared the agents trained by conventional reinforcement learning techniques and those trained by LLM-RS. The experimental results reveal that, under the same equipment training conditions, the LLM-RS technique grounded in the fine-tuning of the LLM and the knowledge base substantially improves the performance of UAVs in air combat simulations, while concurrently diminishing the requisite training duration. Xiancai Xiang, Jian Xue 0002, Ke Lu 0002 |
IJCNN | 2 |
| 2024 | Realistic Full-Body Motion Generation from Sparse Tracking with State Space ModelabstractIn the domain of generative multimedia and interactive experiences, generating realistic and accurate full-body poses from sparse tracking is crucial for many real-world applications, while achieving sequence modeling and efficient motion generation remains challenging. Recently, state space models (SSMs) with efficient hardware-aware designs (i.e., Mamba) have shown great potential for sequence modeling, particularly in temporal contexts. However, processing motion data is still challenging for SSMs. Specifically, the sparsity of input conditions makes motion generation an ill-posed problem. Moreover, the complex structure of the human body further complicates this task. To address these issues, we present Motion Mamba Diffusion (MMD), a novel conditional diffusion model, which effectively utilizes the sequence modeling capability of SSMs and the robust generation ability of diffusion models to track full-body poses accurately. In particular, we design a bidirectional Temporal Mamba Module (TMM) to model motion sequence. Additionally, a Spatial Mamba Module (SMM) is further proposed for feature enhancement within a single frame. Extensive experiments on the large motion capture dataset (AMASS) demonstrate that our proposed approach outperforms the latest methods in terms of accuracy and smoothness, thus providing a crucial advancement for creating realistic virtual avatars in various applications. Kun Dong 0001, Jian Xue 0002, Zehai Niu, Xing Lan, Ke Lu 0002, Qingyuan Liu 0001, Xiaoyu Qin 0001 |
ACM Multimedia | 2 |
| 2024 | Skeleton Cluster Tracking for robust multi-view multi-person 3D human pose estimation
Zehai Niu, Ke Lu 0002, Jian Xue 0002, Jinbao Wang 0001 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Dynamic spatial-temporal topology graph network for skeleton-based action recognition
Lian Chen, Ke Lu 0002, Zehai Niu, Runchen Wei, Jian Xue 0002 |
Multim. Syst. | 5 |
| 2024 | Does Pixel Value Represent Facial Landmark Well in Heatmap?abstractHeatmap-based methods have dominated the face alignment task, yet the maximum response decoding scheme necessitates further reform. While some studies have attempted to compensate for prediction offsets using a post-processing module, the prediction errors induced by the maximum response decoding scheme remain challenging to rectify. In this paper, we assume that using heatmap value to denote the ground-truth probability is not accurate enough. To cure this problem, we propose DISPAL, a novel DIStribution-based Probability for fAcial Landmarks, which signifies the ground-truth probability by the similarity between the pixel’s neighbouring value distribution and Gaussian distribution. This innovative probability enables us to pinpoint the keypoint location more robustly than previous methods that rely solely on the peak score. It also exhibits remarkable generalization to complex decoding methodologies. Furthermore, we propose supervising this probability as an additional task loss to help the model learn better heatmap representation. Extensive empirical results on WFLW, 300W, and COFW datasets demonstrate that our distribution-based probability mechanism significantly surpasses original value-based probability approaches. Xing Lan, Jiayi Lyu, Kun Dong 0001, Hanyu Jiang 0004, Qinghao Hu 0001, Jian Xue 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | From Methods to Applications: A Review of Deep 3D Human Motion CaptureabstractMotion capture technology is crucial in various applications like animation, virtual reality and sports analysis. With the development of deep learning methods, significant progress has been experienced in this field, producing cost-effective and user-friendly solutions for various applications. This paper provides a comprehensive review of deep learning-based human motion capture techniques. Our review aims to bridge the gap between academic research and practical applications, providing valuable insights and guidance for researchers and practitioners in deep learning-based human motion capture. Our study puts forth a new application-oriented taxonomy that comprehensively summarises five fundamental routes of motion capture technology. In addition to that, we also delve into the research priorities linked with each route, following the structure of “hardware requirements - technical routes - datasets - evaluation metrics” and extending the necessary criteria for transferring traditional motion capture systems to deep learning-based ones. Meanwhile, for the motion capture technology, the current state of the art is reviewed, the challenges are identified, and the future directions of the research are outlined. Zehai Niu, Ke Lu 0002, Jian Xue 0002, Xiaoyu Qin 0001, Jinbao Wang 0001, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | BiUNet: Towards More Effective UNet with Bi-Level Routing Attention
Kun Dong 0001, Jian Xue 0002, Xing Lan, Ke Lu 0002 |
BMVC | 2 |
| 2023 | Semi-Supervised Contrastive Learning of Global and Local Representation for 3d Medical Image SegmentationabstractAlthough the application of supervised deep learning in medical image analysis is still very successful, it mainly depends on the quantity and quality of labeled data; and it is time-consuming and labor-intensive to obtain 3D medical image annotation. Recently, contrastive learning has shown its remarkable ability for self-supervised learning and has achieved impressive results on many downstream tasks. In this study, we extend the popular contrastive learning medical image segmentation framework to 3D and design extra reconstruction loss for volumetric medical images to improve the performance of global contrastive learning. We evaluate our method on two public 3D medical image datasets of different modalities. Our proposed method achieves competitive results compared to other methods for different proportions of labeled data. Chuang Jia, Jian Xue 0002, Ke Lu 0002, Zhongqi Wu |
ICIP | 2 |
| 2023 | A Novel Cross-Fusion Method of Different Types of Features for Image CaptioningabstractMulti-modal tasks are receiving more and more attention, including image captioning. Based on X-Linear attention, we simultaneously introduce grid features and region features extracted by Faster RCNN. We obtain a global feature vector of each type of original features through mean pooling. The two types of features are encoded by two parallel encoders. Each encoder has two inputs: a set of feature vectors (region/grid) and the corresponding global feature vector. Each encoding layer outputs an encoded global feature vector and a set of encoded feature vectors. We cross-fuse the global feature vector output by each encoding layer for region features and the set of encoded feature vectors for grid features. In the same way, we cross-fuse another pair of the global feature (grid) and the set of encoded feature vectors (region). Finally, we fuse the two global feature vectors output by the two encoders as the final global features, and the two sets of encoded feature vectors output by the two encoders as the final visual features. Experimental results on the COCO dataset show that our model achieves a new SOTA performance of BLEU-1 81.5%, BLEU-4 40.5%, METEOR 29.6%, and ROUGE 59.5% on the Karpathy test split. Liangshan Lou, Ke Lu 0002, Jian Xue 0002 |
IJCNN | 3 |
| 2023 | Semi-Supervised Medical Image Segmentation With Voxel Stability and Reliability ConstraintsabstractSemi-supervised learning is becoming an effective solution in medical image segmentation because annotations are costly and tedious to acquire. Methods based on the teacher-student model use consistency regularization and uncertainty estimation and have shown good potential in dealing with limited annotated data. Nevertheless, the existing teacher-student model is seriously limited by the exponential moving average algorithm, which leads to the optimization trap. Moreover, the classic uncertainty estimation method calculates the global uncertainty for images but does not consider local region-level uncertainty, which is unsuitable for medical images with blurry regions. In this article, the Voxel Stability and Reliability Constraint (VSRC) model is proposed to address these issues. Specifically, the Voxel Stability Constraint (VSC) strategy is introduced to optimize parameters and exchange effective knowledge between two independent initialized models, which can break through the performance bottleneck and avoid model collapse. Moreover, a new uncertainty estimation strategy, the Voxel Reliability Constraint (VRC), is proposed for use in our semi-supervised model to consider the uncertainty at the local region level. We further extend our model to auxiliary tasks and propose a task-level consistency regularization with uncertainty estimation. Extensive experiments on two 3D medical image datasets demonstrate that our method outperforms other state-of-the-art semi-supervised medical image segmentation methods under limited supervision. Yang Zhao 0028, Ke Lu 0002, Jian Xue 0002 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | A Facial Landmark Detection Method Based on Deep Knowledge TransferabstractFacial landmark detection is a crucial preprocessing step in many applications that process facial images. Deep-learning-based methods have become mainstream and achieved outstanding performance in facial landmark detection. However, accurate models typically have a large number of parameters, which results in high computational complexity and execution time. A simple but effective facial landmark detection model that achieves a balance between accuracy and speed is crucial. To achieve this, a lightweight, efficient, and effective model is proposed called the efficient face alignment network (EfficientFAN) in this article. EfficientFAN adopts the encoder-decoder structure, with a simple backbone EfficientNet-B0 as the encoder and three upsampling layers and convolutional layers as the decoder. Moreover, deep dark knowledge is extracted through feature-aligned distillation and patch similarity distillation on the teacher network, which contains pixel distribution information in the feature space and multiscale structural information in the affinity space of feature maps. The accuracy of EfficientFAN is further improved after it absorbs dark knowledge. Extensive experimental results on public datasets, including 300 Faces in the Wild (300W), Wider Facial Landmarks in the Wild (WFLW), and Caltech Occluded Faces in the Wild (COFW), demonstrate the superiority of EfficientFAN over state-of-the-art methods. Ke Lu 0002, Jian Xue 0002, Jiayi Lyu, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Collision-Risk-Based Event-Triggered Optimal Formation Control for Mobile Multiagent Systems Under Incomplete Information ConditionsabstractThis article deals with collision-risk-based event-triggered optimal formation control problems for mobile multiagent systems. First, several collision-risk-related definitions, such as collision-free margin, moving direction angle, collision risk angle, and collision risk level, are proposed for the moving agents. Then, a collision-risk dependent, time-varying, event-triggered heterogeneous communication network topology is developed, where the agent starts obtaining information of the neighboring agents only when collision risks occur among them. Third, an anti-collision control law, which is composed of a switch function, a control force direction function, and a control strength function, is designed to guarantee the collision avoidance formation of multiagents. Fourth, to ensure the formation quality and save control cost of the multiagent system, an optimal formation control scheme with feedforward compensation is designed. Simulation results illustrate that: 1) by using the collision risk information of mobile agents, the proposed control scheme is effective to realize the collision avoidance optimal formation task and 2) the anti-collision formation controller can be implemented with incomplete information of the agents. Bao-Lin Zhang 0001, Jin Zhou 0003, Jian Xue 0002, Yuanshi Zheng |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | Shape-Adaptive Selection and Measurement for Oriented Object DetectionabstractThe development of detection methods for oriented object detection remains a challenging task. A considerable obstacle is the wide variation in the shape (e.g., aspect ratio) of objects. Sample selection in general object detection has been widely studied as it plays a crucial role in the performance of the detection method and has achieved great progress. However, existing sample selection strategies still overlook some issues: (1) most of them ignore the object shape information; (2) they do not make a potential distinction between selected positive samples; and (3) some of them can only be applied to either anchor-free or anchor-based methods and cannot be used for both of them simultaneously. In this paper, we propose novel flexible shape-adaptive selection (SA-S) and shape-adaptive measurement (SA-M) strategies for oriented object detection, which comprise an SA-S strategy for sample selection and SA-M strategy for the quality estimation of positive samples. Specifically, the SA-S strategy dynamically selects samples according to the shape information and characteristics distribution of objects. The SA-M strategy measures the localization potential and adds quality information on the selected positive samples. The experimental results on both anchor-free and anchor-based baselines and four publicly available oriented datasets (DOTA, HRSC2016, UCAS-AOD, and ICDAR2015) demonstrate the effectiveness of the proposed method. Liping Hou, Ke Lu 0002, Jian Xue 0002, Yuqiu Li |
AAAI | 3 |
| 2022 | Vote-Based Multi-Level Context Attention Network for 3D Point Cloud Object Detectionabstract3D object detection is a challenging task because point clouds are characterized by sparsity and irregularity. Most state-of-the-art detectors recognize objects individually without considering the rich context relationships of objects at different levels. In this paper, we propose an end-to-end vote-based multi-level context attention network. Specifically, a Patch-Context-Module is designed to extract multi-level context features among point patches. Meanwhile, because low-level features contain fine location description information, a Spatial-Context-Module is adopted to combine low-level spatial and semantic features. Furthermore, a Fusion Sampling and Aggregation module is proposed to consider additional semantic information of each vote point, thereby increasing the ratio of positive points and improving detection performance. Finally, the Class-IoU-Guide NMS with an adaptive threshold is implemented to suppress false detection at the inference time. Experiments on the ScanNetV2 and SUN RGB-D datasets demonstrated that our proposed method out-performs current state-of-the-art approaches. Ke Lu 0002, Jian Xue 0002, Liping Hou, Hengsheng Lun |
ICME | 3 |
| 2022 | Improved Transformer with Parallel Encoders for Image CaptioningabstractImage captioning is currently one of the most important multimodal tasks. With Transformer proposed, many Transformer-based models have achieved good performance in image captioning. However, substantial work is still required to improve the performance in the field of image captioning. We propose an improved model that uses the meshed-memory Transformer as its backbone. We propose the use of region features and grid features together. In addition, we use two identical parallel encoders to process region features and grid features separately, and fuse the outputs of each layer of the two encoders to form one of the inputs of the decoder. We comprehensively compare the performance of our model with the existing state-of-the-art models on the official COCO dataset. Experiments show that, on the Karpathy test split, our model outperforms the backbone on all evaluation metrics: for example, it increases BLEU-1 from 80.8% to 81.4%, and CIDEr from 131.2% to 133.5%. Liangshan Lou, Ke Lu 0002, Jian Xue 0002 |
ICPR | 3 |
| 2022 | Dual-Branch Point Cloud Feature Learning for 3D Object DetectionabstractIncomplete feature information is a key problem that limits 3D point cloud object detection and its applications. Many state-of-the-art detectors address this problem from different perspectives, but a comprehensive solution has not yet been obtained. In this paper, a solution is proposed that consists of two branches, one for channel-wise local feature learning and one for spatial-wise global feature learning. The combination of the local features, global contextual features, channel-wise attention features, and spatial attention features of the 3D point cloud is obtained through the two branches. Specifically, a generic spatial self-attention model is proposed that uses skeleton convolution to enhance the extraction of spatial features and combines it with a self-attention mechanism to improve the learning of global features. Further, the proposed skeleton attention mechanism focuses on object contour and rotation invariance of the point cloud. Through sufficient experimental validation, all module proposed in this paper are shown to have a substantial effect on performance and the visualization results demonstrate the importance of our approach. Hengsheng Lun, Jian Xue 0002, Ke Lu 0002 |
SMC | 2 |
| 2022 | Light Attention Embedding for Facial Expression RecognitionabstractFacial expression recognition is important for human–computer interaction and other applications. Several facial expression datasets have been published in recent decades and have enabled improvements in algorithms for classifying emotions. However, recognition of realistic expressions in real-world conditions is still challenging because of uncontrolled conditions, such as lighting, brightness, pose, and occlusion. In this paper, we propose a light attention embedding network based on the spatial attention mechanism (LAENet-SA), which can focus on locations in an image that are relevant to emotion. LAENet-SA allows a small number of attention modules to be embedded and can be constructed from typical convolutional neural networks. The performance of LAENet-SA on facial expression recognition has been validated on three facial expression datasets, including a lab-controlled dataset and two in-the-wild datasets. Experimental results show that LAENet-SA improved the performance on each dataset, compared with state-of-the-art methods, and achieved better generalization when tested on facial images with occlusion. Jian Xue 0002, Ke Lu 0002, Yanfu Yan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Refined One-Stage Oriented Object Detection Method for Remote Sensing ImagesabstractMulti-class object detection in remote sensing images plays an important role in many applications but remains a challenging task because of scale imbalance and arbitrary orientations of the objects with extreme aspect ratios. In this paper, the Asymmetric Feature Pyramid Network (AFPN), Dynamic Feature Alignment (DFA) module, and Area-IoU regression loss are proposed on the basis of a one-stage cascaded detection method for the detection of multi-class objects with arbitrary orientations in remote sensing images. The designed asymmetric convolutional block is embedded into the AFPN for handling objects with extreme aspect ratios and improving the space representation with ignorable increases in calculation. The DFA module is proposed to dynamically align mismatched features, which are caused by the deviation between predefined anchors and arbitrarily oriented predicted boxes. The refined Area-IoU regression loss, which reconciles two new regression loss functions, the area-guided regression loss and IoU-guided regression loss, is proposed to simultaneously solve the scale imbalance problem and angle sensitivity problem. Experiments on three publicly available datasets, DOTA, HRSC2016, and ICDAR2015, show the effectiveness of the proposed method. Liping Hou, Ke Lu 0002, Jian Xue 0002 |
IEEE Trans. Image Process. | 3 |
| 2022 | Fine-Grained Categorization From RGB-D ImagesabstractIn the field of computer vision, fine-grained visual categorization has attracted a lot of attention and made great progress due to convolutional neural networks and a large number of publicly available datasets. With next-generation sensing technology, RGB-D cameras can provide high-quality synchronized RGB and depth images for solving many computer vision problems. Although RGB-D cameras have been used in the context of multi-view object category detection and scene understanding, they have not been widely used in fine-grained classification. In this paper, we introduce a multiview RGB-D dataset RGBD-FG for fine-grained categorization. Currently, the dataset contains 93 051 RGB-D images covering 19 super-categories and 50 sub-categories of common vegetables and fruit, and is organized in a hierarchical manner. We provide extensive experimental results to establish state-of-the-art benchmarks for our dataset, illustrating its diversity and scope for improvement through future work. We also propose a novel modality-specific multimodal network called FS-Multimodal network, which can solve two limitations of multimodal networks trained based on fine-tuning techniques: over-fitting and lack of effective depth-specific features. We hope that our study lays the foundations for fine-grained categorization of RGB-D data. Yanhao Tan, Mohammad Muntasir Rahman, Yanfu Yan, Jian Xue 0002, Ling Shao 0001, Ke Lu 0002 |
IEEE Trans. Multim. | 4 |
| 2021 | Spatiotemporal Features and Local Relationship Learning for Facial Action Unit Intensity RegressionabstractThe action units (AU) encoded by the Facial Action Coding System (FACS) have been widely used in the representation of facial expressions. Although work on automatic facial AU detection has achieved quite good results in recent years, there remains much research potential for more accurate AU detection and intensity regression. Moreover, most work only considers the spatial information and ignores the temporal information. In practice, changes in facial AUs involve both spatial and temporal variation. In this paper, by extracting multi-scale spatial features and corresponding temporal features from the faces in the video image sequence, and learning the local relationship of the spatiotemporal features we propose a method that can obtain robust and accurate regression for AU intensity. The proposed method outperforms the baseline system on FEAFA dataset and obtains comparable performance on DISFA dataset. Ke Lu 0002, Jian Xue 0002 |
ICIP | 4 |
| 2021 | Multi-view 3D Smooth Human Pose Estimation based on Heatmap Filtering and Spatio-temporal InformationabstractThe estimation of 3D human poses from time-synchronized, calibrated multi-view video usually consists of two steps: (1) a 2D detector to locate the 2D coordinate point position of the joint via heatmaps for each frame and (2) a post-processing method such as the recursive pictorial structure model or robust triangulation to obtain 3D coordinate points. However, most existing methods are based on a single frame only. They do not take advantage of the temporal characteristics of the video sequence itself, and must rely on post-processing algorithms. They are also susceptible to human self-occlusion, and the generated sequences suffer from jitter. Therefore, we propose a network model incorporating spatial and temporal features. Using a coarse-to-fine approach, the proposed heatmap temporal network (HTN) generates temporal heatmap information, with an occlusion heatmap filter used to filter low-quality heatmaps before they are sent to the HTN. The heatmap fusion and the triangulation weights are dynamically adjusted, and intermediate supervision is employed to enable better integration of temporal and spatial information. Our network is also end-to-end differentiable. This overcomes the long-standing problem of skeleton jitter being generated and ensures that the sequence is smooth and stable. Zehai Niu, Ke Lu 0002, Jian Xue 0002, Runchen Wei |
ACM Multimedia | 3 |
| 2021 | A Coarse-to-Fine Facial Landmark Detection Method Based on Self-attention MechanismabstractFacial landmark detection in the wild remains a challenging problem in computer vision. Deep learning-based methods currently play a leading role in solving this. However, these approaches generally focus on local feature learning and ignore global relationships. Therefore, in this study, a self-attention mechanism is introduced into facial landmark detection. Specifically, a coarse-to-fine facial landmark detection method is proposed that uses two stacked hourglasses as the backbone, with a new landmark-guided self-attention (LGSA) block inserted between them. The LGSA block learns the global relationships between different positions on the feature map and allows feature learning to focus on the locations of landmarks with the help of a landmark-specific attention map, which is generated in the first-stage hourglass model. A novel attentional consistency loss is also proposed to ensure the generation of an accurate landmark-specific attention map. A new channel transformation block is used as the building block of the hourglass model to improve the model's capacity. The coarse-to-fine strategy is adopted during and between phases to reduce complexity. Extensive experimental results on public datasets demonstrate the superiority of our proposed method against state-of-the-art models. Ke Lu 0002, Jian Xue 0002, Ling Shao 0001, Jiayi Lyu |
IEEE Trans. Multim. | 3 |
| 2020 | Cascade Detector With Feature Fusion For Arbitrary-Oriented Objects In Remote Sensing ImagesabstractDetection of multi-class rotated objects is a challenging task in optical remote sensing images because of large-scale variations, arbitrary orientations and complex backgrounds, etc. Most of the state-of-the-art object detectors for natural images, that use horizontal bounding boxes, are not suitable for oriented objects in remote sensing images. In this paper, we propose an end-to-end cascade detector that can effectively detect rotated objects in complex remote sensing images. Specifically, a feature fusion block is designed to capture features with more details. Meanwhile, a supervised spatial attention mechanism is adopted to improve performance in detecting objects with complex backgrounds by weakening noise and enhancing object regions. Finally, to obtain more accurate object position, a cascade of multi-step detection subnet is implemented to refine anchors. Experiments using a publicly available remote sensing dataset DOTA show that our object detector achieves superior performance over other state-of-the-art approaches. Liping Hou, Ke Lu 0002, Jian Xue 0002 |
ICME | 3 |
| 2020 | Rgbd-Fg: A Large-Scale Rgb-D Dataset For Fine-Grained CategorizationabstractFine-grained visual categorization (FGVC) has received a great deal of attention in recent years. Currently, several public datasets are available for FGVC. However, all these datasets were created with RGB images. RGB-D sensors can provide high-quality synchronized video in terms of both color and depth. In this paper, we introduce a multi-view, large-scale RGB-D dataset called RGBD-FG to establish a novel benchmark for FGVC in RGB-D images. RGBD-FG contains RGB data and the corresponding depth data of vegetables and fruits. Our dataset was captured by a depth sensor and contained 50 categories and a total of 93,051 RGBD images with labels, and organized in a hierarchical manner. Additionally, we used several strategies on our dataset for FGVC, including a multi-modal deep CNN. We present extensive experimental results to create state-of-the-art baselines for the dataset. We hope that this dataset can fill the gap of FGVC among the RGB-D datasets. Yanhao Tan, Ke Lu 0002, Mohammad Muntasir Rahman, Jian Xue 0002 |
ICME | 4 |
| 2020 | EfficientFAN: Deep Knowledge Transfer for Face AlignmentabstractFace alignment plays an important role in many applications that process facial images. At present, deep learning-based methods have achieved excellent results in face alignment. However, these models usually have a large number of parameters, resulting in high computational complexity and execution time. In this paper, a lightweight, efficient, and effective model is proposed and named Efficient Face Alignment Network (EfficientFAN). EfficientFAN adopts the encoder-decoder structure, using a simple backbone Efficient-Net-B0 as the encoder and three deconvolutional layers as the decoder. Compared with state-of-the-art models, it achieves equivalent performance with fewer model parameters, lower computation cost, and higher speed. Moreover, the accuracy of EfficientFAN is further improved by transferring deep knowledge of a complex teacher network through feature-aligned distillation and patch similarity distillation. Extensive experimental results on public data sets demonstrate the superiority of EfficientFAN over state-of-the-art methods. Ke Lu 0002, Jian Xue 0002 |
ICMR | 3 |
| 2020 | A Lightweight Gated Global Module for Global Context Modeling in Neural NetworksabstractGlobal context modeling has been used to achieve better performance in various computer-vision-related tasks, such as classification, detection, segmentation and multimedia retrieval applications. However, most of the existing global mechanisms display problems regarding convergence during training. In this paper, we propose a novel gated global module (GGM) that is lightweight and yet effective in terms of achieving better integration of global information in relation to feature representation. Regarding the original structure of the network as a local block, our module infers global information in parallel with local information, and then a gate function is applied to generate global guidance which is applied to the output of the local module to capture representative information. The proposed GGM can be easily integrated with common CNN architectures and is training friendly. We used a classification task as an example to verify the effectiveness of the proposed GGM, and extensive experiments on ImageNet and CIFAR demonstrated that our method can be widely applied and is conducive to integrating global information into common networks. Liping Hou, Yuantao Song, Ke Lu 0002, Jian Xue 0002 |
ICMR | 5 |
| 2019 | An Automated Lung Nodule Segmentation Method Based On Nodule Detection Network and Region GrowingabstractSegmentation of a specific organ or tissue plays an important role in medical image analysis with the rapid development of clinical decision support systems. With medical imaging equipments, segmenting the lung nodules in the images is able to help physicians diagnose lung cancer diseases and formulate proper schemes. Therefore the research of lung nodule segmentation has attracted a lot of attention these years. However, this task faces some challenges, including the intensity similarity between lung nodules and vessel, inaccurate boundaries and presence of noise in most of the images. In this paper, an automated segmentation method is proposed for lung nodules in CT images. At the first stage, a nodule detection network is used to generate region proposals and locate the bounding boxes of nodules, which are employed as the initial input for the following segmentation. Then the nodules are segmented in the bounding boxes at the second stage. Since the image scale for region growing is reduced by locating the nodule in advance, the efficiency of segmentation can be improved. And due to the localization of nodule before segmentation, some tissues with similar intensity can be excluded from the object region. The proposed method is evaluated on a public lung nodule dataset, and the experimental results indicate the effectiveness and efficiency of the proposed method. Yanhao Tan, Ke Lu 0002, Jian Xue 0002 |
MMAsia | 3 |
| 2019 | Dense Attention Network for Facial Expression Recognition in the WildabstractRecognizing facial expression is significant for human-computer interaction system and other applications. A certain number of facial expression datasets have been published in recent decades and helped with the improvements for emotion classification algorithms. However, recognition of the realistic expressions in the wild is still challenging because of uncontrolled lighting, brightness, pose, occlusion, etc. In this paper, we propose an attention mechanism based module which can help the network focus on the emotion-related locations. Furthermore, we produce two network structures named DenseCANet and DenseSANet by using the attention modules based on the backbone of DenseNet. Then these two networks and original DenseNet are trained on wild dataset AffectNet and lab-controlled dataset CK+. Experimental results show that the DenseSANet has improved the performance on both datasets comparing with the state-of-the-art methods. Ke Lu 0002, Jian Xue 0002, Yanfu Yan |
MMAsia | 3 |
| 2019 | A Single-stage Multi-class Object Detection Method for Remote Sensing ImagesabstractImpressive progresses have been achieved in object detection for images by convolution neural networks. However, a robust multi-class object detection method is still one of the great challenges for remote sensing images. Due to the great diversity of scale, orientation, density and background of objects, most advanced object detection algorithms in natural scenes usually suffer a sharp decline in remote sensing images detection. To solve these problems, we proposed a Single-stage Multi-class Object Detection (SMOD) method, aiming at remote sensing images, which can be trained from scratch and detect multi-class objects quickly and precisely. The proposed method introduces a novel Feature Reuse and Attention (FRA) structure as a key module of feature extraction backbone, which combines SE Attention module and dense Feature Reuse connection. Especially, a multiclass detection structure is proposed to learn from multi-scale, multi-level feature map and get effective attention representation for multi-class remote sensing object detection. SMOD can be trained from scratch without pre-trained network stably and converge well simply by employing batch normalization throughout the network. Experiments show that our trainingfrom-scratch method can obtain better performance compared with some state-of-art algorithms on public multi-class remote sensing dataset AIIA2018-6. Liping Hou, Jian Xue 0002, Ke Lu 0002, Mohammad Muntasir Rahman |
VCIP | 2 |
| 2019 | 3D object detection: Learning 3D bounding boxes from scaled down 2D bounding boxes in RGB-D images
Mohammad Muntasir Rahman, Yanhao Tan, Jian Xue 0002, Ling Shao 0001, Ke Lu 0002 |
Inf. Sci. | 3 |
| 2018 | A fast Cascade Shape Regression Method based on CNN-based InitializationabstractCascade shape regression (CSR) methods predict facial landmarks by iteratively updating an initial shape and are state-of-the-art. The initial shape always limits the result and causes local optimum, which is usually obtained from the average face or by randomly picking a face from the training set. In this paper, we propose a CNN-based initial method for CSR. Convolution neural network provides a highly robust initial shape estimation, while the following CSR algorithm fine-tunes the initialization rapidly to achieve higher accuracy. Furthermore, CNN-based initial approach is proposed to get 68-point initial shape, which is calculated from convolutional network 5-point result by the radial basis function interpolation with thin-plate splines (RBF-TPS). Extensive experiments demonstrate that CSR methods are sensitive to the initialization and proposed approach gets favorable results compared to state-of-the-art algorithms and achieves real-time performance. Jian Xue 0002, Ke Lu 0002, Yanfu Yan |
ICPR | 2 |
| 2018 | Accelerated nonrigid image registration using improved Levenberg-Marquardt method
Jiyang Dong, Ke Lu 0002, Jian Xue 0002, Shuangfeng Dai, Weiguo Pan |
Inf. Sci. | 3 |
| 2018 | Single Image Dehazing Based on the Physical Model and MSRCR AlgorithmabstractTo address the hazy weather image degradation problem, we propose a single image dehazing method based on a physical model and the brightness components of the image by using a multi-scale retinex with color restoration algorithm. The overall dehazing process involves three components, including the atmospheric light value calculation, transmission map estimation, and recovery of the hazy image scene radiance. Our contribution is that we propose a novel algorithm to dehaze a single image by calculating the atmospheric light value and computing the transmission map while considering the dynamic range of the image. Experimental results show that our algorithm can effectively improve the image quality degraded by foggy weather and retain sufficient image details. Jinbao Wang 0001, Ke Lu 0002, Jian Xue 0002, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | RGB-D object recognition with multimodal deep convolutional neural networksabstractObject recognition from RGB-D images has become a hot topic and gained a significant popularity in recent years due to its numerous applications. In this paper, we propose a novel multimodal deep convolutional neural networks architecture for RGB-D object recognition which composed of three streams with two different types of deep CNNs, where each stream can separately learn from each modality. Finally, we propose a combined architecture of joint network of these three streams to classify the objects. Compared to RGB data, RGB-D images provide additional depth information that can be represented as depth colorization methods or surface normals. Our goal is to exploit both colorization and surface normals information to encode depth images. We show that by utilizing both colorization and surface normals of depth images combined with RGB significantly can improves the classification accuracy. We evaluate our model on one of the most challenging RGB-D object dataset and achieves comparable performance to state-of-the-art methods. Mohammad Muntasir Rahman, Yanhao Tan, Jian Xue 0002, Ke Lu 0002 |
ICME | 3 |
| 2016 | Efficient volume rendering methods for out-of-Core datasets by semi-adaptive partitioning
Jian Xue 0002, Ke Lu 0002, Ling Shao 0001, Mohammad Muntasir Rahman |
Inf. Sci. | 1 |
| 2015 | Hybrid architecture for 3D visualization of ultrasonic data
Weiguo Pan, Jian Xue 0002, Ke Lu 0002, Shuangfeng Dai |
Inf. Sci. | 2 |
| 2015 | Learning View-Model Joint Relevance for 3D Object Retrievalabstract3D object retrieval has attracted extensive research efforts and become an important task in recent years. It is noted that how to measure the relevance between 3D objects is still a difficult issue. Most of the existing methods employ just the model-based or view-based approaches, which may lead to incomplete information for 3D object representation. In this paper, we propose to jointly learn the view-model relevance among 3D objects for retrieval, in which the 3D objects are formulated in different graph structures. With the view information, the multiple views of 3D objects are employed to formulate the 3D object relationship in an object hypergraph structure. With the model data, the model-based features are extracted to construct an object graph to describe the relationship among the 3D objects. The learning on the two graphs is conducted to estimate the relevance among the 3D objects, in which the view/model graph weights can be also optimized in the learning process. This is the first work to jointly explore the view-based and model-based relevance among the 3D objects in a graph-based framework. The proposed method has been evaluated in three data sets. The experimental results and comparison with the state-of-the-art methods demonstrate the effectiveness on retrieval accuracy of the proposed 3D object retrieval method. Ke Lu 0002, Jian Xue 0002, Jiyang Dong, Ling Shao 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | 3D model retrieval and classification by semi-supervised learning with content-based similarity
Ke Lu 0002, Jian Xue 0002, Weiguo Pan |
Inf. Sci. | 3 |
| 2008 | A Novel Software Platform for Medical Image Processing and AnalyzingabstractThe design of software platform for medical imaging application has been increasingly prioritized as the sophisticated application of medical imaging. With this demand, we have designed and implemented a novel software platform in traditional object-oriented fashion with some common design patterns. This platform integrates the mainstream algorithms for medical image processing and analyzing within a consistent framework, including reconstruction, segmentation, registration, visualization, etc., and provides a powerful tool for both scientists and engineers. The overall framework and certain key technologies are introduced in detail. Presented experiment examples, numerous downloads, extensive uses, and practical applications commendably demonstrate the validity and flexibility of the platform. Jie Tian 0001, Jian Xue 0002, Yakang Dai, Jian Chen 0014, Jian Zheng 0001 |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2005 | The toolkit and platform for biometric information processingabstractA biometric information processing toolkit (BITK) is designed and implemented to support the biometrics research, e.g. algorithm development and performance evaluation. BITK is an object-oriented C++ software development toolkit (SDK), and it provides a consistent, flexible and reusable framework to integrate algorithms, data structures, and visualization methods. In addition, an application platform based on BITK (BITKAPP) is developed. BITKAPP makes the best of BITK to serve the biometrics researchers with its friendly user interface and the plug-in architecture. The meaningful applications of the toolkit and platform confirm that they effectively support the biometrics research. Xiaoguang He, Jie Tian 0001, Yuliang He, Jian Xue 0002, Xin Yang 0001 |
CSCWD (2) | 4 |
| 2005 | The design and implementation of a novel platform for medical data visualizationabstractAs medical imaging applications become more complex, the design of software platforms for medical imaging is getting greater priority. With this demand, we have designed and implemented a novel software platform for medical data visualization in traditional object-oriented fashion with some common design patterns. This platform integrates the mainstream medical data reconstruction and visualization algorithms and 3D human-computer interaction based on 3D widgets into a consistent framework. The design goals, the overall framework and the implementation of some key technologies are addressed in details, and some application examples are also given to demonstrate the abilities of this platform. The ultimate objective is to provide a flexible, reliable and extensible 3D medical data visualization platform for the medical imaging society. Jian Xue 0002, Jie Tian 0001, Mingchang Zhao, Huiguang He |
CSCWD (2) | 1 |