Xing Lan

dblp:76/11156 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MotionFlow: Efficient Motion Generation With Latent Flow Matching
abstract
In the field of human centric multimedia, text-driven human motion generation is a significant pursuit with wide-ranging applications across diverse scenarios. Despite substantial advancements, existing methods often suffer from a trade-off between inference latency and high-quality generation. To overcome this gap, we propose the Motion Latent Flow Matching model (MotionFlow), a novel and powerful framework for motion generation. It introduces flow matching algorithm in the latent space, which can achieve superior performance with just one-step inference. In addition to the text-driven task, we further extend our method to controllable motion generation. Specifically, we integrate a control encoder into the latent space and further decode the predicted latent code into motion space to support explicit supervision, ensuring the synthesized motion can tightly align with the input signals. Extensive experiments demonstrate that our MotionFlow not only outperforms current leading approaches for the text-driven task, but also delivers remarkable capabilities in controllable motion generation.
Kun Dong 0001, Jian Xue 0002, Xing Lan, Qingyuan Liu 0001, Ke Lu 0002
IEEE Trans. Multim.3
2025 Continuous Action Unit Intensity Modeling for Micro-Expression Recognition
abstract
Micro-Expression Recognition (MER) remains challenging due to the subtle and transient nature of facial muscle movements. While recent methods leverage Action Unit (AU) labels for MER, they often tend to ignore continuous AU intensity variations, which are critical for capturing nuanced facial expressions. To address these limitations, we propose a novel framework integrating continuous AU intensity with hierarchical motion modeling. Our approach begins with a lightweight model that regresses in-frame AU intensity values. These AU intensities are fed into our proposed Continuous AU Transformer (CAUT), which employs a temporal Transformer and a spatial Transformer to model AU evolution across frames and inter-AU dependencies. Simultaneously, a two-stage Transformer architecture extracts hierarchical optical flow features, fused with AU semantics via a multi-scale region-based fusion strategy for enhancing facial motion features. Extensive experiments demonstrate the proposed method’s state-of-the-art performance, validating the effectiveness of continuous AU intensity modeling and hierarchical feature integration for MER.
Hanyu Jiang 0004, Jiayi Lyu, Xing Lan, Jian Xue 0002
ICIP3
2025 One General Plug-In for Facial Heatmap-based Keypoint Detection
abstract
In this paper, we systematically investigate the error distribution in predicted heatmaps for face alignment, and point out that previous works are unreliable in following the rule that decodes coordinates by locating the maximum-score pixel. Our research reveals that the majority of ground-truth positions do not match that pixel but rather lie within a range of a few pixels. Building on this phenomenon, we transform the model’s objective from predicting inaccurate landmarks to identifying precise proposals with that range. We propose a simple but effective module, termed the Response Aware Module (RAM), leveraging response scores in the proposal to regress the proposal offset, which can be used as a plug-and-play layer integrated into public models. Furthermore, we present a novel Heatmap RCNN framework to exploit the distribution of multi-scale heatmaps. Extensive experiments have demonstrated that the trained RAM can be integrated seamlessly as a ready-to-use plugin with the model, yielding impressive improvements. Meanwhile, Heatmap RCNN performs far superior to SOTA results, with 3.82 NME on WFLW, 3.09 on COFW, and 2.90 on 300W.
Hanyu Jiang 0004, Jian Xue 0002, Xing Lan, Ke Lu 0002
ICME3
2025 Multimodal Emotional Talking Face Generation Based on Action Units
abstract
Talking face generation focuses on creating natural facial animations that align with the provided text or audio input. Current methods in this field primarily rely on facial landmarks to convey emotional changes. However, spatial key-points are valuable, yet limited in capturing the intricate dynamics and subtle nuances of emotional expressions due to their restricted spatial coverage. Consequently, this reliance on sparse landmarks can result in decreased accuracy and visual quality, especially when representing complex emotional states. To address this issue, we propose a novel method called Emotional Talking with Action Unit (ETAU), which seamlessly integrates facial Action Units (AUs) into the generation process. Unlike previous works that solely rely on facial landmarks, ETAU employs both Action Units and landmarks to comprehensively represent facial expressions through interpretable representations. Our method provides a detailed and dynamic representation of emotions by capturing the complex interactions among facial muscle movements. Moreover, ETAU adopts a multi-modal strategy by seamlessly integrating emotion prompts, driving videos, and target images, and by leveraging various input data effectively, it generates highly realistic and emotional talking-face videos. Through extensive evaluations across multiple datasets, including MEAD, LRW, GRID and HDTF, ETAU outperforms previous methods, showcasing its superior ability to generate high-quality, expressive talking faces with improved visual fidelity and synchronization. Moreover, ETAU exhibits a significant improvement on the emotion accuracy of the generated results, reaching an impressive average accuracy of 84% on the MEAD dataset.
Jiayi Lyu, Xing Lan, Guohong Hu, Hanyu Jiang 0004, Jinbao Wang 0001, Jian Xue 0002
IEEE Trans. Circuits Syst. Video Technol.2
2025 FoodSAM: Any Food Segmentation
abstract
In this paper, we explore the zero-shot capability of the Segment Anything Model (SAM) for food image segmentation. To address the lack of class-specific information in SAM-generated masks, we propose a novel framework, calledFoodSAM. This innovative approach integrates the coarse semantic mask with SAM-generated masks to enhance semantic segmentation quality. Besides, we recognize that the ingredients in food can be supposed as independent individuals, which motivated us to perform instance segmentation on food images. Furthermore, FoodSAM extends its zero-shot capability to encompass panoptic segmentation by incorporating an object detector, which renders FoodSAM to effectively capture non-food object information. Drawing inspiration from the recent success of promptable segmentation, we also extend FoodSAM to promptable segmentation, supporting various prompt variants. Consequently, FoodSAM emerges as an all-encompassing solution capable of segmenting food items at multiple levels of granularity. Remarkably, this pioneering framework stands as the first-ever work to achieve instance, panoptic, and promptable segmentation on food images. Extensive experiments demonstrate the feasibility and impressing performance of FoodSAM, validating SAM's potential as a prominent and influential tool within the domain of food image segmentation.
Xing Lan, Jiayi Lyu, Hanyu Jiang 0004, Kun Dong 0001, Zehai Niu, Yi Zhang 0162, Jian Xue 0002
IEEE Trans. Multim.1
2025 ExpLLM: Towards Chain of Thought for Facial Expression Recognition
abstract
Facial expression recognition (FER) is a critical task in multimedia with significant implications across various domains. However, analyzing the causes of facial expressions is essential for accurately recognizing them. Current approaches, such as those based on facial action units (AUs), typically provide AU names and intensities but lack insight into the interactions and relationships between AUs and the overall expression. In this paper, we propose a novel method called ExpLLM, which leverages large language models to generate an accurate chain of thought (CoT) for facial expression recognition. Specifically, we have designed the CoT mechanism from three key perspectives: key observations, overall emotional interpretation, and conclusion. The key observations describe the AU's name, intensity, and associated emotions. The overall emotional interpretation provides an analysis based on multiple AUs and their interactions, identifying the dominant emotions and their relationships. Finally, the conclusion presents the final expression label derived from the preceding analysis. Furthermore, we also introduce the Exp-CoT Engine, designed to construct this expression CoT and generate instruction-description data for training our ExpLLM. Extensive experiments on the RAF-DB and AffectNet datasets demonstrate that ExpLLM outperforms current state-of-the-art FER methods. ExpLLM also surpasses the latest GPT-4o in expression CoT generation, particularly in recognizing micro-expressions where GPT-4o frequently fails.
Xing Lan, Jian Xue 0002, Ji Qi 0003, Dongmei Jiang, Ke Lu 0002, Tat-Seng Chua
IEEE Trans. Multim.1
2024 3Dlaneformer: Rethinking Learning Views for 3D Lane Detection
abstract
Accurate 3D lane detection from monocular images is crucial for autonomous driving. Recent advances leverage either front-view (FV) or bird’s-eye-view (BEV) features for prediction, inevitably limiting their ability to perceive driving environments precisely and resulting in suboptimal performance. To overcome the limitations of using features from a single view, we design a novel dual-view cross-attention mechanism, which leverages features from FV and BEV simultaneously. Based on this mechanism, we propose 3DLaneFormer, a powerful framework for 3D lane detection. It outperforms the latest BEV-based or FV-based approaches through extensive experiments on challenging benchmarks and thus verifies the necessity and benefits of utilizing features in both views.
Kun Dong 0001, Jian Xue 0002, Xing Lan, Ke Lu 0002
ICIP3
2024 ETAU: Towards Emotional Talking Head Generation Via Facial Action Unit
abstract
Creating expressive talking heads is crucial for multimedia applications involving virtual human. Existing approaches predominantly rely on facial landmarks to convey emotional changes. However, these spatial keypoints struggle to capture subtle emotional intricacies due to their limited spatial coverage, consequently decreasing accuracy and visual quality, particularly in emotion representation. To address this issue, we introduce a novel method called Emotional Talking with Action Unit (ETAU), which introduces the additional facial Action Units (AUs) to generate talking head video that accurately portray the target emotions. Unlike previous works, ETAU comprehensively quantify facial expressions through Action Units, which provides a detailed and dynamic representation of emotion. To the best of our knowledge, this work pioneers the integration of Action Units for emotional talking head generation. Extensive evaluations on the MEAD dataset showcase ETAU’s state-of-the-art performance with 21.89 PSNR and 0.68 SSIM. Critically, ETAU achieves significant improvement in emotion accuracy of the generated results, reaching 84%, confirming its feasibility in representing emotional expressions.
Jiayi Lyu, Xing Lan, Guohong Hu, Hanyu Jiang 0004, Jian Xue 0002
ICME2
2024 Realistic Full-Body Motion Generation from Sparse Tracking with State Space Model
abstract
In the domain of generative multimedia and interactive experiences, generating realistic and accurate full-body poses from sparse tracking is crucial for many real-world applications, while achieving sequence modeling and efficient motion generation remains challenging. Recently, state space models (SSMs) with efficient hardware-aware designs (i.e., Mamba) have shown great potential for sequence modeling, particularly in temporal contexts. However, processing motion data is still challenging for SSMs. Specifically, the sparsity of input conditions makes motion generation an ill-posed problem. Moreover, the complex structure of the human body further complicates this task. To address these issues, we present Motion Mamba Diffusion (MMD), a novel conditional diffusion model, which effectively utilizes the sequence modeling capability of SSMs and the robust generation ability of diffusion models to track full-body poses accurately. In particular, we design a bidirectional Temporal Mamba Module (TMM) to model motion sequence. Additionally, a Spatial Mamba Module (SMM) is further proposed for feature enhancement within a single frame. Extensive experiments on the large motion capture dataset (AMASS) demonstrate that our proposed approach outperforms the latest methods in terms of accuracy and smoothness, thus providing a crucial advancement for creating realistic virtual avatars in various applications.
Kun Dong 0001, Jian Xue 0002, Zehai Niu, Xing Lan, Ke Lu 0002, Qingyuan Liu 0001, Xiaoyu Qin 0001
ACM Multimedia4
2024 Energy-saving scheduling strategy for variable-speed flexible job-shop problem considering operation-dependent energy consumption
Hongquan Qu, Xiaomeng Tong, Maolin Cai, Yan Shi 0003, Xing Lan
Expert Syst. Appl.5
2024 Does Pixel Value Represent Facial Landmark Well in Heatmap?
abstract
Heatmap-based methods have dominated the face alignment task, yet the maximum response decoding scheme necessitates further reform. While some studies have attempted to compensate for prediction offsets using a post-processing module, the prediction errors induced by the maximum response decoding scheme remain challenging to rectify. In this paper, we assume that using heatmap value to denote the ground-truth probability is not accurate enough. To cure this problem, we propose DISPAL, a novel DIStribution-based Probability for fAcial Landmarks, which signifies the ground-truth probability by the similarity between the pixel’s neighbouring value distribution and Gaussian distribution. This innovative probability enables us to pinpoint the keypoint location more robustly than previous methods that rely solely on the peak score. It also exhibits remarkable generalization to complex decoding methodologies. Furthermore, we propose supervising this probability as an additional task loss to help the model learn better heatmap representation. Extensive empirical results on WFLW, 300W, and COFW datasets demonstrate that our distribution-based probability mechanism significantly surpasses original value-based probability approaches.
Xing Lan, Jiayi Lyu, Kun Dong 0001, Hanyu Jiang 0004, Qinghao Hu 0001, Jian Xue 0002
IEEE Trans. Circuits Syst. Video Technol.1
2024 2-D Magnetotelluric Gradient Prediction With the Transformer + Unet Network Based on Transverse Magnetic Polarization
abstract
The calculation of magnetotelluric (MT) data gradients can be used for data sensitivity analysis, which is of great significance for actual sensitive areas. However, such calculations are highly complex and time consuming and therefore not conducive for the implementation of rapid analysis. To address this issue and improve computational efficiency, we propose a solution based on the Transformer+Unet (T-Unet) neural network model to accelerate the calculation of two-dimensional MT gradients. First, we create a seven-channel dataset corresponding to the gradient label and then obtain the neural network weight model through network training and iteration and predict the gradient value rapidly and accurately on this basis. The experimental results indicate that, compared to traditional gradient computation, the T-Unet network not only significantly reduces computation time but also ensures high gradient prediction accuracy. This research demonstrates the potential of gradient fast prediction in sensitivity analysis and accelerated MT inversion.
Chongxin Yuan, Kunpeng Wang 0003, Xiangpeng Wang, Yunbao Yue, Xing Lan
IEEE Trans. Geosci. Remote. Sens.7
2023 BiUNet: Towards More Effective UNet with Bi-Level Routing Attention
Kun Dong 0001, Jian Xue 0002, Xing Lan, Ke Lu 0002
BMVC3
2023 ATF: An Alternating Training Framework for Weakly Supervised Face Alignment
abstract
In recent years, various face-landmark datasets have been published. Intuitively, it is significant to integrate multiple labeled datasets to achieve higher performance. Due to the different annotation schemes of datasets, it is hard to directly train models using them together. Although numerous efforts have been made in the joint use of datasets, there remain three shortages in previous methods,i.e., additional computation, limitation of the markups scheme, and limited support for the regression method. To solve the above issues, we proposed a novelAlternating Training Framework(ATF), which leverages the similarity and diversity across multiple datasets for a more robust detector. ATF mainly contains two sub-modules:Alternating Training with Decreasing Proportions(ATDP) andMixed Branch Loss($\mathcal {L}_{MB}$). In particular, ATDP trains multiple datasets simultaneously via a weakly supervised way to take advantage of the diversity among them, and$\mathcal {L}_{MB}$utilizes similar landmark pairs to constrain different branches of the corresponding datasets. Besides, we extend the framework to easily handle three situations: single target detector, joint detector, and novel detector. Extensive experiments demonstrate the effectiveness of our framework for both heatmap-based and direct coordinate regression. Moreover, we have achieved a joint detector that outperforms state-of-the-art methods on each benchmark.
Xing Lan, Qinghao Hu 0001, Jian Cheng 0001
IEEE Trans. Multim.1
2022 CSCD: A Cyber Security Community Detection Scheme on Online Social Networks
Yutong Zeng, Honghao Yu, Tiejun Wu, Xing Lan
ICDF2C5
2021 An Evolutionary Study of IoT Malware
abstract
Recent years have witnessed lots of attacks targeted at the widespread Internet of Things (IoT) devices and malicious activities conducted by compromised IoT devices. After some notorious IoT malware released their source code, many new variants emerge, which are usually more powerful and stealthy. Although numerous existing studies have analyzed some exposed families, there is a lack of systematic study to make full use of them, which can be a fundamental step for provenance, triage, labeling, lineage analysis, and authorship attribution. The key challenge of conducting an IoT malware evolutionary study is how to collect sufficient and accurate information about malware and identify the relationships among them. In this article, we take the first step to investigate the IoT malware evolution by leveraging the information from two sources that complement each other. First, we crawl online articles about IoT malware and employ natural language processing techniques to extract the features of malware samples and their relationships with other malware family, which allow us to form the basic lineage graph. Second, we collect real malware samples through our widely deployed honeypots and design a new classifier to group them into families and identify lineage relationships among them. Such results are used to enhance the basic lineage graph. Eventually, we construct the final lineage graph for 72 IoT malware families by correlating the information from the aforementioned sources, which can help the research community better understand and fight IoT malware now and in the future. Our study has been incorporated into the threat awareness system of NSFOCUS company.
Huanran Wang, Weizhe Zhang, Peng Liu 0005, Xiapu Luo, Yang Liu 0039, Yan Li 0075, Wenmao Liu, Runzi Zhang, Xing Lan
IEEE Internet Things J.12
2020 ATF: Towards Robust Face Alignment via Leveraging Similarity and Diversity across Different Datasets
abstract
Face alignment is an important task in the field of multi-media. Together with the impressive progress of algorithms, various benchmark datasets have been released in recent years. Intuitively, it is meaningful to integrate multiple labeled datasets with different annotations to achieve higher performance on a target landmark detector. Although numerous efforts have been made in joint usage, there yet remain three shortages in recent works, e.g., additional computation, limitation of the markups scheme, and limited support for the regression method. To address the above problems, we proposed a novel Alternating Training Framework (ATF), which leverages similarity and diversity across multi-media sources for a more robust detector. Our framework mainly contains two sub-modules: Alternating Training with Decreasing Proportions (ATDP) and Mixed Branch Loss (mathcal LMB). In particular, ATDP trains multiple datasets simultaneously to take advantage of the diversity between them, while mathcal LMB utilizes similar landmark pairs to constrain different branches of corresponding datasets. Extensive experiments on various benchmarks show the effectiveness of our framework, and ATF is feasible for both heatmap-based network and direct coordinate regression. Specifically, the mean error even reaches 3.17 on the experiment on 300W leveraging WFLW, which significantly outperforms state-of-the-art methods. Both in an ordinary convolutional network (OCN) and HRNET, ATF achieves up to 9.96% relative improvement. Our source codes are made publicly available at https://github.com/starhiking/ATF.
Xing Lan, Qinghao Hu 0001, Fangzhou Xiong, Cong Leng, Jian Cheng 0001
ACM Multimedia1
2009 The expected complexity of sphere decoding algorithm in spatial correlated MIMO channels
abstract
The sphere decoding (SD) algorithm is widely considered to be an efficient approach to obtain maximum likelihood (ML) performance in MIMO detection. At present, almost all of the research about the SD algorithm is based on the assumption of independent and identically distributed channel coefficients. However, the channel coefficients are often correlated in practice, which cause the complexity of the SD algorithm to vary. In this paper, we give a theoretical analysis of the complexity of Fincke and Pohst's(FP) SD algorithm in spatial correlated MIMO channels; the exact expression of the expected complexity is derived. We present simulation results obtained from this expression to show the effect of spatial correlation on the complexity of the algorithm, for different Signal-to-Noise Ratios (SNR) and level of spatial correlations.
Jibo Wei, Xing Lan
ISIT2