EDBT 2026 Demo / reviewers in the wild / expert
Jiayi Lyu
dblp:239/4305
· DBLP profile ↗
17ranked-venue papers
2as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-tuning Zero-shot Large Language Models for Patient-reported Outcomes (Student Abstract)abstractRadiotherapy (RT) is a cornerstone of cancer treatment. Following RT, patient-reported outcomes (PROs) collected via standardized questionnaires are crucial for monitoring patients' quality of life and side effects. However, traditional statistical and machine learning methods, which rely on structured numerical data, often fail to capture semantic meaning within patients' health status. To address this, we developed a novel framework using zero- and few-shot large language models (LLMs) to identify patients experiencing mild to severe depression. Furthermore, classification performance is enhanced through parameter-efficient fine-tuning. Experiments on a prostate cancer PRO dataset for depression have demonstrated that our fine-tuned LLMs consistently outperformed other baseline methods across key evaluation metrics. Yang Yan 0003, Matthew W. Chen, Jiayi Lyu, Chen Zhao 0010, Zhong Chen 0003 |
AAAI | 3 |
| 2026 | SeViMatch: A Detector-Based Image Matching Framework with Semantic-Visual Fusion
Yun Liao, Jiayi Lyu, Zongxiao Hu, Qing Duan |
MMM (1) | 4 |
| 2026 | FFMatch: A FilterFormer-Based Network for Accurate Multimodal Image Matching
Yun Liao, Jiayi Lyu, Zongxiao Hu, Qing Duan |
MMM (1) | 2 |
| 2026 | GS2Physics: Semantic-Region-Aware Gaussian Splatting for Physical Property PredictionabstractPredicting the physical properties of reconstructed 3D assets is essential for virtual reality interactions. However, current systems often depend on manually assigning properties such as stiffness and density, which can be inefficient and prone to errors. To address this issue, we present GS2Physics, a novel framework based on 3D Gaussian Splatting. This framework is designed to predict physical properties accurately while maintaining improved consistency in semantic segmentation. Unlike existing approaches, which either struggle with region inconsistency or misalign semantic 3D features, GS2Physics embeds semantic-region-aware features directly into the Gaussian Splatting representation. This allows for region-consistent and accurate physical property prediction, achieving state-of-the-art performance on the ABO-500 mass prediction benchmark. To further evaluate our segmentation capabilities, we introduce PhysSeg-15, a subset dataset of ABO-500 featuring physical property segmentation masks for 15 different 3D objects captured from five viewpoints. Our method significantly outperforms existing approaches in segmentation accuracy. Qualitative results demonstrate more consistent material predictions across different object regions and improved accuracy in physical property prediction. In addition, we showcase the effectiveness of GS2Physics in 3D interaction tasks, where our predicted physical properties result in more realistic object motion. Our dataset and results are available at https://github.com/momaiyc/GS2Physics. Bin Huang 0016, Jiayi Lyu, Zehai Niu, LinLin Shen, Jinbao Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Dynamic Clustering Convolutional Neural NetworkabstractConvolutional neural networks (CNNs) have been playing a dominant role in computer vision. However, the existing approaches of using local window modeling in popular CNNs lack flexibility and hinder their ability to capture long-range dependencies of objects in an image. To overcome these limitations, we propose a novel CNN architecture, termed Dynamic Clustering Convolutional Neural Network (DCCNeXt). The proposed DCCNeXt takes a unique approach by employing global clustering to group image patches with similar semantics into clusters that are then convolved using the shared convolution kernels. To address the high computational complexity of global clustering, the feature vectors from each patch's subspace are extracted for efficient clustering, which makes the proposed model widely compatible with the downstream vision tasks. The extensive experiments of image classification, object detection, instance segmentation, and semantic segmentation on the benchmark datasets demonstrate that the proposed DCCNeXt outperforms the mainstream Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), Vision Multi-layer Perceptrons (MLPs), Vision Graph Neural Networks (GNNs), and Vision Mambas. We anticipate that this study will provide a new perspective and a promising avenue for the design of convolutional neural networks. Tanzhe Li, Baochang Zhang 0001, Jiayi Lyu, Xiawu Zheng, Guodong Guo, Taisong Jin |
AAAI | 3 |
| 2025 | Continuous Action Unit Intensity Modeling for Micro-Expression RecognitionabstractMicro-Expression Recognition (MER) remains challenging due to the subtle and transient nature of facial muscle movements. While recent methods leverage Action Unit (AU) labels for MER, they often tend to ignore continuous AU intensity variations, which are critical for capturing nuanced facial expressions. To address these limitations, we propose a novel framework integrating continuous AU intensity with hierarchical motion modeling. Our approach begins with a lightweight model that regresses in-frame AU intensity values. These AU intensities are fed into our proposed Continuous AU Transformer (CAUT), which employs a temporal Transformer and a spatial Transformer to model AU evolution across frames and inter-AU dependencies. Simultaneously, a two-stage Transformer architecture extracts hierarchical optical flow features, fused with AU semantics via a multi-scale region-based fusion strategy for enhancing facial motion features. Extensive experiments demonstrate the proposed method’s state-of-the-art performance, validating the effectiveness of continuous AU intensity modeling and hierarchical feature integration for MER. Hanyu Jiang 0004, Jiayi Lyu, Xing Lan, Jian Xue 0002 |
ICIP | 2 |
| 2025 | Visual Content Generation in the Era of Large Foundation ModelsabstractThe rapid advancements in large foundation models have significantly transformed the field of visual content generation, impacting domains such as image synthesis, video generation, and 3D modeling. This tutorial will provide an in-depth exploration of the state-of-the-art techniques and methodologies used in visual content generation, emphasizing the role of large-scale generative models. The tutorial will cover fundamental principles, model architectures, recent breakthroughs, and practical applications. We will discuss various generative paradigms, including diffusion models, autoregressive models, and large multimodal models, highlighting their strengths and limitations. Additionally, we will delve into the challenges of controllability, personalization, and realism in generated content, along with open research problems and future directions. By the end of the tutorial, attendees will gain a comprehensive understanding of contemporary visual content generation techniques and their applications, equipping them with the knowledge to leverage these models in their research and projects. Leigang Qu, Fei Shen 0004, Zhenglin Zhou, Jiayi Lyu, Wenjie Wang 0007, Lu Jiang 0004 |
ICMR | 4 |
| 2025 | PMCMatcher: A Parallel Multi-Scale Cascaded Transformer-Based Network for Multimodal Feature MatchingabstractMultimodal image matching is fundamental in computer vision. However, existing methods often struggle to achieve effective cross-modal feature fusion, especially under scale variations and complex scenarios. To this end, we propose PMCMatcher, a Parallel Multi-Scale Cascaded Transformer-Based Network for multimodal feature matching. The core module, the Parallel Multi-Scale Cascaded Transformer, achieves deep interaction and progressive multi-scale fusion through the Selective Multi-Head Linear Attention module and the Cascaded Fusion mechanism. Additionally, the Dynamic Local Feature Enhancement module significantly strengthens the extraction of details by adaptively adjusting convolution kernel weights. To further improve matching accuracy, the Refinement Layer is incorporated to gradually optimize the matching process, enhancing the model’s accuracy and robustness in cross-modal scenarios. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on four representative multimodal image matching datasets, highlighting its superior generalization ability and precise matching accuracy. Yun Liao, Jiayi Lyu, Zongxiao Hu, Qing Duan |
MMAsia | 2 |
| 2025 | DAMamba: Vision State Space Model with Dynamic Adaptive ScanabstractState space models (SSMs) have recently garnered significant attention in computer vision. However, due to the unique characteristics of image data, adapting SSMs from natural language processing to computer vision has not outperformed the state-of-the-art convolutional neural networks (CNNs) and Vision Transformers (ViTs). Existing vision SSMs primarily leverage manually designed scans to flatten image patches into sequences locally or globally. This approach disrupts the original semantic spatial adjacency of the image and lacks flexibility, making it difficult to capture complex image structures. To address this limitation, we propose Dynamic Adaptive Scan (DAS), a data-driven method that adaptively allocates scanning orders and regions. This enables more flexible modeling capabilities while maintaining linear computational complexity and global modeling capacity. Based on DAS, we further propose the vision backbone DAMamba, which significantly outperforms popular vision Mamba models in vision tasks such as image classification, object detection, instance segmentation, and semantic segmentation. Notably, it surpasses some of the latest state-of-the-art CNNs and ViTs. Tanzhe Li, Caoshuo Li, Jiayi Lyu, Hongjuan Pei, Baochang Zhang 0001, Taisong Jin, Rongrong Ji |
NeurIPS | 3 |
| 2025 | Multimodal Emotional Talking Face Generation Based on Action UnitsabstractTalking face generation focuses on creating natural facial animations that align with the provided text or audio input. Current methods in this field primarily rely on facial landmarks to convey emotional changes. However, spatial key-points are valuable, yet limited in capturing the intricate dynamics and subtle nuances of emotional expressions due to their restricted spatial coverage. Consequently, this reliance on sparse landmarks can result in decreased accuracy and visual quality, especially when representing complex emotional states. To address this issue, we propose a novel method called Emotional Talking with Action Unit (ETAU), which seamlessly integrates facial Action Units (AUs) into the generation process. Unlike previous works that solely rely on facial landmarks, ETAU employs both Action Units and landmarks to comprehensively represent facial expressions through interpretable representations. Our method provides a detailed and dynamic representation of emotions by capturing the complex interactions among facial muscle movements. Moreover, ETAU adopts a multi-modal strategy by seamlessly integrating emotion prompts, driving videos, and target images, and by leveraging various input data effectively, it generates highly realistic and emotional talking-face videos. Through extensive evaluations across multiple datasets, including MEAD, LRW, GRID and HDTF, ETAU outperforms previous methods, showcasing its superior ability to generate high-quality, expressive talking faces with improved visual fidelity and synchronization. Moreover, ETAU exhibits a significant improvement on the emotion accuracy of the generated results, reaching an impressive average accuracy of 84% on the MEAD dataset. Jiayi Lyu, Xing Lan, Guohong Hu, Hanyu Jiang 0004, Jinbao Wang 0001, Jian Xue 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | FoodSAM: Any Food SegmentationabstractIn this paper, we explore the zero-shot capability of the Segment Anything Model (SAM) for food image segmentation. To address the lack of class-specific information in SAM-generated masks, we propose a novel framework, calledFoodSAM. This innovative approach integrates the coarse semantic mask with SAM-generated masks to enhance semantic segmentation quality. Besides, we recognize that the ingredients in food can be supposed as independent individuals, which motivated us to perform instance segmentation on food images. Furthermore, FoodSAM extends its zero-shot capability to encompass panoptic segmentation by incorporating an object detector, which renders FoodSAM to effectively capture non-food object information. Drawing inspiration from the recent success of promptable segmentation, we also extend FoodSAM to promptable segmentation, supporting various prompt variants. Consequently, FoodSAM emerges as an all-encompassing solution capable of segmenting food items at multiple levels of granularity. Remarkably, this pioneering framework stands as the first-ever work to achieve instance, panoptic, and promptable segmentation on food images. Extensive experiments demonstrate the feasibility and impressing performance of FoodSAM, validating SAM's potential as a prominent and influential tool within the domain of food image segmentation. Xing Lan, Jiayi Lyu, Hanyu Jiang 0004, Kun Dong 0001, Zehai Niu, Yi Zhang 0162, Jian Xue 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | ETAU: Towards Emotional Talking Head Generation Via Facial Action UnitabstractCreating expressive talking heads is crucial for multimedia applications involving virtual human. Existing approaches predominantly rely on facial landmarks to convey emotional changes. However, these spatial keypoints struggle to capture subtle emotional intricacies due to their limited spatial coverage, consequently decreasing accuracy and visual quality, particularly in emotion representation. To address this issue, we introduce a novel method called Emotional Talking with Action Unit (ETAU), which introduces the additional facial Action Units (AUs) to generate talking head video that accurately portray the target emotions. Unlike previous works, ETAU comprehensively quantify facial expressions through Action Units, which provides a detailed and dynamic representation of emotion. To the best of our knowledge, this work pioneers the integration of Action Units for emotional talking head generation. Extensive evaluations on the MEAD dataset showcase ETAU’s state-of-the-art performance with 21.89 PSNR and 0.68 SSIM. Critically, ETAU achieves significant improvement in emotion accuracy of the generated results, reaching 84%, confirming its feasibility in representing emotional expressions. Jiayi Lyu, Xing Lan, Guohong Hu, Hanyu Jiang 0004, Jian Xue 0002 |
ICME | 1 |
| 2024 | Does Pixel Value Represent Facial Landmark Well in Heatmap?abstractHeatmap-based methods have dominated the face alignment task, yet the maximum response decoding scheme necessitates further reform. While some studies have attempted to compensate for prediction offsets using a post-processing module, the prediction errors induced by the maximum response decoding scheme remain challenging to rectify. In this paper, we assume that using heatmap value to denote the ground-truth probability is not accurate enough. To cure this problem, we propose DISPAL, a novel DIStribution-based Probability for fAcial Landmarks, which signifies the ground-truth probability by the similarity between the pixel’s neighbouring value distribution and Gaussian distribution. This innovative probability enables us to pinpoint the keypoint location more robustly than previous methods that rely solely on the peak score. It also exhibits remarkable generalization to complex decoding methodologies. Furthermore, we propose supervising this probability as an additional task loss to help the model learn better heatmap representation. Extensive empirical results on WFLW, 300W, and COFW datasets demonstrate that our distribution-based probability mechanism significantly surpasses original value-based probability approaches. Xing Lan, Jiayi Lyu, Kun Dong 0001, Hanyu Jiang 0004, Qinghao Hu 0001, Jian Xue 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | IM-IAD: Industrial Image Anomaly Detection Benchmark in ManufacturingabstractImage anomaly detection (IAD) is an emerging and vital computer vision task in industrial manufacturing (IM). Recently, many advanced algorithms have been reported, but their performance deviates considerably with various IM settings. We realize that the lack of a uniform IM benchmark is hindering the development and usage of IAD methods in real-world applications. In addition, it is difficult for researchers to analyze IAD algorithms without a uniform benchmark. To solve this problem, we propose a uniform IM benchmark, for the first time, to assess how well these algorithms perform, which includes various levels of supervision (unsupervised versus fully supervised), learning paradigms (few-shot, continual and noisy label), and efficiency (memory usage and inference speed). Then, we construct a comprehensive IAD benchmark (IM-IAD), which includes 19 algorithms on seven major datasets with a uniform setting. Extensive experiments (17 017 total) on IM-IAD provide in-depth insights into IAD algorithm redesign or selection. Moreover, the proposed IM-IAD benchmark challenges existing algorithms and suggests future research directions. For reproducibility and accessibility, the source code is uploaded to the website: https://github.com/M-3LAB/open-iad. Guoyang Xie, Jinbao Wang 0001, Jiaqi Liu 0004, Jiayi Lyu, Yong Liu 0032, Chengjie Wang 0001, Feng Zheng 0001, Yaochu Jin |
IEEE Trans. Cybern. | 4 |
| 2023 | FedMed-GAN: Federated domain translation on unsupervised cross-modality brain image synthesis
Jinbao Wang 0001, Guoyang Xie, Yawen Huang, Jiayi Lyu, Feng Zheng 0001, Yefeng Zheng 0001, Yaochu Jin |
Neurocomputing | 4 |
| 2023 | A Facial Landmark Detection Method Based on Deep Knowledge TransferabstractFacial landmark detection is a crucial preprocessing step in many applications that process facial images. Deep-learning-based methods have become mainstream and achieved outstanding performance in facial landmark detection. However, accurate models typically have a large number of parameters, which results in high computational complexity and execution time. A simple but effective facial landmark detection model that achieves a balance between accuracy and speed is crucial. To achieve this, a lightweight, efficient, and effective model is proposed called the efficient face alignment network (EfficientFAN) in this article. EfficientFAN adopts the encoder-decoder structure, with a simple backbone EfficientNet-B0 as the encoder and three upsampling layers and convolutional layers as the decoder. Moreover, deep dark knowledge is extracted through feature-aligned distillation and patch similarity distillation on the teacher network, which contains pixel distribution information in the feature space and multiscale structural information in the affinity space of feature maps. The accuracy of EfficientFAN is further improved after it absorbs dark knowledge. Extensive experimental results on public datasets, including 300 Faces in the Wild (300W), Wider Facial Landmarks in the Wild (WFLW), and Caltech Occluded Faces in the Wild (COFW), demonstrate the superiority of EfficientFAN over state-of-the-art methods. Ke Lu 0002, Jian Xue 0002, Jiayi Lyu, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | A Coarse-to-Fine Facial Landmark Detection Method Based on Self-attention MechanismabstractFacial landmark detection in the wild remains a challenging problem in computer vision. Deep learning-based methods currently play a leading role in solving this. However, these approaches generally focus on local feature learning and ignore global relationships. Therefore, in this study, a self-attention mechanism is introduced into facial landmark detection. Specifically, a coarse-to-fine facial landmark detection method is proposed that uses two stacked hourglasses as the backbone, with a new landmark-guided self-attention (LGSA) block inserted between them. The LGSA block learns the global relationships between different positions on the feature map and allows feature learning to focus on the locations of landmarks with the help of a landmark-specific attention map, which is generated in the first-stage hourglass model. A novel attentional consistency loss is also proposed to ensure the generation of an accurate landmark-specific attention map. A new channel transformation block is used as the building block of the hourglass model to improve the model's capacity. The coarse-to-fine strategy is adopted during and between phases to reduce complexity. Extensive experimental results on public datasets demonstrate the superiority of our proposed method against state-of-the-art models. Ke Lu 0002, Jian Xue 0002, Ling Shao 0001, Jiayi Lyu |
IEEE Trans. Multim. | 5 |