EDBT 2026 Demo / reviewers in the wild / expert
Zefang Yu
dblp:272/3989
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2025
0009-0007-7198-3664ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Training-Free Correlation-Weighted Model for Zero-/Few-Shot Industrial Anomaly Detection with Retrieval AugmentationabstractObtaining labeled data in the field of industrial anomaly detection is challenging, which necessitates the development of label-free frameworks. However, current methods mainly focus on the unsupervised paradigm, which uses a large number of normal samples of the same category to train the model, and distinguish anomalies during testing. This training approach necessitates retraining when new datasets or object categories are encountered. Recently, studies have suggested using large pre-trained multimodal vision-language models, such as CLIP, for zero-shot and few-shot anomaly detection, yielding promising outcomes. However, the lack of spatial awareness of these models results in less effectiveness in dense prediction tasks such as anomaly localization. To mitigate this issue, various fine-tuning methods using additional labeled anomaly data have been employed. In other words, substantial data and extensive training efforts are still necessary to ensure optimal model performance on specific datasets. In this paper, we introduce a training-free, CLIP-based model that utilizes patch correlations and prototype guidance to enable zero-shot and few-shot anomaly detection. Specifically, we first use a self-supervised pre-trained model to capture patch correlations within a single image, enhancing the model's regional awareness of defects. Then, we dynamically construct prototypes using a retrieval-enhanced method to alleviate domain gap in general domain models for anomaly detection. Extensive experiments on the popular benchmarks MVTec and VisA demonstrate that our approach achieves state-of-the-art performance across nearly all metrics. Furthermore, we validate the generalization of our method on collected real industrial data. Wei Ran, Zefang Yu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 2 |
| 2025 | GPA: Enhancing Generalizable Physical Adversarial Attacks Across Multiple Vision TasksabstractAdversarial attacks pose a significant challenge in deep learning, as carefully crafted perturbations can severely degrade even the most advanced models. In real-world scenarios, where the target models are often unknown, previous works often focus on creating adversarial patterns for specific known models, with the goal of generalizing these patterns to other models. However, such attacks rely heavily on prior model information, leading to poor generalization. To overcome this, we propose a novel method called GPA. Our solution includes an attention extraction module based on a pre-trained vision encoder, which captures precise and generalizable features of model attention on objects. We also introduce attack loss functions that divert attention away from target objects. Compared to state-of-the-art methods, our approach achieves superior attack performance across various downstream vision tasks, including object detection, instance segmentation, and depth estimation. Moreover, the adversarial patterns generated by GPA maintain their effectiveness in real-world scenarios. Mingye Xie, Suncheng Xiang, Jiacheng Ruan, Zefang Yu, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 5 |
| 2024 | LAMM: Label Alignment for Multi-Modal Prompt LearningabstractWith the success of pre-trained visual-language (VL) models such as CLIP in visual representation tasks, transferring pre-trained models to downstream tasks has become a crucial paradigm. Recently, the prompt tuning paradigm, which draws inspiration from natural language processing (NLP), has made significant progress in VL field. However, preceding methods mainly focus on constructing prompt templates for text and visual inputs, neglecting the gap in class label representations between the VL models and downstream tasks. To address this challenge, we introduce an innovative label alignment method named \textbf{LAMM}, which can dynamically adjust the category embeddings of downstream datasets through end-to-end training. Moreover, to achieve a more appropriate label distribution, we propose a hierarchical loss, encompassing the alignment of the parameter space, feature space, and logits space. We conduct experiments on 11 downstream vision datasets and demonstrate that our method significantly improves the performance of existing multi-modal prompt learning models in few-shot scenarios, exhibiting an average accuracy improvement of 2.31(\%) compared to the state-of-the-art methods on 16 shots. Moreover, our methodology exhibits the preeminence in continual learning compared to other prompt tuning methods. Importantly, our method is synergistic with existing prompt tuning methods and can boost the performance on top of them. Our code and dataset will be publicly available at https://github.com/gaojingsheng/LAMM. Jingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu, Ke Ji, Mingye Xie, Ting Liu 0016, Yuzhuo Fu |
AAAI | 4 |
| 2024 | From Raw Video to Pedagogical Insights: A Unified Framework for Student Behavior AnalysisabstractUnderstanding student behavior in educational settings is critical in improving both the quality of pedagogy and the level of student engagement. While various AI-based models exist for classroom analysis, they tend to specialize in limited tasks and lack generalizability across diverse educational environments. Additionally, these models often fall short in ensuring student privacy and in providing actionable insights accessible to educators. To bridge this gap, we introduce a unified, end-to-end framework by leveraging temporal action detection techniques and advanced large language models for a more nuanced student behavior analysis. Our proposed framework provides an end-to-end pipeline that starts with raw classroom video footage and culminates in the autonomous generation of pedagogical reports. It offers a comprehensive and scalable solution for student behavior analysis. Experimental validation confirms the capability of our framework to accurately identify student behaviors and to produce pedagogically meaningful insights, thereby setting the stage for future AI-assisted educational assessments. Zefang Yu, Mingye Xie, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu |
AAAI | 1 |
| 2024 | Learning to Floorplan like Human Experts via Reinforcement LearningabstractDeep reinforcement learning (RL) has gained popularity for automatically generating placements in modern chip design. However, the visual style of the fioorplans generated by these RL models is significantly different from the manual layouts' style, for RL placers usually only adopt metrics like wirelength and routing congestion as the reward in reinforcement learning, ignoring the complex and fine-grained layout experience of human experts. In this paper, we propose a placement scorer to rate the quality of layouts and apply abnormal detection to the fioorplanning task. In addition, we add the output of this scorer as a part of the reward for reinforcement learning of the placement process. Experimental results on ISPD 2005 benchmark show that our proposed placement quality scorer can evaluate the layouts according to human craft style efficiently, and that adding this scorer into reinforcement learning reward helps generating placements with shorter wirelength than previous methods for some circuit designs. Binjie Yan, Zefang Yu, Mingye Xie, Wei Ran, Jingsheng Gao, Yuzhuo Fu, Ting Liu 0016 |
DATE | 3 |
| 2024 | GIST: Improving Parameter Efficient Fine-Tuning via Knowledge InteractionabstractRecently, the Parameter Efficient Fine-Tuning (PEFT) method, which adjusts or introduces fewer trainable parameters to calibrate pre-trained models on downstream tasks, has been a hot research topic. However, existing PEFT methods within the traditional fine-tuning framework have two main shortcomings: 1) They overlook the explicit association between trainable parameters and downstream knowledge. 2) They neglect the interaction between the intrinsic task-agnostic knowledge of pre-trained models and the task-specific knowledge of downstream tasks. These oversights lead to insufficient utilization of knowledge and suboptimal performance. To address these issues, we propose a novel fine-tuning framework, named GIST, that can be seamlessly integrated into the current PEFT methods in a plug-and-play manner. Specifically, our framework first introduces a trainable token, called the Gist token, when applying PEFT methods on downstream tasks. This token serves as an aggregator of the task-specific knowledge learned by the PEFT methods and builds an explicit association with downstream tasks. Furthermore, to facilitate explicit interaction between task-agnostic and task-specific knowledge, we introduce the concept of knowledge interaction via a Bidirectional Kullback-Leibler Divergence objective. As a result, PEFT methods within our framework can enable the pre-trained model to understand downstream tasks more comprehensively by fully leveraging both types of knowledge. Extensive experiments on the 35 datasets demonstrate the universality and scalability of our framework. Notably, the PEFT method within our GIST framework achieves up to a 2.25% increase on the VTAB-1K benchmark with an addition of just 0.8K parameters (0.009 of ViT-B/16). The code is available at https://github.com/JCruan519/GIST. Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Suncheng Xiang, Zefang Yu, Ting Liu 0016, Yuzhuo Fu, Xiaoye Qu |
ACM Multimedia | 5 |
| 2024 | Deep multimodal representation learning for generalizable person re-identification
Suncheng Xiang, Wei Ran, Zefang Yu, Ting Liu 0016, Dahong Qian, Yuzhuo Fu |
Mach. Learn. | 4 |
| 2023 | AV-TAD: Audio-Visual Temporal Action Detection With TransformerabstractAs an important and challenging task in video understanding, Temporal Action Detection (TAD) has been deeply studied in recent years. However, current works mainly tackle this task with visual information, while neglecting to explore the potential of the audio modality. To address this challenge, in this paper, we propose a simple yet effective AudioVisual Temporal Action Detection Transformer named AV- TAD, which performs early fusion on audio and visual modalities in an end-to-end fashion. On top of it, a novel query formulation is introduced by directly adopting temporal segment coordinates as queries in Transformer decoder, thus allowing us to perform dynamic segment update layer-by-layer. To the best of our knowledge, this is the first attempt to investigate both audio and video feature with a multi-modal Transformer in TAD task. Extensive experiments on THUMOS14 dataset demonstrate that our proposed AV-TAD can outperform the previous methods by a clear margin. Yangcheng Li, Zefang Yu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 2 |
| 2023 | CC-PoseNet: Towards Human Pose Estimation in Crowded ClassroomsabstractHuman pose estimation has long been motivated for its application in human behavior understanding and activity recognition. Despite recent advances in multi-person pose estimation, existing solutions remain challenging in crowded scenes, especially in classroom scenarios where students are extremely overlapped and have different poses. In this paper, we focus on improving human pose estimation in crowded classrooms from the perspective of crowd detection and pose refinement. Specifically, we first follow a top-down strategy to detect persons in a multi-instance prediction manner and perform single-person pose estimation on each detected human region. Then, the pose estimation is refined with Transformer blocks by capturing the interactions among multiple persons in the image. Importantly, we replace self-attention in Transformer with a lightweight attention mechanism to reduce computational complexity. Quantitative and qualitative experiments demonstrate that our method remarkably outperforms previous methods with a clear margin on both standard benchmarks and self-collected classroom images. Zefang Yu, Yanping Hu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 1 |
| 2023 | Colo-SCRL: Self-Supervised Contrastive Representation Learning for Colonoscopic Video RetrievalabstractColonoscopic video retrieval, which is a critical part of polyp treatment, has great clinical significance for the prevention and treatment of colorectal cancer. However, retrieval models trained on action recognition datasets usually produce unsatisfactory retrieval results on colonoscopic datasets due to the large domain gap between them. To seek a solution to this problem, we construct a large-scale colonoscopic dataset named Colo-Pair for medical practice. Based on this dataset, a simple yet effective training method called Colo-SCRL is proposed for more robust representation learning. It aims to refine general knowledge from colonoscopies through masked autoencoder-based reconstruction and momentum contrast to improve retrieval performance. To the best of our knowledge, this is the first attempt to employ the contrastive learning paradigm for medical video retrieval. Empirical results show that our method significantly outperforms current state-of-the-art methods in the colonoscopic video retrieval task. Qingzhong Chen, Shilun Cai, Crystal Cai, Zefang Yu, Dahong Qian, Suncheng Xiang |
ICME | 4 |
| 2022 | Synpose: A Large-Scale and Densely Annotated Synthetic Dataset for Human Pose Estimation in ClassroomabstractDeep learning-based methods for human pose estimation require large volumes of training data to achieve superior performance. However, data acquisition in classroom environments raises privacy concerns, which will undoubtedly hinder the development of the latest deep learning techniques in education domain. Due to the absence of large, richly annotated classroom datasets, research into classroom observation has had to be done by manually collecting and annotating datasets. Unfortunately, the annotation of such data is time-consuming and challenging in over-crowded classrooms. To break through these limitations, we open source SynPose, a large, densely labeled synthetic dataset specifically designed for crowded human pose estimation in classroom and meeting scenarios. Moreover, we propose a novel CTGAN to bridge the domain gap. Comprehensive experiments on real-world classroom images show that our proposed dataset and method deliver important performance benefits compared to existing datasets, revealing the potential of SynPose for future studies. Zefang Yu, Yangcheng Li, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 1 |
| 2020 | Unsupervised person re-identification by hierarchical cluster and domain transfer
Suncheng Xiang, Yuzhuo Fu, Mingye Xie, Zefang Yu, Ting Liu 0016 |
Multim. Tools Appl. | 4 |