EDBT 2026 Demo / reviewers in the wild / expert
Ting Liu 0016
dblp:52/5150-16
· DBLP profile ↗
44ranked-venue papers
2as first author
39since 2021 · last 2026
0000-0003-3489-4578ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 29 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language ModelsabstractRecently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously refined and improved. In this process, benchmarks in the scientific domain have played an important role in assessing the reasoning capabilities of MLLMs. However, existing benchmarks still face three key challenges: 1) Insufficient evaluation of models' reasoning abilities in multilingual scenarios; 2) Inadequate assessment of MLLMs' comprehensive modality coverage; 3) Lack of fine-grained annotation of scientific knowledge points. To address these gaps, we propose MME-SCI, a comprehensive and challenging benchmark. We carefully collected 1,019 high-quality question-answer pairs, which involve 3 distinct evaluation modes. These pairs cover four subjects, namely mathematics, physics, chemistry, and biology, and support five languages: Chinese, English, French, Spanish, and Japanese. We conducted extensive experiments on 16 open-source models and 4 closed-source models, and the results demonstrate that MME-SCI is widely challenging for existing MLLMs. For instance, under the Image-only evaluation mode, o4-mini achieved accuracy of only 52.11%, 24.73%, 36.57%, and 29.80% in mathematics, physics, chemistry, and biology, respectively, indicating a significantly higher difficulty level compared to existing benchmarks. More importantly, using MME-SCI's multilingual and fine-grained knowledge attributes, we analyzed existing models' performance in depth and identified their weaknesses in specific domains. For example, in questions related to "Magnetic Field", o4-mini correctly answered only 5 out of 33 questions, thereby fine-grainedly exposing the model's vulnerabilities. These findings highlight the urgent need to enhance the scientific reasoning capabilities of MLLMs. Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu, Yangyang Kang |
AAAI | 4 |
| 2026 | Unnoticed Yet Effective: A Hybrid Physical Camouflage Framework Against DNNs and Human PerceptionabstractWhile adversarial attacks can effectively deceive deep neural networks, their real-world applicability is often limited by complex and conspicuous patterns that reveal their attack intent to human observers. To overcome this limitation, we propose UYE, a novel camouflage framework designed to simultaneously mislead DNNs and evade human perception. UYE incorporates two key components: an attention refiner leveraging a pre-trained vision encoder to optimize adversarial patterns for robust attacks across diverse environments, and a perception evaluator trained on a preference dataset curated using tailored prompts from human-aligned large multimodal models to ensure natural and unobtrusive camouflage generation. Extensive experiments demonstrate that UYE outperforms state-of-the-art methods in achieving an optimal balance between human stealth and model deception while maintaining effectiveness in real-world scenarios. Mingye Xie, Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu |
AAAI | 4 |
| 2026 | DiverSeed: Integrating Active Learning for Target Domain Data Generation in Instruction Tuning
Jingsheng Gao, Mengnan Qi, Suncheng Xiang, Ke Ji, Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu |
Mach. Learn. | 7 |
| 2025 | MM-CamObj: A Comprehensive Multimodal Dataset for Camouflaged Object ScenariosabstractLarge visual-language models (LVLMs) have achieved great success in multiple applications. However, they still encounter challenges in complex scenes, especially those involving camouflaged objects. This is primarily due to the lack of samples related to camouflaged scenes in the training dataset. To mitigate this issue, we construct the MM-CamObj dataset for the first time, comprising two subsets: CamObj-Align and CamObj-Instruct. Specifically, CamObj-Align contains 11,363 image-text pairs, and it is designed for VL alignment and injecting rich knowledge of camouflaged scenes into LVLMs. CamObj-Instruct is collected for fine-tuning the LVLMs with improved instruction-following capabilities, and it includes 11,363 images and 68,849 conversations with diverse instructions. Based on the MM-CamObj dataset, we propose the CamObj-Llava, an LVLM specifically designed for addressing tasks in camouflaged scenes. To facilitate our model's effective acquisition of knowledge about camouflaged objects and scenes, we introduce a curriculum learning strategy with six distinct modes. Additionally, we construct the CamObj-Bench to evaluate the existing LVLMs' capabilities of understanding, recognition, localization and count in camouflage scenes. This benchmark includes 600 images and 7 tasks, with a total of 9,449 questions. Extensive experiments are conducted on the CamObj-Bench with CamObj-Llava, 8 existing open-source and 3 closed-source LVLMs. Surprisingly, the results indicate that our model achieves a 25.84% improvement in 4 out of 7 tasks compared to GPT-4o. Jiacheng Ruan, Wenzhen Yuan 0002, Zehao Lin, Ning Liao, Feiyu Xiong, Ting Liu 0016, Yuzhuo Fu |
AAAI | 7 |
| 2025 | TTE: Two Tokens Are Enough to Improve Parameter-Efficient TuningabstractExisting fine-tuning paradigms are predominantly characterized by Full Parameter Tuning (FPT) and Parameter-Efficient Tuning (PET). FPT fine-tunes all parameters of a pre-trained model on downstream tasks, whereas PET freezes the pre-trained model and employs only a minimal number of learnable parameters for fine-tuning. However, both approaches face issues of overfitting, especially in scenarios where downstream samples are limited. This issue has been thoroughly explored in FPT, but less so in PET. To this end, this paper investigates overfitting in PET, representing a pioneering study in the field. Specifically, across 19 image classification datasets, we employ three classic PET methods (e.g., VPT, Adapter/Adaptformer, and LoRA) and explore various regularization techniques to mitigate overfitting. Regrettably, the results suggest that existing regularization techniques are incompatible with the PET process and may even lead to performance degradation. Consequently, we introduce a new framework named TTE (Two Tokens are Enough), which effectively alleviates overfitting in PET through a novel constraint function based on the learnable tokens. Experiments conducted on 24 datasets across image and few-shot classification tasks demonstrate that our fine-tuning framework not only mitigates overfitting but also significantly enhances PET's performance. Notably, our TTE framework surpasses the highest-performing FPT framework (DR-Tune), utilizing significantly fewer parameters (0.15M vs. 85.84M) and achieving an improvement of 1%. Jiacheng Ruan, Mingye Xie, Jingsheng Gao, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
AAAI | 6 |
| 2025 | Exploring Generalization Boundaries of Unsupervised Industrial Anomaly Detection Models through Attribute PerturbationabstractIndustrial anomaly detection (IAD) plays a crucial role in large-scale industrial manufacturing. Recently, numerous unsupervised algorithms have been proposed and achieved remarkable performance on benchmark datasets. Given the high homogeneity of samples during training and testing, it appears that the current datasets have been exhaustively addressed. However, it remains uncertain whether these state-of-the-art (SOTA) methods can perform well in more diverse real-world industrial scenarios, or merely overfit to the existing datasets. To address this issue, we propose an attribute perturbing framework utilizing foundation models. Based on it, we quantitatively analyze the impact of attribute perturbation on the anomaly detection system. To the best of our knowledge, we are the first to evaluate the generalization ability of current IAD methods under a shift in testing conditions Systematically. This work helps us have a deeper understanding of the limitations in IAD methods, which also offers valuable insights for future dataset building. Wei Ran, Yuzhuo Fu, Ting Liu 0016 |
ICASSP | 3 |
| 2025 | A Training-Free Correlation-Weighted Model for Zero-/Few-Shot Industrial Anomaly Detection with Retrieval AugmentationabstractObtaining labeled data in the field of industrial anomaly detection is challenging, which necessitates the development of label-free frameworks. However, current methods mainly focus on the unsupervised paradigm, which uses a large number of normal samples of the same category to train the model, and distinguish anomalies during testing. This training approach necessitates retraining when new datasets or object categories are encountered. Recently, studies have suggested using large pre-trained multimodal vision-language models, such as CLIP, for zero-shot and few-shot anomaly detection, yielding promising outcomes. However, the lack of spatial awareness of these models results in less effectiveness in dense prediction tasks such as anomaly localization. To mitigate this issue, various fine-tuning methods using additional labeled anomaly data have been employed. In other words, substantial data and extensive training efforts are still necessary to ensure optimal model performance on specific datasets. In this paper, we introduce a training-free, CLIP-based model that utilizes patch correlations and prototype guidance to enable zero-shot and few-shot anomaly detection. Specifically, we first use a self-supervised pre-trained model to capture patch correlations within a single image, enhancing the model's regional awareness of defects. Then, we dynamically construct prototypes using a retrieval-enhanced method to alleviate domain gap in general domain models for anomaly detection. Extensive experiments on the popular benchmarks MVTec and VisA demonstrate that our approach achieves state-of-the-art performance across nearly all metrics. Furthermore, we validate the generalization of our method on collected real industrial data. Wei Ran, Zefang Yu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 4 |
| 2025 | GPA: Enhancing Generalizable Physical Adversarial Attacks Across Multiple Vision TasksabstractAdversarial attacks pose a significant challenge in deep learning, as carefully crafted perturbations can severely degrade even the most advanced models. In real-world scenarios, where the target models are often unknown, previous works often focus on creating adversarial patterns for specific known models, with the goal of generalizing these patterns to other models. However, such attacks rely heavily on prior model information, leading to poor generalization. To overcome this, we propose a novel method called GPA. Our solution includes an attention extraction module based on a pre-trained vision encoder, which captures precise and generalizable features of model attention on objects. We also introduce attack loss functions that divert attention away from target objects. Compared to state-of-the-art methods, our approach achieves superior attack performance across various downstream vision tasks, including object detection, instance segmentation, and depth estimation. Moreover, the adversarial patterns generated by GPA maintain their effectiveness in real-world scenarios. Mingye Xie, Suncheng Xiang, Jiacheng Ruan, Zefang Yu, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 6 |
| 2025 | VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
Jiacheng Ruan, Wenzhen Yuan 0002, Xiqi Gao 0001, Daoxin Zhang, Yao Hu 0002, Ting Liu 0016, Yuzhuo Fu |
ICCV | 8 |
| 2025 | NLOSdiffuser: Generalized Steady-State Non-Line-of-sight Imaging toward Indoor ScenariosabstractNon-Line-of-sight (NLOS) imaging broadens the field of view to observe objects beyond direct sightlines, attracting considerable research interest. However, previous studies often overlook real-world environmental impacts, focusing on idealized scenarios, which limits their practical applicability. We introduce a steady-state NLOS imaging method tailored for indoor environments, showcasing robust generalization across varied settings despite training on static scenes. Our method employs a synthetic indoor dataset mimicking real-world conditions for training and diverse scenarios for evaluation. A two-stage network initially conducts pixel-level coarse reconstruction to convert diffuse reflection images into intermediate feature maps, followed by semantic-level reconstruction using a pre-trained diffusion model, yielding highly recognizable images. Comprehensive experiments demonstrate superior reconstruction accuracy, image quality, and adaptability compared to conventional methods. Verified in real-world scenarios, our method effectively images various concealed objects, marking significant progress in NLOS imaging. Jiacheng Ruan, Zongyun Zhang, Ting Liu 0016, Yuzhuo Fu |
ICME | 6 |
| 2025 | Bridging the Gap: Balancing Human Perception and Detector Attention in Adversarial AttacksabstractAdversarial attacks on deep neural networks have garnered significant attention, with recent studies focusing on generating intricate, colorful patterns designed to mislead models. However, such patterns are often easily discernible to human observers. To address this limitation, we propose a novel adversarial camouflage framework BHD that simultaneously mitigates human perceptibility and reduces detector attention. Our approach defines a base pattern by leveraging the surroundings of target object, making it less conspicuous to the human eye. We further design a loss function, integrated with a pre-trained vision encoder with fine-tuned projector, to optimize the adversarial pattern. This allows for effective attacks on detectors while ensuring robust camouflage across diverse environments. In addition, we incorporate human-aligned large multimodal models as objective metrics to quantify the impact on human perception. Extensive experimental results demonstrate that our framework achieves a superior balance between human imperceptibility and model deception, outperforming state-of-the-art methods. Mingye Xie, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICME | 4 |
| 2025 | EIAD: Explainable Industrial Anomaly Detection Via Multi-Modal Large Language ModelsabstractIndustrial Anomaly Detection (IAD) is critical to ensure product quality during manufacturing. Although existing zero-shot defect segmentation and detection methods have shown effectiveness, they cannot provide detailed descriptions of the defects. Furthermore, the application of large multi-modal models in IAD remains in its infancy, facing challenges in balancing question-answering (QA) performance and mask-based grounding capabilities, often owing to overfitting during the fine-tuning process. To address these challenges, we propose a novel approach that introduces a dedicated multi-modal defect localization module to decouple the dialog functionality from the core feature extraction. This decoupling is achieved through independent optimization objectives and tailored learning strategies. Additionally, we contribute to the first multi-modal industrial anomaly detection training dataset, named Defect Detection Question Answering (DDQA), encompassing a wide range of defect types and industrial scenarios. Unlike conventional datasets that rely on GPT-generated data, DDQA ensures authenticity and reliability and offers a robust foundation for model training. Experimental results demonstrate that our proposed method, Explainable Industrial Anomaly Detection Assistant (EIAD), achieves outstanding performance in defect detection and localization tasks. It not only significantly enhances accuracy but also improves interpretability. These advancements highlight the potential of EIAD for practical applications in industrial settings. Zongyun Zhang, Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu |
ICME | 4 |
| 2025 | MPI-CD: Multi-Path Information Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language ModelsabstractIn recent years, despite substantial advancements in large vision-language models (LVLMs), they still encounter the issue of ''hallucinations''-where generated results appears reasonable but often deviates from the visual input or actual facts. In contrast, the human cognitive system, when processing visual input, initially relies on visual perception to distinguish between the salient region and non-salient region, integrating relevant information. Subsequently, it recalls pertinent memory details, ultimately generating a comprehensive cognitive outcome. Inspired by this process, we propose a novel, training-free decoding approach, dubbed as Multi-Path Information Contrastive Decoding (MPI-CD). Specifically, to simulate the human information integration process, we design a three-branch structure called the Tri-Branch Integrator (TBI), which contrasts the original, salient region, and non-salient region images to effectively improve the reliability of the LVLMs' output. Furthermore, to mimic the human memory recall mechanism, we further investigate the importance of hidden layer features and propose the Memory Recall Module (MRM). This module adaptively extracts meaningful memory information from the hidden layers and incorporates it into the decoding process, thereby effectively alleviating the hallucination issue. We conduct extensive experiments on three widely used benchmarks (e.g. POPE, AMBER, and MME) using two classic LVLMs. The experimental results demonstrate that our MPI-CD significantly mitigates hallucinations in LVLMs without requiring additional training. Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao, Wenzhen Yuan 0002, Ting Liu 0016, Yuzhuo Fu |
ACM Multimedia | 5 |
| 2024 | LAMM: Label Alignment for Multi-Modal Prompt LearningabstractWith the success of pre-trained visual-language (VL) models such as CLIP in visual representation tasks, transferring pre-trained models to downstream tasks has become a crucial paradigm. Recently, the prompt tuning paradigm, which draws inspiration from natural language processing (NLP), has made significant progress in VL field. However, preceding methods mainly focus on constructing prompt templates for text and visual inputs, neglecting the gap in class label representations between the VL models and downstream tasks. To address this challenge, we introduce an innovative label alignment method named \textbf{LAMM}, which can dynamically adjust the category embeddings of downstream datasets through end-to-end training. Moreover, to achieve a more appropriate label distribution, we propose a hierarchical loss, encompassing the alignment of the parameter space, feature space, and logits space. We conduct experiments on 11 downstream vision datasets and demonstrate that our method significantly improves the performance of existing multi-modal prompt learning models in few-shot scenarios, exhibiting an average accuracy improvement of 2.31(\%) compared to the state-of-the-art methods on 16 shots. Moreover, our methodology exhibits the preeminence in continual learning compared to other prompt tuning methods. Importantly, our method is synergistic with existing prompt tuning methods and can boost the performance on top of them. Our code and dataset will be publicly available at https://github.com/gaojingsheng/LAMM. Jingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu, Ke Ji, Mingye Xie, Ting Liu 0016, Yuzhuo Fu |
AAAI | 7 |
| 2024 | From Raw Video to Pedagogical Insights: A Unified Framework for Student Behavior AnalysisabstractUnderstanding student behavior in educational settings is critical in improving both the quality of pedagogy and the level of student engagement. While various AI-based models exist for classroom analysis, they tend to specialize in limited tasks and lack generalizability across diverse educational environments. Additionally, these models often fall short in ensuring student privacy and in providing actionable insights accessible to educators. To bridge this gap, we introduce a unified, end-to-end framework by leveraging temporal action detection techniques and advanced large language models for a more nuanced student behavior analysis. Our proposed framework provides an end-to-end pipeline that starts with raw classroom video footage and culminates in the autonomous generation of pedagogical reports. It offers a comprehensive and scalable solution for student behavior analysis. Experimental validation confirms the capability of our framework to accurately identify student behaviors and to produce pedagogically meaningful insights, thereby setting the stage for future AI-assisted educational assessments. Zefang Yu, Mingye Xie, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu |
AAAI | 4 |
| 2024 | Learning to Floorplan like Human Experts via Reinforcement LearningabstractDeep reinforcement learning (RL) has gained popularity for automatically generating placements in modern chip design. However, the visual style of the fioorplans generated by these RL models is significantly different from the manual layouts' style, for RL placers usually only adopt metrics like wirelength and routing congestion as the reward in reinforcement learning, ignoring the complex and fine-grained layout experience of human experts. In this paper, we propose a placement scorer to rate the quality of layouts and apply abnormal detection to the fioorplanning task. In addition, we add the output of this scorer as a part of the reward for reinforcement learning of the placement process. Experimental results on ISPD 2005 benchmark show that our proposed placement quality scorer can evaluate the layouts according to human craft style efficiently, and that adding this scorer into reinforcement learning reward helps generating placements with shorter wirelength than previous methods for some circuit designs. Binjie Yan, Zefang Yu, Mingye Xie, Wei Ran, Jingsheng Gao, Yuzhuo Fu, Ting Liu 0016 |
DATE | 8 |
| 2024 | Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition
Suncheng Xiang, Jingsheng Gao, Jiacheng Ruan, Yanping Hu, Ting Liu 0016, Yuzhuo Fu |
ICIC (4) | 7 |
| 2024 | iDAT: inverse Distillation Adapter-TuningabstractAdapter-Tuning (AT) method involves freezing a pre-trained model and introducing trainable adapter modules to acquire downstream knowledge, thereby calibrating the model for better adaptation to downstream tasks. This paper proposes a distillation framework for the AT method instead of crafting a carefully designed adapter module, which aims to improve fine-tuning performance. For the first time, we explore the possibility of combining the AT method with knowledge distillation. Via statistical analysis, we observe significant differences in the knowledge acquisition between adapter modules of different models. Leveraging these differences, we propose a simple yet effective framework called inverse Distillation Adapter-Tuning (iDAT). Specifically, we designate the smaller model as the teacher and the larger model as the student. The two are jointly trained, and online knowledge distillation is applied to inject knowledge of different perspective to student model, and significantly enhance the fine-tuning performance on downstream tasks. Extensive experiments on the VTAB-1K benchmark with 19 image classification tasks demonstrate the effectiveness of iDAT. The results show that using existing AT method within our iDAT framework can further yield a 2.66% performance gain, with only an additional 0.07M trainable parameters. Our approach compares favorably with state-of-the-arts without bells and whistles. Our code is available at https://github.com/JCruan519/iDAT. Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Daize Dong, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICME | 6 |
| 2024 | GIST: Improving Parameter Efficient Fine-Tuning via Knowledge InteractionabstractRecently, the Parameter Efficient Fine-Tuning (PEFT) method, which adjusts or introduces fewer trainable parameters to calibrate pre-trained models on downstream tasks, has been a hot research topic. However, existing PEFT methods within the traditional fine-tuning framework have two main shortcomings: 1) They overlook the explicit association between trainable parameters and downstream knowledge. 2) They neglect the interaction between the intrinsic task-agnostic knowledge of pre-trained models and the task-specific knowledge of downstream tasks. These oversights lead to insufficient utilization of knowledge and suboptimal performance. To address these issues, we propose a novel fine-tuning framework, named GIST, that can be seamlessly integrated into the current PEFT methods in a plug-and-play manner. Specifically, our framework first introduces a trainable token, called the Gist token, when applying PEFT methods on downstream tasks. This token serves as an aggregator of the task-specific knowledge learned by the PEFT methods and builds an explicit association with downstream tasks. Furthermore, to facilitate explicit interaction between task-agnostic and task-specific knowledge, we introduce the concept of knowledge interaction via a Bidirectional Kullback-Leibler Divergence objective. As a result, PEFT methods within our framework can enable the pre-trained model to understand downstream tasks more comprehensively by fully leveraging both types of knowledge. Extensive experiments on the 35 datasets demonstrate the universality and scalability of our framework. Notably, the PEFT method within our GIST framework achieves up to a 2.25% increase on the VTAB-1K benchmark with an addition of just 0.8K parameters (0.009 of ViT-B/16). The code is available at https://github.com/JCruan519/GIST. Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Suncheng Xiang, Zefang Yu, Ting Liu 0016, Yuzhuo Fu, Xiaoye Qu |
ACM Multimedia | 6 |
| 2024 | Deep multimodal representation learning for generalizable person re-identification
Suncheng Xiang, Wei Ran, Zefang Yu, Ting Liu 0016, Dahong Qian, Yuzhuo Fu |
Mach. Learn. | 5 |
| 2024 | Toward an end-to-end implicit addressee modeling for dialogue disentanglement
Jingsheng Gao, Suncheng Xiang, Zhuowei Wang 0003, Ting Liu 0016, Yuzhuo Fu |
Multim. Tools Appl. | 5 |
| 2024 | Editing outdoor scenes with a large annotated synthetic dataset
Mingye Xie, Zongwei Liu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
Multim. Tools Appl. | 4 |
| 2024 | Rethinking Person Re-Identification via Semantic-based PretrainingabstractPretraining is a dominant paradigm in computer vision. Generally, supervised ImageNet pretraining is commonly used to initialize the backbones of person re-identification (Re-ID) models. However, recent works show a surprising result that CNN-based pretraining on ImageNet has limited impacts on Re-ID system due to the large domain gap between ImageNet and person Re-ID data. To seek an alternative to traditional pretraining, here we investigate semantic-based pretraining as another method to utilize additional textual data against ImageNet pretraining. Specifically, we manually construct a diversified FineGPR-C caption dataset for the first time on person Re-ID events. Based on it, a pure semantic-based pretraining approach named VTBR is proposed to adopt dense captions to learn visual representations with fewer images. We train convolutional neural networks from scratch on the captions of FineGPR-C dataset, and then transfer them to downstream Re-ID tasks. Comprehensive experiments conducted on benchmark datasets show that our VTBR can achieve competitive performance compared with ImageNet pretraining—despite using up to 1.4× fewer images, revealing its potential in Re-ID pretraining. Our source code is also publicly available at https://github.com/JeremyXSC/VTBR . Suncheng Xiang, Dahong Qian, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Action Recognition via Fine-Tuned CLIP Model and Temporal Transformer
Yuzhuo Fu, Ting Liu 0016 |
CGI (3) | 3 |
| 2023 | AV-TAD: Audio-Visual Temporal Action Detection With TransformerabstractAs an important and challenging task in video understanding, Temporal Action Detection (TAD) has been deeply studied in recent years. However, current works mainly tackle this task with visual information, while neglecting to explore the potential of the audio modality. To address this challenge, in this paper, we propose a simple yet effective AudioVisual Temporal Action Detection Transformer named AV- TAD, which performs early fusion on audio and visual modalities in an end-to-end fashion. On top of it, a novel query formulation is introduced by directly adopting temporal segment coordinates as queries in Transformer decoder, thus allowing us to perform dynamic segment update layer-by-layer. To the best of our knowledge, this is the first attempt to investigate both audio and video feature with a multi-modal Transformer in TAD task. Extensive experiments on THUMOS14 dataset demonstrate that our proposed AV-TAD can outperform the previous methods by a clear margin. Yangcheng Li, Zefang Yu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 4 |
| 2023 | CC-PoseNet: Towards Human Pose Estimation in Crowded ClassroomsabstractHuman pose estimation has long been motivated for its application in human behavior understanding and activity recognition. Despite recent advances in multi-person pose estimation, existing solutions remain challenging in crowded scenes, especially in classroom scenarios where students are extremely overlapped and have different poses. In this paper, we focus on improving human pose estimation in crowded classrooms from the perspective of crowd detection and pose refinement. Specifically, we first follow a top-down strategy to detect persons in a multi-instance prediction manner and perform single-person pose estimation on each detected human region. Then, the pose estimation is refined with Transformer blocks by capturing the interactions among multiple persons in the image. Importantly, we replace self-attention in Transformer with a lightweight attention mechanism to reduce computational complexity. Quantitative and qualitative experiments demonstrate that our method remarkably outperforms previous methods with a clear margin on both standard benchmarks and self-collected classroom images. Zefang Yu, Yanping Hu, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 4 |
| 2023 | DynaSlim: Dynamic Slimming for Vision TransformersabstractVision transformers (ViTs) have achieved significant performance on various vision tasks. However, high computational and memory costs hinder their edge deployment. Existing compression methods employ static constraints between accuracy and efficiency during sparsification. The static constraints restrict the sparsification efficiency and their initialization relies heavily on human expertise. We propose a dynamic slimming strategy for ViT, DynaSlim, to achieve an adaptive accuracy-efficiency constraint during sparsification. We first equip fine-grained, adjustable sparsity weights, the scaling factor between accuracy and efficiency, for multiple dimensions, including input tokens, Multihead Self-Attention (MSA) and Multilayer Perceptron (MLP). We then employ the heuristic search for these non-differentiable factors and combine the search with regularization-based sparsification to obtain the optimal sparsed model. Finally, we compress and retrain the sparsed model under various budgets to get our resulting submodels. Experiments show that our DynaSlim outperforms previous state-of-the-art methods under different budgets. For example, we reduce both parameters and FLOPs of DeiT-B by 39% while increasing its accuracy by 1.9% on ImageNet-1K. Moreover, we demonstrate the transferability of our compressed models on several downstream datasets. Da Shi, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu |
ICME | 3 |
| 2023 | EGE-UNet: An Efficient Group Enhanced UNet for Skin Lesion Segmentation
Jiacheng Ruan, Mingye Xie, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu |
MICCAI (4) | 4 |
| 2023 | Learning from self-discrepancy via multiple co-teaching for cross-domain person re-identification
Suncheng Xiang, Yuzhuo Fu, Mengyuan Guan, Ting Liu 0016 |
Mach. Learn. | 4 |
| 2023 | Less Is More: Learning from Synthetic Data with Fine-Grained Attributes for Person Re-IdentificationabstractPerson re-identification (ReID) plays an important role in applications such as public security and video surveillance. Recently, learning from synthetic data [ 9 ], which benefits from the popularity of the synthetic data engine, has attracted great attention from the public. However, existing datasets are limited in quantity, diversity, and realisticity, and cannot be efficiently used for the ReID problem. To address this challenge, we manually construct a large-scale person dataset named FineGPR with fine-grained attribute annotations. Moreover, aiming to fully exploit the potential of FineGPR and promote the efficient training from millions of synthetic data, we propose an attribute analysis pipeline called AOST based on the traditional machine learning algorithm, which dynamically learns attribute distribution in a real domain, then eliminates the gap between synthetic and real-world data and thus is freely deployed to new scenarios. Experiments conducted on benchmarks demonstrate that FineGPR with AOST outperforms (or is on par with) existing real and synthetic datasets, which suggests its feasibility for the ReID task and proves the proverbial less-is-more principle. Our synthetic FineGPR dataset is publicly available at https://github.com/JeremyXSC/FineGPR . Suncheng Xiang, Dahong Qian, Mengyuan Guan, Binjie Yan, Ting Liu 0016, Yuzhuo Fu, Guanjie You |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | MALUNet: A Multi-Attention and Light-weight UNet for Skin Lesion SegmentationabstractRecently, some pioneering works have preferred applying more complex modules to improve segmentation performances. However, it is not friendly for actual clinical environments due to limited computing resources. To address this challenge, we propose a light-weight model to achieve competitive performances for skin lesion segmentation at the lowest cost of parameters and computational complexity so far. Briefly, we propose four modules: (1) DGA consists of dilated convolution and gated attention mechanisms to extract global and local feature information; (2) IEA, which is based on external attention to characterize the overall datasets and enhance the connection between samples; (3) CAB is composed of 1D convolution and fully connected layers to perform a global and local fusion of multi-stage features to generate attention maps at channel axis; (4) SAB, which operates on multi-stage features by a shared 2D convolution to generate attention maps at spatial axis. We combine four modules with our U-shape architecture and obtain a light-weight medical image segmentation model dubbed as MALUNet. Compared with UNet, our model improves the mIoU and DSC metrics by 2.39% and 1.49%, respectively, with a 44x and 166x reduction in the number of parameters and computational complexity. In addition, we conduct comparison experiments on two skin lesion segmentation datasets (ISIC2017 and ISIC2018). Experimental results show that our model achieves state-of-the-art in balancing the number of parameters, computational complexity and segmentation performances. Code is available at https://github.com/JCruan519/MALUNet. Jiacheng Ruan, Suncheng Xiang, Mingye Xie, Ting Liu 0016, Yuzhuo Fu |
BIBM | 4 |
| 2022 | GOS: A Large-Scale Annotated Outdoor Scene Synthetic DatasetabstractScene editing has attracted increasing research interests owing to its valuable applications in the field of photography and entertainment. With style-based GAN being proposed, images can be reasonably edited on specific semantics by manipulating in latent space of the generator. However, existing datasets cannot satisfy the demands of large amounts of diverse data and rich semantic annotations at the same time, which makes the existing method difficult to edit on the content of outdoor scene images. To address these problems, we propose a large-scale, diverse synthetic dataset called "GOS dataset" generated based on a video game, which contains fine-grained semantic annotations. Extensive experiments show that utilizing the features obtained from the annotations of our dataset achieves better performance in outdoor scene editing, especially for distance and viewpoint of scenes, which indicates the extracted features have a certain generalization capability. Mingye Xie, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 2 |
| 2022 | Synpose: A Large-Scale and Densely Annotated Synthetic Dataset for Human Pose Estimation in ClassroomabstractDeep learning-based methods for human pose estimation require large volumes of training data to achieve superior performance. However, data acquisition in classroom environments raises privacy concerns, which will undoubtedly hinder the development of the latest deep learning techniques in education domain. Due to the absence of large, richly annotated classroom datasets, research into classroom observation has had to be done by manually collecting and annotating datasets. Unfortunately, the annotation of such data is time-consuming and challenging in over-crowded classrooms. To break through these limitations, we open source SynPose, a large, densely labeled synthetic dataset specifically designed for crowded human pose estimation in classroom and meeting scenarios. Moreover, we propose a novel CTGAN to bridge the domain gap. Comprehensive experiments on real-world classroom images show that our proposed dataset and method deliver important performance benefits compared to existing datasets, revealing the potential of SynPose for future studies. Zefang Yu, Yangcheng Li, Ting Liu 0016, Yuzhuo Fu |
ICASSP | 4 |
| 2022 | Spatial Attention Guided Local Facial Attribute EditingabstractFacial attribute manipulation has attracted great attention from the public due to its wide range of applications. Aiming to smoothly manipulate the attributes of real facial images, it is critical to search for a proper latent code that aligns with the domain of pre-trained GAN for faithful inversion and controls the transformation within the scope of the attribute for precise editing. Previous methods mainly focused on improving the quality of reconstruction but often ignored the editing effect. To address this issue, we first propose a mapping network to manipulate latent code which is effective for diverse situations, and design a spatial attention network to predict binary mask of the certain attribute which encourages to only alter the relevant region of images and suppress irrelevant changes. In addition, we introduce a novel latent space into the GAN inversion framework which achieves high reconstruction quality especially preserving identity features and retains the ability to edit face attributes. Our methods pave the way to semantically meaningful and disentangled manipulations on both generated images and real images. Ex-perimental results indicate a clear improvement over the cur-rent state-of-the-art methods in various metrics. Mingye Xie, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu |
ICME | 4 |
| 2021 | A Hierarchical Assessment Strategy on Soft Error Propagation in Deep Learning ControllerabstractDeep learning techniques have been introduced into the field of intelligent controller design in recent years and become an effective alternative in complex control scenarios. In addition to improve control robustness, deep learning controllers (DLCs) also provide a potential fault tolerance to internal disturbances (such as soft errors) due to the inherent redundant structure of deep neural networks (DNNs). In this paper, we propose a hierarchical assessment to characterize the impact of soft errors on the dependability of a PID controller and its DLC alternative. Single-bit-flip injections in underlying hardware and time series data collection from multiple abstraction layers (ALs) are performed on a virtual prototype system based on an ARM Cortex-A9 CPU, with a PID controller and corresponding recurrent neural network (RNN) implemented DLC deployed on it. We employ generative adversarial networks and Bayesian networks to characterize the local and global dependencies caused by soft errors across the system. By analyzing cross-AL fault propagation paths and component sensitivities, we discover that the parallel data processing pipelines and regular feature size scaling mechanism in DLC can effectively prevent critical failure causing faults from propagating to the control output. Ting Liu 0016, Yuzhuo Fu |
ASP-DAC | 1 |
| 2021 | Taking A Closer Look at Synthesis: Fine-Grained Attribute Analysis for Person Re-IdentificationabstractPerson re-identification (re-ID) plays an important role in applications such as public security and video surveillance. Recently, learning from synthetic data, which benefits from the popularity of synthetic data engine, has achieved remarkable performance. However, in pursuit of high accuracy, researchers in the academic always focus on training with large-scale datasets at a high cost of time and label expenses, while neglect to explore the potential of performing efficient training from millions of synthetic data. To facilitate development in this field, we reviewed the previously developed synthetic dataset GPR and built an improved one (GPR+) with larger number of identities and distinguished attributes. Based on it, we quantitatively analyze the influence of dataset attribute on re-ID system. To our best knowledge, we are among the first attempts to explicitly dissect person re-ID from the aspect of attribute on synthetic dataset. This research helps us have a deeper understanding of the fundamental problems in person re-ID, which also provides useful insights for dataset building and future practical usage. Suncheng Xiang, Yuzhuo Fu, Guanjie You, Ting Liu 0016 |
ICASSP | 4 |
| 2021 | DFFCN: Dual Flow Fusion Convolutional Network for Micro Expression Recognition
Jinjie Chen, Yuzhuo Fu, Yibo Jin 0003, Ting Liu 0016 |
ICONIP (3) | 4 |
| 2021 | Dynamic Channel Pruning for Real-Time Object Detection Networks
Yibo Jin 0003, Ting Liu 0016, Jinjie Chen, Yuzhuo Fu |
ICONIP (5) | 2 |
| 2021 | A cross-layer fault propagation analysis method for edge intelligence systems deployed with DNNs
Ting Liu 0016, Yuzhuo Fu |
J. Syst. Archit. | 1 |
| 2020 | Unsupervised Domain Adaptation Through Synthesis For Person Re-IdentificationabstractPerson re-identification is a hot topic because of its widespread applications in video surveillance and public security. However, it remains a challenging task because of drastic variations in illumination or background across surveillance cameras, which causes the current methods can not work well in real-world scenarios. In addition, due to the scarce dataset, many methods suffer from over-fitting to a different extent. To remedy the above two problems, firstly, we develop a data collector and labeler, which can generate the synthetic random scenes and simultaneously annotate them without any manpower. Based on it, we build a large-scale, diverse synthetic dataset. Secondly, we propose a novel unsupervised Re-ID method via domain adaptation, which can exploit the synthetic data to boost the performance of re-identification in a completely unsupervised way, and free humans from heavy data annotations. Extensive experiments show that our proposed method achieves the state-of-the-art performance on two benchmark datasets, and is very competitive with current cross-domain Re-ID method. Suncheng Xiang, Yuzhuo Fu, Guanjie You, Ting Liu 0016 |
ICME | 4 |
| 2020 | Progressive learning with style transfer for distant domain adaptationabstractThis article studies a novel transfer learning problem termed distant domain transfer learning. Different from traditional transfer learning which assumes there is a close relation between source and target data, in this study, the objective is to execute an unseen and unrelated task based on a labelled data set training previously without any samples from intermediate domains. To this end, the authors propose deep unsupervised progressive learning (DUPL) framework and its upgraded version, end‐to‐end DUPL (eDUPL). eDUPL consists of two components, i.e. (i) translating the style of labelled images from irrelevant source domain to the target domain and (ii) learning a domain adaptation model with progressive learning for testing on the target domain. In comparison, eDUPL can integrate the two components of the framework seamlessly. In general, the proposed method is easy to be implemented and can be viewed as a strong convolutional baseline for distant domain adaptation task. Comprehensive experiments based on VeRi Vehicle, CUB‐200‐2011 Birds and Oxford5k Buildings data sets are conducted and the results indicate that the proposed method robustly achieves state‐of‐the‐art performances compared with existing approaches, which demonstrates the effectiveness and superiority of the proposed algorithm. Suncheng Xiang, Yuzhuo Fu, Ting Liu 0016 |
IET Image Process. | 3 |
| 2020 | Multi-level feature learning with attention for person re-identification
Suncheng Xiang, Yuzhuo Fu, Wei Ran, Ting Liu 0016 |
Multim. Tools Appl. | 5 |
| 2020 | Unsupervised person re-identification by hierarchical cluster and domain transfer
Suncheng Xiang, Yuzhuo Fu, Mingye Xie, Zefang Yu, Ting Liu 0016 |
Multim. Tools Appl. | 5 |
| 2019 | Deep Unsupervised Progressive Learning for Distant Domain AdaptationabstractThe superiority of deeply learned representation has been reported in very recent literature of re-identification (Re-ID) task. In this paper, we study a novel transfer learning problem termed Distant Domain Transfer Learning (DDTL) for Re-ID task. Different from existing transfer learning problems which assume that there is a close relation between source domain and target domain, in the DDTL problem, target domain can be totally different from source domain. For example, the source domain classifies pedestrian images but the target domain distinguishes vehicle images. In this work, our goal is to execute an unseen and unrelated task based on a labeled dataset training previously without any samples from intermediate domains. Particularly, we consider the more pragmatic issue of learning a deep feature with no labels, and propose a Deep Unsupervised Progressive Learning (DUPL) method to transfer pretrained deep representations to unseen domains. Specifically, our work performs clustering and fine-tuning of the CNN to improve the performance of original model trained on the irrelevant labeled dataset. Empirical studies on distant domain adaptation task (pedestrian -> vehicle) demonstrate the effectiveness of the proposed method, and the improvement in terms of the mAP accuracy is up to 15% over "non-transfer" methods. Suncheng Xiang, Yuzhuo Fu, Ting Liu 0016 |
ICTAI | 3 |