EDBT 2026 Demo / reviewers in the wild / expert
Yufan Liu 0001
dblp:51/10612-1
· DBLP profile ↗
36ranked-venue papers
10as first author
28since 2021 · last 2026
0000-0002-8426-9335ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 17 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Burst Image Quality Assessment: A New Benchmark and Unified Framework for Multiple Downstream TasksabstractIn recent years, the development of burst imaging technology has improved the capture and processing capabilities of visual data, enabling a wide range of applications. However, the redundancy in burst images leads to the increased storage and transmission demands, as well as reduced efficiency of downstream tasks. To address this, we propose a new task of Burst Image Quality Assessment (BuIQA), to evaluate the task-driven quality of each frame within a burst sequence, providing reasonable cues for burst image selection. Specifically, we establish the first benchmark dataset for BuIQA, consisting of 7,346 burst sequences with 45,827 images and 191,572 annotated quality scores for multiple downstream scenarios. Inspired by the data analysis, a unified BuIQA framework is proposed to achieve an efficient adaption for BuIQA under diverse downstream scenarios. Specifically, a task-driven prompt generation network is developed with heterogeneous knowledge distillation, to learn the priors of the downstream task. Then, the task-aware quality assessment network is introduced to assess the burst image quality based on the task prompt. Extensive experiments across 10 downstream scenarios demonstrate the impressive BuIQA performance of the proposed approach, outperforming the state-of-the-art. Furthermore, it can achieve 0.33 dB PSNR improvement in the downstream tasks of denoising and super-resolution, by applying our approach to select the high-quality burst frames. Xiaoye Liang, Lai Jiang 0004, Minglang Qiao, Yue Zhang 0082, Xin Deng 0002, Shengxi Li, Yufan Liu 0001, Mai Xu |
AAAI | 8 |
| 2026 | Compressed image super-resolution based on invertible degradation and restoration
Mai Xu, Lai Jiang 0004, Xin Deng 0002, Yue Zhang 0082, Yufan Liu 0001 |
Pattern Recognit. | 6 |
| 2026 | Multi-modal face anti-spoofing via self-supervised learning
Yufan Liu 0001, Lai Jiang 0004, Shengxi Li, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jinlong Lin |
Pattern Recognit. Lett. | 2 |
| 2026 | Deepfake Detection via Exploring Degradation InconsistencyabstractThe detection of face forgery has become increasingly vital due to the severe security concerns posed by face manipulation techniques. While recent studies on forgery detection have demonstrated promising results when the training and testing samples come from the same domains, the problem remains challenging when attempting to extend the detector to unseen methods. In this work, we propose an innovative approach to enhance the generalization capability of forgery detection methods by exploring degradation inconsistency clues interspersed between the background and the manipulated face regions. Our motivation stems from the observation that digital photos undergo different degradation during acquisition and transmission, resulting in backgrounds and faces from different sources containing distinct degradation patterns in the forged faces. The proposed framework, termed the Degradation Consistency Learning Framework, integrates two core components: a data generation network that modulates degradation transformations to obtain tampered facial images, and a detection network that mines degradation inconsistency clues from both spatial and frequency domains. These two components are tightly coupled through adversarial training, forming a dynamic architecture akin to a Generative Adversarial Network (GAN). Experimental results on different benchmark and evaluation protocols (i.e., indataset and cross-dataset) have demonstrated the effectiveness of our method. Weiming Bai, Yufan Liu 0001, Aixi Zhang, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Visual-Instructed Degradation Diffusion for All-in-One Image RestorationabstractImage restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that utilizes visual instruction-guided degradation diffusion. Unlike existing methods that rely on task-specific models or ambiguous text-based priors, Defusion constructs explicit visual instructions that align with the visual degradation patterns. These instructions are grounded by applying degradations to standardized visual elements, capturing intrinsic degradation features while agnostic to image semantics. Defusion then uses these visual instructions to guide a diffusion-based model that operates directly in the degradation space, where it reconstructs high-quality images by denoising the degradation effects with enhanced stability and generalizability. Comprehensive experiments demonstrate that Defusion outperforms state-of-the-art methods across diverse image restoration tasks, including complex and real-world degradations. Wenyang Luo, Haina Qin, Zewen Chen, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
CVPR | 7 |
| 2025 | Noise-Optimized Distribution Distillation for Dataset Condensation
Tongfei Liu, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Chenguang Ma |
ACM Multimedia | 2 |
| 2025 | Each Complexity Deserves a Pruning PolicyabstractThe established redundancy in visual tokens within large vision–language models (LVLMs) allows for pruning to effectively reduce their substantial computational demands. Empirical evidence from previous works indicates that visual tokens in later decoder stages receive less attention than shallow layers. Then, previous methods typically employ heuristics layer-specific pruning strategies where, although the number of tokens removed may differ across decoder layers, the overall pruning schedule is fixed and applied uniformly to all input samples and tasks, failing to align token elimination with the model’s holistic reasoning trajectory. Cognitive science indicates that human visual processing often begins with broad exploration to accumulate evidence before narrowing focus as the target becomes distinct. Our experiments reveal an analogous pattern in LVLMs. This observation strongly suggests that neither a fixed pruning schedule nor a heuristics layer-wise strategy can optimally accommodate the diverse complexities inherent in different inputs. To overcome this limitation, we introduce Complexity-Adaptive Pruning (AutoPrune), which is a training-free, plug-and-play framework that tailors pruning policies to varying sample and task complexities. Specifically, AutoPrune quantifies the mutual information between visual and textual tokens, and then projects this signal to a budget-constrained logistic retention curve. Each such logistic curve, defined by its unique shape, is shown to effectively correspond with the specific complexity of different tasks, and can easily guarantee adherence to a pre-defined computational constraints. We evaluate AutoPrune not only on standard vision-language tasks but also on Vision-Language-Action (VLA) models for autonomous driving. Notably, when applied to LLaVA-1.5-7B, our method prunes 89% of visual tokens and reduces inference FLOPs by 76.8%, but still retaining 96.7% of the original accuracy averaged over all tasks. This corresponds to a 9.1% improvement over the recent work PDrop (CVPR'2025), demonstrating the effectivenes. Code is available at https://github.com/AutoLab-SAI-SJTU/AutoPrune. Hanshi Wang, Yufan Liu 0001, Weiming Hu 0004 |
NeurIPS | 5 |
| 2025 | MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural NetworksabstractBrain-inspired spiking neural networks (SNNs) provide energy-efficient computation through event-driven processing. However, the shared weights across multiple timesteps lead to serious temporal feature redundancy, limiting both efficiency and performance. This issue is further aggravated when processing static images due to the duplicated input. To mitigate this problem, we propose a parameter-free and plug-and-play module named Mutual Information-based Temporal Redundancy Quantification and Reduction (MI-TRQR), constructing energy-efficient SNNs. Specifically, Mutual Information (MI) is properly introduced to quantify redundancy between discrete spike features at different timesteps on two spatial scales: pixel (local) and the entire spatial features (global). Based on the multi-scale redundancy quantification, we apply a probabilistic masking strategy to remove redundant spikes. The final representation is subsequently recalibrated to account for the spike removal. Extensive experimental results demonstrate that our MI-TRQR achieves sparser spiking firing, higher energy efficiency, and better performance concurrently with different SNN architectures in tasks of neuromorphic data classification, static data classification, and time-series forecasting. Notably, MI-TRQR increases accuracy by \textbf{1.7\%} on CIFAR10-DVS with 4 timesteps while reducing energy cost by \textbf{37.5\%}. Our codes are available at https://github.com/dfxue/MI-TRQR. Dengfeng Xue, Yifan Lu 0001, Chunfeng Yuan, Yufan Liu 0001, Wei Liu 0153, Man Yao, Li Yang 0014, Bing Li 0001, Stephen J. Maybank, Weiming Hu 0004, Zhetao Li |
NeurIPS | 5 |
| 2025 | FiGVCL: Fine-Grained Benchmark and Method for Video Copy LocalizationabstractContent-based video copy localization (VCL) aims to detect and locate copied segments in pairs of videos. VCL requires fine-grained video analysis to robustly identify copied segments that have been edited. Despite recent progress, the prohibitive cost of annotating copied segments and the lack of a fine-grained benchmark hinder the development of effective VCL systems. In this work, we annotate a new real-world dataset, FiGVCL, with challenging scenarios designed to evaluate VCL methods. FiGVCL is carefully annotated to preserve the temporal correspondences observed in copied segments. Moreover, we propose a novel fine-grained VCL benchmark metric based on temporal correspondences to improve discriminability. Finally, we design a simple but effective baseline model that uses fine-grained local embeddings for accurate copied segment localization. We also present an unsupervised training strategy that outperforms previous supervised VCL methods. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | iESTA: Instance-Enhanced Spatial-Temporal Alignment for Video Copy LocalizationabstractVideo copy Segment Localization (VSL) requires the identification of the temporal segments within a pair of videos that contain copied content. Current methods primarily focus on global temporal modeling, overlooking the complementarity of global semantic and local fine-grained features, which limits their effectiveness. Some related methods attempt to incorporate local spatial information but often disrupt spatial semantic structures, resulting in less accurate matching. To address these issues, we propose the Instance-Enhanced Spatial-Temporal Alignment Framework (iESTA), based on a proper representation granularity that integrates instance-level local features and semantic global features. Specifically, the Instance-relation Graph (IRG) is constructed to capture instance-level features and fine-grained interactions, preserving local information integrity and better representing the video feature space in a proper granularity. An instance-GNN structure is designed to refine these graph representations. For global features, we enhance the representation of semantic information, capturing temporal relationships within videos using a Transformer framework. Additionally, we design a Complementarity-perception Alignment Module (CAM) to effectively process and integrate complementary spatial-temporal information, producing accurate frame-to-frame alignment maps. Our approach also incorporates a differentiable Dynamic Time Warping (DTW) method to utilize latent temporal alignments as weak supervisory signals, improving the accuracy of the matching process. Experimental results indicate that our proposed iESTA outperforms state-of-the-art methods on both the small-scale dataset VCDB and the large-scale dataset VCSL. Xinmiao Ding, Jinming Lou, Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Task-Aware Attentional Dynamic Alignment for Few-Shot Compressed Video ClassificationabstractWe present a novel Task-aware Attentional Dynamic Alignment (TADA) framework for visual-based few-shot video classification (FSVC) that addresses two key challenges in this field: efficiency and nuanced spatio-temporal reasoning. Existing methods are often hindered by computationally expensive video decoding processes and neglect the temporal order of videos. In contrast, our method harnesses compressed domain data to extract rich spatio-temporal cues at a fraction of the cost of traditional video processing methods. Specifically, we propose an embedding module to extract informative features from compressed domain data while minimizing computational overheads. Furthermore, to exploit the temporal order of frames, we develop a prototypical ADA module to align and classify videos with an explicit temporal order constraint. Our framework also incorporates a contextual mixer to enrich video embeddings with task-specific context. Extensive experiments on multiple datasets demonstrate that TADA achieves state-of-the-art performance and outperforms existing methods in accuracy and efficiency. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Hierarchical Semantic Compression for Consistent Image Semantic RestorationabstractThe emerging semantic compression has been receiving increasing research efforts most recently, capable of achieving high fidelity restoration during compression, even at extremely low bitrates. However, existing semantic compression methods typically combine standard pipelines with either pre-defined or high-dimensional semantics, thus suffering from deficiency in compression. To address this issue, we propose a novel hierarchical semantic compression (HSC) framework that purely operates within intrinsic semantic spaces from generative models, which is able to achieve efficient compression for consistent semantic restoration. More specifically, we first analyse the entropy models for the semantic compression, which motivates us to employ a hierarchical architecture based on a newly developed general inversion encoder. Then, we propose the feature compression network (FCN) and semantic compression network (SCN), such that the middle-level semantic feature and core semantics are hierarchically compressed to restore both accuracy and consistency of image semantics, via an entropy model progressively shared by channel-wise context. Experimental results demonstrate that the proposed HSC framework achieves the state-of-the-art performance on subjective quality and consistency for human vision, together with superior performances on machine vision tasks given compressed bitstreams. This essentially coincides with human visual system in understanding images, thus providing a new framework for future image/video compression paradigms. The source code and trained models are available at https://github.com/bblgbr/HSC-TIP2025. Shengxi Li, Zifu Zhang, Mai Xu, Lai Jiang 0004, Yufan Liu 0001, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2024 | Learn from Noise: Detecting Deepfakes via Regional Noise ConsistencyabstractFace forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Various methods primarily concentrate on the features specific to certain generation techniques, potentially resulting in overfitting to the distinctive fingerprint characteristics of those manipulation techniques, thus undermining their generalizability. In contrast, our investigation reveals a prevalent phenomenon wherein regional noise consistency is disrupted during the integration of synthesized faces into source images, regardless of specific manipulation techniques. Motivated by this observation, we introduce the Regional Noise Consistency Learning Framework (RNCL), a novel approach designed to discern manipulated faces. Central to RNCL are two pivotal modules: Noise Consistency Enhancement (NCE) and Pyramidal Noise Consistency Learning (PNCL). The NCE module facilitates channel-wise and spatial-wise feature enhancement by exploiting noise inconsistencies between facial and non-facial regions. Complementarily, the PNCL module constructs a noise consistency pyramid to analyze enhanced features across multiple scales, enabling adaptive multi-scale feature integration. Leveraging the NCE and PNCL modules, our framework effectively transforms noise information into useful forgery cues, significantly enhancing forgery detection performance. Experimental results demonstrate that our method achieves state-of-the-art performance on standard benchmarks. The code will be publicly available. Weiming Bai, Yufan Liu 0001, Bo Wang 0147, Chengwei Peng, Weiming Hu 0004, Bing Li 0001 |
IJCNN | 2 |
| 2024 | Cross-Architecture Knowledge Distillation
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jingting Ding, Liang Li 0006, Stephen J. Maybank |
Int. J. Comput. Vis. | 1 |
| 2024 | Joint Learning of Audio-Visual Saliency Prediction and Sound Source Localization on Multi-face Videos
Minglang Qiao, Yufan Liu 0001, Mai Xu, Xin Deng 0002, Bing Li 0001, Weiming Hu 0004, Ali Borji |
Int. J. Comput. Vis. | 2 |
| 2023 | AUNet: Learning Relations Between Action Units for Face Forgery DetectionabstractFace forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Recent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same domain. However, the problem remains challenging when one tries to generalize the detector to forgeries created by unseen methods during training. Observing that face manipulation may alter the relation between different facial action units (AU), we propose the Action-Units Relation Learning framework to improve the generality of forgery detection. In specific, it consists of the Action Units Relation Transformer (ART) and the Tampered AU Prediction (TAP). The ART constructs the relation between different AUs with AU-agnostic Branch and AU-specific Branch, which complement each other and work together to exploit forgery clues. In the Tampered AU Prediction, we tamper AU-related regions at the image level and develop challenging pseudo samples at the feature level. The model is then trained to predict the tampered AU regions with the generated location-specific supervision. Experimental results demonstrate that our method can achieve state-of-the-art performance in both the in-dataset and cross-dataset evaluations. Weiming Bai, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
CVPR | 2 |
| 2023 | Nasty-SFDA: Source Free Domain Adaptation from a Nasty ModelabstractA challenging problem called Nasty Source Free Domain Adaptation (Nasty-SFDA) is proposed in this work, where only a nasty source model and unlabeled target samples are available for DA. Further, after DA, the target model is expected to be a nasty model. In order to deal with Nasty-SFDA, Nasty HypOthesis Transfer (NHOT) with an improved version of Information Maximization (IM) loss called Multi-Peak Constraints (MPC) and several Label Generation (LG) techniques is proposed. Experiments on four popular datasets show the superiority of NHOT for both Nasty-SFDA and SFDA. In addition, the target model obtained via NHOT is proven to be a nasty model. Jiajiong Cao, Yufan Liu 0001, Weiming Bai, Jingting Ding, Liang Li 0006 |
ICASSP | 2 |
| 2023 | Learning from the Raw Domain: Cross Modality Distillation for Compressed Video Action RecognitionabstractVideo action recognition is faced with the challenges of both huge computation burden and performance requirements. Using compressed domain data, which saves much decoding computation, is a possible solution. Unfortunately, existing compressed-domain-based (CD) methods fail to obtain high performance, compared with state-of-the-art (SOTA) raw-domain-based (RD) methods. In order to solve the problem, we propose a cross-modality knowledge distillation method to force the CD model to learn the knowledge from the RD model. In particular, spatial knowledge and temporal knowledge are first constructed to align feature space between the raw domain and the compressed domain. Then, an adaptively multi-path knowledge learning scheme is presented to help the CD model learn in a more efficient way. Experiments verify the effectiveness of the proposed method in large-scale and small-scale datasets. Yufan Liu 0001, Jiajiong Cao, Weiming Bai, Bing Li 0001, Weiming Hu 0004 |
ICASSP | 1 |
| 2023 | Learning to Explore Distillability and Sparsability: A Joint Framework for Model CompressionabstractDeep learning shows excellent performance usually at the expense of heavy computation. Recently, model compression has become a popular way of reducing the computation. Compression can be achieved using knowledge distillation or filter pruning. Knowledge distillation improves the accuracy of a lightweight network, while filter pruning removes redundant architecture in a cumbersome network. They are two different ways of achieving model compression, but few methods simultaneously consider both of them. In this paper, we revisit model compression and define two attributes of a model: distillability and sparsability, which reflect how much useful knowledge can be distilled and how many pruned ratios can be obtained, respectively. Guided by our observations and considering both accuracy and model size, a dynamically distillability-and-sparsability learning framework (DDSL) is introduced for model compression. DDSL consists of teacher, student and dean. Knowledge is distilled from the teacher to guide the student. The dean controls the training process by dynamically adjusting the distillation supervision and the sparsity supervision in a meta-learning framework. An alternating direction method of multiplier (ADMM)-based knowledge distillation-with-pruning (KDP) joint optimization algorithm is proposed to train the model. Extensive experimental results show that DDSL outperforms 24 state-of-the-art methods, including both knowledge distillation and filter pruning methods. Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Cross-Architecture Knowledge Distillation
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jingting Ding, Liang Li 0006 |
ACCV (5) | 1 |
| 2022 | Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video ClassificationabstractCompared with image few-shot learning, most of the existing few-shot video classification methods perform worse on feature matching, because they fail to sufficiently exploit the temporal information and relation. Specifically, frames are usually evenly sampled, which may miss important frames. On the other hand, the heuristic model simply encodes the equally treated frames in sequence, which results in the lack of both long-term and short-term temporal modeling and interaction. To alleviate these limitations, we take advantage of the compressed domain knowledge and propose a long-short term Cross-Transformer (LSTC) for few-shot video classification. For short terms, the motion vector (MV) contains temporal cues and reflects the importance of each frame. For long terms, a video can be natively divided into a sequence of GOPs (Group Of Picture). Using this compressed domain knowledge helps to obtain a more accurate spatial-temporal feature space. Consequently, we design the long-short term selection module, short-term module, and long-term module to comprise the LSTC. Long-short term selection is performed to select informative compressed domain data. Long/short-term modules are utilized to sufficiently exploit the temporal information so that the query and support can be well-matched by cross-attention. Experimental results show the superiority of our method on various datasets. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Yanan Miao, Yangxi Li |
IJCAI | 2 |
| 2022 | Self-supervised Face Anti-spoofing via Anti-contrastive Learning
Jiajiong Cao, Yufan Liu 0001, Jingting Ding, Liang Li 0006 |
PRCV (2) | 2 |
| 2022 | SDTP: Semantic-Aware Decoupled Transformer Pyramid for Dense Image PredictionabstractAlthough transformer has achieved great progress on computer vision tasks, the scale variation in dense image prediction is still the key challenge. Few effective multi-scale techniques are applied in transformer and there are two main limitations in the current methods. On the one hand, self-attention module in vanilla transformer fails to sufficiently exploit the diversity of semantic information because of its rigid mechanism. On the other hand, it is difficult to build attention and interaction among different levels due to the heavy computational burden. To alleviate this problem, we first revisit multi-scale problem in dense prediction, verifying the significance of diverse semantic representation and multi-scale interaction, and exploring the adaptation of transformer to pyramidal structure. Inspired by these findings, we propose a novel Semantic-aware Decoupled Transformer Pyramid (SDTP) for dense image prediction, consisting of Intra-level Semantic Promotion (ISP), Cross-level Decoupled Interaction (CDI) and Attention Refinement Function (ARF). ISP explores the semantic diversity in different receptive space through more flexible self-attention strategy. CDI builds the global attention and interaction among different levels in decoupled space which also solves the problem of heavy computation. Besides, ARF is further added to refine the attention in transformer. Experimental results demonstrate the validity and generality of the proposed method, which outperforms the state-of-the-art by a significant margin in dense image prediction tasks. Furthermore, the proposed components are all plug-and-play, which can be embedded in other methods. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Bailan Feng, Kebin Wu, Chengwei Peng, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | DPFPS: Dynamic and Progressive Filter Pruning for Compressing Convolutional Neural Networks from ScratchabstractFilter pruning is a commonly used method for compressing Convolutional Neural Networks (ConvNets), due to its friendly hardware supporting and flexibility. However, existing methods mostly need a cumbersome procedure, which brings many extra hyper-parameters and training epochs. This is because only using sparsity and pruning stages cannot obtain a satisfying performance. Besides, many works do not consider the difference of pruning ratio across different layers. To overcome these limitations, we propose a novel dynamic and progressive filter pruning (DPFPS) scheme that directly learns a structured sparsity network from Scratch. In particular, DPFPS imposes a new structured sparsity-inducing regularization specifically upon the expected pruning parameters in a dynamic sparsity manner. The dynamic sparsity scheme determines sparsity allocation ratios of different layers and a Taylor series based channel sensitivity criteria is presented to identify the expected pruning parameters. Moreover, we increase the structured sparsity-inducing penalty in a progressive manner. This helps the model to be sparse gradually instead of forcing the model to be sparse at the beginning. Our method solves the pruning ratio based optimization problem by an iterative soft-thresholding algorithm (ISTA) with dynamic sparsity. At the end of the training, we only need to remove the redundant parameters without other stages, such as fine-tuning. Extensive experimental results show that the proposed method is competitive with 11 state-of-the-art methods on both small-scale and large-scale datasets (i.e., CIFAR and ImageNet). Specifically, on ImageNet, we achieve a 44.97% pruning ratio of FLOPs by compressing ResNet-101, even with an increase of 0.12% Top-5 accuracy. Our pruned models and codes are released at https://github.com/taoxvzi/DPFPS. Xiaofeng Ruan, Yufan Liu 0001, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004 |
AAAI | 2 |
| 2021 | DSIC: Dynamic Sample-Individualized Connector for Multi-Scale Object DetectionabstractAlthough object detection has reached a milestone recently, the scale variation is still the key challenge. Integrating multilevel features is presented to alleviate the problems, like Feature Pyramid Network (FPN) and its improvements. However, the specifically designed architectures and fixed data flow paths of these methods are not flexible for feature fusion, especially when fed with various samples. To overcome the limitations, we propose a Dynamic Sample-Individualized Connector (DSIC) for multi-scale object detection, which dynamically adjusts network connections to fit different samples. In particular, DSIC consists of two components: Intra-scale Selection Gate (ISG) and Cross-scale Selection Gate (CSG). With the help of the presented gate operator, ISG adaptively extracts proper multi-level features from backbone as the inputs of feature integration. CSG automatically activates informative data flow paths based on the extracted multi-level features. These two components are both plug-and-play and can be embedded in any backbone. Experimental results demonstrate that the proposed method outperforms the state-of-the- arts. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Yanan Miao |
ICME | 2 |
| 2021 | Adaptive Coarse-to-Fine Interactor for Multi-Scale Object DetectionabstractScale variation is one of the key challenges of object detection. Multi-level feature fusion is presented to alleviate the problems, e.g., Feature Pyramid Network (FPN) and its extended methods. However, the input features fed into these methods and the interaction among features from different levels are insufficient and rigid. To fully exploit the features of multi-scale objects and enhance the feature interaction, we propose a novel and effective framework called Adaptive Coarse-to-Fine Interactor (ACFI). Specifically, ACFI consists of three cascaded components: Multi-Resolution Fusion (MRF), Fine-Grained Interaction (FGI), and Edge-aware Enhancement (EAE). MRF adaptively extracts multi-level features from multi-resolution images and multi-stage features, and then these features are fed into FGI to have a fine-grained interaction utilizing bottom-up guidance. After that, EAE further refines the features obtained by FGI, and enhances the detailed edge information and suppresses the redundant noise. After the coarse-to-fine process, we can obtain powerful multiscale representations of various objects. Each component can be embedded into any backbones, separately. Experimental results show the superiority of our method and verify the effectiveness of each proposed module. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
IJCNN | 2 |
| 2021 | Toward Accurate Pixelwise Object Tracking via Attention RetrievalabstractPixelwise single object tracking is challenging due to the competition of running speeds and segmentation accuracy. Current state-of-the-art real-time approaches seamlessly connect tracking and segmentation by sharing computation of the backbone network, e.g., SiamMask and D3S fork a light branch from the tracking model to predict segmentation mask. Although efficient, directly reusing features from tracking networks may harm the segmentation accuracy, since background clutter in the backbone feature tends to introduce false positives in segmentation. To mitigate this problem, we propose a unified tracking-retrieval-segmentation framework consisting of an attention retrieval network (ARN) and an iterative feedback network (IFN). Instead of segmenting the target inside the bounding box, the proposed framework performs soft spatial constraints on backbone features to obtain an accurate global segmentation map. Concretely, in ARN, a look-up-table (LUT) is first built by sufficiently using the information of the first frame. By retrieving it, a target-aware attention map is generated to suppress the negative influence of background clutter. To ulteriorly refine the contour of the segmentation, IFN iteratively enhances the features at different resolutions by taking the predicted mask as feedback guidance. Our framework sets a new state of the art on the recent pixelwise tracking benchmark VOT2020 and runs at 40 fps. Notably, the proposed model surpasses SiamMask by 11.7/4.2/5.5 points on VOT2020, DAVIS2016, and DAVIS2017, respectively. Code is available at https://github.com/JudasDie/SOTS. Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Houwen Peng |
IEEE Trans. Image Process. | 2 |
| 2021 | EDP: An Efficient Decomposition and Pruning Scheme for Convolutional Neural Network CompressionabstractModel compression methods have become popular in recent years, which aim to alleviate the heavy load of deep neural networks (DNNs) in real-world applications. However, most of the existing compression methods have two limitations: 1) they usually adopt a cumbersome process, including pretraining, training with a sparsity constraint, pruning/decomposition, and fine-tuning. Moreover, the last three stages are usually iterated multiple times. 2) The models are pretrained under explicit sparsity or low-rank assumptions, which are difficult to guarantee wide appropriateness. In this article, we propose an efficient decomposition and pruning (EDP) scheme via constructing a compressed-aware block that can automatically minimize the rank of the weight matrix and identify the redundant channels. Specifically, we embed the compressed-aware block by decomposing one network layer into two layers: a new weight matrix layer and a coefficient matrix layer. By imposing regularizers on the coefficient matrix, the new weight matrix learns to become a low-rank basis weight, and its corresponding channels become sparse. In this way, the proposed compressed-aware block simultaneously achieves low-rank decomposition and channel pruning by only one single data-driven training stage. Moreover, the network of architecture is further compressed and optimized by a novel Pruning & Merging (PM) module which prunes redundant channels and merges redundant decomposed layers. Experimental results (17 competitors) on different data sets and networks demonstrate that the proposed EDP achieves a high compression ratio with acceptable accuracy degradation and outperforms state-of-the-arts on compression rate, accuracy, inference time, and run-time memory. Xiaofeng Ruan, Yufan Liu 0001, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Yangxi Li, Stephen J. Maybank |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model
Yufan Liu 0001, Minglang Qiao, Mai Xu, Bing Li 0001, Weiming Hu 0004, Ali Borji |
ECCV (20) | 1 |
| 2019 | Knowledge Distillation via Instance Relationship GraphabstractThe key challenge of knowledge distillation is to extract general, moderate and sufficient knowledge from a teacher network to guide a student network. In this paper, a novel Instance Relationship Graph (IRG) is proposed for knowledge distillation. It models three kinds of knowledge, including instance features, instance relationships and feature space transformation, while the latter two kinds of knowledge are neglected by previous methods. Firstly, the IRG is constructed to model the distilled knowledge of one network layer, by considering instance features and instance relationships as vertexes and edges respectively. Secondly, an IRG transformation is proposed to models the feature space transformation across layers. It is more moderate than directly mimicking the features at intermediate layers. Finally, hint loss functions are designed to force a student's IRGs to mimic the structures of a teacher's IRGs. The proposed method effectively captures the knowledge along the whole network via IRGs, and thus shows stable convergence and strong robustness to different network architectures. In addition, the proposed method shows superior performance over existing methods on datasets of various scales. Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004, Yangxi Li, Yunqiang Duan |
CVPR | 1 |
| 2018 | Rate control schemes for panoramic video coding
Yufan Liu 0001, Li Yang 0014, Mai Xu, Zulin Wang |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Find Who to Look at: Turning From Action to SaliencyabstractThe past decade has witnessed the use of highlevel features in saliency prediction for both videos and images. Unfortunately, the existing saliency prediction methods only handle high-level static features, such as face. In fact, high-level dynamic features (also called actions), such as speaking or head turning, are also extremely attractive to visual attention in videos. Thus, in this paper, we propose a data-driven method for learning to predict the saliency of multiple-face videos, by leveraging both static and dynamic features at high-level. Specifically, we introduce an eye-tracking database, collecting the fixations of 39 subjects viewing 65 multiple-face videos. Through analysis on our database, we find a set of high-level features that cause a face to receive extensive visual attention. These high-level features include the static features of face size, center-bias and head pose, as well as the dynamic features of speaking and head turning. Then, we present the techniques for extracting these high-level features. Afterwards, a novel model, namely multiple hidden Markov model (M-HMM), is developed in our method to enable the transition of saliency among faces. In our MHMM, the saliency transition takes into account both the state of saliency at previous frames and the observed high-level features at the current frame. The experimental results show that the proposed method is superior to other state-of-the-art methods in predicting visual attention on multiple-face videos. Finally, we shed light on a promising implementation of our saliency prediction method in locating the region-of-interest (ROI), for video conference compression with high efficiency video coding (HEVC). Mai Xu, Yufan Liu 0001, Haoji Hu, Feng He 0007 |
IEEE Trans. Image Process. | 2 |
| 2017 | Predicting Salient Face in Multiple-Face VideosabstractAlthough the recent success of convolutional neural network (CNN) advances state-of-the-art saliency prediction in static images, few work has addressed the problem of predicting attention in videos. On the other hand, we find that the attention of different subjects consistently focuses on a single face in each frame of videos involving multiple faces. Therefore, we propose in this paper a novel deep learning (DL) based method to predict salient face in multiple-face videos, which is capable of learning features and transition of salient faces across video frames. In particular, we first learn a CNN for each frame to locate salient face. Taking CNN features as input, we develop a multiple-stream long short-term memory (M-LSTM) network to predict the temporal transition of salient faces in video sequences. To evaluate our DL-based method, we build a new eye-tracking database of multiple-face videos. The experimental results show that our method outperforms the prior state-of-the-art methods in predicting visual attention on faces in multiple-face videos. Yufan Liu 0001, Songyang Zhang 0001, Mai Xu, Xuming He 0001 |
CVPR | 1 |
| 2017 | A novel rate control scheme for panoramic video codingabstractThe popularity of multi-view panoramic videos has been considerably increased for producing Virtual Reality (VR) content, due to its immersive visual experience. We argue in this paper that PSNR is less effective in assessing visual quality of compressed panoramic videos than Sphere-based PSNR (S-PNSR), in which sphere-to-plain mapping of panoramic videos is considered. Thus, the conventional rate control (R-C) schemes of 2-Dimensional (2D) video coding, which optimize on PSNR, are not suitable for panoramic video coding. To optimize S-PSNR, we propose in this paper a novel RC scheme for panoramic video coding. Specifically, we develop an S-PSNR optimization formulation with constraint on bit-rate. Then, a solution is provided to the developed formulation, such that bits can be allocated to each coding block for achieving optimal S-PSNR in panoramic video coding. Finally, the experiment results validate the effectiveness of the proposed RC scheme in improving S-PSNR of panoramic video coding. Yufan Liu 0001, Mai Xu, Chen Li 0049, Shengxi Li, Zulin Wang |
ICME | 1 |
| 2017 | A subjective visual quality assessment method of panoramic videosabstractDifferent from 2-dimensional (2D) videos, panoramic videos contain spherical viewing direction with the support of head-mounted displays, thus improving immersive and interactive visual experience. Unfortunately, to our best knowledge, there are few subjective visual quality assessment (VQA) methods for panoramic videos. In this paper, we therefore propose a subjective VQA method for assessing quality loss of impaired panoramic videos. Specifically, we first establish a database containing viewing direction data of several subjects on watching panoramic videos. Then, we find out that there exists high consistency of viewing direction on panoramic videos across different subjects. Upon this finding, we present a procedure of subjective test in measuring quality of panoramic videos by different subjects, yielding different mean opinion score (DMOS). To couple with inconsistency of viewing directions on panoramic videos, we further propose a vectorized DMOS metric. Finally, experimental results verify that our subjective VQA method, in the forms of both overall and vectorized DMOS metrics, is effective in measuring subjective quality of panoramic videos. Mai Xu, Chen Li 0049, Yufan Liu 0001, Xin Deng 0002, Jiaxin Lu 0003 |
ICME | 3 |
| 2015 | Subjective rate-distortion optimization in HEVC with perceptual model of multiple facesabstractThis paper proposes a novel perceptual video coding approach with a perceptual model of multiple faces, to improve the coding efficiency of HEVC in video conferencing scenarios. For the perceptual model, a latest active appearance model (AAM) is used to detect multiple faces in a video frame. Then, the perceptual model of multiple faces can be established on the basis of the detected multiple faces. With the established perceptual model of multiple faces, all faces in a video frame can be taken into account for subjective rate-distortion optimization, which is based on the state-of-the-art r-λ rate control scheme of HEVC. As such, the perceptual video coding can be achieved for HEVC of video conferencing scenarios. Finally, the experimental results validate the effectiveness of the proposed perceptual video coding approach, in terms of subjective quality. Yufan Liu 0001, Haoji Hu, Mai Xu |
VCIP | 1 |