VLDB 2026 Research / reviewers in the wild / expert
Hanbin Zhao
dblp:222/7871
· DBLP profile ↗
37ranked-venue papers
5as first author
36since 2021 · last 2026
0000-0001-8906-4534ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 4 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 23 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bring Your Dreams to Life: Continual Text-to-Video CustomizationabstractCustomized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that personalized concepts remain static and do not expand incrementally over time. Additionally, they struggle with forgetting and concept neglect when continuously learning new concepts, including subjects and motions. To resolve the above challenges, we develop a novel Continual Customized Video Diffusion (CCVD) model, which can continuously learn new concepts to generate videos across various text-to-video generation tasks by tackling forgetting and concept neglect. To address catastrophic forgetting, we introduce a concept-specific attribute retention module and a task-aware concept aggregation strategy. They can capture the unique characteristics and identities of old concepts during training, while combining all subject and motion adapters of old concepts based on their relevance during testing. Besides, to tackle concept neglect, we develop a controllable conditional synthesis to enhance regional features and align video contexts with user conditions, by incorporating layer-specific region attention-guided noise estimation. Extensive experimental comparisons demonstrate that our CCVD outperforms existing CTVG models. Jiahua Dong 0001, Wenqi Liang, Zongyan Han, Meng Cao 0002, Duzhen Zhang, Hanbin Zhao, Zhi Han, Salman Khan 0001, Fahad Shahbaz Khan |
AAAI | 7 |
| 2026 | Mass Concept Erasure in Diffusion Models with Concept HierarchyabstractThe success of diffusion models has raised concerns about the generation of unsafe or harmful content, prompting concept erasure approaches that fine-tune modules to suppress specific concepts while preserving general generative capabilities. However, as the number of erased concepts grows, these methods often become inefficient and ineffective, since each concept requires a separate set of fine-tuned parameters and may degrade the overall generation quality. In this work, we propose a supertype-subtype concept hierarchy that organizes erased concepts into a parent–child structure. Each erased concept is treated as a child node, and semantically related concepts (e.g., macaw, and bald eagle) are grouped under a shared parent node, referred to as a supertype concept (e.g., bird). Rather than erasing concepts individually, we introduce an effective and efficient group-wise suppression method, where semantically similar concepts are grouped and erased jointly by sharing a single set of learnable parameters. During the erasure phase, standard diffusion regularization is applied to preserve denoising process in unmasked regions. To mitigate the degradation of supertype generation caused by excessive erasure of semantically related subtypes, we propose a novel method called Supertype-Preserving Low-Rank Adaptation (SuPLoRA), which encodes the supertype concept information in the frozen down-projection matrix and updates only the up-projection matrix during erasure. Theoretical analysis demonstrates the effectiveness of SuPLoRA in mitigating generation performance degradation. We construct a more challenging benchmark that requires simultaneous erasure of concepts across diverse domains, including celebrities, objects, and pornographic content. Comprehensive experiments demonstrate that our method achieves a superior balance between effective multi-concept erasure and the preservation of desirable generative performance. Jiahang Tu, Ye Li 0043, Hanbin Zhao, Chao Zhang 0029, Hui Qian 0001 |
AAAI | 4 |
| 2026 | CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word Vocabulary
Jiahang Tu, Jiahua Dong 0001, Hanbin Zhao, Chao Zhang 0001, Nicu Sebe, Hui Qian 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware PromptingabstractRecent pre-trained vision-language models (PT-VLMs) often face a Multi-Domain Task Incremental Learning (MTIL) scenario in practice, where several classes and domains of multi-modal tasks are arrive incrementally. Without access to previously seen tasks and unseen tasks, memory-constrained MTIL suffers from forward and backward forgetting. To alleviate the above challenges, parameter-efficient fine-tuning techniques (PEFT), such as prompt tuning, are employed to adapt the PT-VLM to the diverse incrementally learned tasks. To achieve effective new task adaptation, existing methods only consider the effect of PEFT strategy selection, but neglect the influence of PEFT parameter setting (e.g., prompting). In this paper, we tackle the challenge of optimizing prompt designs for diverse tasks in MTIL and propose an Instance-Aware Prompting (IAP) framework. Specifically, our Instance-Aware Gated Prompting (IA-GP) strategy enhances adaptation to new tasks while mitigating forgetting by adaptively assigning prompts across transformer layers at the instance level. Our Instance-Aware Class-Distribution-Driven Prompting (IA-CDDP) improves the task adaptation process by determining an accurate task-label-related confidence score for each instance. Experimental evaluations across 11 datasets, using three performance metrics, demonstrate the effectiveness of our proposed method. The source codes are available at https://github.com/FerdinandZJU/IAP. Hao Fu 0023, Hanbin Zhao, Jiahua Dong 0001, Henghui Ding, Chao Zhang 0001, Hui Qian 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Hyperbolic Active Learning for Label-Efficient Action SegmentationabstractRecent advances in action segmentation have greatly enhanced our understanding of complex and dynamic scenes in video content. Despite these improvements, the field continues to face persistent challenges, particularly in terms of model efficiency and the substantial cost associated with manual annotation. In this work, we introduce a novel framework that integrates active learning within hyperbolic space to effectively address these issues. By leveraging the hierarchical representational capacity of hyperbolic space, which is naturally suited for modeling structured data, and combining it with the selective efficiency of active learning, our method introduces hyperbolic uncertainty metrics to guide the targeted selection of the most informative video frames and sequences for annotation. This enables the model to prioritize annotation efforts where they are most impactful. Furthermore, the model iteratively refines pseudo labels using all available annotations, significantly reducing the need for exhaustive labeling while preserving high segmentation accuracy. To further mitigate reliance on precise annotations, we enhance the MS-TCN model by incorporating soft pseudo labels and a weighting mechanism that dynamically adjusts learning based on label confidence, allowing for more robust training in the presence of noisy or weakly labeled data. Extensive experiments conducted on two widely used action segmentation benchmark datasets validate the effectiveness of our approach, demonstrating that it can substantially reduce annotation effort while maintaining overall segmentation performance. Jingqiao Xiu, Wei Ji 0008, Menglin Yang 0001, Hanbin Zhao, Roger Zimmermann |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental LearningabstractClass-Incremental Learning (CIL) requires models to continually acquire knowledge of new classes without forgetting old ones. Despite Pre-trained Models (PTMs) have shown excellent performance in CIL, catastrophic forgetting still occurs as the model learns new concepts. Existing work seeks to utilize lightweight components to adjust the PTM, while the forgetting phenomenon still comes from parameter and retrieval levels. Specifically, iterative updates of the model result in parameter drift, while mistakenly retrieving irrelevant modules leads to the mismatch during inference. To this end, we propose MOdel Surgery (MOS) to rescue the model from forgetting previous knowledge. By training task-specific adapters, we continually adjust the PTM to downstream tasks. To mitigate parameter-level forgetting, we present an adapter merging approach to learn task-specific adapters, which aims to bridge the gap between different components while reserve task-specific information. Besides, to address retrieval-level forgetting, we introduce a training-free self-refined adapter retrieval mechanism during inference, which leverages the model's inherent ability for better adapter retrieval. By jointly rectifying the model with those steps, MOS can robustly resist catastrophic forgetting in the learning process. Extensive experiments on seven benchmark datasets validate MOS's state-of-the-art performance. Hai-Long Sun, Da-Wei Zhou 0001, Hanbin Zhao, Le Gan, De-Chuan Zhan, Han-Jia Ye |
AAAI | 3 |
| 2025 | TextToucher: Fine-Grained Text-to-Touch GenerationabstractTactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tactile images by vision-to-touch image translation. However, compared to text modality, visual modality-driven tactile generation cannot accurately depict human tactile sensation. In this work, we analyze the characteristics of tactile images in detail from two granularities: object-level (tactile texture, tactile shape), and sensor-level (gel status). We model these granularities of information through text descriptions and propose a fine-grained Text-to-Touch generation method (TextToucher) to generate high-quality tactile samples. Specifically, we introduce a multimodal large language model to build the text sentences about object-level tactile information and employ a set of learnable text prompts to represent the sensor-level tactile information. To better guide the tactile generation process with the built text information, we fuse the dual grains of text information and explore various dual-grain text conditioning methods within the diffusion transformer architecture. Furthermore, we propose a Contrastive Text-Touch Pre-training (CTTP) metric to precisely evaluate the quality of text-driven generated tactile data. Extensive experiments demonstrate the superiority of our TextToucher method. Jiahang Tu, Hao Fu 0023, Fengyu Yang 0003, Hanbin Zhao, Chao Zhang 0029, Hui Qian 0001 |
AAAI | 4 |
| 2025 | Hierarchical Visual Prompt Learning for Continual Video Instance SegmentationabstractVideo instance segmentation (VIS) has gained significant attention for its capability in tracking and segmenting object instances across video frames. However, most of the existing VIS approaches unrealistically assume that the categories of object instances remain fixed over time. Moreover, they experience catastrophic forgetting of old classes when required to continuously learn object instances belonging to new categories. To resolve these challenges, we develop a novel Hierarchical Visual Prompt Learning (HVPL) model that overcomes catastrophic forgetting of previous categories from both frame-level and video-level perspectives. Specifically, to mitigate forgetting at the frame level, we devise a task-specific frame prompt and an orthogonal gradient correction (OGC) module. The OGC module helps the frame prompt encode task-specific global instance information for new classes in each individual frame by projecting its gradients onto the orthogonal feature space of old classes. Furthermore, to address forgetting at the video level, we design a task-specific video prompt and a video context decoder. This decoder first embeds structural inter-class relationships across frames into the frame prompt features, and then propagates task-specific global video contexts from the frame prompt features to the video prompt. Through rigorous comparisons, our HVPL model proves to be more effective than baseline approaches. The code is available at https://github.com/JiahuaDong/HVPL. Jiahua Dong 0001, Wenqi Liang, Hanbin Zhao, Henghui Ding, Nicu Sebe, Salman Khan 0001, Fahad Shahbaz Khan |
ICCV | 4 |
| 2025 | FG-OrIU: Towards Better Forgetting via Feature-Gradient Orthogonality for Incremental Unlearning
Jiahang Tu, Mintong Kang, Hanbin Zhao, Chao Zhang 0001, Hui Qian 0001 |
ICCV | 4 |
| 2025 | Unleashing High-Quality Image Generation in Diffusion Sampling Using Second-Order Levenberg-Marquardt-LangevinabstractThe diffusion models (DMs) have demonstrated the remarkable capability of generating images via learning the noised score function of data distribution. Current DM sampling techniques typically rely on first-order Langevin dynamics at each noise level, with efforts concentrated on refining inter-level denoising strategies. While leveraging additional second-order Hessian geometry to enhance the sampling quality of Langevin is a common practice in Markov chain Monte Carlo (MCMC), the naive attempts to utilize Hessian geometry in high-dimensional DMs lead to quadratic-complexity computational costs, rendering them non-scalable. In this work, we introduce a novel Levenberg-Marquardt-Langevin (LML) method that approximates the diffusion Hessian geometry in a training-free manner, drawing inspiration from the celebrated Levenberg-Marquardt optimization algorithm. Our approach introduces two key innovations: (1) A low-rank approximation of the diffusion Hessian, leveraging the DMs' inherent structure and circumventing explicit quadratic-complexity computations; (2) A damping mechanism to stabilize the approximated Hessian. This LML approximated Hessian geometry enables the diffusion sampling to execute more accurate steps and improve the image generation quality. We further conduct a theoretical analysis to substantiate the approximation error bound of low-rank approximation and the convergence property of the damping mechanism. Extensive experiments across multiple pretrained DMs validate that the LML method significantly improves image generation quality, with negligible computational overhead. Fangyikang Wang, Hubery Yin, Shaobin Zhuang, Huminhao Zhu, Yanlong Tang, Chao Zhang 0029, Hanbin Zhao, Hui Qian 0001, Chen Li 0031 |
ICCV | 10 |
| 2025 | Efficiently Access Diffusion Fisher: Within the Outer Product Span SpaceabstractRecent Diffusion models (DMs) advancements have explored incorporating the second-order diffusion Fisher information (DF), defined as the negative Hessian of log density, into various downstream tasks and theoretical analysis.
However, current practices typically approximate the diffusion Fisher by applying auto-differentiation to the learned score network. This black-box method, though straightforward, lacks any accuracy guarantee and is time-consuming.
In this paper, we show that the diffusion Fisher actually resides within a space spanned by the outer products of score and initial data.
Based on the outer-product structure, we develop two efficient approximation algorithms to access the trace and matrix-vector multiplication of DF, respectively.
These algorithms bypass the auto-differentiation operations with time-efficient vector-product calculations.
Furthermore, we establish the approximation error bounds for the proposed algorithms.
Experiments in likelihood evaluation and adjoint optimization demonstrate the superior accuracy and reduced computational cost of our proposed algorithms.
Additionally, based on the novel outer-product formulation of DF, we design the first numerical verification experiment for the optimal transport property of the general PF-ODE deduced map. Fangyikang Wang, Hubery Yin, Shaobin Zhuang, Huminhao Zhu, Chao Zhang 0029, Hanbin Zhao, Hui Qian 0001, Chen Li 0031 |
ICML | 8 |
| 2025 | A Timestep-Adaptive Frequency-Enhancement Framework for Diffusion-based Image Super-ResolutionabstractImage super-resolution (ISR) is a classic and challenging problem in computer vision because of complex and unknown degradation patterns in the data collection process. Leveraging powerful generative priors, diffusion-based methods have recently established new state-of-the-art ISR performance, but their characteristics in the frequency domain are still underexplored. In this paper, we innovatively investigate their frequency-domain behaviors from a sampling timestep perspective. Experimentally, we find that current diffusion-based ISR algorithms exhibit insufficiency in different frequency components in distinct groups of timesteps during the sampling. To address this, we first propose a Timestep Division Controller that is able to adaptively divide the timesteps into groups based on the performance gradient across different components. Next, we design two dedicated modules --- the Amplitude and Phase Enhancement Module (APEM) and the High- and Low-Frequency Enhancement Module (HLEM), to regulate the information flow of distinct frequency-domain features. By adaptively enhancing specific frequency components at different stages of the sampling process, the two modules effectively compensate for the insufficient frequency-domain perception of diffusion-based ISR models. Extensive experiments on three benchmark datasets verify the superior ISR performance of our method, e.g., achieving an average 5.40% improvement on CLIP-IQA compared to the best diffusion-based ISR baseline. Hanbin Zhao, Jiaqing Zhou, Guozhi Xu, Tianlei Hu, Gang Chen 0001, Haobo Wang 0001 |
IJCAI | 2 |
| 2025 | Few-Shot Incremental Multi-modal Learning via Touch Guidance and Imaginary Vision SynthesisabstractMultimodal perception, which integrates vision and touch, is increasingly demonstrating its significance in domains such as embodied intelligence and human-computer interaction. However, in open-world scenarios, multimodal data streams face significant challenges, including catastrophic forgetting and overfitting, during few-shot class incremental learning (FSCIL), leading to a severe degradation in model performance. In this work, we propose a novel approach named Few-Shot Incremental Multi-modal Learning via Touch Guidance and Imaginary Vision Synthesis (TIFS). Our method leverages vision imagination synthesis to enhance the semantic understanding and integrates touch and vision fusion to improve the problem of modal imbalance. Specifically, we introduce a framework that employs touch-guided vision information for cross-modal contrastive learning to address the challenges of few-shot learning. Additionally, we incorporate multiple learning mechanisms, including regularization, memory mechanisms, and attention mechanisms, to mitigate catastrophic forgetting during multi-incremental step learning. Experimental results on the Touch and Go and VisGel datasets demonstrate that the TIFS framework exhibits robust continuous learning capabilities and strong generalization performance in touch-vision few-shot incremental learning tasks. Our code is available at https://github.com/Vision-Multimodal-Lab-HZCU/TIFS. Lina Wei, Zhongsheng Lin, Canghong Jin, Hanbin Zhao, Dapeng Chen |
IJCAI | 6 |
| 2025 | PECTP: Parameter-Efficient Cross-Task Prompts for Incremental Vision TransformerabstractIncremental Learning (IL) aims to learn deep models on sequential tasks continually, where each new task includes a batch of new classes and deep models have no access to task ID information at the inference time. Recent vast pre-trained models (PTMs) have achieved outstanding performance by prompt technique in practical IL without the old samples (rehearsal-free) and with a memory constraint (memory-constrained): Prompt-extending and Prompt-fixed methods. However, prompt-extending methods need a large memory buffer to maintain an ever-expanding prompt pool and meet an extra challenging prompt selection problem. Prompt-fixed methods only learn a single set of prompts on one of the incremental tasks and can not handle all the incremental tasks effectively. To achieve a good balance between the memory cost and the performance on all the tasks, we propose a Parameter-Efficient Cross-Task Prompt (PECTP) framework with Prompt Retention Module (PRM) and classifier Head Retention Module (HRM). To make the final learned prompts effective on all incremental tasks, PRM constrains the evolution of cross-task prompts’ parameters from Outer Prompt Granularity and Inner Prompt Granularity. Besides, we employ HRM to inherit old knowledge in the previously learned classifier heads to facilitate the cross-task prompts’ generalization ability. Extensive experiments show the effectiveness of our method. The source codes are available at https://github.com/RAIAN08/PECTP. Hanbin Zhao, Chao Zhang 0001, Jiahua Dong 0001, Henghui Ding, Yu-Gang Jiang 0001, Hui Qian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | GCSTG: Generating Class-Confusion-Aware Samples With a Tree-Structure Graph for Few-Shot Object DetectionabstractFew-Shot Object Detection (FSOD) aims to detect the objects of novel classes using only a few manually annotated samples. With the few novel class samples, learning the inter-class relationships among foreground and constructing the corresponding class hierarchy in FSOD is a challenging task. The poor construction of the class hierarchy will result in the inter-class confusion problem, which has been identified as a primary cause of inferior performance in novel classes by recent FSOD methods. In this work, we further find that the intra-super-class confusion, where samples are misclassified as classes within their associated super-classes, is the main challenge in solving the confusion problem. To solve this issue, this work generates class-confusion-aware samples with a pre-defined tree-structure graph, for helping models to construct a precise class hierarchy. In precise, for generating class-confusion-aware samples, we add the noise into available samples and update the noise to maximize confidence scores on associated confusion categories of samples. Then, a confusion-aware curriculum learning strategy is proposed to make generated samples gradually participate in the training, which benefits the model convergence while learning the generated samples. Experimental results show that our method can be used as a plug-in in recent FSOD methods and consistently improve the model performance. Longrong Yang, Hanbin Zhao, Hongliang Li 0001, Liang Qiao 0001, Xi Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Visuo-Tactile Class-Incremental LearningabstractThe ability to associate sight with touch is essential for human and robot agents to understand material properties and to interact with the physical world. In the real-world scenarios, the robot agents often operate in a dynamically changing environment where new classes of objects are continually collected by visual and tactile sensors. In this article, we define this scenario as V isuo- T actile C lass- I ncremental L earning (VT-CIL). In practical VT-CIL, the robot needs to adapt to a new environment with constrained storage and computing resources, and suffers from the severe forgetting of vision and touch knowledge about old environments. To alleviate this problem, we consider visuo-tactile correlations in VT-CIL and propose a novel framework. It efficiently incorporates the Visuo-Tactile Cross-Modal Pseudo-Label-Consistent (VT-CMPLC) constraint, Dual-Visuo-Tactile Exemplars (DVT-E), and the Dual-Visuo-Tactile-Compatible (DVT-C) constraint. The old visual–tactile classes are preserved by the VT-CMPLC constraint and DVT-E, while the visuo-tactile correlations and the VT-CMPLC and DVT-E capabilities are enhanced by the DVT-C constraint. We built two benchmarks, the Touch-and-Go Class-Incremental (TaG-CI) benchmark and the ObjectFolder-Real Class-Incremental (OFR-CI) benchmark. Experimental results on TaG-CI and OFR-CI benchmarks demonstrate the effectiveness of our method against previous state-of-the-art class-incremental learning methods in VT-CIL. Hao Fu 0023, Fengyu Yang 0003, Boyang Wang 0008, Wei Ji 0008, Hanbin Zhao, Chao Zhang 0029, Roger Zimmermann, Hui Qian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | DriveDiTFit: Fine-tuning Diffusion Transformers for Autonomous Driving Data GenerationabstractIn autonomous driving, deep models have shown remarkable performance across various visual perception tasks with the demand of high-quality and huge-diversity training datasets. Such datasets are expected to cover various driving scenarios with adverse weather, lighting conditions, and diverse moving objects. However, manually collecting these data presents huge challenges and is expensive. With the rapid development of large generative models, we propose DriveDiTFit, a novel method for efficiently generating autonomous Driv ing data by Fi ne- t uning pre-trained Di ffusion T ransformers (DiTs). Specifically, DriveDiTFit utilizes a gap-driven modulation technique to carefully select and efficiently fine-tune a few parameters in DiTs according to the discrepancy between the pre-trained source data and the target driving data. Additionally, DriveDiTFit develops an effective weather and lighting condition embedding module to ensure diversity in the generated data, which is initialized by a nearest-semantic-similarity initialization approach. Through progressive tuning scheme to refine the process of detail generation in early diffusion process and enlarging the weights corresponding to small objects in training loss, DriveDiTFit ensures high-quality generation of small moving objects in the generated data. Extensive experiments conducted on driving datasets confirm that our method could efficiently produce diverse real driving data. Jiahang Tu, Wei Ji 0008, Hanbin Zhao, Chao Zhang 0029, Roger Zimmermann, Hui Qian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | GAD-PVI: A General Accelerated Dynamic-Weight Particle-Based Variational Inference FrameworkabstractParticle-based Variational Inference (ParVI) methods approximate the target distribution by iteratively evolving finite weighted particle systems. Recent advances of ParVI methods reveal the benefits of accelerated position update strategies and dynamic weight adjustment approaches. In this paper, we propose the first ParVI framework that possesses both accelerated position update and dynamical weight adjustment simultaneously, named the General Accelerated Dynamic-Weight Particle-based Variational Inference (GAD-PVI) framework. Generally, GAD-PVI simulates the semi-Hamiltonian gradient flow on a novel Information-Fisher-Rao space, which yields an additional decrease on the local functional dissipation. GAD-PVI is compatible with different dissimilarity functionals and associated smoothing approaches under three information metrics. Experiments on both synthetic and real-world data demonstrate the faster convergence and reduced approximation error of GAD-PVI methods over the state-of-the-art. Fangyikang Wang, Huminhao Zhu, Chao Zhang 0029, Hanbin Zhao, Hui Qian 0001 |
AAAI | 4 |
| 2024 | APISR: Anime Production Inspired Real-World Anime Super-ResolutionabstractWhile real-world anime super-resolution (SR) has gained increasing attention in the SR community, existing methods still adopt techniques from the photorealistic domain. In this paper, we analyze the anime production workflow and rethink how to use characteristics of it for the sake of the real-world anime SR. First, we argue that video networks and datasets are not necessary for anime SR due to the repetition use of hand-drawing frames. Instead, we propose an anime image collection pipeline by choosing the least compressed and the most informative frames from the video sources. Based on this pipeline, we introduce the Anime Production-oriented Image (API) dataset. In addition, we identify two anime-specific challenges of distorted and faint hand-drawn lines and unwanted color artifacts. We address the first issue by introducing a prediction-oriented compression module in the image degradation model and a pseudo-ground truth preparation with enhanced hand-drawn lines. In addition, we introduce the balanced twin perceptual loss combining both anime and photorealistic high-level features to mitigate unwanted color artifacts and increase visual clarity. We evaluate our method through extensive experiments on the public benchmark, showing our method outperforms state-of-the-art anime dataset-trained approaches. The code is available at https://github.com/Kiteretsu77/APISR. Boyang Wang 0008, Fengyu Yang 0003, Xihang Yu, Chao Zhang 0029, Hanbin Zhao |
CVPR | 5 |
| 2024 | RCS-Prompt: Learning Prompt to Rearrange Class Space for Prompt-Based Continual Learning
Longrong Yang, Hanbin Zhao, Yunlong Yu 0001, Xiaodong Zeng, Xi Li 0001 |
ECCV (47) | 2 |
| 2024 | Hierarchical Debiasing and Noisy Correction for Cross-domain Video Tube RetrievalabstractVideo Tube Retrieval (VTR) has attracted wide attention in the multi-modal domain, aiming to accurately localize the spatial-temporal tube in videos based on the natural language description. Despite the remarkable progress, existing VTR models trained on a specific domain (source domain) often perform unsatisfactory in another domain (target domain), due to the domain gap. Toward this issue, we introduce the learning strategy, Unsupervised Domain Adaptation, into the VTR task (UDA-VTR), which enables the knowledge transfer from the labeled source domain to the unlabeled target domain without additional manual annotations. An intuitive solution is generating the pseudo labels for the target domain samples with the fully trained source model and fine-tuning the source model on the target domain with pseudo labels. However, the existing domain gap gives rise to two problems for this process: (1) The transfer of model parameters across domains may introduce source domain bias into target domain features, significantly impacting the feature-based prediction for target domain samples. (2) The pseudo labels tend to identify video tubes that are widely present in the source domain, rather than accurately localizing the correct video tubes specific to the target domain samples. To address the above issues, we propose the unsupervised domain adaptation model via Hierarchical dEbiAsing and noisy corRecTion (HEART) for cross-domain video tube retrieval, which contains two characteristic modules: Layered Feature Debiasing (including the adversarial feature alignment and the graph based alignment) and Pseudo Label Refinement. Extensive experiments prove the effectiveness of our HEART model by significantly surpassing the state-of-the-arts. Jingqiao Xiu, Mengze Li 0001, Wei Ji 0008, Jingyuan Chen 0003, Hanbin Zhao, Shin'ichi Satoh 0001, Roger Zimmermann |
ACM Multimedia | 5 |
| 2024 | D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language ModelsabstractLarge language models have shown an impressive societal impact owing to their excellent understanding and logical reasoning skills. However, such strong ability relies on a huge amount of computing resources, which makes it difficult to deploy LLMs on computing resource-constrained platforms. Currently, LLMs process each token equivalently, but we argue that not every word is equally important. Some words should not be allocated excessive computing resources, particularly for dispensable terms in simple questions. In this paper, we propose a novel dynamic inference paradigm for LLMs, namely D-LLMs, which adaptively allocate computing resources in token processing. We design a dynamic decision module for each transformer layer that decides whether a network unit should be executed or skipped. Moreover, we tackle the issue of adapting D-LLMs to real-world applications, specifically concerning the missing KV-cache when layers are skipped. To overcome this, we propose a simple yet effective eviction policy to exclude the skipped layers from subsequent attention calculations. The eviction policy not only enables D-LLMs to be compatible with prevalent applications but also reduces considerable storage resources. Experimentally, D-LLMs show superior performance, in terms of computational cost and KV storage utilization. It can reduce up to 45\% computational cost and KV storage on Q\&A, summarization, and math solving tasks, 50\% on commonsense reasoning tasks. Yikun Jiang, Hanbin Zhao, Zhang Chao, Hui Qian 0001, John C. S. Lui |
NeurIPS | 4 |
| 2024 | BELM: Bidirectional Explicit Linear Multi-step Sampler for Exact Inversion in Diffusion ModelsabstractThe inversion of diffusion model sampling, which aims to find the corresponding initial noise of a sample, plays a critical role in various tasks.
Recently, several heuristic exact inversion samplers have been proposed to address the inexact inversion issue in a training-free manner.
However, the theoretical properties of these heuristic samplers remain unknown and they often exhibit mediocre sampling quality.
In this paper, we introduce a generic formulation, \emph{Bidirectional Explicit Linear Multi-step} (BELM) samplers, of the exact inversion samplers, which includes all previously proposed heuristic exact inversion samplers as special cases.
The BELM formulation is derived from the variable-stepsize-variable-formula linear multi-step method via integrating a bidirectional explicit constraint. We highlight this bidirectional explicit constraint is the key of mathematically exact inversion.
We systematically investigate the Local Truncation Error (LTE) within the BELM framework and show that the existing heuristic designs of exact inversion samplers yield sub-optimal LTE.
Consequently, we propose the Optimal BELM (O-BELM) sampler through the LTE minimization approach.
We conduct additional analysis to substantiate the theoretical stability and global convergence property of the proposed optimal sampler.
Comprehensive experiments demonstrate our O-BELM sampler establishes the exact inversion property while achieving high-quality sampling.
Additional experiments in image editing and image interpolation highlight the extensive potential of applying O-BELM in varying applications. Fangyikang Wang, Hubery Yin, Yuejiang Dong, Huminhao Zhu, Zhang Chao, Hanbin Zhao, Hui Qian 0001, Chen Li 0031 |
NeurIPS | 6 |
| 2024 | Solving Zero-Sum Markov Games with Continuous State via Spectral Dynamic EmbeddingabstractIn this paper, we propose a provably efficient natural policy gradient algorithm called Spectral Dynamic Embedding Policy Optimization (\SDEPO) for two-player zero-sum stochastic Markov games with continuous state space and finite action space.
In the policy evaluation procedure of our algorithm, a novel kernel embedding method is employed to construct a finite-dimensional linear approximations to the state-action value function.
We explicitly analyze the approximation error in policy evaluation, and show that \SDEPO\ achieves an $\tilde{O}(\frac{1}{(1-\gamma)^3\epsilon})$ last-iterate convergence to the $\epsilon-$optimal Nash equilibrium, which is independent of the cardinality of the state space.
The complexity result matches the best-known results for global convergence of policy gradient algorithms for single agent setting.
Moreover, we also propose a practical variant of \SDEPO\ to deal with continuous action space and empirical results demonstrate the practical superiority of the proposed method. Chenhao Zhou 0003, Zebang Shen, Zhang Chao, Hanbin Zhao, Hui Qian 0001 |
NeurIPS | 4 |
| 2024 | MgSvF: Multi-Grained Slow versus Fast Framework for Few-Shot Class-Incremental LearningabstractAs a challenging problem, few-shot class-incremental learning (FSCIL) continually learns a sequence of tasks, confronting the dilemma between slow forgetting of old knowledge and fast adaptation to new knowledge. In this paper, we concentrate on this "slow versus fast" (SvF) dilemma to determine which knowledge components to be updated in a slow fashion or a fast fashion, and thereby balance old-knowledge preservation and new-knowledge adaptation. We propose a multi-grained SvF learning strategy to cope with the SvF dilemma from two different grains: intra-space (within the same feature space) and inter-space (between two different feature spaces). The proposed strategy designs a novel frequency-aware regularization to boost the intra-space SvF capability, and meanwhile develops a new feature space composition operation to enhance the inter-space SvF learning performance. With the multi-grained SvF learning strategy, our method outperforms the state-of-the-art approaches by a large margin. Hanbin Zhao, Yongjian Fu 0002, Mintong Kang, Qi Tian 0001, Fei Wu 0001, Xi Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Epoch-Evolving Gaussian Process Guided Learning for ClassificationabstractThe conventional mini-batch gradient descent algorithms are usually trapped in the local batch-level distribution information, resulting in the ``zig-zag'' effect in the learning process. To characterize the correlation information between the batch-level distribution and the global data distribution, we propose a novel learning scheme called epoch-evolving Gaussian process guided learning (GPGL) to encode the global data distribution information in a non-parametric way. Upon a set of class-aware anchor samples, our GP model is built to estimate the class distribution for each sample in mini-batch through label propagation from the anchor samples to the batch samples. The class distribution, also named the context label, is provided as a complement for the ground-truth one-hot label. Such a class distribution structure has a smooth property and usually carries a rich body of contextual information that is capable of speeding up the convergence process. With the guidance of the context label and ground-truth label, the GPGL scheme provides a more efficient optimization through updating the model parameters with a triangle consistency loss. Furthermore, our GPGL scheme can be generalized and naturally applied to the current deep models, outperforming the state-of-the-art optimization methods on six benchmark datasets. Jiabao Cui, Xuewei Li 0003, Hanbin Zhao, Hui Wang 0107, Bin Li 0038, Xi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Unsupervised Domain Adaptation With Class-Aware Memory AlignmentabstractUnsupervised domain adaptation (UDA) is to make predictions on unlabeled target domain by learning the knowledge from a label-rich source domain. In practice, existing UDA approaches mainly focus on minimizing the discrepancy between different domains by mini-batch training, where only a few instances are accessible at each iteration. Due to the randomness of sampling, such a batch-level alignment pattern is unstable and may lead to misalignment. To alleviate this risk, we propose class-aware memory alignment (CMA) that models the distributions of the two domains by two auxiliary class-aware memories and performs domain adaptation on these predefined memories. CMA is designed with two distinct characteristics: class-aware memories that create two symmetrical class-aware distributions for different domains and two reliability-based filtering strategies that enhance the reliability of the constructed memory. We further design a unified memory-based loss to jointly improve the transferability and discriminability of features in the memories. State-of-the-art (SOTA) comparisons and careful ablation studies show the effectiveness of our proposed CMA. Hui Wang 0107, Liangli Zheng, Hanbin Zhao, Shijian Li, Xi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Elastic Knowledge Distillation by Learning From RecollectionabstractModel performance can be further improved with the extra guidance apart from the one-hot ground truth. To achieve it, recently proposed recollection-based methods utilize the valuable information contained in the past training history and derive a "recollection" from it to provide data-driven prior to guide the training. In this article, we focus on two fundamental aspects of this method, i.e., recollection construction and recollection utilization. Specifically, to meet the various demands of models with different capacities and at different training periods, we propose to construct a set of recollections with diverse distributions from the same training history. After that, all the recollections collaborate together to provide guidance, which is adaptive to different model capacities, as well as different training periods, according to our similarity-based elastic knowledge distillation (KD) algorithm. Without any external prior to guide the training, our method achieves a significant performance gain and outperforms the methods of the same category, even as well as KD with well-trained teacher. Extensive experiments and further analysis are conducted to demonstrate the effectiveness of our method. Yongjian Fu 0002, Hanbin Zhao, Wenfu Wang, Weihao Fang, Yueting Zhuang, Xi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | RBC: Rectifying the Biased Context in Continual Semantic Segmentation
Hanbin Zhao, Fengyu Yang 0003, Xinghe Fu, Xi Li 0001 |
ECCV (34) | 1 |
| 2022 | Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image GenerationabstractAs a challenging task, text-to-image generation aims to generate photo-realistic and semantically consistent images according to the given text descriptions. Existing methods mainly extract the text information from only one sentence to represent an image and the text representation effects the quality of the generated image well. However, directly utilizing the limited information in one sentence misses some key attribute descriptions, which are the crucial factors to describe an image accurately. To alleviate the above problem, we propose an effective text representation method with the complements of attribute information. Firstly, we construct an attribute memory to jointly control the text-to-image generation with sentence input. Secondly, we explore two update mechanisms, sample-aware and sample-joint mechanisms, to dynamically optimize a generalized attribute memory. Furthermore, we design an attribute-sentence-joint conditional generator learning scheme to align the feature embeddings among multiple representations, which promotes the cross-modal network training. Experimental results illustrate that the proposed method obtains substantial performance improvements on both the CUB (FID from 14.81 to 8.57) and COCO (FID from 21.42 to 12.39) datasets. Xintian Wu, Hanbin Zhao, Liangli Zheng, Shouhong Ding, Xi Li 0001 |
ACM Multimedia | 2 |
| 2022 | PcmNet: Position-sensitive context modeling network for temporal action localization
Hanbin Zhao, Guangchen Lin, Songcen Xu, Xi Li 0001 |
Neurocomputing | 2 |
| 2022 | Structure-conditioned adversarial learning for unsupervised domain adaptation
Hui Wang 0107, Hanbin Zhao, Fei Wu 0001, Xi Li 0001 |
Neurocomputing | 4 |
| 2022 | Memory-Efficient Class-Incremental Learning for Image ClassificationabstractWith the memory-resource-limited constraints, class-incremental learning (CIL) usually suffers from the "catastrophic forgetting" problem when updating the joint classification model on the arrival of newly added classes. To cope with the forgetting problem, many CIL methods transfer the knowledge of old classes by preserving some exemplar samples into the size-constrained memory buffer. To utilize the memory buffer more efficiently, we propose to keep more auxiliary low-fidelity exemplar samples, rather than the original real-high-fidelity exemplar samples. Such a memory-efficient exemplar preserving scheme makes the old-class knowledge transfer more effective. However, the low-fidelity exemplar samples are often distributed in a different domain away from that of the original exemplar samples, that is, a domain shift. To alleviate this problem, we propose a duplet learning scheme that seeks to construct domain-compatible feature extractors and classifiers, which greatly narrows down the above domain gap. As a result, these low-fidelity auxiliary exemplar samples have the ability to moderately replace the original exemplar samples with a lower memory cost. In addition, we present a robust classifier adaptation scheme, which further refines the biased classifier (learned with the samples containing distillation label knowledge about old classes) with the help of the samples of pure true class labels. Experimental results demonstrate the effectiveness of this work against the state-of-the-art approaches. We will release the code, baselines, and training statistics for all models to facilitate future research. Hanbin Zhao, Hui Wang 0107, Yongjian Fu 0002, Fei Wu 0001, Xi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | What and Where: Learn to Plug Adapters via NAS for Multidomain LearningabstractAs an important and challenging problem, multidomain learning (MDL) typically seeks a set of effective lightweight domain-specific adapter modules plugged into a common domain-agnostic network. Usually, existing ways of adapter plugging and structure design are handcrafted and fixed for all domains before model learning, resulting in learning inflexibility and computational intensiveness. With this motivation, we propose to learn a data-driven adapter plugging strategy with neural architecture search (NAS), which automatically determines where to plug for those adapter modules. Furthermore, we propose an NAS-adapter module for adapter structure design in an NAS-driven learning scheme, which automatically discovers effective adapter module structures for different domains. Experimental results demonstrate the effectiveness of our MDL model against existing approaches under the conditions of comparable performance. Hanbin Zhao, Yongjian Fu 0002, Hui Wang 0107, Omar El Farouk Bourahla, Xi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | When Video Classification Meets Incremental ClassesabstractWith the rapid development of social media, tremendous videos with new classes are generated daily, which raise an urgent demand for video classification methods that can continuously update new classes while maintaining the knowledge of old videos with limited storage and computing resources. In this paper, we summarize this task as Class-Incremental Video Classification (CIVC) and propose a novel framework to address it. As a subarea of incremental learning tasks, the challenge of catastrophic forgetting is unavoidable in CIVC. To better alleviate it, we utilize some characteristics of videos. First, we decompose the spatio-temporal knowledge before distillation rather than treating it as a whole in the knowledge transfer process; trajectory is also used to refine the decomposition. Second, we propose a dual granularity exemplar selection method to select and store representative video instances of old classes and key-frames inside videos under a tight storage budget. We benchmark our method and previous SOTA class-incremental learning methods on Something-Something V2 and Kinetics datasets, and our method outperforms previous methods significantly. Hanbin Zhao, Shihao Su, Yongjian Fu 0002, Zibo Lin, Xi Li 0001 |
ACM Multimedia | 1 |
| 2021 | Progressive Class-Based Expansion Learning for Image ClassificationabstractIn this paper, we propose a novel image process scheme called class-based expansion learning for image classification, which aims at improving the supervision-stimulation frequency for the samples of the confusing classes. Class-based expansion learning takes a bottom-up growing strategy in a class-based expansion optimization fashion, which pays more attention to the quality of learning the fine-grained classification boundaries for the preferentially selected classes. Besides, we develop a class confusion criterion to select the confusing class preferentially for training. In this way, the classification boundaries of the confusing classes are frequently stimulated, resulting in a fine-grained form. Experimental results demonstrate the effectiveness of the proposed scheme on several benchmarks. Hui Wang 0107, Hanbin Zhao, Xi Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2018 | Progressive Blockwise Knowledge Distillation for Neural Network AccelerationabstractAs an important and challenging problem in machine learning and computer vision, neural network acceleration essentially aims to enhance the computational efficiency without sacrificing the model accuracy too much. In this paper, we propose a progressive blockwise learning scheme for teacher-student model distillation at the subnetwork block level. The proposed scheme is able to distill the knowledge of the entire teacher network by locally extracting the knowledge of each block in terms of progressive blockwise function approximation. Furthermore, we propose a structure design criterion for the student subnetwork block, which is able to effectively preserve the original receptive field from the teacher network. Experimental results demonstrate the effectiveness of the proposed scheme against the state-of-the-art approaches. Hui Wang 0107, Hanbin Zhao, Xi Li 0001, Xu Tan 0003 |
IJCAI | 2 |