Han Fang 0002

dblp:209/7867-2 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-4379-2971ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval
abstract
In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in aligning with complex or ambiguous moments. To overcome these limitations, we explore Deep Evidential Regression (DER) to construct a vanilla Evidential baseline. However, this approach encounters two major issues: the inability to effectively handle modality imbalance and the structural differences in DER's heuristic uncertainty regularizer, which adversely affect uncertainty estimation. This misalignment results in high uncertainty being incorrectly associated with accurate samples rather than challenging ones. Our observations indicate that existing methods lack the adaptability required for complex video scenarios. In response, we propose Debiased Evidential Learning for Moment Retrieval (DEMR), a novel framework that incorporates a Reflective Flipped Fusion (RFF) block for cross-modal alignment and a query reconstruction task to enhance text sensitivity, thereby reducing bias in uncertainty estimation. Additionally, we introduce a Geom-regularizer to refine uncertainty predictions, enabling adaptive alignment with difficult moments and improving retrieval accuracy. Extensive testing on standard datasets and debiased datasets ActivityNet-CD and Charades-CD demonstrates significant enhancements in effectiveness, robustness, and interpretability, positioning our approach as a promising solution for temporal-semantic robustness in moment retrieval.
Haojian Huang, Kaijing Ma, Xianghao Zang, Han Fang 0002, Chao Ban, Hao Sun 0038, Mulin Chen, Zhongjiang He
AAAI7
2025 Trusted Unified Feature-Neighborhood Dynamics for Multi-View Classification
abstract
Multi-view classification (MVC) faces inherent challenges due to domain gaps and inconsistencies across different views, often resulting in uncertainties during the fusion process. While Evidential Deep Learning (EDL) has been effective in addressing view uncertainty, existing methods predominantly rely on the Dempster-Shafer combination rule, which is sensitive to conflicting evidence and often neglects the critical role of neighborhood structures within multi-view data. To address these limitations, we propose a Trusted Unified Feature-NEighborhood Dynamics (TUNED) model for robust MVC. This method effectively integrates local and global feature-neighborhood (F-N) structures for robust decision-making. Specifically, we begin by extracting local F-N structures within each view. To further mitigate potential uncertainties and conflicts in multi-view fusion, we employ a selective Markov random field that adaptively manages cross-view neighborhood dependencies. Additionally, we employ a shared parameterized evidence extractor that learns global consensus conditioned on local F-N structures, thereby enhancing the global integration of multi-view features. Experiments on benchmark datasets show that our method improves accuracy and robustness over existing approaches, particularly in scenarios with high uncertainty and conflicting views.
Haojian Huang, Chuanyu Qin, Zhe Liu 0041, Kaijing Ma, Han Fang 0002, Chao Ban, Hao Sun 0038, Zhongjiang He
AAAI6
2025 FASTER: Face Attribute Sliders with Semantic Rewards
abstract
Large-scale text-to-image generative models have demonstrated remarkable success in generating diverse and high-quality faces. However, current methods for face editing often unintentionally modify facial features that are intended to be preserved. Multi-step denoising methods necessitate storing multi-step gradients, leading to considerable time and memory consumption. In this study, we propose FASTER(Face Attribute Sliders wiTh sEmantic Rewards), an effective method that employs stable diffusion models for face attribute editing. The key idea is to identify a low-rank attribute editing direction by leveraging attribute reward and S-CLIP reward between the original face and the edited face. This process helps to establish the desired face attribute slider. To acquire the edited face, we introduce an efficient one-step reward technique by utilizing denoised results at random timesteps for learning. This technique reduces training time by 6x. FASTER achieves 98.67% editing accuracy while simultaneously improving attribute preservation by nearly 10% compared to other methods on the CelebA-HQ dataset, all without compromising identity information.
Jingyan Chen, Lanxiang Zhou, Han Fang 0002, Zerun Feng, Chao Ban, Hao Sun 0038, Jiani Hu
ICASSP3
2025 ViCo: A Multitask Video-enhanced and Cognition-preserving Modality Alignment Training Framework
abstract
The rapid development of multimodal large language models (MLLMs) has brought significant breakthroughs to this field. However, current MLLMs typically rely on vision instruction tuning based on large language models (LLMs) to endow them with multimodal capabilities, which may lead to low video utilization and catastrophic forgetting due to the task-specific nature of vision instruction tuning. Take case of multiple-choice Video Question Answering (VideoQA) as an example, requiring only selecting correct answers makes it easier for LLMs to shortcut rather than fully comprehend video content. Moreover, the fixed task format of ’’predicting the correct option" risks catastrophic forgetting, which can cause LLMs to lose their origin cognitive abilities. To address this, we propose ViCo, a training framework that enables LLMs to perform a specific auxiliary tasks during vision instruction tuning, which helps the model leverage video information more thoroughly while mitigating catastrophic forgetting, thereby improving modality alignment and enhancing performance. We evaluate our method on several strong VideoQA benchmarks, where it outperforms all state-of-the-art methods, demonstrating its effectiveness and generality. Additionally, we also provide extensive analysis on the video information utilization and catastrophic forgetting resulting. The code will be made available at https://github.com/dunknsabsw/ViCo.
Zhenda Yu, Jiayu Shen, Lanxiang Zhou, Han Fang 0002, Xianghao Zang, Chao Ban, Jingfeng Chen, Zhongjiang He, Hao Sun 0038, Yanmei Kang
ICASSP5
2025 SelectVision: Adaptive Vision Resolution Selection for Visual Document Understanding
Zhongjiang He, Han Fang 0002, Hao Sun 0015, Kongming Liang, Zhanyu Ma
ICDAR (4)4
2025 DDL: Dynamic Direction Learning for Semi-Supervised Facial Expression Recognition
abstract
Most semi-supervised facial expression recognition (FER) algorithms leverage pseudo-labeling to mine additional information from unlabeled samples. Despite its good performance, two critical issues persist: class imbalance and domain shift. The former is a typical challenge due to the significant variation in sample numbers across different FER classes, resulting in highly imbalanced pseudo labels in existing semi-supervised methods. For the latter, given that labeled and unlabeled data usually come from different sources, a considerable domain gap might exist, leading the model to generate low-quality pseudo labels. To tackle these issues, we introduce a novel semi-supervised FER algorithm called Dynamic Direction Learning (DDL), which consists of adaptive balance learning (ABL) and adaptive alignment learning (AAL). ABL allows a balanced training process by dynamically adjusting the constraints of self-training based on the performance of a balanced validation dataset. Moreover, AAL adaptively aligns the feature distribution of labeled and unlabeled data by minimizing their distance in feature space. Additionally, a role rotation mechanism (RRM) is proposed to avoid confirmation bias, which further improves self-training. Extensive experiments demonstrate that DDL achieves state-of-the-art performance on different FER datasets.
Yuhang Zhang 0016, Han Fang 0002, Jiani Hu, Weihong Deng
IEEE Trans. Affect. Comput.4
2024 ProTA: Probabilistic Token Aggregation for Text-Video Retrieval
abstract
Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (ProTA) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, ProTA achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%).
Han Fang 0002, Xianghao Zang, Chao Ban, Zerun Feng, Lanxiang Zhou, Zhongjiang He, Hao Sun 0038
ICME1
2024 GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment Tuning
abstract
Recent text-to-image (T2I) synthesis models have demonstrated intriguing abilities to produce high-quality images based on text prompts. However, current models still face Text-Image Misalignment problem (e.g., attribute errors and relation mistakes) for compositional generation. Existing models attempted to condition T2I models on grounding inputs to improve controllability while ignoring the explicit supervision from the layout conditions. To tackle this issue, we propose Grounded jOint lAyout aLignment (GOAL), an effective framework for T2I synthesis. Two novel modules, discriminative semantic alignment (DSAlign) and masked attention alignment (MAAlign), are proposed and incorporated in this framework to improve the text-image alignment. DSAlign leverages discriminative tasks at the region-wise level to ensure low-level semantic alignment. MAAlign provides high-level attention alignment by guiding the model to focus on the target object. We also build a dataset GOAL2K for model fine-tuning, which composes 2000 semantically accurate image-text pairs and their layout annotations. Comprehensive evaluations on T2I-Compbench, NSR-1K, and Drawbench demonstrate the superior generation performance of our method. Especially, there are improvements of 19%, 13%, and 12% in color, shape, and texture metrics for T2I-Compbench. Additionally, Q-Align metrics demonstrate that our method can generate images of higher quality.
Han Fang 0002, Zerun Feng, Kaijing Ma, Chao Ban, Xianghao Zang, Lanxiang Zhou, Zhongjiang He, Jingyan Chen, Jiani Hu, Hao Sun 0038
ACM Multimedia2
2023 Mask to Reconstruct: Cooperative Semantics Completion for Video-text Retrieval
abstract
Recently, masked video modeling has been widely explored and improved the model's understanding ability of visual regions at a local level. However, existing methods usually adopt random masking and follow the same reconstruction paradigm to complete the masked regions, which do not leverage the correlations between cross-modal content. In this paper, we present MAsk for Semantics COmpleTion (MASCOT) based on semantic-based masked modeling. Specifically, after applying attention-based video masking to generate high-informed and low-informed masks, we propose Informed Semantics Completion to recover masked semantics information. The recovery mechanism is achieved by aligning the masked content with the unmasked visual regions and corresponding textual context, which makes the model capture more text-related details at a patch level. Additionally, we shift the emphasis of reconstruction from irrelevant backgrounds to discriminative parts to ignore regions with low-informed masks. Furthermore, we design co-learning to incorporate video cues under different masks and learn more aligned representation. Our MASCOT performs state-of-the-art performance on four text-video retrieval benchmarks, including MSR-VTT, LSMDC, ActivityNet, and DiDeMo.
Han Fang 0002, Zhifei Yang 0004, Xianghao Zang, Chao Ban, Zhongjiang He, Hao Sun 0038, Lanxiang Zhou
ACM Multimedia1
2023 A Baseline Investigation: Transformer-based Cross-view Baseline for Text-based Person Search
abstract
This paper investigates a baseline approach for text-based person search by using a transformer-based framework. Existing methods usually treat the visual and textual features as independent entities for speeding up the model inference process. However, the attention to the same images should be changed according to different texts. In this paper, we use a commonly employed framework with a fused feature as the baseline, which overcomes the misalignment problem introduced by fixed features. A thorough investigation is conducted in this paper. Moreover, we propose Cross-View Matching (CVM) to provide challenging, positive text-image pairs that enable the model to learn cross-view meta-information. Furthermore, we suggest a novel evaluation process to reduce the inference time and GPU memory demand. The experiments are conducted on CUHK-PEDES, ICFG-PEDES, and RSTPReid benchmarks. Through extensive parameter analysis, the potentials of a transformer-based framework are fully explored. Although the proposed scheme is a simple framework, it achieves significant performance improvements compared with other state-of-the-art methods.
Xianghao Zang, Wei Gao 0003, Ge Li 0002, Han Fang 0002, Chao Ban, Zhongjiang He, Hao Sun 0038
ACM Multimedia4
2023 Transferring Image-CLIP to Video-Text Retrieval via Temporal Relations
abstract
We present a novel network to transfer the image-language pre-trained model to video-text retrieval in an end-to-end manner. Leading approaches in the domain of video-and-language learning try to distill the spatio-temporal video features and multi-modal interaction between videos and language from a large-scale video-text dataset. Differently, we leverage the pre-trained image-language model, and simplify it as a two-stage framework including co-learning of image and text, and enhancing temporal relations between video frames and video-text respectively. Specifically, based on the spatial semantics captured by Contrastive Language-Image Pre-training (CLIP) model, our model involves a Temporal Difference Block (TDB) to capture motions at fine temporal video frames, and a Temporal Alignment Block (TAB) to re-align the tokens of video clips and phrases and enhance the cross-modal correlation. These two temporal blocks efficiently realize video-language learning and enable the proposed model to scale well on comparatively small datasets. We conduct extensive experimental studies including ablation studies and comparisons with existing SOTA methods, and our proposed approach outperforms them on the popularly-employed text-to-video and video-to-text retrieval benchmarks, including MSR-VTT, MSVD, LSMDC, and VATEX.
Han Fang 0002, Pengfei Xiong, Luhui Xu, Wenhan Luo
IEEE Trans. Multim.1
2022 Dynamic Training Data Dropout for Robust Deep Face Recognition
abstract
Learning with noise is a practically challenging problem in deep face recognition. Despite the success of large margin softmax loss functions, these methods are designed for clean face databases. Considering the inevitable noise in the large scale databases, we first analyze the performance of noise in the training databases. For noise-robust deep face recognition, we propose a dynamic training data dropout (DTDD) method to dynamically filter the noise in the training database and gradually form a stable refined database for model learning. Specifically, we leverage the information provided by the model predictions of accumulated training epochs, which can distinguish regular samples and noise effectively and accurately. The proposed DTDD method is easy and stable for implementation, and can be combined with existing state-of-the-art loss functions and network architectures. Extensive experiments on CASIA-WebFace, VGGFace2, and MS-Celeb-1 M databases empirically demonstrate that our proposed method can robustly train deep face recognition models in the presence of label noise and low quality images.
Yaoyao Zhong, Weihong Deng, Han Fang 0002, Jiani Hu, Dongyue Zhao, Dongchao Wen
IEEE Trans. Multim.3
2021 Augmented Face Representation Learning via Transitive Distillation
abstract
The wild face of large variations is hard to recognize in unconstrained scenarios. To tackle this issue, existing works synthesize and augment the variation-specific faces for recognition. However, directly feeding generated samples results in negative transfer, because the feature spaces are shifted compared with normal samples. Instead, we propose a transitive distillation network (TDNet) that introduces a transitive domain to transfer cross-variation representations, which alleviates the negative influence of synthesized data. Specifically, data of diverse variations are firstly synthesized. Then we construct distributions from different variations as teachers to distill student. The negative transfer is mitigated by adopting adaptor as a bridge to break large domain distance. To handle faces of different quality, we propose a novel strategy to define easy and hard samples, which are utilized to select specific transitive status. Meanwhile, bilateral classification with curriculum learning is proposed to improve confidence of synthesized data gradually, enhancing the robustness of representation learning. Experiments show that our method achieves superiority on unconstrained face benchmarks such as IJB-C and SCface, while maintaining competence on general test sets.
Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu, Dongyue Zhao, Dongchao Wen
FG1
2021 Adaptive Re-Balancing Network with Gate Mechanism for Long-Tailed Visual Question Answering
abstract
Visual Question Answering (VQA) is a challenging task which requires a fine-grained semantic understanding of visual and textual contents. Existing works focus on better modality representations. However, these methods give little consideration to the long-tailed data distribution in common VQA datasets. The extreme class imbalance causes training bias to behave well in head class, but fail in tail class. Therefore, we propose a unified Adaptive Re-balancing Network (ARN) to take care of classification in both head and tail classes, exhaustively improving performance for VQA. Specifically, two training branches are introduced to per-form their own duty iteratively, which learn the universal representations first and then emphasize the tail data progressively by the re-balancing branch with adaptive learning. Meanwhile, contextual information in the question is vital for guiding accurate visual attention. Thus our network is further equipped with a novel gate mechanism to give higher weight to contextual information. The Experimental results on common benchmarks such as VQA-v2 have demonstrated the superiority of our method compared with state of the art.
Hongyu Chen 0005, Ruifang Liu, Han Fang 0002
ICASSP3
2020 Generate to Adapt: Resolution Adaption Network for Surveillance Face Recognition
Han Fang 0002, Weihong Deng, Yaoyao Zhong, Jiani Hu
ECCV (15)1