EDBT 2026 Demo / reviewers in the wild / expert
Kaihao Zhang
dblp:179/6089
· DBLP profile ↗
89ranked-venue papers
19as first author
75since 2021 · last 2026
0000-0002-4317-660XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 10 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 14 first-author · 40 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-powered assistant with electrotactile feedback to assist blind and low vision people with maps and routes preview
Chutian Jiang, Yinan Fan, Junan Xie, Emily Kuang, Kaihao Zhang, Mingming Fan 0001 |
Int. J. Hum. Comput. Stud. | 5 |
| 2026 | A guard against ambiguous sentiment for multimodal aspect-level sentiment classification
Yanjing Wang 0008, Bin Shi 0003, Kaihao Zhang, Bo Dong 0001 |
Inf. Process. Manag. | 5 |
| 2026 | Condition-Guided Diffusion for Multi-Modal Pedestrian Trajectory Prediction Incorporating Intention and Interaction PriorsabstractPedestrian behavior exhibits inherent multi-modality, necessitating predictions that balance accuracy and diversity to adapt effectively to various complex scenarios. However, conventional noise addition in diffusion models is often aimless and unguided, leading to redundant noise reduction steps and the generation of uncontrollable samples. To address these issues, we propose a Prior Condition-Guided Diffusion Model (CGD-TraP) for multi-modal pedestrian trajectory prediction. Instead of directly adding Gaussian noise to trajectories at each timestep during the forward process, our approach leverages internal intention and external interaction to guide noise estimation. Specifically, we design two specialized modules to extract and aggregate intention and interaction features. These features are then adaptively fused through a spatial-temporal fusion based on selective state space, which estimates a controllable noisy trajectory distribution. By optimizing the noise addition process in a more controlled and efficient manner, our method ensures that the denoising process is effectively guided, resulting in predictions that are both accurate and diverse. Extensive experiments on the ETH-UCY, SDD, and NBA datasets demonstrate that CGD-TraP surpasses state-of-the-art diffusion-based and other generative methods, achieving superior efficiency, accuracy, and diversity. Yanghong Liu, Xingping Dong, Yutian Lin, Mang Ye, Kaihao Zhang, Bo Du 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Enhancing outdoor vision: Binocular desnowing with dual-stream temporal transformerabstractVideo desnowing, aimed at removing snowflakes and enhancing the quality of videos, is a crucial yet intricate task essential for improving the effectiveness of outdoor vision systems. Compared to rain and haze, the inherent opacity and diverse morphology of snowflakes result in more pronounced background occlusions, thereby challenging the efficacy of current desnowing techniques, particularly those focusing solely on images or videos captured from a monocular perspective. To address these challenges, this paper proposes a Dual-Stream Temporal Transformer (DSTT) to advance snow removal and visual enhancement by leveraging comprehensive information from stereo views and spatial-temporal cues. More specifically, it incorporates a Dual-Stream Weight-shared Transformer (DSWT) module to exploit spatial information from different views. This module employs a hierarchical weight-sharing strategy to extract fused spatial features across different views from low-level to high-level layers. Subsequently, the Dual-Stream ConvLSTM (DS-CLSTM) module is introduced to capture temporal correlations across streaming frames. By combining temporal-spatial cues and complementary details from diverse views, videos can be effectively restored while preserving the original content’s details. In addition, two binocular snowy datasets – SnowKITTI2012 and SnowKITTI 2015 – are presented, providing a valuable resource for evaluating the binocular desnowing task. Comprehensive experiments evaluated on both synthetic and real-world snowy datasets demonstrate that our proposed method outperforms the state-of-the-art baselines. En Yu, Jie Lu 0001, Kaihao Zhang, Guangquan Zhang 0001 |
Pattern Recognit. | 3 |
| 2026 | LoHi-SSL: A Multi-Level Synergistic Learning Model for Integrating Single-Cell Multi-Omics Data via Low- and High-Order Information FusionabstractRecent advancements in single-cell sequencing technologies have enabled researchers to identify cell subpopulations and their functional states with greater accuracy, thereby uncovering cellular heterogeneity. However, due to the heterogeneity across different single-cell multi-omics datasets and the intrinsic variability among cells, effectively integrating data from multiple molecular layers remains a significant challenge. To address this issue, a Single-cell Multi-level Synergistic Learning model (LoHi-SSL) is proposed, which integrates low-order and high-order information to achieve efficient multi-omics data fusion. LoHi-SSL consists of three key modules: low-order information learning, high-order information learning, and feature integration. The low-order information learning module focuses on addressing intra-omics cellular heterogeneity. It first extracts features from each omics dataset, then constructs a graph structure to capture intercellular relationships. A Graph Autoencoder is employed to extract local neighborhood information, effectively preserving intra-omics cellular similarity. The high-order information learning module is designed to eliminate cross-omics heterogeneity and align data in a unified latent representation space. To achieve this, multi-omics hypergraph learning is introduced to model complex cellular relationships across different omics, enhancing feature interactions.In the feature integration module, contrastive learning is utilized to guide the model in learning more discriminative feature representations by constructing positive and negative sample pairs. For different omics data from the same cell, LoHi-SSL encourages feature alignment within a shared latent space, reducing cross-omics heterogeneity. Meanwhile, for different cell types, a contrastive loss function is applied to increase the separation between their representations, thereby enhancing cellular distinguishability and achieving efficient single-cell multi-omics integration. Experimental results demonstrate that LoHi-SSL outperforms existing methods on six publicly available datasets, achieving superior performance in clustering tasks, particularly in terms of NMI (Normalized Mutual Information), ARI (Adjusted Rand Index), AMI (Adjusted Mutual Information), and ACC (Clustering Accuracy). Furthermore, robustness analysis shows that LoHi-SSL exhibits strong resistance to noise. Additionally, cell trajectory analysis using the latent representations learned by LoHi-SSL accurately reflects biological evolutionary pathways. In summary, LoHi-SSL provides an efficient and robust approach for single-cell multi-omics data integration, offering a powerful tool for studying cellular state transitions, heterogeneity, and regulatory mechanisms. Xiaoyun Xiong, Kaihao Zhang, Chengdong Zhang, Yuanyuan Zhang 0008 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2026 | Virtual Consistency Model for All-in-One Image RestorationabstractAll-in-one Image Restoration (AIR) seeks to address diverse degradations using a unified model trained only once. Existing methods often rely on degradation-specific guidance, leading to conflicting gradients during training. In contrast, diffusion models offer a promising alternative by operating in a high-noise space where diverse degradations exhibit a homogeneous Gaussian distribution. This characteristic alleviates gradient conflicts associated with task-specific degradations. However, existing diffusion-based AIR methods often suffer from a lack of direct supervision in the image space, leading to error accumulation during the iterative denoising process and image fidelity compromisation. This highlights a fundamental dilemma for AIR: the optimal space for modeling degradations is inherently suboptimal for preserving image fidelity. To address this issue, we propose a Virtual Consistency Model for AIR (VCMAIR), which restores images in the high-noise space while employing a novel consistency function to enforce accurate supervision in the image space. Extensive experiments demonstrate that the proposed method outperforms existing state-of-the-art methods across a comprehensive benchmark of diverse degradation scenarios, including both standard AIR tasks and challenging real-world image restoration tasks. Jiawei Wu 0001, Luwei Tu, Zhi Jin 0002, Kaihao Zhang, Wenqi Ren, Xiaochun Cao |
IEEE Trans. Image Process. | 5 |
| 2026 | Dual Alignment-Enhanced Fashion Vision-Language Pre-TrainingabstractFashion vision-language pre-training (VLP) models have demonstrated remarkable capabilities in excelling at a wide range of fashion cross-modal tasks. However, current models still face three notable limitations (1) an inability to discern varying levels of consistency between textual descriptions and multi-view images, (2) a deficiency in explicit fine-grained alignment between images and text, and (3) a lack of specific supervision mechanisms for facilitating global joint embedding learning. To address these limitations, we propose a novel dual alignment-enhanced fashion VLP model. This model delves deeply into the rich resources of multi-view images and semantic attributes associated with each fashion item. Notably, we introduce two novel pre-training tasks: Multi-grained Adaptive Image-Text Alignment (MAITA) and Joint Embedding-oriented Alignment (JEA). MAITA focuses on optimizing the text/image encoder by orchestrating adaptive alignment between multi-view images and input text. This encompasses both coarse-grained and fine-grained alignment strategies to enrich semantic understanding, while JEA is devised to supervise the fine-grained semantic learning process of the global multimodal joint embedding. Experimental results spanning four diverse downstream tasks, including cross-modal retrieval, text-guided image retrieval, category recognition, and subcategory recognition, substantiate the significant performance superiority of our model over prior state-of-the-art fashion VLP models. Weili Guan, Kejie Wang, Xuemeng Song, Kaihao Zhang, Xiaojun Chang, Shengping Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language ModelsabstractJiaqi Zhao, Miao Zhang, Ming Wang, Yuzhang Shang, Kaihao Zhang, Weili Guan, Yaowei Wang, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Miao Zhang 0022, Yuzhang Shang, Kaihao Zhang, Weili Guan, Yaowei Wang 0001, Min Zhang 0005 |
ACL (1) | 5 |
| 2025 | DBAANet: Dual-Branch Attention Aggregation Network for Medical Image SegmentationabstractRecent hybrid architectures combining Transformer encoders with U-Net have advanced medical image segmentation by modeling global dependencies, but they often suffer from semantic misalignment between local CNN features and global Transformer representations, leading to inefficient multi-scale fusion and boundary detail loss in resourceconstrained clinical settings. To overcome these limitations, we propose DBAANet, a novel and efficient architecture featuring a dual-branch encoder that synergistically combines an enhanced Vision Transformer (ViT) with multi-scale convolutional blocks to optimize feature extraction. In the decoder, we employ channel and spatial attention mechanisms to capture complex inter-feature relationships and introduce a gated attention mechanism to fuse multi-source features from different stages, thereby fully leveraging diverse information sources. To evaluate its feasibility, we conducted extensive experiments on two 3D medical image segmentation tasks and five polyp segmentation datasets. The results demonstrate that DBAANet achieves strong performance while maintaining excellent computational efficiency, providing an accurate and efficient solution for medical image segmentation. Qihao Wang, Guiling Shi, Kaihao Zhang, Yuanyuan Zhang 0008 |
BIBM | 4 |
| 2025 | MMIF-COAL: Multi-Task Multi-Omics Integration Framework with Clustering Optimization and Adversarial LearningabstractSingle-cell multi-omics technologies enable simultaneous profiling of transcriptomic and epigenomic features, providing rich information for dissecting cellular heterogeneity and functional states. However, high dimensionality, modality-specific noise, and cross-modal heterogeneity pose significant challenges for integration and robust latent representation learning. In this study, we present MMIF-COAL, a multi-task framework for single-cell multiomics integration that jointly leverages clustering optimization and adversarial learning to achieve cross-modal feature alignment, enhanced latent space separability, and improved representation of rare cell types. We systematically evaluated MMIF-COAL on five publicly available single-cell multi-omics datasets. Experimental results demonstrate that MMIF-COAL consistently outperforms existing multi-task and single-task methods in clustering concordance and cell type classification accuracy, maintaining stable performance on large-scale and highly heterogeneous datasets. Ablation studies further confirm the contributions of KL divergence, entropy regularization, generator, discriminator, and data augmentation modules to the overall model performance. Analysis of cell type-specific latent features shows that MMIF-COAL effectively captures functional signals in highly active immune cells while maintaining low expression in resting or inactive cell types, demonstrating the biological interpretability of the learned latent representations. Overall, MMIF-COAL provides an efficient and robust solution for multi-task, multi-modal single-cell integration, offering valuable support for downstream analyses of cellular heterogeneity and multi-omics representations. Kaihao Zhang, Xiaoyun Xiong, Chengdong Zhang |
BIBM | 1 |
| 2025 | Designing LLM-Powered Multimodal Instructions to Support Rich Hands-on Skills Remote Learning: A Case Study with Massage Instructors and LearnersabstractAlthough remote learning is widely used for delivering and capturing knowledge, it has limitations in teaching hands-on skills that require nuanced instructions and demonstrations of precise actions, such as massage. Furthermore, scheduling conflicts between instructors and learners often limit the availability of real-time feedback, reducing learning efficiency. To address these challenges, we developed a synthesis tool utilizing an LLM-powered Virtual Teaching Assistant (VTA). This tool integrates multimodal instructions that convey precise data, such as stroke patterns and pressure control, while providing real-time feedback for learners and summarizing their performance for instructors. Our case study with instructors and learners demonstrated the effectiveness of these multimodal instructions and the VTA in enhancing massage teaching and learning. We then discuss the tools' use in other hands-on skills instruction and cognitive process differences in various courses. Chutian Jiang, Yinan Fan, Junan Xie, Emily Kuang, Baichuan Feng, Kaihao Zhang, Mingming Fan 0001 |
CHI | 6 |
| 2025 | CorrMoE: Mixture of Experts with De-Stylization Learning for Cross-Scene and Cross-Domain Correspondence PruningabstractEstablishing reliable correspondences between image pairs is a fundamental task in computer vision, underpinning applications such as 3D reconstruction and visual localization. Although recent methods have made progress in pruning outliers from dense correspondence sets, they often hypothesize consistent visual domains and overlook the challenges posed by diverse scene structures. In this paper, we propose CorrMoE, a novel correspondence pruning framework that enhances robustness under cross-domain and cross-scene variations. To address domain shift, we introduce a De-stylization Dual Branch, performing style mixing on both implicit and explicit graph features to mitigate the adverse influence of domain-specific representations. For scene diversity, we design a Bi-Fusion Mixture of Experts module that adaptively integrates multi-perspective features through linear-complexity attention and dynamic expert routing. Extensive experiments on benchmark datasets demonstrate that CorrMoE achieves superior accuracy and generalization compared to state-of-the-art methods. The code and pre-trained models are available at https://github.com/peiwenxia/CorrMoE. Peiwen Xia, Tangfei Liao, Danhuai Zhao, Jianjun Ke, Kaihao Zhang, Tong Lu 0002, Tao Wang 0052 |
ECAI | 6 |
| 2025 | LLFA: Fusing Global Illumination and Local Priors for Low-Light Face Image Enhancement with AdaptorabstractLow-light image enhancement problem has been widely studied. However, most existing methods do not perform well on low-light face images due to no specific facial characteristic considerations. We first create large-scale low-light face datasets with synthesized and real-world images to address the absence of suitable datasets. Our experiments show that existing LLIE and face restoration methods are limited in enhancing low-light face images. To overcome these challenges, we propose a novel framework, the Low-Light Face Adaptor (LLFA), featuring an auxiliary encoder and an adaptor module. The encoder captures global illumination information, while the adaptor module adaptively fuses this information with high-quality priors. We also introduce a joint learning strategy that optimizes the model by simultaneously learning face priors and the enhancement process. Comprehensive experiments demonstrate that LLFA significantly outperforms state-of-the-art methods. Ziqian Shao, Tao Wang 0052, Kaihao Zhang, Danhuai Zhao, Tong Lu 0002 |
ICASSP | 3 |
| 2025 | Segmentation-Guided Sparse Transformer for Under-Display Camera Image RestorationabstractUnder-display Camera is an emerging technology for full-screen display with a camera under the display. However, the current implementation of UDC causes serious image degradation. Incident light required for camera imaging undergoes attenuation and diffraction when passing through the display. Current UDC image restoration methods predominantly utilize convolutional networks, whereas transformer-based methods with superior performance are lacking. This paper proposes a Segmentation-Guided Sparse Transformer method (SGSFormer) for restoring images from UDC degraded images. Specifically, we utilize sparse self-attention to filter out redundant information and noise, directing the model’s attention to focus on the features more relevant to the degraded regions in need of reconstruction. Moreover, we integrate an instance segmentation map as prior information to guide sparse self-attention in filtering and focusing on the correct regions. Extensive experiments exhibit the superior performance of our model over the state-of-the-art methods. Jingyun Xue, Tao Wang 0052, Pengwen Dai, Kaihao Zhang |
ICASSP | 4 |
| 2025 | MaterialMVP: Illumination-Invariant Material Generation via Multi-View PBR Diffusion
Zebin He, Mingxin Yang, Tao Wang 0052, Kaihao Zhang, Guanying Chen, Jie Jiang 0015, Chunchao Guo, Wenhan Luo |
ICCV | 6 |
| 2025 | Cross-View Isolated Sign Language Recognition via View Synthesis and Feature Disentanglement
Kaihao Zhang, Xin Yu 0002 |
ICCV | 4 |
| 2025 | MOERL: When Mixture-Of-Experts Meet Reinforcement Learning for Adverse Weather Image Restoration
Tao Wang 0052, Peiwen Xia, Peng-Tao Jiang, Zhe Kong, Kaihao Zhang, Tong Lu 0002, Wenhan Luo |
ICCV | 6 |
| 2025 | LDPose: Towards Inclusive Human Pose Estimation for Limb-Deficient Individuals in the Wild
Jiaying Ying, Heming Du, Kaihao Zhang, Lincheng Li, Xin Yu 0002 |
ICCV | 3 |
| 2025 | Towards Multiple Character Image Animation Through Enhancing Implicit DecouplingabstractControllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. To address these challenges, we propose a novel multi-condition guided framework for character image animation, employing several well-designed input modules to enhance the implicit decoupling capability of the model. First, the optical flow guider calculates the background optical flow map as guidance information, which enables the model to implicitly learn to decouple the background motion into background constants and background momentum during training, and generate a stable background by setting zero background momentum during inference. Second, the depth order guider calculates the order map of the characters, which transforms the depth information into the positional information of multiple characters. This facilitates the implicit learning of decoupling different characters, especially in accurately separating the occluded body parts of multiple characters. Third, the reference pose map is input to enhance the ability to decouple character texture and pose information in the reference image. Furthermore, to fill the gap of fair evaluation of multi-character image animation, we propose a new benchmark comprising about 4,000 frames. Extensive qualitative and quantitative evaluations demonstrate that our method excels in generating high-quality character animations, especially in scenarios of complex backgrounds and multiple characters. Jingyun Xue, Hongfa Wang, Qi Tian 0003, Yue Ma 0016, Andong Wang, Zhiyuan Zhao 0002, Shaobo Min, Kaihao Zhang, Harry Shum, Wei Liu 0005, Mengyang Liu, Wenhan Luo |
ICLR | 9 |
| 2025 | Gradient-Based Adversarial Attacks on Deep LiDAR OdometryabstractAdversarial attacks have been recently investigated in LiDAR perception problems for autonomous driving, where a small perturbation of source inputs can result in incorrect predictions. However, most previous studies focus on attacks on single-frame perception modules, lacking explorations of attacks on consecutive-frame tasks, i.e. the LiDAR odometry. In this paper, we propose a gradient optimization-based adversarial attack towards deep LiDAR odometry networks. To generate point clouds consistent with real-world scenarios, we constrain adversarial points within the range of a small object, e.g. a traffic cone, and render new points to simulate real LiDAR measurements. By incorporating such adversarial points in consecutive frames, we demonstrate a significant decrease in pose estimation accuracy of current popular LiDAR odometry networks. In addition, we also evaluate traditional geometric odometry approaches and report their robustness against adversarial points. Extensive experiments on the KITTI and Waymo datasets illustrate the effectiveness of the proposed attack method and the vulnerability of deep LiDAR odometry networks against adversarial points. Zhenbo Song, Xuanzhu Chen, Zhenyuan Zhang 0001, Kaihao Zhang, Jianfeng Lu 0003 |
ICRA | 4 |
| 2025 | AuslanWeb: A Scalable Web-Based Australian Sign Language Communication System for Deaf and Hearing IndividualsabstractEffective communication between the deaf community and hearing individuals facilitates social inclusion, equal opportunities, and the dignity of vulnerable populations. However, existing region-specific sign language systems are constrained by limited training datasets and narrow topic domains, rendering them ineffective for bridging the linguistic gaps between sign languages and spoken languages. Auslan, as the sign language specific to Australia, still lacks a reliable bidirectional translation tool for effective communication. To address these challenges, we propose AuslanWeb, a web-based system for bidirectional translation of both isolated and successive sign language. For the former, AuslanWeb achieves high-precision mapping between isolated signs (glosses) and spoken language words or phrases through a multimodal recognition system and a versatile Auslan dictionary. For the latter, it leverages the advanced contextual understanding and text generation capabilities of Large Language Models (LLMs) to support bidirectional translation between successive sign language videos and long-form spoken language. By integrating linguistic structure with advanced AI capabilities, AuslanWeb overcomes the limitations of dataset dependency and enhances the scalability of sign language translation systems. The effectiveness of the system is further validated through user feedback, receiving consistent praise from Auslan experts, Australian deaf individuals, and volunteers. The demo video of AuslanWeb is provided here. Heming Du, Hongwei Sheng, Lincheng Li, Kaihao Zhang |
WWW | 5 |
| 2025 | DM-HAP: Diffusion model for accurate hand pose predictionabstractForecasting hand poses is a challenging task due to inherent uncertainties, occlusions, and inaccuracies in 3D pose estimation. Diffusion models provide a promising direction for predicting precise 3D hand poses under noisy conditions. In this work, we introduce Dual-diffusion, an innovative framework for the precise prediction of future hand poses. Our approach leverages the strengths of diffusion models by framing hand pose forecasting as a reverse diffusion process, effectively addressing the complexities of noisy hand joint movements and their subtleties. To enhance the learning of hand pose representations, we propose a unique neural architecture that simultaneously captures both local and global features. This is achieved through the deployment of Global and Local Diffusion (GLD) blocks within our network, which facilitate the exchange of information between local and global features. This diffusion-based interaction enables the integration of global hand actions and local finger actions, leading to a more powerful representation learning approach. We evaluate the effectiveness of our Dual-diffusion method on three public 3D hand pose estimation datasets (MSRA, F-PHAB, and BigHand2.2M), and our approach outperforms previous methods. Zhifeng Wang 0004, Kaihao Zhang, Ramesh S. Sankaranarayana |
Neurocomputing | 2 |
| 2025 | MB-TaylorFormer V2: Improved Multi-Branch Linear Transformer Expanded by Taylor Formula for Image RestorationabstractRecently, Transformer networks have demonstrated outstanding performance in the field of image restoration due to the global receptive field and adaptability to input. However, the quadratic computational complexity of Softmax-attention poses a significant limitation on its extensive application in image restoration tasks, particularly for high-resolution images. To tackle this challenge, we propose a novel variant of the Transformer. This variant leverages the Taylor expansion to approximate the Softmax-attention and utilizes the concept of norm-preserving mapping to approximate the remainder of the first-order Taylor expansion, resulting in a linear computational complexity. Moreover, we introduce a multi-branch architecture featuring multi-scale patch embedding into the proposed Transformer, which has four distinct advantages: 1) various sizes of the receptive field; 2) multi-level semantic information; 3) flexible shapes of the receptive field; 4) accelerated training and inference speed. Hence, the proposed model, named the second version of Taylor formula expansion-based Transformer (for short MB-TaylorFormer V2) has the capability to concurrently process coarse-to-fine features, capture long-distance pixel interactions with limited computational cost, and improve the approximation of the Taylor expansion remainder. Experimental results across diverse image restoration benchmarks demonstrate that MB-TaylorFormer V2 achieves state-of-the-art performance in multiple image restoration tasks, such as image dehazing, deraining, desnowing, motion deblurring, and denoising, with very little computational overhead. Zhi Jin 0002, Yuwei Qiu, Kaihao Zhang, Hongdong Li, Wenhan Luo |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Visual style prompt learning using diffusion models for blind face restoration
Wanglong Lu, Tao Wang 0052, Kaihao Zhang, Xianta Jiang, Hanli Zhao |
Pattern Recognit. | 4 |
| 2025 | LLDiffusion: Learning degradation representations in diffusion models for low-light image enhancement
Tao Wang 0052, Kaihao Zhang, Yong Zhang 0034, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005 |
Pattern Recognit. | 2 |
| 2025 | Visual and Textual Prompts in VLLMs for Enhancing Emotion RecognitionabstractVision Large Language Models (VLLMs) exhibit promising potential for multi-modal understanding, yet their application to video-based emotion recognition remains limited by insufficient spatial and contextual awareness. Traditional approaches, which prioritize isolated facial features, often neglect critical non-verbal cues such as body language, environmental context, and social interactions, leading to reduced robustness in real-world scenarios. To address this gap, we propose Set-of-Vision-Text Prompting (SoVTP), a novel framework that enhances zero-shot emotion recognition by integrating spatial annotations (e.g., bounding boxes, facial landmarks), physiological signals (facial action units), and contextual cues (body posture, scene dynamics, others’ emotions) into a unified prompting strategy. SoVTP preserves holistic scene information while enabling fine-grained analysis of facial muscle movements and interpersonal dynamics. Extensive experiments show that SoVTP achieves substantial improvements over existing visual prompting methods, demonstrating its effectiveness in enhancing VLLMs’ video emotion recognition capabilities. Zhifeng Wang 0004, Qixuan Zhang, Peter Zhang, Wenjia Niu, Kaihao Zhang, Ramesh S. Sankaranarayana, Sabrina B. Caldwell, Tom Gedeon |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | All-in-One Weather-Degraded Image Restoration Via Adaptive Degradation-Aware Self-Prompting ModelabstractExisting approaches for all-in-one weather-degraded image restoration suffer from inefficiencies in leveraging degradation-aware priors, resulting in sub-optimal performance in adapting to different weather conditions. To this end, we develop an adaptive degradation-aware self-prompting model (ADSM) for all-in-one weather-degraded image restoration. Specifically, our model employs the contrastive language-image pre-training model (CLIP) to facilitate the training of our proposed latent prompt generators (LPGs), which represent three types of latent prompts to characterize the degradation type, degradation property and image caption. Moreover, we integrate the acquired degradation-aware prompts into the time embedding of diffusion model to improve degradation perception. Meanwhile, we employ the latent caption prompt to guide the reverse sampling process using the cross-attention mechanism, thereby guiding the accurate image reconstruction. Furthermore, to accelerate the reverse sampling procedure of diffusion model and address the limitations of frequency perception, we introduce a wavelet-oriented noise estimating network (WNE-Net). Extensive experiments conducted on eight publicly available datasets demonstrate the effectiveness of our proposed approach in both task-specific and all-in-one applications. Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Kaihao Zhang, Ting Chen 0003 |
IEEE Trans. Multim. | 5 |
| 2025 | Multiprior Learning Via Neural Architecture Search for Blind Face RestorationabstractBlind face restoration (BFR) aims to recover high-quality (HQ) face images from low-quality (LQ) ones and usually resorts to facial priors for improving restoration performance. However, current methods still suffer from two major difficulties: 1) how to derive a powerful network architecture without extensive hand tuning and 2) how to capture complementary information from multiple facial priors in one network to improve restoration performance. To this end, we propose a face restoration searching network (FRSNet) to adaptively search the suitable feature extraction architecture within our specified search space, which can directly contribute to the restoration quality. On the basis of FRSNet, we further design our multiple facial prior searching network (MFPSNet) with a multiprior learning scheme. MFPSNet optimally extracts information from diverse facial priors and fuses the information into image features, ensuring that both external guidance and internal features are reserved. In this way, MFPSNet takes full advantage of semantic-level (parsing maps), geometric-level (facial heat maps), reference-level (facial dictionaries), and pixel-level (degraded images) information and, thus, generates faithful and realistic images. Quantitative and qualitative experiments show that the MFPSNet performs favorably on both synthetic and real-world datasets against the state-of-the-art (SOTA) BFR methods. The codes are publicly available at: https://github.com/YYJ1anG/MFPSNet. Yanjiang Yu, Puyang Zhang, Kaihao Zhang, Wenhan Luo |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Designing Unobtrusive Modulated Electrotactile Feedback on Fingertip Edge to Assist Blind and Low Vision (BLV) People in Comprehending ChartsabstractCharts are crucial in conveying information across various fields but are inaccessible to blind and low vision (BLV) people without assistive technology. Chart comprehension tools leveraging haptic feedback have been used widely but are often bulky, expensive, and static, rendering them inefficient for conveying chart data. To increase device portability, enable multitasking, and provide efficient assistance in chart comprehension, we introduce a novel system that delivers unobtrusive modulated electrotactile feedback directly to the fingertip edge. Our three-part study with twelve participants confirmed the effectiveness of this system, demonstrating that electrotactile feedback, when applied for 0.5 seconds with a 0.12-second interval, provides the most accurate position and direction recognition. Furthermore, our electrotactile device has proven valuable in assisting BLV participants in comprehending four commonly used charts: line charts, scatterplots, bar charts, and pie charts. We also delve into the implications of our findings on recognition enhancement, presentation modes, and function synergy. Chutian Jiang, Yinan Fan, Junan Xie, Emily Kuang, Kaihao Zhang, Mingming Fan 0001 |
CHI | 5 |
| 2024 | View from Above: Orthogonal-View Aware Cross-View LocalizationabstractThis paper presents a novel aerial-to-ground feature ag-gregation strategy, tailored for the task of cross- view image-based geo-localization. Conventional vision-based methods heavily rely on matching ground-view image features with a pre-recorded image database, often through establishing planar homography correspondences via a planar ground assumption. As such, they tend to ignore features that are off-ground and not suited for handling visual occlusions, leading to unreliable localization in challenging scenarios. We propose a Top-to-Ground Aggregation (T2GA) module that capitalizes aerial orthographic views to aggregate features down to the ground level, leveraging reliable off-ground information to improve feature alignment. Furthermore, we introduce a Cycle Domain Adaptation (CycDA) loss that ensures feature extraction robustness across do-main changes. Additionally, an Equidistant Re-projection (ERP) loss is introduced to equalize the impact of all key-points on orientation error, leading to a more extended distribution of keypoints which benefits orientation estimation. On both KITTI and Ford Multi-AV datasets, our method consistently achieves the lowest mean longitudinal and lateral translations across different settings and obtains the smallest orientation error when the initial pose is less ac-curate, a more challenging setting. Further, it can complete an entire route through continual vehicle pose estimation with initial vehicle pose given only at the starting point.11Code is available at https://github.com/ShanWang-Shan/View FromAbove. Shan Wang 0010, Jiawei Liu 0005, Yanhao Zhang 0003, Sundaram Muthu, Fahira A. Maken, Kaihao Zhang, Hongdong Li |
CVPR | 7 |
| 2024 | OMG: Occlusion-Friendly Personalized Multi-concept Generation in Diffusion Models
Zhe Kong, Yong Zhang 0034, Tianyu Yang 0003, Tao Wang 0052, Kaihao Zhang, Bizhu Wu, Guanying Chen, Wei Liu 0005, Wenhan Luo |
ECCV (31) | 5 |
| 2024 | Prompting Future Driven Diffusion Model for Hand Motion Prediction
Bowen Tang 0002, Kaihao Zhang, Wenhan Luo, Wei Liu 0005, Hongdong Li |
ECCV (7) | 2 |
| 2024 | Lrdif: Diffusion Models For Under-Display Camera Emotion RecognitionabstractThis study introduces LRDif, a novel diffusion-based framework designed specifically for facial expression recognition (FER) within the context of under-display cameras (UDC). To address the inherent challenges posed by UDC’s image degradation, such as reduced sharpness and increased noise, LRDif employs a two-stage training strategy that integrates a condensed preliminary extraction network (FPEN) and an agile transformer network (UDCformer) to effectively identify emotion labels from UDC images. By harnessing the robust distribution mapping capabilities of Diffusion Models (DMs) and the spatial dependency modeling strength of transformers, LRDif effectively overcomes the obstacles of noise and distortion inherent in UDC environments. Comprehensive experiments on standard FER datasets including RAFDB, KDEF, and FERPlus, LRDif demonstrate state-of-the-art performance, underscoring its potential in advancing FER applications. This work not only addresses a significant gap in the literature by tackling the UDC challenge in FER but also sets a new benchmark for future research in the field. Zhifeng Wang 0004, Kaihao Zhang, Ramesh S. Sankaranarayana |
ICIP | 2 |
| 2024 | LLDif: Diffusion Models for Low-Light Facial Expression Recognition
Zhifeng Wang 0004, Kaihao Zhang, Ramesh S. Sankaranarayana |
ICPR (13) | 2 |
| 2024 | Blind Face Video Restoration with Temporal Consistent Generative Prior and Degradation-Aware PromptabstractWithin the domain of blind face restoration (BFR), approaches lacking facial priors frequently result in excessively smoothed visual outputs. Exiting BFR methods predominantly utilize generative facial priors to achieve realistic and authentic details. However, these methods, primarily designed for images, encounter challenges in maintaining temporal consistency when applied to face video restoration. To tackle this issue, we introduce StableBFVR, an innovative Blind Face Video Restoration method based on Stable Diffusion that incorporates temporal information into the generative prior. This is achieved through the introduction of temporal layers in the diffusion process. These temporal layers consider both long-term and short-term information aggregation. Moreover, to improve generalizability, BFR methods employ complex, large-scale degradation during training, but it often sacrifices accuracy. Addressing this, StableBFVR features a novel mixed-degradation-aware prompt module, capable of encoding specific degradation information to dynamically steer the restoration process. Comprehensive experiments demonstrate that our proposed StableBFVR outperforms state-of-the-art methods. Jingfan Tan, Hyunhee Park, Tao Wang 0052, Kaihao Zhang, Pengwen Dai, Zikun Liu 0001, Wenhan Luo |
ACM Multimedia | 5 |
| 2024 | Fast Ultra High-Definition Video Deblurring via Multi-scale Separable Network
Wenqi Ren, Senyou Deng, Kaihao Zhang, Fenglong Song, Xiaochun Cao, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | GridFormer: Residual Dense Transformer with Grid Structure for Image Restoration in Adverse Weather Conditions
Tao Wang 0052, Kaihao Zhang, Ziqian Shao, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005, Hongdong Li |
Int. J. Comput. Vis. | 2 |
| 2024 | HTNet for micro-expression recognitionabstractFacial expression is related to facial muscle contractions and different muscle movements correspond to different emotional states. For micro-expression recognition, the muscle movements are usually subtle, which has a negative impact on the performance of current facial emotion recognition algorithms. Most existing methods use self-attention mechanisms to capture relationships between tokens in a sequence, but they do not take into account the inherent spatial relationships between facial landmarks. This can result in sub-optimal performance on micro-expression recognition tasks. Therefore, learning to recognize facial muscle movements is a key challenge in the area of micro-expression recognition. In this paper, we propose a Hierarchical Transformer Network (HTNet) to identify critical areas of facial muscle movement. HTNet includes two major components: a transformer layer that leverages the local temporal features and an aggregation layer that extracts local and global semantical facial features. Specifically, HTNet divides the face into four different facial areas: left lip area, left eye area, right eye area and right lip area. The transformer layer is used to focus on representing local minor muscle movement with local self-attention in each area. The aggregation layer is used to learn the interactions between eye areas and lip areas. The experiments on four publicly available micro-expression datasets show that the proposed approach outperforms previous methods by a large margin. The codes and models are available at: https://github.com/wangzhifengharrison/HTNet. Zhifeng Wang 0004, Kaihao Zhang, Wenhan Luo, Ramesh S. Sankaranarayana |
Neurocomputing | 2 |
| 2024 | Blind face restoration: Benchmark datasets and a baseline model
Puyang Zhang, Kaihao Zhang, Wenhan Luo, Guoren Wang |
Neurocomputing | 2 |
| 2024 | Aircraft type recognition in 3D-view optical image with contour segmentation
Zhixiang Liang, Yanshan Li, Rui Yu 0004, Kaihao Zhang |
Multim. Tools Appl. | 4 |
| 2024 | Restoring vision in hazy weather with hierarchical contrastive learning
Tao Wang 0052, Guangpin Tao, Wanglong Lu, Kaihao Zhang, Wenhan Luo, Xiaoqin Zhang 0002, Tong Lu 0002 |
Pattern Recognit. | 4 |
| 2024 | Restoring vision in rain-by-snow weather with simple attention-based sampling cross-hierarchy Transformer
Yuanbo Wen 0002, Tao Gao 0001, Kaihao Zhang, Peng Cheng 0002, Ting Chen 0003 |
Pattern Recognit. | 3 |
| 2024 | From heavy rain removal to detail restoration: A faster and better network
Yuanbo Wen 0002, Tao Gao 0001, Jing Zhang 0052, Kaihao Zhang, Ting Chen 0003 |
Pattern Recognit. | 4 |
| 2024 | Toward Real-World Blind Face Restoration With Generative Diffusion PriorabstractBlind face restoration is an important task in computer vision and has gained significant attention due to its wide-range applications. Previous works mainly exploit facial priors to restore face images and have demonstrated high-quality results. However, generating faithful facial details remains a challenging problem due to the limited prior knowledge obtained from finite data. In this work, we delve into the potential of leveraging the pretrained Stable Diffusion for blind face restoration. We propose BFRffusion which is thoughtfully designed to effectively extract features from low-quality face images and could restore realistic and faithful facial details with the generative prior of the pretrained Stable Diffusion. In addition, we build a privacy-preserving face dataset called PFHQ with balanced attributes like race, gender, and age. This dataset can serve as a viable alternative for training blind face restoration networks, effectively addressing privacy and bias concerns usually associated with the real face datasets. Through an extensive series of experiments, we demonstrate that our BFRffusion achieves state-of-the-art performance on both synthetic and real-world public testing datasets for blind face restoration and our PFHQ dataset is an available resource for training blind face restoration networks. The codes, pretrained models, and dataset are released at https://github.com/chenxx89/BFRffusion. Jingfan Tan, Tao Wang 0052, Kaihao Zhang, Wenhan Luo, Xiaochun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Frequency-Oriented Efficient Transformer for All-in-One Weather-Degraded Image RestorationabstractAdverse weather conditions, such as rain, raindrop, snow and haze, consistently degrade images in an unpredictable manner, thereby rendering existing task-specific and task-aligned methods inadequate in addressing this formidable problem. To this end, we investigate the application of Transformer in image restoration and introduce an efficient frequency-oriented method called AIRFormer, which is designed to restore weather-degraded images comprehensively and holistically. Specifically, we identify that the initial self-attention mechanism exhibits distinctive properties akin to a low-pass filter. Therefore, we construct a frequency-guided Transformer encoder by incorporating wavelet-based prior information to guide the extraction of image features. Additionally, considering the non-specific frequency characteristics of self-attention in the later stages, we develop a frequency-refined Transformer decoder that incorporates learnable task-specific queries across spatial dimensions, channel dimensions, and wavelet domains. To facilitate the training of our proposed method, we curate a comprehensive benchmark dataset named AIR40K that, encompasses a wide range of challenging scenarios. Extensive experimental evaluations demonstrate the superiority of our AIRFormer over both task-aligned and all-in-one methods across 15 publicly available datasets. Notably, AIRFormer achieves the best trade-off between the inference time and quality of reconstructed image, comparing with existing methods such as TransWeather and Restormer. The source code, dataset and pre-trained models will be available at https://github.com/chdwyb/AIRFormer. Tao Gao 0001, Yuanbo Wen 0002, Kaihao Zhang, Jing Zhang 0052, Ting Chen 0003, Lidong Liu, Wenhan Luo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Dual Teacher Knowledge Distillation With Domain Alignment for Face Anti-SpoofingabstractFace recognition systems have raised concerns due to their vulnerability to different presentation attacks, and system security has become an increasingly critical concern. Although many face anti-spoofing (FAS) methods perform well in intra-dataset scenarios, their generalization remains a challenge. To address this issue, some methods adopt domain adversarial training (DAT) to extract domain-invariant features. Differently, in this paper, we propose a domain adversarial attack (DAA) method by adding perturbations to the input images, which makes them indistinguishable across domains and enables domain alignment. Moreover, since models trained on limited data and types of attacks cannot generalize well to unknown attacks, we propose a dual perceptual and generative knowledge distillation framework for face anti-spoofing that utilizes pre-trained face-related models containing rich face priors. Specifically, we adopt two different face-related models as teachers to transfer knowledge to the target student model. The pre-trained teacher models are not from the task of face anti-spoofing but from perceptual and generative tasks, respectively, which implicitly augment the data. By combining both DAA and dual-teacher knowledge distillation, we develop a dual teacher knowledge distillation with domain alignment framework (DTDA) for face anti-spoofing. The advantage of our proposed method has been verified through extensive ablation studies and comparison with state-of-the-art methods on public datasets across multiple protocols. Zhe Kong, Wentian Zhang, Tao Wang 0052, Kaihao Zhang, Yuexiang Li, Xiaoying Tang 0001, Wenhan Luo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Blind Face Restoration for Under-Display Camera via Dictionary Guided TransformerabstractBy hiding the front-facing camera below the display panel, Under-Display Camera (UDC) provides users with a full-screen experience. However, due to the characteristics of the display, images taken by UDC suffer from significant quality degradation. Methods have been proposed to tackle UDC image restoration and advances have been achieved. There are still no specialized methods and datasets for restoring UDC face images, which may be the most common problem in the UDC scene. To this end, considering color filtering, brightness attenuation, and diffraction in the imaging process of UDC, we propose a two-stage network UDC Degradation Model Network named UDC-DMNet to synthesize UDC images by modeling the processes of UDC imaging. Then we use UDC-DMNet and high-quality face images from FFHQ and CelebA-Test to create UDC face training datasets FFHQ-P/T and testing datasets CelebA-Test-P/T for UDC face restoration. We propose a novel dictionary-guided transformer network named DGFormer. Introducing the facial component dictionary and the characteristics of the UDC image in the restoration makes DGFormer capable of addressing blind face restoration in UDC scenarios. Experiments show that our DGFormer and UDC-DMNet achieve state-of-the-art performance. Jingfan Tan, Tao Wang 0052, Kaihao Zhang, Wenhan Luo, Xiaochun Cao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | MC-Blur: A Comprehensive Benchmark for Image DeblurringabstractBlur artifacts can seriously degrade the visual quality of images, and numerous deblurring methods have been proposed for specific scenarios. However, in most real-world images, blur is caused by different factors, e.g., motion, and defocus. In this paper, we address how other deblurring methods perform in the case of multiple types of blur. For in-depth performance evaluation, we construct a new large-scale multi-cause image deblurring dataset (MC-Blur), including real-world and synthesized blurry images with different blur factors. The images in the proposed MC-Blur dataset are collected using other techniques: averaging sharp images captured by a 1000-fps high-speed camera, convolving Ultra-High-Definition (UHD) sharp images with large-size kernels, adding defocus to images, and real-world blurry images captured by various camera models. Based on the MC-Blur dataset, we conduct extensive benchmarking studies to compare SOTA methods in different scenarios, analyze their efficiency, and investigate the buildataset’s capacity. These benchmarking results provide a comprehensive overview of the advantages and limitations of current deblurring methods, revealing our dataset’s advances. The dataset is available to the public athttps://github.com/HDCVLab/MC-Blur-Dataset. Kaihao Zhang, Tao Wang 0052, Wenhan Luo, Wenqi Ren, Björn Stenger, Wei Liu 0005, Hongdong Li, Ming-Hsuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Ultra-High-Definition Low-Light Image Enhancement: A Benchmark and Transformer-Based MethodabstractAs the quality of optical sensors improves, there is a need for processing large-scale images. In particular, the ability of devices to capture ultra-high definition (UHD) images and video places new demands on the image processing pipeline. In this paper, we consider the task of low-light image enhancement (LLIE) and introduce a large-scale database consisting of images at 4K and 8K resolution. We conduct systematic benchmarking studies and provide a comparison of current LLIE algorithms. As a second contribution, we introduce LLFormer, a transformer-based low-light enhancement method. The core components of LLFormer are the axis-based multi-head self-attention and cross-layer attention fusion block, which significantly reduces the linear complexity. Extensive experiments on the new dataset and existing public datasets show that LLFormer outperforms state-of-the-art methods. We also show that employing existing LLIE methods trained on our benchmark as a pre-processing step significantly improves the performance of downstream tasks, e.g., face detection in low-light conditions. The source code and pre-trained models are available at https://github.com/TaoWangzj/LLFormer. Tao Wang 0052, Kaihao Zhang, Tianrun Shen, Wenhan Luo, Björn Stenger, Tong Lu 0002 |
AAAI | 2 |
| 2023 | Robust Single Image Reflection Removal Against Adversarial AttacksabstractThis paper addresses the problem of robust deep single-image reflection removal (SIRR) against adversarial attacks. Current deep learning based SIRR methods have shown significant performance degradation due to unnoticeable distortions and perturbations on input images. For a comprehensive robustness study, we first conduct diverse adversarial attacks specifically for the SIRR problem, i.e. towards different attacking targets and regions. Then we propose a robust SIRR model, which integrates the cross-scale attention module, the multi-scale fusion module, and the adversarial image discriminator. By exploiting the multi-scale mechanism, the model narrows the gap between features from clean and adversarial images. The image discriminator adaptively distinguishes clean or noisy inputs, and thus further gains reliable robustness. Extensive experiments on Nature, SIR2, and Real datasets demonstrate that our model remarkably improves the robustness of SIRR across disparate scenes. Zhenbo Song, Zhenyuan Zhang 0001, Kaihao Zhang, Wenhan Luo, Zhaoxin Fan, Wenqi Ren, Jianfeng Lu 0003 |
CVPR | 3 |
| 2023 | Model Calibration in Dense Classification with Adaptive Label PerturbationabstractFor safety-related applications, it is crucial to produce trustworthy deep neural networks whose prediction is associated with confidence that can represent the likelihood of correctness for subsequent decision-making. Existing dense binary classification models are prone to being over-confident. To improve model calibration, we propose Adaptive Stochastic Label Perturbation (ASLP) which learns a unique label perturbation level for each training image. ASLP employs our proposed Self-Calibrating Binary Cross Entropy (SC-BCE) loss, which unifies label perturbation processes including stochastic approaches (like DisturbLabel), and label smoothing, to correct calibration while maintaining classification rates. ASLP follows Maximum Entropy Inference of classic statistical mechanics to maximise prediction entropy with respect to missing information. It performs this while: (1) preserving classification accuracy on known data as a conservative solution, or (2) specifically improves model calibration degree by minimising the gap between the prediction accuracy and expected confidence of the target training label. Extensive results demonstrate that ASLP can significantly improve calibration degrees of dense binary classification models on both in-distribution and out-of-distribution data. The code is available on https://github.com/Carlisle-Liu/ASLP. Jiawei Liu 0005, Changkun Ye, Shan Wang 0010, Ruikai Cui, Jing Zhang 0052, Kaihao Zhang, Nick Barnes |
ICCV | 6 |
| 2023 | MB-TaylorFormer: Multi-branch Efficient Transformer Expanded by Taylor Formula for Image DehazingabstractIn recent years, Transformer networks are beginning to replace pure convolutional neural networks (CNNs) in the field of computer vision due to their global receptive field and adaptability to input. However, the quadratic computational complexity of softmax-attention limits the wide application in image dehazing task, especially for high-resolution images. To address this issue, we propose a new Transformer variant, which applies the Taylor expansion to approximate the softmax-attention and achieves linear computational complexity. A multi-scale attention refinement module is proposed as a complement to correct the error of the Taylor expansion. Furthermore, we introduce a multi-branch architecture with multi-scale patch embedding to the proposed Transformer, which embeds features by overlapping deformable convolution of different scales. The design of multi-scale patch embedding is based on three key ideas: 1) various sizes of the receptive field; 2) multi-level semantic information; 3) flexible shapes of the receptive field. Our model, named Multi-branch Transformer expanded by Taylor formula (MB-TaylorFormer), can em-bed coarse to fine features more flexibly at the patch embedding stage and capture long-distance pixel interactions with limited computational cost. Experimental results on several dehazing benchmarks show that MB-TaylorFormer achieves state-of-the-art (SOTA) performance with a light computational burden. The source code and pre-trained models are available at https://github.com/FVL2020/ICCV-2023-MB-TaylorFormer. Yuwei Qiu, Kaihao Zhang, Wenhan Luo, Hongdong Li, Zhi Jin 0002 |
ICCV | 2 |
| 2023 | Homography Guided Temporal Fusion for Road Line and Marking SegmentationabstractReliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance and overall high appearance consistency. To solve these issues, we propose a Homography Guided Fusion (HomoFusion) module to exploit temporally-adjacent video frames for complementary cues facilitating the correct classification of the partially occluded road lines or markings. To reduce computational complexity, a novel surface normal estimator is proposed to establish spatial correspondences between the sampled frames, allowing the HomoFusion module to perform a pixel-to-pixel attention mechanism in updating the representation of the occluded road lines or markings. Experiments on ApolloScape, a large-scale lane mark segmentation dataset, and ApolloScape Night with artificial simulated night-time road conditions, demonstrate that our method outperforms other existing SOTA lane mark segmentation models with less than 9% of their parameters and computational complexity. We show that exploiting available camera intrinsic data and ground plane assumption for cross-frame correspondence can lead to a light-weight network with significantly improved performances in speed and accuracy. We also prove the versatility of our HomoFusion approach by applying it to the problem of water puddle segmentation and achieving SOTA performance1. Shan Wang 0010, Jiawei Liu 0005, Kaihao Zhang, Wenhan Luo, Yanhao Zhang 0003, Sundaram Muthu, Fahira A. Maken, Hongdong Li |
ICCV | 4 |
| 2023 | F&F Attack: Adversarial Attack against Multiple Object Trackers by Inducing False Negatives and False PositivesabstractMulti-object tracking (MOT) aims to build moving trajectories for number-agnostic objects. Modern multi-object trackers commonly follow the tracking-by-detection strategy. Therefore, fooling detectors can be an effective solution but it usually requires attacks in multiple successive frames, resulting in low efficiency. Attacking association processes improves efficiency but may require model-specific design, leading to poor generalization. In this paper, we propose a novel False negative and False positive attack (F&F attack) mechanism: it perturbs the input image to erase original detections and to inject deceptive false alarms around original ones while integrating the association attack implicitly. The mechanism can produce effective identity switches against multi-object trackers by only fooling detectors in a few frames. To demonstrate the flexibility of the mechanism, we deploy it to three multi-object trackers (ByteTrack, SORT, and CenterTrack) which are enabled by two representative detectors (YOLOX and CenterNet). Comprehensive experiments on MOT17 and MOT20 datasets show that our method significantly outperforms existing attackers, revealing the vulnerability of the tracking-by-detection paradigm to detection attacks. Qi Ye 0001, Wenhan Luo, Kaihao Zhang, Zhiguo Shi 0001, Jiming Chen 0001 |
ICCV | 4 |
| 2023 | InterTracker: Discovering and Tracking General Objects Interacting with Hands in the WildabstractUnderstanding human interaction with objects is an important research topic for embodied Artificial Intelligence and identifying the objects that humans are interacting with is a primary problem for interaction understanding. Existing methods rely on frame-based detectors to locate interacting objects. However, this approach is subjected to heavy occlusions, background clutter, and distracting objects. To address the limitations, in this paper, we propose to leverage spatio-temporal information of hand-object interaction to track interactive objects under these challenging cases. Without prior knowledge of the general objects to be tracked like object tracking problems, we first utilize the spatial relation between hands and objects to adaptively discover the interacting objects from the scene. Second, the consistency and continuity of the appearance of objects between successive frames are exploited to track the objects. With this tracking formulation, our method also benefits from training on large-scale general object-tracking datasets. We further curate a video-level hand-object interaction dataset for testing and evaluation from 100DOH. The quantitative results demonstrate that our proposed method outperforms the state-of-the-art methods. Specifically, in scenes with continuous interaction with different objects, we achieve an impressive improvement of about 10% as evaluated using the Average Precision (AP) metric. Our qualitative findings also illustrate that our method can produce more continuous trajectories for interacting objects. Yanyan Shao, Qi Ye 0001, Wenhan Luo, Kaihao Zhang, Jiming Chen 0001 |
IROS | 4 |
| 2023 | Punctuation-level Attack: Single-shot and Single Punctuation Can Fool Text ModelsabstractThe adversarial attacks have attracted increasing attention in various fields including natural language processing. The current textual attacking models primarily focus on fooling models by adding character-/word-/sentence-level perturbations, ignoring their influence on human perception. In this paper, for the first time in the community, we propose a novel mode of textual attack, punctuation-level attack. With various types of perturbations, including insertion, displacement, deletion, and replacement, the punctuation-level attack achieves promising fooling rates against SOTA models on typical textual tasks and maintains minimal influence on human perception and understanding of the text by mere perturbation of single-shot single punctuation. Furthermore, we propose a search method named Text Position Punctuation Embedding and Paraphrase (TPPEP) to accelerate the pursuit of optimal position to deploy the attack, without exhaustive search, and we present a mathematical interpretation of TPPEP. Thanks to the integrated Text Position Punctuation Embedding (TPPE), the punctuation attack can be applied at a constant cost of time. Experimental results on public datasets and SOTA models demonstrate the effectiveness of the punctuation attack and the proposed TPPE. We additionally apply the single punctuation attack to summarization, semantic-similarity-scoring, and text-to-image tasks, and achieve encouraging results. Chongyang Du, Tao Wang 0052, Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wei Liu 0005, Xiaochun Cao |
NeurIPS | 4 |
| 2023 | Vicinity Vision TransformerabstractVision transformers have shown great success on numerous computer vision tasks. However, their central component, softmax attention, prohibits vision transformers from scaling up to high-resolution images, due to both the computational complexity and memory footprint being quadratic. Linear attention was introduced in natural language processing (NLP) which reorders the self-attention mechanism to mitigate a similar issue, but directly applying existing linear attention to vision may not lead to satisfactory results. We investigate this problem and point out that existing linear attention methods ignore an inductive bias in vision tasks, i.e., 2D locality. In this article, we propose Vicinity Attention, which is a type of linear attention that integrates 2D locality. Specifically, for each image patch, we adjust its attention weight based on its 2D Manhattan distance from its neighbouring patches. In this case, we achieve 2D locality in a linear complexity where the neighbouring image patches receive stronger attention than far away patches. In addition, we propose a novel Vicinity Attention Block that is comprised of Feature Reduction Attention (FRA) and Feature Preserving Connection (FPC) in order to address the computational bottleneck of linear attention approaches, including our Vicinity Attention, whose complexity grows quadratically with respect to the feature dimension. The Vicinity Attention Block computes attention in a compressed feature space with an extra skip connection to retrieve the original feature distribution. We experimentally validate that the block further reduces computation without degenerating the accuracy. Finally, to validate the proposed methods, we build a linear vision transformer backbone named Vicinity Vision Transformer (VVT). Targeting general vision tasks, we build VVT in a pyramid structure with progressively reduced sequence length. We perform extensive experiments on CIFAR-100, ImageNet-1 k, and ADE20 K datasets to validate the effectiveness of our method. Our method has a slower growth rate in terms of computational overhead than previous transformer-based and convolution-based networks when the input resolution increases. In particular, our approach achieves state-of-the-art image classification accuracy with 50% fewer parameters than previous approaches. Weixuan Sun, Zhen Qin 0003, Yi Zhang 0137, Kaihao Zhang, Nick Barnes, Stanley T. Birchfield, Lingpeng Kong, Yiran Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | EDFace-Celeb-1 M: Benchmarking Face Hallucination With a Million-Scale DatasetabstractRecent deep face hallucination methods show stunning performance in super-resolving severely degraded facial images, even surpassing human ability. However, these algorithms are mainly evaluated on non-public synthetic datasets. It is thus unclear how these algorithms perform on public face hallucination datasets. Meanwhile, most of the existing datasets do not well consider the distribution of races, which makes face hallucination methods trained on these datasets biased toward some specific races. To address the above two problems, in this paper, we build a public Ethnically Diverse Face dataset, EDFace-Celeb-1 M, and design a benchmark task for face hallucination. Our dataset includes 1.7 million photos that cover different countries, with relatively balanced race composition. To the best of our knowledge, it is the largest-scale and publicly available face hallucination dataset in the wild. Associated with this dataset, this paper also contributes various evaluation protocols and provides comprehensive analysis to benchmark the existing state-of-the-art methods. The benchmark evaluations demonstrate the performance and limitations of state-of-the-art algorithms. https://github.com/HDCVLab/EDFace-Celeb-1M. Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Jingyu Liu 0004, Jiankang Deng, Wei Liu 0005, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Enhanced Spatio-Temporal Interaction Learning for Video Deraining: Faster and BetterabstractVideo deraining is an important task in computer vision as the unwanted rain hampers the visibility of videos and deteriorates the robustness of most outdoor vision systems. Despite the significant success which has been achieved for video deraining recently, two major challenges remain: 1) how to exploit the vast information among successive frames to extract powerful spatio-temporal features across both the spatial and temporal domains, and 2) how to restore high-quality derained videos with a high-speed approach. In this paper, we present a new end-to-end video deraining framework, dubbed Enhanced Spatio-Temporal Interaction Network (ESTINet), which considerably boosts current state-of-the-art video deraining quality and speed. The ESTINet takes the advantage of deep residual networks and convolutional long short-term memory, which can capture the spatial features and temporal correlations among successive frames at the cost of very little computational resource. Extensive experiments on three public datasets show that the proposed ESTINet can achieve faster speed than the competitors, while maintaining superior performance over the state-of-the-art methods. https://github.com/HDCVLab/Enhanced-Spatio-Temporal-Interaction-Learning-for-Video-Deraining. Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Wenqi Ren, Wei Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Event-Aware Video Deraining via Multi-Patch Progressive LearningabstractIn this paper, we address the problem of video-based rain streak removal by developing an event-aware multi-patch progressive neural network. Rain streaks in video exhibit correlations in both temporal and spatial dimensions. Existing methods have difficulties in modeling the characteristics. Based on the observation, we propose to develop a module encoding events from neuromorphic cameras to facilitate deraining. Events are captured asynchronously at pixel-level only when intensity changes by a margin exceeding a certain threshold. Due to this property, events contain considerable information about moving objects including rain streaks passing though the camera across adjacent frames. Thus we suggest that utilizing it properly facilitates deraining performance non-trivially. In addition, we develop a multi-patch progressive neural network. The multi-patch manner enables various receptive fields by partitioning patches and the progressive learning in different patch levels makes the model emphasize each patch level to a different extent. Extensive experiments show that our method guided by events outperforms the state-of-the-art methods by a large margin in synthetic and real-world datasets. Shangquan Sun, Wenqi Ren, Jingzhi Li 0002, Kaihao Zhang, Meiyu Liang, Xiaochun Cao |
IEEE Trans. Image Process. | 4 |
| 2023 | T-Net: Deep Stacked Scale-Iteration Network for Image DehazingabstractHaze reduces the visibility of image content and leads to failure in handling subsequent computer vision tasks. In this paper, we address the problem of single image dehazing by proposing a dehazing network named T-Net, which consists of a backbone network based on the U-Net architecture and a dual attention module. Multi-scale feature fusion can be achieved by using skip connections with a new fusion strategy. Furthermore, by repeatedly unfolding the plain T-Net, Stack T-Net is proposed to take advantage of the dependence of deep features across stages via a recursive strategy. To reduce network parameters, the intra-stage recursive computation of ResNet is adopted in our Stack T-Net. We take both the stage-wise result and the original hazy image as input to each T-Net and finally output the prediction of the clean image. Experimental results on both synthetic and real-world images demonstrate that our plain T-Net and the advanced Stack T-Net perform favorably against state-of-the-art dehazing algorithms and show that our Stack T-Net could further improve the dehazing effect, demonstrating the effectiveness of the recursive strategy. Lirong Zheng 0003, Yanshan Li, Kaihao Zhang, Wenhan Luo |
IEEE Trans. Multim. | 3 |
| 2022 | Beyond Monocular Deraining: Parallel Stereo Deraining Network Via Semantic Prior
Kaihao Zhang, Wenhan Luo, Yanjiang Yu, Wenqi Ren, Fang Zhao 0006, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
Int. J. Comput. Vis. | 1 |
| 2022 | Deep Image Deblurring: A Survey
Kaihao Zhang, Wenqi Ren, Wenhan Luo, Wei-Sheng Lai, Björn Stenger, Ming-Hsuan Yang 0001, Hongdong Li |
Int. J. Comput. Vis. | 1 |
| 2022 | Displacement-Invariant Cost Computation for Stereo MatchingabstractAbstract Although deep learning-based methods have dominated stereo matching leaderboards by yielding unprecedented disparity accuracy, their inference time is typically slow, i.e., less than 4 FPS for a pair of 540p images. The main reason is that the leading methods employ time-consuming 3D convolutions applied to a 4D feature volume. A common way to speed up the computation is to downsample the feature volume, but this loses high-frequency details. To overcome these challenges, we propose a displacement-invariant cost computation module to compute the matching costs without needing a 4D feature volume. Rather, costs are computed by applying the same 2D convolution network on each disparity-shifted feature map pair independently. Unlike previous 2D convolution-based methods that simply perform context mapping between inputs and disparity maps, our proposed approach learns to match features between the two images. We also propose an entropy-based refinement strategy to refine the computed disparity map, which further improves the speed by avoiding the need to compute a second disparity map on the right image. Extensive experiments on standard datasets (SceneFlow, KITTI, ETH3D, and Middlebury) demonstrate that our method achieves competitive accuracy with much less inference time. On typical image sizes (e.g., $$540\times 960$$ 540 × 960 ), our method processes over 100 FPS on a desktop GPU, making our method suitable for time-critical applications such as autonomous driving. We also show that our approach generalizes well to unseen datasets, outperforming 4D-volumetric methods. We will release the source code to ensure the reproducibility. Yiran Zhong, Charles T. Loop, Wonmin Byeon, Stanley T. Birchfield, Yuchao Dai, Kaihao Zhang, Alexey Kamenev, Thomas M. Breuel, Hongdong Li, Jan Kautz |
Int. J. Comput. Vis. | 6 |
| 2022 | Four-player GroupGAN for weak expression recognition via latent expression magnification
Wenjia Niu, Kaihao Zhang, Dongxu Li 0003, Wenhan Luo |
Knowl. Based Syst. | 2 |
| 2022 | Disentangled Feature Networks for Facial Portrait and Caricature GenerationabstractFacial portrait is an artistic form which draws faces by emphasizing discriminative or prominent parts of faces via various kinds of drawing tools. However, the complex interplay between the different facial factors, such as facial parts, background, and drawing styles, and the significant domain gap between natural facial images and their portrait counterparts makes the task challenging. In this paper, a flexible four-stream Disentangled Feature Networks (DFN) is proposed to learn disentangled feature representation of different facial factors and generate plausible portraits with reasonable exaggerations and richness in style. Four factors are encoded as embedding features, and combined to reconstruct facial portraits. Meanwhile, to make the process fully automatic (without manually specifying either portrait style or exaggerating form), we propose a new Adversarial Portrait Mapping Module (APMM) to map noise to the embedding feature space, as proxies for portrait style and exaggerating. Thanks to the proposedDFNandAPMM, we are able to manipulate the portrait style and facial geometric structures to generate a large number of portraits. Extensive experiments on two public datasets show that our proposed methods can generate a diverse set of artistic portraits. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wenqi Ren, Hongdong Li |
IEEE Trans. Multim. | 1 |
| 2021 | ARVo: Learning All-Range Volumetric Correspondence for Video DeblurringabstractVideo deblurring models exploit consecutive frames to remove blurs from camera shakes and object motions. In order to utilize neighboring sharp patches, typical methods rely mainly on homography or optical flows to spatially align neighboring blurry frames. However, such explicit approaches are less effective in the presence of fast motions with large pixel displacements. In this work, we propose a novel implicit method to learn spatial correspondence among blurry frames in the feature space. To construct distant pixel correspondences, our model builds a correlation volume pyramid among all the pixel-pairs between neigh-boring frames. To enhance the features of the reference frame, we design a correlative aggregation module that maximizes the pixel-pair correlations with its neighbors based on the volume pyramid. Finally, we feed the aggregated features into a reconstruction module to obtain the restored frame. We design a generative adversarial paradigm to optimize the model progressively. Our proposed method is evaluated on the widely-adopted DVD dataset, along with a newly collected High-Frame-Rate (1000 fps) Dataset for Video Deblurring (HFR-DVD). Quantitative and qualitative experiments show that our model performs favorably on both datasets against previous state-of-the-art methods, confirming the benefit of modeling all-range spatial correspondence for video deblurring. Dongxu Li 0003, Kaihao Zhang, Xin Yu 0002, Yiran Zhong, Wenqi Ren, Hanna Suominen, Hongdong Li |
CVPR | 3 |
| 2021 | Deep Two-View Structure-From-Motion RevisitedabstractTwo-view structure-from-motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM. Existing deep learning-based approaches formulate the problem by either recovering absolute pose scales from two consecutive frames or predicting a depth map from a single image, both of which are ill-posed problems. In contrast, we propose to revisit the problem of deep two-view SfM by leveraging the well-posedness of the classic pipeline. Our method consists of 1) an optical flow estimation network that predicts dense correspondences between two frames; 2) a normalized pose estimation module that computes relative camera poses from the 2D optical flow correspondences, and 3) a scale-invariant depth estimation network that leverages epipolar geometry to reduce the search space, refine the dense correspondences, and estimate relative depth maps. Extensive experiments show that our method outperforms all state-of-the-art two-view SfM methods by a clear margin on KITTI depth, KITTI VO, MVS, Scenes11, and SUN3D datasets in both relative pose and depth estimation. Yiran Zhong, Yuchao Dai, Stanley T. Birchfield, Kaihao Zhang, Nikolai Smolyanskiy, Hongdong Li |
CVPR | 5 |
| 2021 | Pyramid Architecture Search for Real-Time Image DeblurringabstractMulti-scale and multi-patch deep models have been shown effective in removing blurs of dynamic scenes. However, these methods still suffer from one major obstacle: manually designing a lightweight and high-efficiency network is challenging and time-consuming. To tackle this obstacle, we propose a novel deblurring method, dubbed PyNAS (pyramid neural architecture search network), towards automatically designing hyper-parameters including the scales, patches, and standard cell operators. The proposed PyNAS adopts gradient-based search strategies and innovatively searches the hierarchy patch and scale scheme not limited to cell searching. Specifically, we introduce a hierarchical search strategy tailored to the multi-scale and multi-patch deblurring task. The strategy follows the principle that the first distinguishes between the top-level (pyramid-scales and pyramid-patches) and bottom-level variables (cell operators) and then searches multi-scale variables using the top-to-bottom principle. During the search stage, PyNAS employs an early stopping strategy to avoid the collapse and computational issues. Furthermore, we use a path-level binarization mechanism for multi-scale cell searching to save the memory consumption. Our primary contribution is a real-time deblurring algorithm (around 58 fps) for 720p images while achieves state-of-the-art deblurring performance on the GoPro and Video Deblurring datasets. Xiaobin Hu, Wenqi Ren, Kaicheng Yu, Kaihao Zhang, Xiaochun Cao, Wei Liu 0005, Bjoern Menze |
ICCV | 4 |
| 2021 | Benchmarking Ultra-High-Definition Image Super-resolutionabstractIncreasingly, modern mobile devices allow capturing images at Ultra-High-Definition (UHD) resolution, which includes 4K and 8K images. However, current single image super-resolution (SISR) methods focus on super-resolving images to ones with resolution up to high definition (HD) and ignore higher-resolution UHD images. To explore their performance on UHD images, in this paper, we first introduce two large-scale image datasets, UHDSR4K and UHDSR8K, to benchmark existing SISR methods. With 70,000 V100 GPU hours of training, we benchmark these methods on 4K and 8K resolution images under seven different settings to provide a set of baseline models. Moreover, we propose a baseline model, called Mesh Attention Network (MANet) for SISR. The MANet applies the attention mechanism in both different depths (horizontal) and different levels of receptive field (vertical). In this way, correlations among feature maps are learned, enabling the network to focus on more important features. Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Wenqi Ren, Björn Stenger, Wei Liu 0005, Hongdong Li, Ming-Hsuan Yang 0001 |
ICCV | 1 |
| 2021 | Deep robust image deblurring via blur distilling and information comparison in latent space
Wenjia Niu, Kaihao Zhang, Wenhan Luo, Yiran Zhong, Hongdong Li |
Neurocomputing | 2 |
| 2021 | Blind Motion Deblurring Super-Resolution: When Dynamic Spatio-Temporal Learning Meets Static Image UnderstandingabstractSingle-image super-resolution (SR) and multi-frame SR are two ways to super resolve low-resolution images. Single-Image SR generally handles each image independently, but ignores the temporal information implied in continuing frames. Multi-frame SR is able to model the temporal dependency via capturing motion information. However, it relies on neighbouring frames which are not always available in the real world. Meanwhile, slight camera shake easily causes heavy motion blur on long-distance-shot low-resolution images. To address these problems, a Blind Motion Deblurring Super-Reslution Networks, BMDSRNet, is proposed to learn dynamic spatio-temporal information from single static motion-blurred images. Motion-blurred images are the accumulation over time during the exposure of cameras, while the proposed BMDSRNet learns the reverse process and uses three-streams to learn Bidirectional spatio-temporal information based on well designed reconstruction loss functions to recover clean high-resolution images. Extensive experiments demonstrate that the proposed BMDSRNet outperforms recent state-of-the-art methods, and has the ability to simultaneously deal with image deblurring and SR. Wenjia Niu, Kaihao Zhang, Wenhan Luo, Yiran Zhong |
IEEE Trans. Image Process. | 2 |
| 2021 | Angular-Driven Feedback Restoration Networks for Imperfect Sketch RecognitionabstractAutomatic hand-drawn sketch recognition is an important task in computer vision. However, the vast majority of prior works focus on exploring the power of deep learning to achieve better accuracy on complete and clean sketch images, and thus fail to achieve satisfactory performance when applied to incomplete or destroyed sketch images. To address this problem, we first develop two datasets that contain different levels of scrawl and incomplete sketches. Then, we propose an angular-driven feedback restoration network (ADFRNet), which first detects the imperfect parts of a sketch and then refines them into high quality images, to boost the performance of sketch recognition. By introducing a novel "feedback restoration loop" to deliver information between the middle stages, the proposed model can improve the quality of generated sketch images while avoiding the extra memory cost associated with popular cascading generation schemes. In addition, we also employ a novel angular-based loss function to guide the refinement of sketch images and learn a powerful discriminator in the angular space. Extensive experiments conducted on the proposed imperfect sketch datasets demonstrate that the proposed model is able to efficiently improve the quality of sketch images and achieve superior performance over the current state-of-the-art methods. Jia Wan 0001, Kaihao Zhang, Hongdong Li, Antoni B. Chan |
IEEE Trans. Image Process. | 2 |
| 2021 | Dual Attention-in-Attention Model for Joint Rain Streak and Raindrop RemovalabstractRain streaks and raindrops are two natural phenomena, which degrade image capture in different ways. Currently, most existing deep deraining networks take them as two distinct problems and individually address one, and thus cannot deal adequately with both simultaneously. To address this, we propose a Dual Attention-in-Attention Model (DAiAM) which includes two DAMs for removing both rain streaks and raindrops. Inside the DAM, there are two attentive maps - each of which attends to the heavy and light rainy regions, respectively, to guide the deraining process differently for applicable regions. In addition, to further refine the result, a Differential-driven Dual Attention-in-Attention Model (D-DAiAM) is proposed with a "heavy-to-light" scheme to remove rain via addressing the unsatisfying deraining regions. Extensive experiments on one public raindrop dataset, one public rain streak and our synthesized joint rain streak and raindrop (JRSRD) dataset have demonstrated that the proposed method not only is capable of removing rain streaks and raindrops simultaneously, but also achieves the state-of-the-art performance on both tasks. Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Wenqi Ren |
IEEE Trans. Image Process. | 1 |
| 2021 | Deep Dense Multi-Scale Network for Snow Removal Using Semantic and Depth PriorsabstractImages captured in snowy days suffer from noticeable degradation of scene visibility, which degenerates the performance of current vision-based intelligent systems. Removing snow from images thus is an important topic in computer vision. In this paper, we propose a Deep Dense Multi-Scale Network (DDMSNet) for snow removal by exploiting semantic and depth priors. As images captured in outdoor often share similar scenes and their visibility varies with depth from camera, such semantic and depth information provides a strong prior for snowy image restoration. We incorporate the semantic and depth maps as input and learn the semantic-aware and geometry-aware representation to remove snow. In particular, we first create a coarse network to remove snow from the input images. Then, the coarsely desnowed images are fed into another network to obtain the semantic and depth labels. Finally, we design a DDMSNet to learn semantic-aware and geometry-aware representation via a self-attention mechanism to produce the final clean images. Experiments evaluated on public synthetic and real-world snowy images verify the superiority of the proposed method, offering better results both quantitatively and qualitatively. https://github.com/HDCVLab/Deep-Dense-Multi-scale-Network https://github.com/HDCVLab/Deep-Dense-Multi-scale-Network. Kaihao Zhang, Rongqing Li, Yanjiang Yu, Wenhan Luo |
IEEE Trans. Image Process. | 1 |
| 2020 | Deblurring by Realistic BlurringabstractExisting deep learning methods for image deblurring typically train models using pairs of sharp images and their blurred counterparts. However, synthetically blurring images does not necessarily model the blurring process in real-world scenarios with sufficient accuracy. To address this problem, we propose a new method which combines two GAN models, i.e., a learning-to-Blur GAN (BGAN) and learning-to-DeBlur GAN (DBGAN), in order to learn a better model for image deblurring by primarily learning how to blur images. The first model, BGAN, learns how to blur sharp images with unpaired sharp and blurry image sets, and then guides the second model, DBGAN, to learn how to correctly deblur such images. In order to reduce the discrepancy between real blur and synthesized blur, a relativistic blur loss is leveraged. As an additional contribution, this paper also introduces a Real-World Blurred Image (RWBI) dataset including diverse blurry images. Our experiments show that the proposed method achieves consistently superior quantitative performance as well as higher perceptual quality on both the newly proposed dataset and the public GOPRO dataset. Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma 0002, Björn Stenger, Wei Liu 0005, Hongdong Li |
CVPR | 1 |
| 2020 | Single Image Super-Resolution via a Holistic Attention Network
Weilei Wen, Wenqi Ren, Xiangde Zhang, Lianping Yang, Shuzhen Wang, Kaihao Zhang, Xiaochun Cao, Haifeng Shen |
ECCV (12) | 7 |
| 2020 | Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding
Kaihao Zhang, Wenhan Luo, Wenqi Ren, Jingwen Wang 0003, Fang Zhao 0006, Lin Ma 0002, Hongdong Li |
ECCV (27) | 1 |
| 2020 | Unsupervised Domain Adaptation with Noise Resistible Mutual-Training for Person Re-identification
Fang Zhao 0006, Shengcai Liao, Guosen Xie, Jian Zhao 0006, Kaihao Zhang, Ling Shao 0001 |
ECCV (11) | 5 |
| 2020 | Every Moment Matters: Detail-Aware Networks to Bring a Blurry Image AliveabstractMotion-blurred images are the result of light accumulation over the period of camera exposure time, during which the camera and objects in the scene are in relative motion to each other. The inverse process of extracting an image sequence from a single motion-blurred image is an ill-posed vision problem. One key challenge is that the motions across frames are subtle, which makes the generating networks difficult to capture them and thus the recovery sequences lack motion details. In order to alleviate this problem, we propose a detail-aware network with three consecutive stages to improve the reconstruction quality by addressing specific aspects in the recovery process. The detail-aware network firstly models the dynamics using a cycle flow loss, resolving the temporal ambiguity of the reconstruction in the first stage. Then, a GramNet is proposed in the second stage to refine subtle motion between continuous frames using Gram matrices as motion representation. Finally, we introduce a HeptaGAN in the third stage to bridge the continuous and discrete nature of exposure time and recovered frames, respectively, in order to maintain rich detail. Experiments show that the proposed detail-aware networks produce sharp image sequences with rich details and subtle motion, outperforming the state-of-the-art methods. Kaihao Zhang, Wenhan Luo, Björn Stenger, Wenqi Ren, Lin Ma 0002, Hongdong Li |
ACM Multimedia | 1 |
| 2020 | TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationabstractSign language translation (SLT) aims to interpret sign video sequences into text-based natural language sentences. Sign videos consist of continuous sequences of sign gestures with no clear boundaries in between. Existing SLT models usually represent sign visual features in a frame-wise manner so as to avoid needing to explicitly segmenting the videos into isolated signs. However, these methods neglect the temporal information of signs and lead to substantial ambiguity in translation. In this paper, we explore the temporal semantic structures of sign videos to learn more discriminative features. To this end, we first present a novel sign video segment representation which takes into account multiple temporal granularities, thus alleviating the need for accurate video segmentation. Taking advantage of the proposed segment representation, we develop a novel hierarchical sign video feature learning method via a temporal semantic pyramid network, called TSPNet. Specifically, TSPNet introduces an inter-scale attention to evaluate and enhance local semantic consistency of sign segments and an intra-scale attention to resolve semantic ambiguity by using non-local video context. Experiments show that our TSPNet outperforms the state-of-the-art with significant improvements on the BLEU score (from 9.58 to 13.41) and ROUGE score (from 31.80 to 34.96) on the largest commonly used SLT dataset. Our implementation is available at https://github.com/verashira/TSPNet. Dongxu Li 0003, Xin Yu 0002, Kaihao Zhang, Ben Swift, Hanna Suominen, Hongdong Li |
NeurIPS | 4 |
| 2020 | Displacement-Invariant Matching Cost Learning for Accurate Optical Flow EstimationabstractLearning matching costs has been shown to be critical to the success of the state-of-the-art deep stereo matching methods, in which 3D convolutions are applied on a 4D feature volume to learn a 3D cost volume. However, this mechanism has never been employed for the optical flow task. This is mainly due to the significantly increased search dimension in the case of optical flow computation, \ie, a straightforward extension would require dense 4D convolutions in order to process a 5D feature volume, which is computationally prohibitive. This paper proposes a novel solution that is able to bypass the requirement of building a 5D feature volume while still allowing the network to learn suitable matching costs from data. Our key innovation is to decouple the connection between 2D displacements and learn the matching costs at each 2D displacement hypothesis independently, \ie, displacement-invariant cost learning. Specifically, we apply the same 2D convolution-based matching net independently on each 2D displacement hypothesis to learn a 4D cost volume. Moreover, we propose a displacement-aware projection layer to scale the learned cost volume, which reconsiders the correlation between different displacement candidates and mitigates the multi-modal problem in the learned cost volume. The cost volume is then projected to optical flow estimation through a 2D soft-argmin layer. Extensive experiments show that our approach achieves state-of-the-art accuracy on various datasets, and outperforms all published optical flow methods on the Sintel benchmark. The code is available at https://github.com/jytime/DICL-Flow. Yiran Zhong, Yuchao Dai, Kaihao Zhang, Pan Ji, Hongdong Li |
NeurIPS | 4 |
| 2020 | Human Parsing Based Texture Transfer from Single Image to 3D Human via Cross-View ConsistencyabstractThis paper proposes a human parsing based texture transfer model via cross-view consistency learning to generate the texture of 3D human body from a single image. We use the semantic parsing of human body as input for providing both the shape and pose information to reduce the appearance variation of human image and preserve the spatial distribution of semantic parts. Meanwhile, in order to improve the prediction for textures of invisible parts, we explicitly enforce the consistency across different views of the same subject by exchanging the textures predicted by two views to render images during training. The perception loss and total variation regularization are optimized to maximize the similarity between rendered and input images, which does not necessitate extra 3D texture supervision. Experimental results on pedestrian images and fashion photos demonstrate that our method can produce higher quality textures with convincing details than other texture generation methods. Fang Zhao 0006, Shengcai Liao, Kaihao Zhang, Ling Shao 0001 |
NeurIPS | 3 |
| 2019 | Cousin Network Guided Sketch Recognition via Latent Attribute WarehouseabstractWe study the problem of sketch image recognition. This problem is plagued with two major challenges: 1) sketch images are often scarce in contrast to the abundance of natural images, rendering the training task difficult, and 2) the significant domain gap between sketch image and its natural image counterpart makes the task of bridging the two domains challenging. In order to overcome these challenges, in this paper we propose to transfer the knowledge of a network learned from natural images to a sketch network - a new deep net architecture which we term as cousin network. This network guides a sketch-recognition network to extract more relevant features that are close to those of natural images, via adversarial training. Moreover, to enhance the transfer ability of the classification model, a sketch-to-image attribute warehouse is constructed to approximate the transformation between the sketch domain and the real image domain. Extensive experiments conducted on the TU-Berlin dataset show that the proposed model is able to efficiently distill knowledge from natural images and achieves superior performance than the current state of the art. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Hongdong Li |
AAAI | 1 |
| 2019 | Learning Joint Gait Representation via Quintuplet Loss MinimizationabstractGait recognition is an important biometric method popularly used in video surveillance, where the task is to identify people at a distance by their walking patterns from video sequences. Most of the current successful approaches for gait recognition either use a pair of gait images to form a cross-gait representation or rely on a single gait image for unique-gait representation. These two types of representations emperically complement one another. In this paper, we propose a new Joint Unique-gait and Cross-gait Network (JUCNet), to combine the advantages of unique-gait representation with that of cross-gait representation, leading to an significantly improved performance. Another key contribution of this paper is a novel quintuplet loss function, which simultaneously increases the inter-class differences by pushing representations extracted from different subjects apart and decreases the intra-class variations by pulling representations extracted from the same subject together. Experiments show that our method achieves the state-of-the-art performance tested on standard benchmark datasets, demonstrating its superiority over existing methods. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
CVPR | 1 |
| 2019 | Adversarial Spatio-Temporal Learning for Video DeblurringabstractCamera shake or target movement often leads to undesired blur effects in videos captured by a hand-held camera. Despite significant efforts having been devoted to video-deblur research, two major challenges remain: 1) how to model the spatio-temporal characteristics across both the spatial domain (i.e., image plane) and the temporal domain (i.e., neighboring frames) and 2) how to restore sharp image details with respect to the conventionally adopted metric of pixel-wise errors. In this paper, to address the first challenge, we propose a deblurring network (DBLRNet) for spatial-temporal learning by applying a 3D convolution to both the spatial and temporal domains. Our DBLRNet is able to capture jointly spatial and temporal information encoded in neighboring frames, which directly contributes to the improved video deblur performance. To tackle the second challenge, we leverage the developed DBLRNet as a generator in the generative adversarial network (GAN) architecture and employ a content loss in addition to an adversarial loss for efficient adversarial training. The developed network, which we name as deblurring GAN, is tested on two standard benchmarks and achieves the state-of-the-art performance. Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
IEEE Trans. Image Process. | 1 |
| 2017 | Facial Expression Recognition Based on Deep Evolutional Spatial-Temporal NetworksabstractOne key challenging issue of facial expression recognition is to capture the dynamic variation of facial physical structure from videos. In this paper, we propose a part-based hierarchical bidirectional recurrent neural network (PHRNN) to analyze the facial expression information of temporal sequences. Our PHRNN models facial morphological variations and dynamical evolution of expressions, which is effective to extract "temporal features" based on facial landmarks (geometry information) from consecutive frames. Meanwhile, in order to complement the still appearance information, a multi-signal convolutional neural network (MSCNN) is proposed to extract "spatial features" from still frames. We use both recognition and verification signals as supervision to calculate different loss functions, which are helpful to increase the variations of different expressions and reduce the differences among identical expressions. This deep evolutional spatial-temporal network (composed of PHRNN and MSCNN) extracts the partial-whole, geometry-appearance, and dynamic-still information, effectively boosting the performance of facial expression recognition. Experimental results show that this method largely outperforms the state-of-the-art ones. On three widely used facial expression databases (CK+, Oulu-CASIA, and MMI), our method reduces the error rates of the previous best ones by 45.5%, 25.8%, and 24.4%, respectively. Kaihao Zhang, Yongzhen Huang, Liang Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Localize heavily occluded human faces via deep segmentationabstractLocalizing heavily occluded human faces is a challenging problem in facial detection. Previous methods mainly employ sliding windows by determining whether windows include human faces. In this paper, we provide a novel segmentation-based perspective for heavily occluded face localization with deep convolutional neural networks (CNN). Our model takes an image as input without complicated pre-processing. After several convolutional layers, fully-connected layers and a softmax classifier, we can predict the labels of all pixels in an image, which is the key to localize heavily occluded human faces. Finally, we search a minimal rectangle to localize the human face. Our detector needs neither complex pre-processing nor the time-consuming sliding window. Besides, we use a single model to localize faces to further alleviate computational complexity. Experimental results show that our proposed method is a very effective way to localize heavily occluded human face. Kaihao Zhang, Yongzhen Huang, Ran He 0001, Liang Wang 0001 |
ICIP | 1 |
| 2015 | Kinship Verification with Deep Convolutional Neural Networks
Kaihao Zhang, Yongzhen Huang, Chunfeng Song, Liang Wang 0001 |
BMVC | 1 |