VLDB 2026 Research / reviewers in the wild / expert
Hao Li 0058
dblp:17/5705-58
· DBLP profile ↗
18ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-7915-582XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable Artificial Intelligence for Deepfake Detection: Pipeline, Open Source and ComparisonsabstractABSTRACT Deepfake detection models achieve high accuracy, yet their interpretability remains underexplored. This study presents a unified evaluation pipeline for post hoc visual explanations, grounded in the Co‐12 attributes of explanation quality and operationalised through a structured framework that assesses Coherence, Composition, Correctness, Completeness, Compactness, Covariate completeness and Continuity. Applying this pipeline to 16 representative explanation methods reveals systematic differences across methodological categories. CAM‐based approaches demonstrate strong spatial coherence and temporal continuity; gradient‐based techniques such as Guided Backprop and LRP yield compact and accurate attributions; and redistribution‐based methods including ExcitationBP and Deep Taylor maintain high consistency across evaluation conditions. In contrast, perturbation‐based approaches such as SHAP and LIME exhibit weaker localisation and reduced temporal stability. By enabling controlled, attribute‐level comparison of explanation methods, the proposed pipeline bridges conceptual interpretability frameworks and empirical analysis, offering practical guidance for the development and deployment of interpretable deepfake detectors in forensic and auditing applications. The source code is publicly available at: https://github.com/junxinchenieee/EAI‐Deepfake‐Detection . Hao Li 0058, Junxin Chen 0001, Bo Wang 0024, David Camacho |
Expert Syst. J. Knowl. Eng. | 1 |
| 2026 | GlareLane: A real-world benchmark and method for lane detection under challenging illumination
Xiaojie Yu, Qiankun Li 0004, Hao Li 0058, Ben-Guo He, Qiang He 0002, Junxin Chen 0001 |
Image Vis. Comput. | 3 |
| 2026 | Collaborative Feedback Discriminative Propagation for Video Super-ResolutionabstractThe key success of existing video super-resolution (VSR) methods stems mainly from exploring spatial and temporal information that is usually achieved by a temporal propagation with alignment strategies. However, inaccurate alignment usually leads to significant artifacts that will be accumulated during propagation and thus affect video restoration. Moreover, only propagating the same timestep features forward or backward does not handle the videos with complex motion or occlusion. To address these issues, we propose a collaborative feedback discriminative (CFD) method to correct inaccurate aligned features and better model spatial and temporal information for VSR. Specifically, we first develop a discriminative alignment correction (DAC) method to reduce the influences of the artifacts caused by inaccurate alignment. Then, we propose a collaborative feedback propagation (CFP) module based on feedback and gating mechanisms to explore spatial and temporal information of different timestep features from forward and backward propagation simultaneously. Finally, we embed the proposed DAC and CFP into commonly used VSR networks to verify the effectiveness of our method. Experimental results demonstrate that our method improves the performance of existing VSR models while maintaining a lower model complexity. Hao Li 0058, Xiang Chen 0015, Jiangxin Dong, Jinhui Tang 0001, Jinshan Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | FoundIR: Unleashing Million-Scale Training Data to Advance Foundation Models for Image RestorationabstractDespite the significant progress made by all-in-one models in universal image restoration, existing methods suffer from a generalization bottleneck in real-world scenarios, as they are mostly trained on small-scale synthetic datasets with limited degradations. Therefore, large-scale high-quality real-world training data is urgently needed to facilitate the emergence of foundational models for image restoration. To advance this field, we spare no effort in contributing a million-scale dataset with two notable advantages over existing training data: real-world samples with larger-scale, and degradation types with higher diversity. By adjusting internal camera settings and external imaging conditions, we can capture aligned image pairs using our well-designed data acquisition system over multiple rounds and our data alignment criterion. Moreover, we propose a robust model, FoundIR, to better address a broader range of restoration tasks in real-world scenarios, taking a further step toward foundation models. Specifically, we first utilize a diffusion-based generalist model to remove degradations by learning the degradation-agnostic common representations from diverse inputs, where incremental learning strategy is adopted to better guide model training. To refine the model's restoration capability in complex scenarios, we introduce degradation-aware specialist models for achieving final high-quality results. Extensive experiments show the value of our dataset and the effectiveness of our method. Hao Li 0058, Xiang Chen 0015, Jiangxin Dong, Jinhui Tang 0001, Jinshan Pan |
ICCV | 1 |
| 2025 | Medical image translation with deep learning: Advances, datasets and perspectives
Junxin Chen 0001, Zhiheng Ye, Renlong Zhang, Hao Li 0058, Bo Fang 0005, Li-bo Zhang 0004, Wei Wang 0077 |
Medical Image Anal. | 4 |
| 2024 | Federated continual representation learning for evolutionary distributed intrusion detection in Industrial Internet of Things
Zhao Zhang 0001, Hao Li 0058, Shenbo Liu, Lijun Tang |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Hybrid CNN-Transformer Feature Fusion for Single Image DerainingabstractSince rain streaks exhibit diverse geometric appearances and irregular overlapped phenomena, these complex characteristics challenge the design of an effective single image deraining model. To this end, rich local-global information representations are increasingly indispensable for better satisfying rain removal. In this paper, we propose a lightweight Hybrid CNN-Transformer Feature Fusion Network (dubbed as HCT-FFN) in a stage-by-stage progressive manner, which can harmonize these two architectures to help image restoration by leveraging their individual learning strengths. Specifically, we stack a sequence of the degradation-aware mixture of experts (DaMoE) modules in the CNN-based stage, where appropriate local experts adaptively enable the model to emphasize spatially-varying rain distribution features. As for the Transformer-based stage, a background-aware vision Transformer (BaViT) module is employed to complement spatially-long feature dependencies of images, so as to achieve global texture recovery while preserving the required structure. Considering the indeterminate knowledge discrepancy among CNN features and Transformer features, we introduce an interactive fusion branch at adjacent stages to further facilitate the reconstruction of high-quality deraining results. Extensive evaluations show the effectiveness and extensibility of our developed HCT-FFN. The source code is available at https://github.com/cschenxiang/HCT-FFN. Xiang Chen 0015, Jinshan Pan, Jiyang Lu, Zhentao Fan, Hao Li 0058 |
AAAI | 5 |
| 2023 | Learning A Sparse Transformer Network for Effective Image DerainingabstractTransformers-based methods have achieved significant performance in image deraining as they can model the non-local information which is vital for high-quality image reconstruction. In this paper, we find that most existing Transformers usually use all similarities of the tokens from the query-key pairs for the feature aggregation. However, if the tokens from the query are different from those of the key, the self-attention values estimated from these tokens also involve in feature aggregation, which accordingly interferes with the clear image restoration. To overcome this problem, we propose an effective DeRaining network, Sparse Transformer (DRSformer) that can adaptively keep the most useful self-attention values for feature aggregation so that the aggregated features better facilitate high-quality image reconstruction. Specifically, we develop a learnable top-k selection operator to adaptively retain the most crucial attention scores from the keys for each query for better feature aggregation. Simultaneously, as the naive feed-forward network in Transformers does not model the multi-scale information that is important for latent clear image restoration, we develop an effective mixed-scale feed-forward network to generate better features for image deraining. To learn an enriched set of hybrid features, which combines local context from CNN operators, we equip our model with mixture of experts feature compensator to present a cooperation refinement deraining scheme. Extensive experimental results on the commonly used benchmarks demonstrate that the proposed method achieves favorable performance against state-of-the-art approaches. The source code and trained models are available at https://github.com/cschenxiang/DRSformer. Xiang Chen 0015, Hao Li 0058, Mingqiang Li, Jinshan Pan |
CVPR | 2 |
| 2023 | A Closer Look at Classifier in Adversarial Domain GeneralizationabstractThe task of domain generalization is to learn a classification model from multiple source domains and generalize it to unknown target domains. The key to domain generalization is learning discriminative domain-invariant features. Invariant representations are achieved using adversarial domain generalization as one of the primary techniques. For example, generative adversarial networks have been widely used, but suffer from the problem of low intra-class diversity, which can lead to poor generalization ability. To address this issue, we propose a new method called auxiliary classifier in adversarial domain generalization (CloCls). CloCls improve the diversity of the source domain by introducing auxiliary classifier. Combining typical task-related losses, e.g., cross-entropy loss for classification and adversarial loss for domain discrimination, our overall goal is to guarantee the learning of condition-invariant features for all source domains while increasing the diversity of source domains. Further, inspired by smoothing optima have improved generalization for supervised learning tasks like classification. We leverage that converging to a smooth minima with respect task loss stabilizes the adversarial training leading to better performance on unseen target domain which can effectively enhances the performance of domain adversarial methods. We have conducted extensive image classification experiments on benchmark datasets in domain generalization, and our model exhibits sufficient generalization ability and outperforms state-of-the-art DG methods. Ye Wang 0023, Junyang Chen 0001, Mengzhu Wang, Hao Li 0058, Wei Wang 0335, Houcheng Su, Zhihui Lai 0001, Wei Wang 0077, Zhenghan Chen |
ACM Multimedia | 4 |
| 2023 | Real-World Image Super-Resolution by Exclusionary Dual-LearningabstractReal-world image super-resolution is a practical image restoration problem that aims to obtain high-quality images from in-the-wild input, has recently received considerable attention with regard to its tremendous application potentials. Although deep learning-based methods have achieved promising restoration quality on real-world image super-resolution datasets, they ignore the relationship between L1- and perceptual- minimization and roughly adopt auxiliary large-scale datasets for pre-training. In this paper, we discuss the image types within a corrupted image and the property of perceptual- and Euclidean- based evaluation protocols. Then we propose a method, Real-World image Super-Resolution by Exclusionary Dual-Learning (RWSR-EDL) to address the feature diversity in perceptual- and L1- based cooperative learning. Moreover, a noise-guidance data collection strategy is developed to address the training time consumption in multiple datasets optimization. When an auxiliary dataset is incorporated, RWSR-EDL achieves promising results and repulses any training time increment by adopting the noise-guidance data collection strategy. Extensive experiments show that RWSR-EDL achieves competitive performance over state-of-the-art methods on four in-the-wild image super-resolution datasets. Hao Li 0058, Jinghui Qin, Zhijing Yang, Pengxu Wei, Jinshan Pan, Liang Lin 0004, Yukai Shi |
IEEE Trans. Multim. | 1 |
| 2023 | OccluMix: Towards De-Occlusion Virtual Try-on by Semantically-Guided MixupabstractImage Virtual try-on aims at replacing the cloth on a personal image with a garment image (in-shop clothes), which has attracted increasing attention from the multimedia and computer vision communities. Prior methods successfully preserve the character of clothing images, however, occlusion remains a pernicious effect for realistic virtual try-on. In this work, we first present a comprehensive analysis of the occlusions and categorize them into two aspects: i) Inherent-Occlusion: the ghost of the former cloth still exists in the try-on image; ii) Acquired-Occlusion: the target cloth warps to the unreasonable body part. Based on the in-depth analysis, we find that the occlusions can be simulated by a novel semantically-guided mixup module, which can generate semantic-specific occluded images that work together with the try-on images to facilitate training a de-occlusion try-on (DOC-VTON) framework. Specifically, DOC-VTON first conducts a sharpened semantic parsing on the try-on person. Aided by semantics guidance and pose prior, various complexities of texture are selectively blending with human parts in a copy-and-paste manner. Then, the Generative Module (GM) is utilized to take charge of synthesizing the final try-on image and learning to de-occlusion jointly. In comparison to the state-of-the-art methods, DOC-VTON achieves better perceptual quality by reducing occlusion effects. Zhijing Yang, Junyang Chen 0002, Yukai Shi, Hao Li 0058, Tianshui Chen, Liang Lin 0004 |
IEEE Trans. Multim. | 4 |
| 2022 | Progressive representation recalibration for lightweight super-resolution
Ruimian Wen, Zhijing Yang, Tianshui Chen, Hao Li 0058 |
Neurocomputing | 4 |
| 2022 | DnSwin: Toward real-world denoising via a continuous Wavelet Sliding Transformer
Hao Li 0058, Zhijing Yang, Ziying Zhao, Junyang Chen 0001, Yukai Shi, Jinshan Pan |
Knowl. Based Syst. | 1 |
| 2022 | Criteria Comparative Learning for Real-Scene Image Super-ResolutionabstractReal-scene image super-resolution aims to restore real-world low-resolution images into their high-quality versions. A typical RealSR framework usually includes the optimization of multiple criteria which are designed for different image properties, by making the implicit assumption that the ground-truth images can provide a good trade-off between different criteria. However, this assumption could be easily violated in practice due to the inherent contrastive relationship between different image properties. Contrastive learning (CL) provides a promising recipe to relieve this problem by learning discriminative features using the triplet contrastive losses. Though CL has achieved significant success in many computer vision tasks, it is non-trivial to introduce CL to RealSR due to the difficulty in defining valid positive image pairs in this case. Inspired by the observation that the contrastive relationship could also exist between the criteria, in this work, we propose a novel training paradigm for RealSR, named Criteria Comparative Learning (Cria-CL), by developing contrastive losses defined on criteria instead of image patches. In addition, a spatial projector is proposed to obtain a good view for Cria-CL in RealSR. Our experiments demonstrate that compared with the typical weighted regression strategy, our method achieves a significant improvement under similar parameter settings. Yukai Shi, Hao Li 0058, Sen Zhang 0006, Zhijing Yang, Xiao Wang 0014 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Local Spatial Constraint and Total Variation for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection, which is aimed at locating anomaly, has received widespread attention. In this article, a new anomaly detector, named local spatial constraint and total variation (LSC-TV), is proposed for hyperspectral imagery. In anomaly detection methods based on low-rank representation, background pixels are usually considered to have a global low-dimensional structure. However, the complex background distribution in hyperspectral images (HSIs) means that this global low-dimensional structure rarely occurs. In LSC-TV, the effective local spatial information is extracted by superpixel segmentation, and the regularization based on the F-norm is used to force the background within the same superpixel to show uniform spectral features. Moreover, each pixel is given a penalty based on the degree of anomaly determined during model iteration, while the anomaly is not considered by the background constraint. In addition, the background pixels in the neighborhood often show a high correlation, whereas the anomaly does not possess this feature. Nonisotropic TV is introduced into the proposed LSC model using the correlation of first-order neighborhoods to make it easier for anomalies to be separated. The proposed LSC-TV method and current state-of-the-art methods are tested on a set of simulated data and four sets of real data. The experimental results demonstrate that the proposed method is superior to the comparative method in terms of both color map detection and quantitative evaluation. Ruyi Feng, Hao Li 0058, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Tieyong Zeng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Low-Rank Representation Incorporating Local Spatial Constraint for Hyperspectral Anomaly DetectionabstractRecently, hyperspectral anomaly detection methods based on low-rank representation(LRR) have been widely studied. However, the assumption of global low dimension of background may ignore the local structure information of hyperspectral image. In this paper, a novel LRR incorporating local spatial constraint method is proposed for hyperspectral anomaly detection. Different from LRR detector, the proposed method considers the spatial information based on the supe pixel in the background part. The proposed method and current state-of-the-art methods are tested on two sets of real data. The experimental results demonstrate that the proposed method is superior to the comparative method in terms of both colour map detection and quantitative evaluation. Hao Li 0058, Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Lifei Wei |
IGARSS | 1 |
| 2021 | Superpixel-Based Reweighted Low-Rank and Total Variation Sparse Unmixing for Hyperspectral Remote Sensing ImageryabstractSparse unmixing, as a semisupervised unmixing method, has attracted extensive attention. The process of sparse unmixing involves treating the mixed pixels of hyperspectral imagery as a linear combination of a small number of spectral signatures (endmembers) in a standard spectral library, associated with fractional abundances. Over the past ten years, to achieve a better performance, sparse unmixing algorithms have begun to focus on the spatial information of hyperspectral images. However, less accurate spatial information greatly limits the performance of the spatial-regularization-based sparse unmixing algorithms. In this article, to overcome this limitation and obtain more reliable spatial information, a novel sparse unmixing algorithm named superpixel-based reweighted low-rank and total variation (SUSRLR-TV) is proposed to enhance the performance of the traditional spatial-regularization-based sparse unmixing approaches. In the proposed approach, superpixel segmentation is adopted to consider both the spatial proximity and the spectral similarity. In addition, a low-rank constraint is enforced on the objective function as pixels within each superpixel have the same endmembers and similar abundance values, and they naturally satisfy the low-rank constraint. Differing from the traditional nuclear norm, a reweighted nuclear norm is used to achieve a more efficient and accurate low-rank constraint. Meanwhile, low-rank consideration is also used to enhance the spatial continuity and suppress the effects of random noise. Furthermore, TV regularization is introduced to promote the smoothness of the abundance maps. Experiments on three simulated data sets, as well as a well-known real hyperspectral imagery data set, confirm the superior performance of the proposed method in both the qualitative assessment and the quantitative evaluation, compared with the state-of-the-art sparse unmixing methods. Hao Li 0058, Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Semi-Supervised Hyperspectral Unmixing with Very Deep Convolutional Neural NetworksabstractHyperspectral unmixing is an essential task in hyperspectral imagery applications. Deep learning methods have been taken into hyperspectral unmixing because of its great feature extraction ability and better performance. However, there are several problems in existing deep learning based spectral unmixing methods. The networks are not deep enough to exploit their feature extraction capabilities in these unsupervised autoencoders based methods, and their effects are not stable. The main reason may be the limited prior information limited the ability of conducting the supervised method. In this manuscript, a semi-supervised deep learning based unmixing method is proposed. Unlike the existing methods, our model uses deeper neural networks without pooling layers, and the endmember spectrum are selected supervised from the original data, which uses nature and nurture cooperatively. The experimental results show that the proposed method achieves better performance and produces more accurate abundance maps, as well as higher quantitative results, compared with the current state-of-the-art deep learning unmixing algorithms. Jiayu Bai, Ruyi Feng, Lizhe Wang 0001, Hao Li 0058, Fengpeng Li, Yanfei Zhong, Liangpei Zhang 0001 |
IGARSS | 4 |