EDBT 2026 Demo / reviewers in the wild / expert
Zhangkai Ni
dblp:185/7403
· DBLP profile ↗
43ranked-venue papers
15as first author
35since 2021 · last 2026
0000-0003-3682-6288ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 12 first-author · 23 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Contrastive Mean Teacher for Robust Low-Light Image Enhancement
Zhangkai Ni, Menglin Han, Wenhan Yang, Hanli Wang, Lin Ma 0002, Sam Kwong |
Int. J. Comput. Vis. | 1 |
| 2026 | A Consensus Resistance-Based Autonomous Vehicle Social Group Self-Adaption MethodabstractThe advancement of autonomous driving technology has brought significant benefits to modern transportation systems. However, individual autonomous vehicles face challenges such as limited perception range and insufficient autonomous capabilities. Cooperative groups of autonomous vehicles, enabled by advanced communication technologies, can enhance traffic efficiency through information exchange. Existing research primarily focuses on centralized autonomous vehicle groups, where the leading node suffers from weak resilience and high computational load, making it difficult to maintain group collaboration over time. To address these issues, this article proposes a decentralized formation and self-adaptation method for autonomous vehicle social groups based on consensus resistance in closed scenes. First, we introduceconsensus resistanceas a metric to evaluate social group and member consistency, and develop a decentralized formation approach. Second, we present a self-adaption model for autonomous vehicle social groups, incorporating four evolutionary events: 1) expansion; 2) merging; 3) reduction; and 4) splitting, to ensure the stability of moving social groups. Simulation results demonstrate the proposed method effectively constructs social groups in both real-world and simulated environments, exhibiting robust consistency throughout the self-adaption process. Jiujun Cheng, Lu Yang 0019, Zhangkai Ni, Guangtao Zhou, Zhenhua Huang 0001, Shangce Gao |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2026 | Integrated Perception, Communication, and Computation for Autonomous Vehicle and Road Infrastructure NetworkabstractVehicle-to-Infrastructure (V2I) collaboration constitutes an emerging paradigm for advancing autonomous driving. However, the integrated collaboration of perception, communication, and computation within V2I system remains a critical challenge. To address it, we propose a Software-Defined Network (SDN)-based collaborative approach for Autonomous Vehicle and Road Infrastructure Network (AVRIN). The architecture designates road infrastructures as road nodes and autonomous vehicles as dynamic vehicle nodes, establishing AVRIN through SDN. The control plane dynamically maintains global network topology and distributed flow tables by continuously evaluating node accessibility, while the forwarding plane is responsible for packet transmission via the OpenFlow protocol. In the perception module, road nodes divide the perception range into spatial units, whereas vehicle nodes dynamically align these units with their drivable areas across temporal sequences. Through coordinated communication and computation modules, road nodes strategically allocate dedicated bandwidth and computational resources. Building on this approach, we develop a particle swarm-based multi-objective optimization algorithm to achieve balanced co-optimization across perception, communication, and computation. Experimental validation demonstrates its superior collaborative Bird's Eye View (BEV) detection performance on the V2X-Sim 2.0 dataset, outperforming existing approaches by 10.37% in mean Average Precision. Furthermore, evaluations on the newly collected Jiading dataset, from a real-world urban roadway, confirm the approach's robustness with 1.823-second computation time under dynamic network conditions. Lu Yang 0019, Jiujun Cheng, MengChu Zhou, Cong Liu 0012, Zhangkai Ni, Mande Xie, Shangce Gao |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | AFUNet: Cross-Iterative Alignment-Fusion Synergy for HDR Reconstruction via Deep Unfolding ParadigmabstractExisting learning-based methods effectively reconstruct HDR images from multi-exposure LDR inputs with extended dynamic range and improved detail, but they rely more on empirical design rather than theoretical foundation, which can impact their reliability. To address these limitations, we propose the cross-iterative Alignment and Fusion deep Unfolding Network (AFUNet), where HDR reconstruction is systematically decoupled into two interleaved subtasks -- alignment and fusion -- optimized through alternating refinement, achieving synergy between the two subtasks to enhance the overall performance. Our method formulates multi-exposure HDR reconstruction from a Maximum A Posteriori (MAP) estimation perspective, explicitly incorporating spatial correspondence priors across LDR images and naturally bridging the alignment and fusion subproblems through joint constraints. Building on the mathematical foundation, we reimagine traditional iterative optimization through unfolding -- transforming the conventional solution process into an end-to-end trainable AFUNet with carefully designed modules that work progressively. Specifically, each iteration of AFUNet incorporates an Alignment-Fusion Module (AFM) that alternates between a Spatial Alignment Module (SAM) for alignment and a Channel Fusion Module (CFM) for adaptive feature fusion, progressively bridging misaligned content and exposure discrepancies. Extensive qualitative and quantitative evaluations demonstrate AFUNet's superior performance, consistently surpassing state-of-the-art methods. Our code is available at: https://github.com/eezkni/AFUNet Zhangkai Ni, Wenhan Yang |
ICCV | 2 |
| 2025 | D2AD: Diffusion Distillation for Unsupervised Image Anomaly DetectionabstractImage anomaly detection identifies images that deviate from normal patterns. Knowledge Distillation has been extensively studied in unsupervised anomaly detection, where it is assumed that the student model learns and reconstructs the multi-scale normal representations of the pre-trained teacher model, with the representation discrepancies between the student and teacher identified as anomalies. However, the over-generalization of the student model causes it to reconstruct anomalies, leading to the failure of anomaly detection. To address this issue, we propose a novel Diffusion Distillation method for Anomaly Detection (D2AD), which utilizes a lightweight diffusion pipeline to prevent student from imitating abnormal behavior. Specifically, we use the diffusion process to confuse the abnormal student representation with noisy normal representations and restore them to normal, effectively avoiding the over-generalization issue and enabling the student to efficiently transfer the teacher’s refined knowledge of normality. Moreover, we propose the Anomaly-Guided Diffusion Process (AGDP), which dynamically adjusts the noise intensity based on the likelihood of anomalies to balance anomaly confusion and representation recovery. Extensive experiments on the MVTec AD, VisA, and Real-IAD datasets show significant improvements in anomaly detection and localization, while maintaining high efficiency. Yuheng Shao, Zhangkai Ni, Qinyuan Liu |
ICME | 2 |
| 2025 | Perceptual-GS: Scene-adaptive Perceptual Densification for Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis. However, existing methods struggle to adaptively optimize the distribution of Gaussian primitives based on scene characteristics, making it challenging to balance reconstruction quality and efficiency. Inspired by human perception, we propose scene-adaptive perceptual densification for Gaussian Splatting (Perceptual-GS), a novel framework that integrates perceptual sensitivity into the 3DGS training process to address this challenge. We first introduce a perception-aware representation that models human visual sensitivity while constraining the number of Gaussian primitives. Building on this foundation, we develop a perceptual sensitivity-adaptive distribution to allocate finer Gaussian granularity to visually critical regions, enhancing reconstruction quality and robustness. Extensive evaluations on multiple datasets, including BungeeNeRF for large-scale scenes, demonstrate that Perceptual-GS achieves state-of-the-art performance in reconstruction quality, efficiency, and robustness. The code is publicly available at: https://github.com/eezkni/Perceptual-GS Hongbi Zhou, Zhangkai Ni |
ICML | 2 |
| 2025 | Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report GenerationabstractVision-grounded medical report generation aims to produce clinically accurate descriptions of medical images, anchored in explicit visual evidence to improve interpretability and facilitate integration into clinical workflows. However, existing methods often rely on separately trained detection modules that require extensive expert annotations, introducing high labeling costs and limiting generalizability due to pathology distribution bias across datasets. To address these challenges, we propose Self-Supervised Anatomical Consistency Learning (SS-ACL)-a novel and annotation-free framework that aligns generated reports with corresponding anatomical regions using simple textual prompts. SS-ACL constructs a hierarchical anatomical graph inspired by the invariant top-down inclusion structure of human anatomy, organizing entities by spatial location. It recursively reconstructs fine-grained anatomical regions to enforce intra-sample spatial alignment, inherently guiding attention maps toward visually relevant areas prompted by text. To further enhance inter-sample semantic alignment for abnormality recognition, SS-ACL introduces a region-level contrastive learning based on anatomical consistency. These aligned embeddings serve as priors for report generation, enabling attention maps to provide interpretable visual evidence. Extensive experiments demonstrate that SS-ACL, without relying on expert annotations, (i) generates accurate and visually grounded reports-outperforming state-of-the-art methods by 10% in lexical accuracy and 25% in clinical efficacy, and (ii) achieves competitive performance on various downstream visual tasks, surpassing current leading visual foundation models by 8% in zero-shot visual grounding. Our code is available at https://github.com/kaelsunkiller/ssacl. Longzhen Yang, Zhangkai Ni, Ying Wen 0003, Lianghua He, Heng Tao Shen |
ACM Multimedia | 2 |
| 2025 | Semantic Masking with Curriculum Learning for Robust HDR Image Reconstruction
Zhangkai Ni, Kerui Ren, Wenhan Yang, Hanli Wang, Sam Kwong |
Int. J. Comput. Vis. | 1 |
| 2025 | Rethinking Artifact Mitigation in HDR Reconstruction: From Detection to OptimizationabstractArtifact remains a long-standing challenge in High Dynamic Range (HDR) reconstruction. Existing methods focus on model designs for artifact mitigation but ignore explicit detection and suppression strategies. Because artifact lacks clear boundaries, distinct shapes, and semantic consistency, and there is no existing dedicated dataset for HDR artifact, progress in direct artifact detection and recovery is impeded. To bridge the gap, we propose a unified HDR reconstruction framework that integrates artifact detection and model optimization. Firstly, we build the first HDR artifact dataset (HADataset), comprising 1,213 diverse multi-exposure Low Dynamic Range (LDR) image sets and 1,765 HDR image pairs with per-pixel artifact annotations. Secondly, we develop an effective HDR artifact detector (HADetector), a robust artifact detection model capable of accurately localizing HDR reconstruction artifact. HADetector plays two pivotal roles: (1) enhancing existing HDR reconstruction models through fine-tuning, and (2) serving as a non-reference image quality assessment (NR-IQA) metric, the Artifact Score (AS), which aligns closely with human visual perception for reliable quality evaluation. Extensive experiments validate the effectiveness and generalizability of our framework, including the HADataset, HADetector, fine-tuning paradigm, and AS metric. The code and datasets are available at: https://github.com/xinyueliii/hdr-artifact-detect-optimize. Zhangkai Ni, Wenhan Yang, Hanli Wang, Lianghua He, Sam Kwong |
IEEE Trans. Image Process. | 2 |
| 2025 | Structural Similarity-Inspired Unfolding for Lightweight Image Super-ResolutionabstractMajor efforts in data-driven image super-resolution (SR) primarily focus on expanding the receptive field of the model to better capture contextual information. However, these methods are typically implemented by stacking deeper networks or leveraging transformer-based attention mechanisms, which consequently increases model complexity. In contrast, model-driven methods based on the unfolding paradigm show promise in improving performance while effectively maintaining model compactness through sophisticated module design. Based on these insights, we propose a Structural Similarity-Inspired Unfolding (SSIU) method for efficient image SR. This method is designed through unfolding an SR optimization function constrained by structural similarity, aiming to combine the strengths of both data-driven and model-driven approaches. Our model operates progressively following the unfolding paradigm. Each iteration consists of multiple Mixed-Scale Gating Modules (MSGM) and an Efficient Sparse Attention Module (ESAM). The former implements comprehensive constraints on features, including a structural similarity constraint, while the latter aims to achieve sparse activation. In addition, we design a Mixture-of-Experts-based Feature Selector (MoE-FS) that fully utilizes multi-level feature information by combining features from different steps. Extensive experiments validate the efficacy and efficiency of our unfolding-inspired network. Our model outperforms current state-of-the-art models, boasting lower parameter counts and reduced memory consumption. Our code will be available at: https://github.com/eezkni/SSIU. Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2025 | Breaking Boundaries: Unifying Imaging and Compression for HDR Image CompressionabstractHigh Dynamic Range (HDR) images present unique challenges for Learned Image Compression (LIC) due to their complex domain distribution compared to Low Dynamic Range (LDR) images. In coding practice, HDR-oriented LIC typically adopts preprocessing steps (e.g., perceptual quantization and tone mapping operation) to align the distributions between LDR and HDR images, which inevitably comes at the expense of perceptual quality. To address this challenge, we rethink the HDR imaging process which involves fusing multiple exposure LDR images to create an HDR image and propose a novel HDR image compression paradigm, Unifying Imaging and Compression (HDR-UIC). The key innovation lies in establishing a seamless pipeline from image capture to delivery and enabling end-to-end training and optimization. Specifically, a Mixture-ATtention (MAT)-based compression backbone merges LDR features while simultaneously generating a compact representation. Meanwhile, the Reference-guided Misalignment-aware feature Enhancement (RME) module mitigates ghosting artifacts caused by misalignment in the LDR branches, maintaining fidelity without introducing additional information. Furthermore, we introduce an Appearance Redundancy Removal (ARR) module to optimize coding resource allocation among LDR features, thereby enhancing the final HDR compression performance. Extensive experimental results demonstrate the efficacy of our approach, showing significant improvements over existing state-of-the-art HDR compression schemes. Our code is available at: https://github.com/plf1999/HDR-UIC. Xuelin Shen, Linfeng Pan, Zhangkai Ni, Yu-Lin He, Wenhan Yang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 3 |
| 2025 | Shell-Guided Compression of Voxel Radiance FieldsabstractIn this paper, we address the challenge of significant memory consumption and redundant components in large-scale voxel-based model, which are commonly encountered in real-world 3D reconstruction scenarios. We propose a novel method called Shell-guided compression of Voxel Radiance Fields (SVRF), aimed at optimizing voxel-based model into a shell-like structure to reduce storage costs while maintaining rendering accuracy. Specifically, we first introduce a Shell-like Constraint, operating in two main aspects: 1) enhancing the influence of voxels neighboring the surface in determining the rendering outcomes, and 2) expediting the elimination of redundant voxels both inside and outside the surface. Additionally, we introduce an Adaptive Thresholds to ensure appropriate pruning criteria for different scenes. To prevent the erroneous removal of essential object parts, we further employ a Dynamic Pruning Strategy to conduct smooth and precise model pruning during training. The compression method we propose does not necessitate the use of additional labels. It merely requires the guidance of self-supervised learning based on predicted depth. Furthermore, it can be seamlessly integrated into any voxel-grid-based method. Extensive experimental results demonstrate that our method achieves comparable rendering quality while compressing the original number of voxel grids by more than 70%. Our code will be available at: https://github.com/eezkni/SVRF. Peiqi Yang, Zhangkai Ni, Hanli Wang, Wenhan Yang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 2 |
| 2025 | M2Trans: Multi-Modal Regularized Coarse-to-Fine Transformer for Ultrasound Image Super-ResolutionabstractUltrasound image super-resolution (SR) aims to transform low-resolution images into high-resolution ones, thereby restoring intricate details crucial for improved diagnostic accuracy. However, prevailing methods relying solely on image modality guidance and pixel-wise loss functions struggle to capture the distinct characteristics of medical images, such as unique texture patterns and specific colors harboring critical diagnostic information. To overcome these challenges, this paper introduces the Multi-Modal Regularized Coarse-to-fine Transformer (M2Trans) for Ultrasound Image SR. By integrating the text modality, we establish joint image-text guidance during training, leveraging the medical CLIP model to incorporate richer priors from text descriptions into the SR optimization process, enhancing detail, structure, and semantic recovery. Furthermore, we propose a novel coarse-to-fine transformer comprising multiple branches infused with self-attention and frequency transforms to efficiently capture signal dependencies across different scales. Extensive experimental results demonstrate significant improvements over state-of-the-art methods on benchmark datasets, including CCA-US, US-CASE, and our newly created dataset MMUS1K, with a minimum improvement of 0.17dB, 0.30dB, and 0.28dB in terms of PSNR. Zhangkai Ni, Runyu Xiao, Wenhan Yang, Hanli Wang, Zhihua Wang 0002, Lihua Xiang |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Edge Computing-Based Contributed Perception and Autonomous Vehicle Groups in Open Scenes
Qichao Mao, Jiujun Cheng, MengChu Zhou, Zhangkai Ni, Shangce Gao, Chuanhuang Li |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Contributed Perception-Based Dynamic Evolution Method for Autonomous Vehicle Groups in Open Scenes
Qichao Mao, Jiujun Cheng, MengChu Zhou, Zhangkai Ni, Guiyuan Yuan, Shangce Gao, Chuanhuang Li |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Similarity Shuffled Criss-Cross Transformer With Angle Loss for Image-Text MatchingabstractImage-text matching aims to retrieve images from the guidance of textual queries or retrieve text expressions with the help of images. Existing Transformer-based methods compute attention for all tokens and thus suffer from redundant information, resulting in inadequate focus on salient features. On the other hand, the widely adopted bidirectional ranking loss overlooks the importance of expanding the distance between positive and negative samples, leading to the misclassification of negative samples as positive ones. In this work, we propose similarity shuffled criss-cross Transformer (SSCT) with angle loss for image-text matching. Specifically, a grouping-shuffling operation is introduced to better distinguish salient features from redundant information, bypassing the need for fully connected mapping. The grouping-shuffling operation establishes channel dependencies across different groups of feature representations, enhancing salient features while suppressing unimportant ones. Then, a criss-cross attention mechanism that equips self-attention with a novel criss-cross convolution is designed to make isolated information cooperatively express integral semantics. Moreover, a novel angle loss is introduced to expand the distances between positive and negative samples. Extensive experiments on the benchmark datasets of MSCOCO and Flickr30K demonstrate that the proposed methods achieve superior performances compared to state-of-the-art methods. Taiyi Su, Hanli Wang, Zhangkai Ni |
IEEE Trans. Multim. | 4 |
| 2025 | Vision-Language Relational Transformer for Video-to-Text GenerationabstractVideo-to-text generation is a challenging task that involves translating video contents into accurate and expressive sentences. Existing methods often ignore the importance of establishing fine-grained semantics within visual representations and exploring textual knowledge implied by video contents, leading to difficulty in generating satisfactory sentences. To address these problems, a vision-language relational transformer model is proposed for video-to-text generation. Three key novel aspects are investigated. First, a visual relation modeling block is designed to obtain higher-order feature representations and establish semantic relationships between regional and global features. Second, a knowledge attention block is developed to explore hierarchical textual information and capture cross-modal dependencies. Third, a video-centric conversation system is constructed to complete multi-round dialogues by incorporating the proposed modules including visual relation modeling, knowledge attention and text generation. Extensive experiments on five benchmark datasets including MSVD, MSRVTT, ActivityNet, Charades and EMVPC demonstrate that the proposed scheme achieves remarkable performance compared with the state-of-the-art methods. Besides, the qualitative experiment reveals the system's favorable conversation capability and provides a valuable exemplar for future video understanding works. Tengpeng Li, Hanli Wang, Qinyu Li, Zhangkai Ni |
IEEE Trans. Multim. | 4 |
| 2025 | Transferring From Distortion to Perception-Oriented Optimization: Just-Noticeable-Distortion-Based Domain AdaptationabstractTheperception-distortion- tradeoffreveals the limitation of current low-level deep learning paradigms,i.e., minimizing reconstruction distortion does not guarantee improved perceptual quality. Acknowledging the lack of a reliableperception-oriented optimization function, we are motivated to explore a flexible approach for enhancing perceptual quality by steering thetradeoffto prioritizeperception. To this end, we reconsider theperception-distortionfunction by incorporating the Just-Noticeable-Distortion (JND) mechanism. We mathematically demonstrate that in the common image restoration process, altering the optimization target from natural images to distorted images—where the distortion intensity is constrained by the JND threshold and the distortion type aligns with that arising from the restorer itself—effectively obtained improvedperceptionindices without any changes to the restorer or optimization function. Accordingly, to facilitate various low-level learning models, we are motivated to construct the first large-scale CNN-oriented JND image dataset. Our dataset comprises 500 natural images and 4,500 degraded versions generated by a series of autoencoders, as well as the actual JND judgment results collected through rigorous subjective testing from twenty volunteers. Finally, a learning-based JND inference model is established on the proposed dataset and employed in the proposed JND-based adaptation scheme, where the inferred JND images serve as pseudo-ground truth for the training or fine-tuning processes of low-level vision models. Extensive experiments on image super-resolution and end-to-end image compression across multiple models have shown encouraging improvements in perceptual quality, demonstrating the effectiveness of the proposed scheme. Our dataset is available at:https://github.com/ohq17/CNN-Oriented-JND-Dataset. Xuelin Shen, Haoqiao Ou, Zhangkai Ni, Wenhan Yang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 3 |
| 2025 | EIN: Exposure-Induced Network for Single-Image HDR ReconstructionabstractReconstructing high dynamic range (HDR) images from standard dynamic range (SDR) ones has received growing attention in recent years. A predominant problem of this task lies in the absence of texture and structural information in under/over-exposed regions. In this article, we propose an efficient and stable single-image HDR reconstruction method, namely exposure-induced network (EIN). More specifically, a dynamic range expansion branch (DB) is designed to expand the global dynamic range of the input SDR image. Moreover, two exposure-gated detail recovering branches for local over- (OB) and under- (UB) exposed regions are proposed to interact with the DB to progressively infer the texture and structural details with the learned confidence maps to resolve challenging ambiguities in such regions. The features from these three interactional branches are adaptively fused in the joint global–local decoder to reconstruct the final HDR image. The proposed network is trained based upon a large-scale dataset constructed with diverse content. Extensive experimental results demonstrate that the proposed model achieves consistent visual quality improvement for input SDR images with different exposures compared with state-of-the-art methods. The source code is available at: https://github.com/Yliu724/EIN . Zhangkai Ni, Peilin Chen 0001, Shiqi Wang 0001, Xinfeng Zhang 0001, Hanli Wang, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance FieldabstractNeural Radiance Fields (NeRF) have demonstrated impressive potential in synthesizing novel views from dense input, however, their effectiveness is challenged when dealing with sparse input. Existing approaches that incorporate additional depth or semantic supervision can alleviate this issue to an extent. However, the process of supervision collection is not only costly but also potentially inaccurate. In our work, we introduce a novel model: the Collaborative Neural Radiance Fields (ColNeRF) designed to work with sparse input. The collaboration in ColNeRF includes the cooperation among sparse input source images and the cooperation among the output of the NeRF. Through this, we construct a novel collaborative module that aligns information from various views and meanwhile imposes self-supervised constraints to ensure multi-view consistency in both geometry and appearance. A Collaborative Cross-View Volume Integration module (CCVI) is proposed to capture complex occlusions and implicitly infer the spatial location of objects. Moreover, we introduce self-supervision of target rays projected in multiple directions to ensure geometric and color consistency in adjacent regions. Benefiting from the collaboration at the input and output ends, ColNeRF is capable of capturing richer and more generalized scene representation, thereby facilitating higher-quality results of the novel view synthesis. Our extensive experimental results demonstrate that ColNeRF outperforms state-of-the-art sparse input generalizable NeRF methods. Furthermore, our approach exhibits superiority in fine-tuning towards adapting to new scenes, achieving competitive performance compared to per-scene optimized NeRF-based methods while significantly reducing computational costs. Our code is available at: https://github.com/eezkni/ColNeRF. Zhangkai Ni, Peiqi Yang, Wenhan Yang, Hanli Wang, Lin Ma 0002, Sam Kwong |
AAAI | 1 |
| 2024 | Misalignment-Robust Frequency Distribution Loss for Image TransformationabstractThis paper aims to address a common challenge in deep learning-based image transformation methods, such as im-age enhancement and super-resolution, which heavily rely on precisely aligned paired datasets with pixel-level align-ments. However, creating precisely aligned paired images presents significant challenges and hinders the advance-ment of methods trained on such data. To overcome this challenge, this paper introduces a novel and simple frequency Distribution Loss (FDL) for computing distribution distance within the frequency domain. Specifically, we transform image features into the frequency domain using Discrete Fourier Transformation (DFT). Subsequently, frequency components (amplitude and phase) are processed separately to form the FDL loss function. Our method is empirically proven effective as a training constraint due to the thoughtful utilization of global information in the frequency domain. Extensive experimental evaluations, fo-cusing on image enhancement and super-resolution tasks, demonstrate that FDL outperforms existing misalignment-robust loss functions. Furthermore, we explore the poten-tial of our FDL for image style transfer that relies solely on completely misaligned data. Our code is available at: https://github.com/eezkni/FDL Zhangkai Ni, Juncheng Wu, Wenhan Yang, Hanli Wang, Lin Ma 0002 |
CVPR | 1 |
| 2024 | Unrolled Decomposed Unpaired Learning for Controllable Low-Light Video Enhancement
Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Hanwei Zhu, Zhangkai Ni, Qi Mao 0002, Shiqi Wang 0001 |
ECCV (23) | 5 |
| 2024 | DDR: Exploiting Deep Degradation Response as Flexible Image DescriptorabstractImage deep features extracted by pre-trained networks are known to contain rich and informative representations. In this paper, we present Deep Degradation Response (DDR), a method to quantify changes in image deep features under varying degradation conditions. Specifically, our approach facilitates flexible and adaptive degradation, enabling the controlled synthesis of image degradation through text-driven prompts. Extensive evaluations demonstrate the versatility of DDR as an image descriptor, with strong correlations observed with key image attributes such as complexity, colorfulness, sharpness, and overall quality. Moreover, we demonstrate the efficacy of DDR across a spectrum of applications. It excels as a blind image quality assessment metric, outperforming existing methodologies across multiple datasets. Additionally, DDR serves as an effective unsupervised learning objective in image restoration tasks, yielding notable advancements in image deblurring and single-image super-resolution. Our code is available at: https://github.com/eezkni/DDR. Juncheng Wu, Zhangkai Ni, Hanli Wang, Wenhan Yang, Yuyin Zhou, Shiqi Wang 0001 |
NeurIPS | 2 |
| 2024 | CD-iNet: Deep Invertible Network for Perceptual Image Color Difference Measurement
Zhihua Wang 0002, Keshuo Xu, Keyan Ding, Qiuping Jiang, Yifan Zuo 0001, Zhangkai Ni, Yuming Fang 0001 |
Int. J. Comput. Vis. | 6 |
| 2024 | A Dynamic Evolution Model for Decentralized Autonomous Car Clusters in a Highway SceneabstractCluster evolution is a challenging problem for vehicular ad hoc network (VANET) in a highway scene with fast moving autonomous vehicles and frequent cluster topology changes. Most of the existing studies analyze the cluster evolution behavior of cluster heads (CHs), and these approaches lead to frequent changes in vehicle structure when CHs change, which easily makes the cluster unstable. In this work, we propose a decentralized autonomous car cluster dynamic evolution model. First, we define a decentralized cluster structure. Then, we analyze the cluster evolution behavior and propose a maintenance method. Next, we define eight vehicle states and their transitions. Finally, we introduce the cluster dynamic evolution model and the collaboration model. The results of extensive simulation experiments show that our method can effectively maintain the consistency of cluster consensus and improve the stability of the cluster structure compared with the centralized cluster maintenance method. Jiujun Cheng, Huiyu Sun, Zhangkai Ni, Aiguo Zhou |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Opinion-Unaware Blind Image Quality Assessment Using Multi-Scale Deep Feature StatisticsabstractDeep learning-based methods have significantly influenced the blind image quality assessment (BIQA) field, however, these methods often require training using large amounts of human rating data. In contrast, traditional knowledge-based methods are cost-effective for training but face challenges in effectively extracting features aligned with human visual perception. To bridge these gaps, we propose integrating deep features from pre-trained visual models with a statistical analysis model into a Multi-scale Deep Feature Statistics (MDFS) model for achieving opinion-unaware BIQA (OU-BIQA), thereby eliminating the reliance on human rating data and significantly improving training efficiency. Specifically, we extract patch-wise multi-scale features from pre-trained vision models, which are subsequently fitted into a multivariate Gaussian (MVG) model. The final quality score is determined by quantifying the distance between the MVG model derived from the test image and the benchmark MVG model derived from the high-quality image set. A comprehensive series of experiments conducted on various datasets show that our proposed model exhibits superior consistency with human visual perception compared to state-of-the-art BIQA models. Furthermore, it shows improved generalizability across diverse target-specific BIQA tasks. Our code is available at:https://github.com/eezkni/MDFS Zhangkai Ni, Keyan Ding, Wenhan Yang, Hanli Wang, Shiqi Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Glow in the Dark: Low-Light Image Enhancement With External MemoryabstractDeep learning-based methods have achieved remarkable success with powerful modeling capabilities. However, the weights of these models are learned over the entire training dataset, which inevitably leads to the ignorance of sample specific properties in the learned enhancement mapping. This situation causes ineffective enhancement in the testing phase for the samples that differ significantly from the training distribution. In this paper, we introduce external memory to form an external memory-augmented network (EMNet) for low-light image enhancement. The external memory aims to capture the sample specific properties of the training dataset to guide the enhancement in the testing phase. Benefiting from the learned memory, more complex distributions of reference images in the entire dataset can be “remembered” to facilitate the adjustment of the testing samples more adaptively. To further augment the capacity of the model, we take the transformer as our baseline network, which specializes in capturing long-range spatial redundancy. Experimental results demonstrate that our proposed method has a promising performance and outperforms state-of-the-art methods. It is noted that, the proposed external memory is a plug-and-play mechanism that can be integrated with any existing method to further improve the enhancement quality. More practices of integrating external memory with other image enhancement methods are qualitatively and quantitatively analyzed. The results further confirm that the effectiveness of our proposed memory mechanism when combing with existing enhancement methods. Dongjie Ye, Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 2 |
| 2023 | Adaptive Token Excitation with Negative Selection for Video-Text Retrieval
Juntao Yu, Zhangkai Ni, Taiyi Su, Hanli Wang |
ICANN (7) | 2 |
| 2023 | Structure-Aware Generative Adversarial Network for Text-to-Image GenerationabstractText-to-image generation aims at synthesizing photo-realistic images from textual descriptions. Existing methods typically align images with the corresponding texts in a joint semantic space. However, the presence of the modality gap in the joint semantic space leads to misalignment. Meanwhile, the limited receptive field of the convolutional neural network leads to structural distortions of generated images. In this work, a structure-aware generative adversarial network (SaGAN) is proposed for (1) semantically aligning multimodal features in the joint semantic space in a learnable manner; and (2) improving the structure and contour of generated images by the designed content-invariant negative samples. Experimental results show that SaGAN achieves over 30.1% and 8.2% improvements in terms of FID on the datasets of CUB and COCO when compared with the state-of-the-art approaches. Zhangkai Ni, Hanli Wang |
ICIP | 2 |
| 2023 | A CTU-Level Screen Content Rate Control for Low-Delay Versatile Video CodingabstractIn this paper, a rate control scheme for screen content video coding is proposed for the Versatile Video Coding (VVC) standard. In view of the critical challenges arising from the spatial and temporal unnaturalness of screen content sequences, the proposed method relies on the specifically designed pre-analysis such that the content information regarding the scene complexity can be obtained. As such, the estimated residual complexity is then incorporated into the proposed complexity-aware rate models and distortion models, leading to the optimal bit allocations for each frame and coding tree unit (CTU). In particular, the optimization problem can be analytically solved with the proposed models, and the coding parameters such as Lagrangian multiplier$\lambda $and quantization parameter of each frame and CTU could be delicately derived according to the allocated bits through the proposed analytical models. Extensive experiments have been conducted to evaluate the effectiveness of the proposed method. Compared to the default hierarchical$\lambda $-domain rate control and other screen content rate control algorithms, the proposed method could achieve obvious RD performance gain, and the bit-rate accuracy could be improved. Yi Chen 0028, Meng Wang 0017, Shiqi Wang 0001, Zhangkai Ni, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | High Dynamic Range Image Quality Assessment Based on Frequency DisparityabstractIn this paper, a novel and effective image quality assessment (IQA) algorithm based on frequency disparity for high dynamic range (HDR) images is proposed, termed as local-global frequency feature-based model (LGFM). Motivated by the assumption that the human visual system (HVS) is highly adapted for extracting structural information and partial frequencies when perceiving the visual scene, the Gabor and the Butterworth filters are applied to the luminance component of the HDR image to extract the local and global frequency features, respectively. The similarity measurement and feature pooling strategy are sequentially performed on the frequency features to obtain the predicted single quality score. The experiments evaluated on four widely used benchmarks demonstrate that the proposed LGFM can provide a higher consistency with the subjective perception compared with the state-of-the-art HDR IQA methods. Our code is available at:https://github.com/eezkni/LGFM. Zhangkai Ni, Shiqi Wang 0001, Hanli Wang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Neural Network Based Rate Control for Versatile Video CodingabstractIn this work, we propose a neural network based rate control algorithm for Versatile Video Coding (VVC). The proposed method relies on the modeling of the Rate-Quantization (R-Q) and Distortion-Quantization (D-Q) relationships in a data driven manner based upon the characteristics of prediction residuals. In particular, a pre-analysis framework is adopted, in an effort to obtain the prediction residuals which govern the Rate-Distortion (R-D) behaviors. By inferring from the prediction residuals with deep neural networks, the Coding Tree Unit (CTU) level R-Q and D-Q model parameters are derived, which could efficiently guide the optimal bit allocation. Subsequently, the coding parameters, including Quantization Parameter (QP) and$\lambda $, at both frame and CTU levels, are obtained according to allocated bit-rates. We implement the proposed rate control algorithm on VVC Test Model (VTM-13.0). Experimental results exhibit that the proposed rate control algorithm achieves 0.77% BD-Rate savings under Low Delay B (LDB) configurations when compared to the default rate control algorithm used in VTM-13.0. For Random Access (RA) configurations, 1.77% BD-Rate savings can be observed. Furthermore, with better bit-rate estimation, more stable buffer status can be observed, further demonstrating the advantages of the proposed rate control method. Yunhao Mao, Meng Wang 0017, Zhangkai Ni, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | CSformer: Bridging Convolution and Transformer for Compressive SensingabstractConvolutional Neural Networks (CNNs) dominate image processing but suffer from local inductive bias, which is addressed by the transformer framework with its inherent ability to capture global context through self-attention mechanisms. However, how to inherit and integrate their advantages to improve compressed sensing is still an open issue. This paper proposes CSformer, a hybrid framework to explore the representation capacity of local and global features. The proposed approach is well-designed for end-to-end compressive image sensing, composed of adaptive sampling and recovery. In the sampling module, images are measured block-by-block by the learned sampling matrix. In the reconstruction stage, the measurements are projected into an initialization stem, a CNN stem, and a transformer stem. The initialization stem mimics the traditional reconstruction of compressive sensing but generates the initial reconstruction in a learnable and efficient manner. The CNN stem and transformer stem are concurrent, simultaneously calculating fine-grained and long-range features and efficiently aggregating them. Furthermore, we explore a progressive strategy and window-based transformer block to reduce the parameters and computational complexity. The experimental results demonstrate the effectiveness of the dedicated transformer-based architecture for compressive sensing, which achieves superior performance compared to state-of-the-art methods on different datasets. Our codes is available at: https://github.com/Lineves7/CSformer. Dongjie Ye, Zhangkai Ni, Hanli Wang, Jian Zhang 0018, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 2 |
| 2022 | Cycle-Interactive Generative Adversarial Network for Robust Unsupervised Low-Light EnhancementabstractGetting rid of the fundamental limitations in fitting to the paired training data, recent unsupervised low-light enhancement methods excel in adjusting illumination and contrast of images. However, for unsupervised low light enhancement, the remaining noise suppression issue due to the lacking of supervision of detailed signal largely impedes the wide deployment of these methods in real-world applications. Herein, we propose a novel Cycle-Interactive Generative Adversarial Network (CIGAN) for unsupervised low-light image enhancement, which is capable of not only better transferring illumination distributions between low/normal-light images but also manipulating detailed signals between two domains, e.g., suppressing/synthesizing realistic noise in the cyclic enhancement/degradation process. In particular, the proposed low-light guided transformation feed-forwards the features of low-light images from the generator of enhancement GAN (eGAN) into the generator of degradation GAN (dGAN). With the learned information of real low-light images, dGAN can synthesize more realistic diverse illumination and contrast in low-light images. Moreover, the feature randomized perturbation module in dGAN learns to increase the feature randomness to produce diverse feature distributions, persuading the synthesized low-light images to contain realistic noise. Extensive experiments demonstrate both the superiority of the proposed method and the effectiveness of each module in CIGAN. Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
ACM Multimedia | 1 |
| 2021 | Just Noticeable Distortion Profile Inference: A Patch-Level Structural Visibility Learning ApproachabstractIn this paper, we propose an effective approach to infer the just noticeable distortion (JND) profile based on patch-level structural visibility learning. Instead of pixel-level JND profile estimation, the image patch, which is regarded as the basic processing unit to better correlate with the human perception, can be further decomposed into three conceptually independent components for visibility estimation. In particular, to incorporate the structural degradation into the patch-level JND model, a deep learning-based structural degradation estimation model is trained to approximate the masking of structural visibility. In order to facilitate the learning process, a JND dataset is further established, including 202 pristine images and 7878 distorted images generated by advanced compression algorithms based on the upcoming Versatile Video Coding (VVC) standard. Extensive experimental results further show the superiority of the proposed approach over the state-of-the-art. Our dataset is available at: https://github.com/ShenXuelin-CityU/PWJNDInfer. Xuelin Shen, Zhangkai Ni, Wenhan Yang, Xinfeng Zhang 0001, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 2 |
| 2020 | Unpaired Image Enhancement with Quality-Attention Generative Adversarial NetworkabstractIn this work, we aim to learn an unpaired image enhancement model, which can enrich low-quality images with the characteristics of high-quality images provided by users. We propose a quality attention generative adversarial network (QAGAN) trained on unpaired data based on the bidirectional Generative Adversarial Network (GAN) embedded with a quality attention module (QAM). The key novelty of the proposed QAGAN lies in the injected QAM for the generator such that it learns domain-relevant quality attention directly from the two domains. More specifically, the proposed QAM allows the generator to effectively select semantic-related characteristics from the spatial-wise and adaptively incorporate style-related attributes from the channel-wise, respectively. Therefore, in our proposed QAGAN, not only discriminators but also the generator can directly access both domains which significantly facilitate the generator to learn the mapping function. Extensive experimental results show that, compared with the state-of-the-art methods based on unpaired learning, our proposed method achieves better performance in both objective and subjective evaluations. Zhangkai Ni, Wenhan Yang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
ACM Multimedia | 1 |
| 2020 | Color Image Demosaicing Using Progressive Collaborative RepresentationabstractIn this paper, a progressive collaborative representation (PCR) framework is proposed that is able to incorporate any existing color image demosaicing method for further boosting its demosaicing performance. Our PCR consists of two phases: (i) offline training and (ii) online refinement. In phase (i), multiple training-and-refining stages will be performed. In each stage, a new dictionary will be established through the learning of a large number of feature-patch pairs, extracted from the demosaicked images of the current stage and their corresponding original full-color images. After training, a projection matrix will be generated and exploited to refine the current demosaicked image. The updated image with improved image quality will be used as the input for the next training-and-refining stage and performed the same processing likewise. At the end of phase (i), all the projection matrices generated as above-mentioned will be exploited in phase (ii) to conduct online demosaicked image refinement of the test image. Extensive simulations conducted on two commonly-used test datasets (i.e., the IMAX and Kodak) for evaluating the demosaicing algorithms have clearly demonstrated that our proposed PCR framework is able to constantly boost the performance of any image demosaicing method we experimented, in terms of the objective and subjective performance evaluations. Zhangkai Ni, Kai-Kuang Ma, Huanqiang Zeng, Baojiang Zhong |
IEEE Trans. Image Process. | 1 |
| 2020 | Towards Unsupervised Deep Image Enhancement With Generative Adversarial NetworkabstractImproving the aesthetic quality of images is challenging and eager for the public. To address this problem, most existing algorithms are based on supervised learning methods to learn an automatic photo enhancer for paired data, which consists of low-quality photos and corresponding expert-retouched versions. However, the style and characteristics of photos retouched by experts may not meet the needs or preferences of general users. In this paper, we present an unsupervised image enhancement generative adversarial network (UEGAN), which learns the corresponding image-to-image mapping from a set of images with desired characteristics in an unsupervised manner, rather than learning on a large number of paired images. The proposed model is based on single deep GAN which embeds the modulation and attention mechanisms to capture richer global and local features. Based on the proposed model, we introduce two losses to deal with the unsupervised image enhancement: (1) fidelity loss, which is defined as a l2 regularization in the feature domain of a pre-trained VGG network to ensure the content between the enhanced image and the input image is the same, and (2) quality loss that is formulated as a relativistic hinge adversarial loss to endow the input image the desired characteristics. Both quantitative and qualitative results show that the proposed model effectively improves the aesthetic quality of images. Zhangkai Ni, Wenhan Yang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
IEEE Trans. Image Process. | 1 |
| 2018 | Screen Content Image Quality Assessment Using Multi-Scale Difference of GaussianabstractIn this paper, a novel image quality assessment (IQA) model for the screen content images (SCIs) is proposed by using multi-scale difference of Gaussian (MDOG). Motivated by the observation that the human visual system (HVS) is sensitive to the edges while the image details can be better explored in different scales, the proposed model exploits MDOG to effectively characterize the edge information of the reference and distorted SCIs at two different scales, respectively. Then, the degree of edge similarity is measured in terms of the smaller-scale edge map. Finally, the edge strength computed based on the larger-scale edge map is used as the weighting factor to generate the final SCI quality score. Experimental results have shown that the proposed IQA model for the SCIs produces high consistency with human perception of the SCI quality and outperforms the state-of-the-art quality models. Ying Fu 0004, Huanqiang Zeng, Lin Ma 0002, Zhangkai Ni, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | A Gabor Feature-Based Quality Assessment Model for the Screen Content ImagesabstractIn this paper, an accurate and efficient full-reference image quality assessment (IQA) model using the extracted Gabor features, called Gabor feature-based model (GFM), is proposed for conducting objective evaluation of screen content images (SCIs). It is well-known that the Gabor filters are highly consistent with the response of the human visual system (HVS), and the HVS is highly sensitive to the edge information. Based on these facts, the imaginary part of the Gabor filter that has odd symmetry and yields edge detection is exploited to the luminance of the reference and distorted SCI for extracting their Gabor features, respectively. The local similarities of the extracted Gabor features and two chrominance components, recorded in the LMN color space, are then measured independently. Finally, the Gabor-feature pooling strategy is employed to combine these measurements and generate the final evaluation score. Experimental simulation results obtained from two large SCI databases have shown that the proposed GFM model not only yields a higher consistency with the human perception on the assessment of SCIs but also requires a lower computational complexity, compared with that of classical and state-of-the-art IQA models. The source code for the proposed GFM will be available at http://smartviplab.org/pubilcations/GFM.html. Zhangkai Ni, Huanqiang Zeng, Lin Ma 0002, Junhui Hou, Jing Chen 0001, Kai-Kuang Ma |
IEEE Trans. Image Process. | 1 |
| 2017 | ESIM: Edge Similarity for Screen Content Image Quality AssessmentabstractIn this paper, an accurate full-reference image quality assessment (IQA) model developed for assessing screen content images (SCIs), called the edge similarity (ESIM), is proposed. It is inspired by the fact that the human visual system (HVS) is highly sensitive to edges that are often encountered in SCIs; therefore, essential edge features are extracted and exploited for conducting IQA for the SCIs. The key novelty of the proposed ESIM lies in the extraction and use of three salient edge features-i.e., edge contrast, edge width, and edge direction. The first two attributes are simultaneously generated from the input SCI based on a parametric edge model, while the last one is derived directly from the input SCI. The extraction of these three features will be performed for the reference SCI and the distorted SCI, individually. The degree of similarity measured for each above-mentioned edge attribute is then computed independently, followed by combining them together using our proposed edge-width pooling strategy to generate the final ESIM score. To conduct the performance evaluation of our proposed ESIM model, a new and the largest SCI database (denoted as SCID) is established in our work and made to the public for download. Our database contains 1800 distorted SCIs that are generated from 40 reference SCIs. For each SCI, nine distortion types are investigated, and five degradation levels are produced for each distortion type. Extensive simulation results have clearly shown that the proposed ESIM model is more consistent with the perception of the HVS on the evaluation of distorted SCIs than the multiple state-of-the-art IQA methods. Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Jing Chen 0001, Canhui Cai, Kai-Kuang Ma |
IEEE Trans. Image Process. | 1 |
| 2016 | Screen content image quality assessment using edge modelabstractSince the human visual system (HVS) is highly sensitive to edges, a novel image quality assessment (IQA) metric for assessing screen content images (SCIs) is proposed in this paper. The turnkey novelty lies in the use of an existing parametric edge model to extract two types of salient attributes - namely, edge contrast and edge width, for the distorted SCI under assessment and its original SCI, respectively. The extracted information is subject to conduct similarity measurements on each attribute, independently. The obtained similarity scores are then combined using our proposed edge-width pooling strategy to generate the final IQA score. Hopefully, this score is consistent with the judgment made by the HVS. Experimental results have shown that the proposed IQA metric produces higher consistency with that of the HVS on the evaluation of the image quality of the distorted SCI than that of other state-of-the-art IQA metrics. Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma |
ICIP | 1 |
| 2016 | Gradient Direction for Screen Content Image Quality AssessmentabstractIn this letter, we make the first attempt to explore the usage of the gradient direction to conduct the perceptual quality assessment of the screen content images (SCIs). Specifically, the proposed approach first extracts the gradient direction based on the local information of the image gradient magnitude, which not only preserves gradient direction consistency in local regions, but also demonstrates sensitivities to the distortions introduced to the SCI. A deviation-based pooling strategy is subsequently utilized to generate the corresponding image quality index. Moreover, we investigate and demonstrate the complementary behaviors of the gradient direction and magnitude for SCI quality assessment. By jointly considering them together, our proposed SCI quality metric outperforms the state-of-the-art quality metrics in terms of correlation with human visual system perception. Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 1 |