VLDB 2026 Research / reviewers in the wild / expert
Renjie Wan
dblp:191/2619
· DBLP profile ↗
59ranked-venue papers
10as first author
45since 2021 · last 2026
0000-0002-0161-0367ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 7 first-author · 27 since 2021Artificial intelligence and machine learning · 33 · 6 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GS-Checker: Tampering Localization for 3D Gaussian SplattingabstractRecent advances in editing technologies for 3D Gaussian Splatting (3DGS) have made it simple to manipulate 3D scenes. However, these technologies raise concerns about potential malicious manipulation of 3D content. To avoid such malicious applications, localizing tampered regions becomes crucial. In this paper, we propose GS-Checker, a novel method for locating tampered areas in 3DGS models. Our approach integrates a 3D tampering attribute into the 3D Gaussian parameters to indicate whether the Gaussian has been tampered. Additionally, we design a 3D contrastive mechanism by comparing the similarity of key attributes between 3D Gaussians to seek tampering cues at 3D level. Furthermore, we introduce a cyclic optimization strategy to refine the 3D tampering attribute, enabling more accurate tampering localization. Notably, our approach does not require expensive 3D labels for supervision. Extensive experimental results demonstrate the effectiveness of our proposed method to locate the tampered 3DGS area. Haoliang Han, Ziyuan Luo, Anderson Rocha 0001, Renjie Wan |
AAAI | 5 |
| 2026 | Creating Blank Canvas Against AI-enabled Image ForgeryabstractAIGC-based image editing technology has greatly simplified the realistic-level image modification, causing serious potential risks of image forgery. This paper introduces a new approach to tampering detection using the Segment Anything Model (SAM). Instead of training SAM to identify tampered areas, we propose a novel strategy. The entire image is transformed into a blank canvas from the perspective of neural models. Any modifications to this blank canvas would be noticeable to the models. To achieve this idea, we introduce adversarial perturbations to prevent SAM from seeing anything, allowing it to identify forged regions when the image is tampered with. Due to SAM's powerful perceiving capabilities, naive adversarial attacks cannot completely tame SAM. To thoroughly deceive SAM and make it blind to the image, we introduce a frequency-aware optimization strategy, which further enhances the capability of tamper localization. Extensive experimental results demonstrate the effectiveness of our method. Qi Song 0003, Ziyuan Luo, Renjie Wan |
AAAI | 3 |
| 2026 | Naturalistic Typographic Attacks on VLM-Based Image Quality Assessment
Ziyuan Luo, Qi Song 0003, Renjie Wan |
QoMEX | 4 |
| 2026 | Adversarially robust multimedia watermarking via data-centric optimization
Ziyuan Luo, Qi Song 0003, Haoliang Li, Anderson Rocha 0001, Renjie Wan |
Pattern Recognit. | 5 |
| 2026 | Open-Set Deepfake Detection: A Parameter-Efficient Adaptation Method With Forgery Style MixtureabstractOpen-set face forgery detection poses significant security threats and presents substantial challenges for existing detection models. These detectors primarily have two limitations: they cannot generalize across unknown forgery domains or inefficiently adapt to new data. To address these issues, we introduce an approach that is both general and parameter-efficient for face forgery detection. Our method builds on the assumption that different forgery source domains exhibit distinct style statistics. Specifically, we design a forgery-style-mixture formulation that augments the diversity of forgery source domains, enhancing the model’s generalizability across unseen domains. In addition, previous methods typically require fully fine-tuning pretrained networks, consuming substantial time and computational resources. Drawing on recent advancements in vision transformers (ViT) for face forgery detection, we develop a parameter-efficient ViT-based detection model that includes lightweight forgery feature extraction modules and enables the model to extract global and local forgery clues simultaneously. We only optimize the inserted lightweight modules during training, maintaining the original ViT structure with its pre-trained weights. This training strategy effectively preserves the informative pre-trained knowledge while flexibly adapting the model to the task of Deepfake detection. Extensive experimental results demonstrate that the designed model achieves state-of-the-art generalizability with significantly reduced trainable parameters, representing an important step toward open-set Deepfake detection in the wild. Chenqi Kong, Anwei Luo, Peijun Bao, Haoliang Li, Renjie Wan, Zengwei Zheng, Anderson Rocha 0001, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | MantleMark: Migrating Watermarks From Multi-View Images to Radiance Fields via Frequency ModulationabstractMulti-view images are essential for modern radiance field reconstruction methods like Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). While image watermarking is a crucial data protection and ownership verification technique, it faces unprecedented challenges in multi-view scenarios. Traditional 2D watermarking techniques often fail to maintain detectability in rendered views, while existing 3D watermarking methods are typically limited to specific reconstruction methods and require access to the reconstruction process. To address these limitations, we propose MantleMark, a watermarking framework that migrates watermarks from multi-view images to radiance fields via frequency modulation. Our key insight is constructing a mantle-like Frequency-domain Watermarking Representation in 3D frequency space, which can be projected to create view-dependent watermarking patterns. Relying upon the Fourier Projection-Slice Theorem, we embed these patterns through magnitude spectrum modulation in the image frequency domain, enabling watermarks to migrate into 3D representations. This approach ensures watermark detectability in rendered views regardless of the reconstruction methods used by adversaries. Extensive experiments demonstrate that our method achieves robust watermark detection while maintaining high visual quality across various radiance field-based reconstruction methods. Ziyuan Luo, Jun Liu 0036, Haoliang Li, Anderson Rocha 0001, Renjie Wan |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Generalizable Dynamic Representation Learning for Source Identification in Sequential DataabstractSource identification is a foundational task in multimedia forensics, enabling the attribution and verification of digital content. While existing methods have achieved significant progress for static data, they often fail to generalize effectively on sequential data, which exhibit unique challenges such as temporal dependencies and dynamic variations caused by environmental and transmission factors. These challenges are further exacerbated in real-world scenarios, where crossdomain variations-spanning devices, software, and transmission protocols-significantly degrade the performance of traditional approaches. To address these limitations, we propose VoVAE, a probabilistic variational framework tailored for generalizable source identification in sequential data. VoVAE explicitly models temporal dependencies while disentangling dynamic variations (e.g., transmission distortions) from static source-specific features (e.g., device patterns) within a decoupled but complementary feature space. By separating these factors, VoVAE enables the extraction of robust and transferable representations, ensuring accurate source attribution across diverse and unseen conditions. We evaluate VoVAE on two challenging forensic applications: cross-domain VoIP phone call identification and cross-domain video source camera identification, using the VPCID and QUFVD datasets. Experimental results demonstrate that VoVAE outperforms state-of-the-art methods, achieving significant improvements in generalization across cross-device, cross-software, and cross-brand scenarios. Comprehensive ablation studies further highlight the importance of dynamic representation learning and feature disentanglement in capturing temporal patterns and enhancing robustness to domain shifts. These findings establish VoVAE as a scalable and robust solution for source identification in sequential data across diverse forensic scenarios. Bo Ding 0006, Tiexin Qin, Renjie Wan, Haoliang Li |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal TriggersabstractHateful meme detection aims to prevent the proliferation of hateful memes on various social media platforms. Considering its impact on social environments, this paper introduces a previously ignored but significant threat to hateful meme detection: backdoor attacks. By injecting specific triggers into meme samples, backdoor attackers can manipulate the detector to output their desired outcomes. To explore this, we propose the Meme Trojan framework to initiate backdoor attacks on hateful meme detection. Meme Trojan involves creating a novel Cross-Modal Trigger (CMT) and a learnable trigger augmentor to enhance the trigger pattern according to each input sample. Due to the cross-modal property, the proposed CMT can effectively initiate backdoor attacks on hateful meme detectors under an automatic application scenario. Additionally, the injection position and size of our triggers are adaptive to the texts contained in the meme, which ensures that the trigger is seamlessly integrated with the meme content. Our approach outperforms the state-of-the-art backdoor attack methods, showing significant improvements in effectiveness and stealthiness. We believe that this paper will draw more attention to the potential threat posed by backdoor attacks on hateful meme detection. Ruofei Wang, Hongzhan Lin 0001, Ziyuan Luo, Ka Chun Cheung, Simon See, Jing Ma 0004, Renjie Wan |
AAAI | 7 |
| 2025 | Asynchronous Event Error-Minimizing Noise for Safeguarding Event DatasetabstractWith more event datasets being released online, safeguarding the event dataset against unauthorized usage has become a serious concern for data owners. Unlearnable Examples are proposed to prevent the unauthorized exploitation of image datasets. However, it's unclear how to create unlearnable asynchronous event streams to prevent event misuse. In this work, we propose the first unlearnable event stream generation method to prevent unauthorized training from event datasets. A new form of asynchronous event error-minimizing noise is proposed to perturb event streams, tricking the unauthorized model into learning embedded noise instead of realistic features. To be compatible with the sparse event, a projection strategy is presented to sparsify the noise to render our unlearnable event streams (UEvs). Extensive experiments demonstrate that our method effectively protects event data from unauthorized exploitation, while preserving their utility for legitimate use. We hope our UEvs contribute to the advancement of secure and trustworthy event dataset sharing. Code is available at: https://github.com/rfww/uevs. Ruofei Wang, Peiqi Duan 0002, Boxin Shi, Renjie Wan |
ICCV | 4 |
| 2025 | Align 3D Representation and Text Embedding for 3D Content PersonalizationabstractRecent advances in NeRF and 3DGS have significantly enhanced the efficiency and quality of 3D content synthesis. However, efficient personalization of generated 3D content remains a critical challenge. Current 3D personalization approaches predominantly rely on knowledge distillation-based methods, which require computationally expensive retraining procedures. To address this challenge, we propose Invert3D, a novel framework for convenient 3D content personalization. Nowadays, vision-language models such as CLIP enable direct image personalization through aligned vision-text embedding spaces. However, the inherent structural differences between 3D content and 2D images preclude direct application of these techniques to 3D personalization. Our approach bridges this gap by establishing alignment between 3D representations and text embedding spaces. Specifically, we develop a camera-conditioned 3D-to-text inverse mechanism that projects 3D contents into a 3D embedding aligned with text embeddings. This alignment enables efficient manipulation and personalization of 3D content through natural language prompts, eliminating the need for computationally retraining procedures. Extensive experiments demonstrate that Invert3D achieves effective personalization of 3D content. Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan |
ACM Multimedia | 5 |
| 2025 | Stereo-GS: Multi-View Stereo Vision Model for Generalizable 3D Gaussian Splatting ReconstructionabstractGeneralizable 3D Gaussian Splatting reconstruction showcases advanced Image-to-3D content creation but requires substantial computational resources and large datasets, posing challenges to training models from scratch. Current methods usually entangle the prediction of 3D Gaussian geometry and appearance, which rely heavily on data-driven priors and result in slow regression speeds. To address this, we propose Stereo-GS, a disentangled framework for efficient 3D Gaussian prediction. Our method extracts features from local image pairs using a stereo vision backbone and fuses them via global attention blocks. Dedicated point and Gaussian prediction heads generate multi-view point-maps for geometry and Gaussian features for appearance, combined as GS-maps to represent the 3DGS object. A refinement network enhances these GSmaps for high-quality reconstruction. Unlike existing methods that depend on camera parameters, our approach achieves pose-free 3D reconstruction, improving robustness and practicality. By reducing resource demands while maintaining high-quality outputs, Stereo- GS provides an efficient, scalable solution for real-world 3D content generation. Project page: https://kevinhuangxf.github.io/stereo-gs. Xiufeng Huang, Ka Chun Cheung, Runmin Cong, Simon See, Renjie Wan |
ACM Multimedia | 5 |
| 2025 | MarkSplatter: Generalizable Watermarking for 3D Gaussian Splatting Model via Splatter Image StructureabstractThe growing popularity of 3D Gaussian Splatting (3DGS) has intensified the need for effective copyright protection. Current 3DGS watermarking methods rely on computationally expensive fine-tuning procedures for each predefined message. We propose the first generalizable watermarking framework that enables efficient protection of Splatter Image-based 3DGS models through a single forward pass. We introduce GaussianBridge that transforms unstructured 3D Gaussians into Splatter Image format, enabling direct neural processing for arbitrary message embedding. To ensure imperceptibility, we design a Gaussian-Uncertainty-Perceptual heatmap prediction strategy for preserving visual quality. For robust message recovery, we develop a dense segmentation-based extraction mechanism that maintains reliable extraction even when watermarked objects occupy minimal regions in rendered views. Project page: https://kevinhuangxf.github.io/marksplatter. Xiufeng Huang, Ziyuan Luo, Qi Song 0003, Ruofei Wang, Renjie Wan |
ACM Multimedia | 5 |
| 2025 | ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image GenerationabstractThe widespread adoption of Retrieval-Augmented Image Generation (RAIG) has raised significant concerns about the unauthorized use of private image datasets. While these systems have shown remarkable capabilities in enhancing generation quality through reference images, protecting visual datasets from unauthorized use in such systems remains a challenging problem. Traditional digital watermarking approaches face limitations in RAIG systems, as the complex feature extraction and recombination processes fail to preserve watermark signals during generation. To address these challenges, we propose ImageSentinel, a novel framework for protecting visual datasets in RAIG. Our framework synthesizes sentinel images that maintain visual consistency with the original dataset. These sentinels enable protection verification through randomly generated character sequences that serve as retrieval keys. To ensure seamless integration, we leverage vision-language models to generate the sentinel images. Experimental results demonstrate that ImageSentinel effectively detects unauthorized dataset usage while preserving generation quality for authorized applications. Ziyuan Luo, Yangyi Zhao, Ka Chun Cheung, Simon See, Renjie Wan |
NeurIPS | 5 |
| 2025 | The NeRF Signature: Codebook-Aided Watermarking for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have been gaining attention as a significant form of 3D content representation. With the proliferation of NeRF-based creations, the need for copyright protection has emerged as a critical issue. Although some approaches have been proposed to embed digital watermarks into NeRF, they often neglect essential model-level considerations and incur substantial time overheads, resulting in reduced imperceptibility and robustness, along with user inconvenience. In this paper, we extend the previous criteria for image watermarking to the model level and propose NeRF Signature, a novel watermarking method for NeRF. We employ a Codebook-aided Signature Embedding (CSE) that does not alter the model structure, thereby maintaining imperceptibility and enhancing robustness at the model level. Furthermore, after optimization, any desired signatures can be embedded through the CSE, and no fine-tuning is required when NeRF owners want to use new binary signatures. Then, we introduce a joint pose-patch encryption watermarking strategy to hide signatures into patches rendered from a specific viewpoint for higher robustness. In addition, we explore a Complexity-Aware Key Selection (CAKS) scheme to embed signatures in high visual complexity patches to enhance imperceptibility. The experimental results demonstrate that our method outperforms other baseline methods in terms of imperceptibility and robustness. Ziyuan Luo, Anderson Rocha 0001, Boxin Shi, Qing Guo 0005, Haoliang Li, Renjie Wan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Reference-Based Iterative Interaction With P2-Matching for Stereo Image Super-ResolutionabstractStereo Image Super-Resolution (SSR) holds great promise in improving the quality of stereo images by exploiting the complementary information between left and right views. Most SSR methods primarily focus on the inter-view correspondences in low-resolution (LR) space. The potential of referencing a high-quality SR image of one view benefits the SR for the other is often overlooked, while those with abundant textures contribute to accurate correspondences. Therefore, we propose Reference-based Iterative Interaction (RIISSR), which utilizes reference-based iterative pixel-wise and patch-wise matching, dubbed $P^{2}$ -Matching, to establish cross-view and cross-resolution correspondences for SSR. Specifically, we first design the information perception block (IPB) cascaded in parallel to extract hierarchical contextualized features for different views. Pixel-wise matching is embedded between two parallel IPBs to exploit cross-view interaction in LR space. Iterative patch-wise matching is then executed by utilizing the SR stereo pair as another mutual reference, capitalizing on the cross-scale patch recurrence property to learn high-resolution (HR) correspondences for SSR performance. Moreover, we introduce the supervised side-out modulator (SSOM) to re-weight local intra-view features and produce intermediate SR images, which seamlessly bridge two matching mechanisms. Experimental results demonstrate the superiority of RIISSR against existing state-of-the-art methods. Runmin Cong, Rongxin Liao, Feng Li 0037, Ronghui Sheng, Huihui Bai 0001, Renjie Wan, Sam Kwong, Wei Zhang 0021 |
IEEE Trans. Image Process. | 6 |
| 2025 | Unsupervised Domain Adaptation for Low-Dose CT Reconstruction via Bayesian Uncertainty AlignmentabstractLow-dose computed tomography (LDCT) image reconstruction techniques can reduce patient radiation exposure while maintaining acceptable imaging quality. Deep learning (DL) is widely used in this problem, but the performance of testing data (also known as target domain) is often degraded in clinical scenarios due to the variations that were not encountered in training data (also known as source domain). Unsupervised domain adaptation (UDA) of LDCT reconstruction has been proposed to solve this problem through distribution alignment. However, existing UDA methods fail to explore the usage of uncertainty quantification, which is crucial for reliable intelligent medical systems in clinical scenarios with unexpected variations. Moreover, existing direct alignment for different patients would lead to content mismatch issues. To address these issues, we propose to leverage a probabilistic reconstruction framework to conduct a joint discrepancy minimization between source and target domains in both the latent and image spaces. In the latent space, we devise a Bayesian uncertainty alignment to reduce the epistemic gap between the two domains. This approach reduces the uncertainty level of target domain data, making it more likely to render well-reconstructed results on target domains. In the image space, we propose a sharpness-aware distribution alignment (SDA) to achieve a match of second-order information, which can ensure that the reconstructed images from the target domain have similar sharpness to normal-dose CT (NDCT) images from the source domain. Experimental results on two simulated datasets and one clinical low-dose imaging dataset show that our proposed method outperforms other methods in quantitative and visualized performance. Kecheng Chen, Jie Liu 0044, Renjie Wan, Victor Ho-fun Lee, Varut Vardhanabhuti, Hong Yan 0001, Haoliang Li |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Colorizing Monochromatic Radiance FieldsabstractThough Neural Radiance Fields (NeRF) can produce colorful 3D representations of the world by using a set of 2D images, such ability becomes non-existent when only monochromatic images are provided. Since color is necessary in representing the world, reproducing color from monochromatic radiance fields becomes crucial. To achieve this goal, instead of manipulating the monochromatic radiance fields directly, we consider it as a representation-prediction task in the Lab color space. By first constructing the luminance and density representation using monochromatic images, our prediction stage can recreate color representation on the basis of an image colorization module. We then reproduce a colorful implicit model through the representation of luminance, density, and color. Extensive experiments have been conducted to validate the effectiveness of our approaches. Our project page: https://liquidammonia.github.io/color-nerf. Yean Cheng, Renjie Wan, Shuchen Weng, Chengxuan Zhu, Yakun Chang, Boxin Shi |
AAAI | 2 |
| 2024 | CogSimulator: A Model for Simulating User Cognition & Behavior with Minimal Data for Tailored Cognitive Enhancement
Weizhen Bian, Yubo Zhou, Yuanhang Luo, Ming Mo, Siyan Liu 0001, Yikai Gong, Ziyuan Luo, Aobo Wang, Renjie Wan |
CogSci | 9 |
| 2024 | Neural Underwater Scene RepresentationabstractAmong the numerous efforts towards digitally recovering the physical world, Neural Radiance Fields (NeRFs) have proved effective in most cases. However, underwater scene introduces unique challenges due to the absorbing water medium, the local change in lighting and the dynamic contents in the scene. We aim at developing a neural under-water scene representation for these challenges, modeling the complex process of attenuation, unstable in-scattering and moving objects during light transport. The proposed method can reconstruct the scenes from both established datasets and in-the-wild videos with outstanding fidelity. Yunkai Tang, Chengxuan Zhu, Renjie Wan, Boxin Shi |
CVPR | 3 |
| 2024 | GeometrySticker: Enabling Ownership Claim of Recolorized Neural Radiance Fields
Xiufeng Huang, Ka Chun Cheung, Simon See, Renjie Wan |
ECCV (9) | 4 |
| 2024 | Imaging Interiors: An Implicit Solution to Electromagnetic Inverse Scattering Problems
Ziyuan Luo, Boxin Shi, Haoliang Li, Renjie Wan |
ECCV (7) | 4 |
| 2024 | Protecting NeRFs' Copyright via Plug-And-Play Watermarking Base Model
Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan |
ECCV (11) | 5 |
| 2024 | Event Trojan: Asynchronous Event-Based Backdoor Attacks
Ruofei Wang, Qing Guo 0005, Haoliang Li, Renjie Wan |
ECCV (7) | 4 |
| 2024 | SPY-Watermark: Robust Invisible Watermarking for Backdoor AttackabstractBackdoor attack aims to deceive a victim model when facing backdoor instances while maintaining its performance on benign data. Current methods use manual patterns or special perturbations as triggers, while they often overlook the robustness against data corruption, making backdoor attacks easy to defend in practice. To address this issue, we propose a novel backdoor attack method named Spy-Watermark, which remains effective when facing data collapse and backdoor defense. Therein, we introduce a learnable watermark embedded in the latent domain of images, serving as the trigger. Then, we search for a watermark that can withstand collapse during image decoding, cooperating with several anti-collapse operations to further enhance the resilience of our trigger against data corruption. Extensive experiments are conducted on CIFAR10, GTSRB, and ImageNet datasets, demonstrating that Spy-Watermark overtakes ten state-of-the-art methods in terms of robustness and stealthiness. Ruofei Wang, Renjie Wan, Zongyu Guo, Qing Guo 0005, Rui Huang 0006 |
ICASSP | 2 |
| 2024 | Geometry Cloak: Preventing TGS-based 3D Reconstruction from Copyrighted ImagesabstractSingle-view 3D reconstruction methods like Triplane Gaussian Splatting (TGS) have enabled high-quality 3D model generation from just a single image input within seconds. However, this capability raises concerns about potential misuse, where malicious users could exploit TGS to create unauthorized 3D models from copyrighted images. To prevent such infringement, we propose a novel image protection approach that embeds invisible geometry perturbations, termed ``geometry cloaks'', into images before supplying them to TGS. These carefully crafted perturbations encode a customized message that is revealed when TGS attempts 3D reconstructions of the cloaked image. Unlike conventional adversarial attacks that simply degrade output quality, our method forces TGS to fail the 3D reconstruction in a specific way - by generating an identifiable customized pattern that acts as a watermark. This watermark allows copyright holders to assert ownership over any attempted 3D reconstructions made from their protected images. Extensive experiments have verified the effectiveness of our geometry cloak. Qi Song 0003, Ziyuan Luo, Ka Chun Cheung, Simon See, Renjie Wan |
NeurIPS | 5 |
| 2024 | GaussianMarker: Uncertainty-Aware Copyright Protection of 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has become a crucial method for acquiring 3D assets. To protect the copyright of these assets, digital watermarking techniques can be applied to embed ownership information discreetly within 3DGS mod- els. However, existing watermarking methods for meshes, point clouds, and implicit radiance fields cannot be directly applied to 3DGS models, as 3DGS models use explicit 3D Gaussians with distinct structures and do not rely on neural networks. Naively embedding the watermark on a pre-trained 3DGS can cause obvious distortion in rendered images. In our work, we propose an uncertainty- based method that constrains the perturbation of model parameters to achieve invisible watermarking for 3DGS. At the message decoding stage, the copyright messages can be reliably extracted from both 3D Gaussians and 2D rendered im- ages even under various forms of 3D and 2D distortions. We conduct extensive experiments on the Blender, LLFF, and MipNeRF-360 datasets to validate the effectiveness of our proposed method, demonstrating state-of-the-art performance on both message decoding accuracy and view synthesis quality. Xiufeng Huang, Yiu-Ming Cheung, Ka Chun Cheung, Simon See, Renjie Wan |
NeurIPS | 6 |
| 2024 | Scenedoor: An Environmental Backdoor Attack for Face RecognitionabstractFace recognition is often used for biometric validation, which has become a significant technique in our society. Due to its sensitive applications, security vulnerabilities posed by backdoor attacks have attracted considerable focus. Current backdoor attack methods use digital perturbations or physical objects as triggers, while these additional requirements make existing backdoor attacks less viable in real-world applications. To address this issue, we propose a novel backdoor attack method named Scene Backdoor (Scenedoor), which injects a 3D scene as the trigger that effectively simplifies the backdoor activation. Any person who appears in this scene will be attacked as the attacker-desired identity. Specifically, we reconstruct a 3D scene from several 2D images and then blend the facial part extracted from the input sample with the reconstructed scene to generate the poisoned image. Extensive experiments are conducted on CelenDF (v2), CelebA-HQ, and PinsFace datasets, demonstrating that Scenedoor overtakes five state-of-the-art methods in terms of effectiveness, stealthiness, and robustness. Ruofei Wang, Ziyuan Luo, Haoliang Li, Renjie Wan |
VCIP | 4 |
| 2024 | Cross-Image Disentanglement for Low-Light Enhancement in Real WorldabstractImages captured in the low-light condition suffer from low visibility and various imaging artifacts, e.g., real noise. Existing supervised algorithms for low-light image enhancement require a large set of pixel-aligned training image pairs, which are hard to prepare in practice. Though some recent unsupervised methods can alleviate such data challenges, many real world artifacts inevitably get falsely amplified in the enhanced results due to the lack of corresponding supervision. In this paper, instead of using perfectly aligned images for training, we creatively employ the misaligned real world images as the guidance, which are considerably easier to collect. Specifically, we propose a Cross-Image Disentanglement Network (CIDN) with weakly supervised learning, to separately extract cross-image brightness and image-specific content features from low/normal-light images. Based on that, CIDN can simultaneously correct the brightness and suppress image artifacts in the feature domain, which largely increases the robustness of the pixel shifts between training pairs. By considering real world corruptions, we propose a new training dataset with misaligned and noisy image pairs and its corresponding evaluation dataset. Experimental results show that our model achieves state-of-the-art performances on both the newly proposed dataset and other popular low-light datasets. The code implementation is publicly available at:https://github.com/GuoLanqing/CIDN. Lanqing Guo, Renjie Wan, Wenhan Yang, Alex Chichung Kot, Bihan Wen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Auto Diagnosis of Parkinson's Disease Via a Deep Learning Model Based on Mixed Emotional Facial ExpressionsabstractParkinson's disease (PD) is a common degenerative disease of the nervous system in the elderly. The early diagnosis of PD is very important for potential patients to receive prompt treatment and avoid the aggravation of the disease. Recent studies have found that PD patients always suffer from emotional expression disorder, thus forming the characteristics of "masked faces". Based on this, we thus propose an auto PD diagnosis method based on mixed emotional facial expressions in the paper. Specifically, the proposed method is cast into four steps: Firstly, we synthesize virtual face images containing six basic expressions (i.e., anger, disgust, fear, happiness, sadness, and surprise) via generative adversarial learning, in order to approximate the premorbid expressions of PD patients; Secondly, we design an effective screening scheme to assess the quality of the above synthesized facial expression images and then shortlist the high-quality ones; Thirdly, we train a deep feature extractor accompanied with a facial expression classifier based on the mixture of the original facial expression images of the PD patients, the high-quality synthesized facial expression images of PD patients, and the normal facial expression images from other public face datasets; Finally, with the well-trained deep feature extractor, we thus adopt it to extract the latent expression features for six facial expression images of a potential PD patient to conduct PD/non-PD prediction. To show real-world impacts, we also collected a new facial expression dataset of PD patients in collaboration with a hospital. Extensive experiments are conducted to validate the effectiveness of the proposed method for PD diagnosis and facial expression recognition. Wei Huang 0013, Renjie Wan, Peng Zhang 0005, Yufei Zha |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Occlusion-Free Scene Recovery via Neural Radiance FieldsabstractOur everyday lives are filled with occlusions that we strive to see through. By aggregating desired background information from different viewpoints, we can easily eliminate such occlusions without any external occlusion-free supervision. Though several occlusion removal methods have been proposed to empower machine vision systems with such ability, their performances are still unsatisfactory due to reliance on external supervision. We propose a novel method for occlusion removal by directly building a mapping between position and viewing angles and the corresponding occlusion-free scene details leveraging Neural Radiance Fields (NeRF). We also develop an effective scheme to jointly optimize camera parameters and scene reconstruction when occlusions are present. An additional depth constraint is applied to supervise the entire optimizaion without labeled external data for training. The experimental results on existing and newly collected datasets validate the effectiveness of our method. Our project page: https://freebutuselesssoul.github.io/occnerf. Chengxuan Zhu, Renjie Wan, Yunkai Tang, Boxin Shi |
CVPR | 2 |
| 2023 | CopyRNeRF: Protecting the CopyRight of Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have the potential to be a major representation of media. Since training a NeRF has never been an easy task, the protection of its model copyright should be a priority. In this paper, by analyzing the pros and cons of possible copyright protection solutions, we propose to protect the copyright of NeRF models by replacing the original color representation in NeRF with a watermarked color representation. Then, a distortion-resistant rendering scheme is designed to guarantee robust message extraction in 2D renderings of NeRF. Our proposed method can directly protect the copyright of NeRF models while maintaining high rendering quality and bit accuracy when compared among optional solutions. Project page: https://luo-ziyuan.github.io/copyrnerf. Ziyuan Luo, Qing Guo 0005, Ka Chun Cheung, Simon See, Renjie Wan |
ICCV | 5 |
| 2023 | Enhancing Low-Light Images Using Infrared Encoded ImagesabstractLow-light image enhancement task is essential yet challenging as it is ill-posed intrinsically. Previous arts mainly focus on the low-light images captured in the visible spectrum using pixel-wise loss, which limits the capacity of recovering the brightness, contrast, and texture details due to the small number of income photons. In this work, we propose a novel approach to increase the visibility of images captured under low-light environments by removing the in-camera infrared (IR) cut-off filter, which allows for the capture of more photons and results in improved signal-to-noise ratio due to the inclusion of information from the IR spectrum. To verify the proposed strategy, we collect a paired dataset of low-light images captured without the IR cut-off filter, with corresponding long-exposure reference images with an external filter. The experimental results on the proposed dataset demonstrate the effectiveness of the proposed method, showing better performance quantitatively and qualitatively. The dataset and code are publicly available at https://wyf0912.github.io/ELIEI/ Shulin Tian, Yufei Wang 0006, Renjie Wan, Wenhan Yang, Alex Chichung Kot, Bihan Wen |
ICIP | 3 |
| 2023 | Removing Image Artifacts From Scratched Lens ProtectorsabstractA protector is placed in front of the camera lens for mobile devices to avoid damage, while the protector itself can be easily scratched accidentally, especially for plastic ones. The artifacts appear in a wide variety of patterns, making it difficult to see through them clearly. Removing image artifacts from the scratched lens protector is inherently challenging due to the occasional flare artifacts and the co-occurring interference within mixed artifacts. Though different methods have been proposed for some specific distortions, they seldom consider such inherent challenges. In our work, we consider the inherent challenges in a unified framework with two cooperative modules, which facilitate the performance boost of each other. We also collect a new dataset from the real world to facilitate training and evaluation purposes. The experimental results demonstrate that our method outperforms the baselines qualitatively and quantitatively. The code and datasets will be released at https://github.com/wyf0912/flare-removal Yufei Wang 0006, Renjie Wan, Wenhan Yang, Bihan Wen, Lap-Pui Chau, Alex Chichung Kot |
ISCAS | 2 |
| 2023 | Gene-Induced Multimodal Pre-training for Image-Omic Classification
Xingran Xie, Renjie Wan, Qingli Li, Yan Wang 0033 |
MICCAI (6) | 3 |
| 2023 | Benchmarking Single-Image Reflection Removal AlgorithmsabstractReflection removal has been discussed for more than decades. This paper aims to provide the analysis for different reflection properties and factors that influence image formation, an up-to-date taxonomy for existing methods, a benchmark dataset, and the unified benchmarking evaluations for state-of-the-art (especially learning-based) methods. Specifically, this paper presents a SIngle-image Reflection Removal Plus dataset “SIR$^{2+}$” with the new consideration for in-the-wild scenarios and glass with diverse color and unplanar shapes. We further perform quantitative and visual quality comparisons for state-of-the-art single-image reflection removal algorithms. Open problems for improving reflection removal algorithms are discussed at the end. Our dataset and follow-up update can be found athttps://reflectionremoval.github.io/sir2data/. Renjie Wan, Boxin Shi, Haoliang Li, Yuchen Hong, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Coarse-to-fine Disentangling Demoiréing Framework for Recaptured Screen ImagesabstractRemoving the undesired moiré patterns from images capturing the contents displayed on screens is of increasing research interest, as the need for recording and sharing the instant information conveyed by the screens is growing. Previous demoiréing methods provide limited investigations into the formation process of moiré patterns to exploit moiré-specific priors for guiding the learning of demoiréing models. In this paper, we investigate the moiré pattern formation process from the perspective of signal aliasing, and correspondingly propose a coarse-to-fine disentangling demoiréing framework. In this framework, we first disentangle the moiré pattern layer and the clean image with alleviated ill-posedness based on the derivation of our moiré image formation model. Then we refine the demoiréing results exploiting both the frequency domain features and edge attention, considering moiré patterns' property on spectrum distribution and edge intensity revealed in our aliasing based analysis. Experiments on several datasets show that the proposed method performs favorably against state-of-the-art methods. Besides, the proposed method is validated to adapt well to different data sources and scales, especially on the high-resolution moiré images. Ce Wang 0007, Shengsen Wu, Renjie Wan, Boxin Shi, Ling-Yu Duan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Purifying Low-Light Images via Near-Infrared Enlightened ImageabstractCameras usually produce low-quality images under low-light conditions. Though many methods have been proposed to enhance the visibility of low-light images, they are mainly designed for illumination correction and less capable of sup-pressing the artifacts. In this paper, we propose to enhance the visibility and suppress artifacts by purifying low-light images under the guidance of the NIR enlightened image captured by using the near-infrared light as compensation. Specifically, we introduce a disentanglement framework to disentangle the structure and color components from the NIR enlightened and RGB images, respectively. Correspondingly, we introduce a new dataset with the RGB and NIR enlightened images for training and evaluation purposes. The experimental results show that our proposed method achieves promising results. Renjie Wan, Boxin Shi, Wenhan Yang, Bihan Wen, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Multim. | 1 |
| 2023 | Background Scene Recovery From an Image Looking Through Colored GlassabstractColored glass, which is commonly seen in modern city life, often degrades images taken through it with co-occurring reflection and color bias due to its optical property of simultaneous transmission, reflection, and wavelength-selective absorption. Recovering the clean background behind colored glass is inherently challenging due to the mutual interference of two degradations within a single mixture observation, and has barely been specifically considered by existing image restoration methods. In this paper, we aim at realizing faithful background scene recovery for an image taken in front of colored glass. We first analyze the formation model of mixed degradations caused by colored glass, and propose a cooperative framework to address the mutual interference problem, featuring a novel glass color invariant loss and progressive refinement. Besides, we propose a data synthesis strategy for network training. Experimental results on our newly collected real-world dataset show that our proposed method achieves state-of-the-art performance. Ce Wang 0007, Dejia Xu, Renjie Wan, Boxin Shi, Ling-Yu Duan |
IEEE Trans. Multim. | 3 |
| 2022 | Low-Light Image Enhancement with Normalizing FlowabstractTo enhance low-light images to normally-exposed ones is highly ill-posed, namely that the mapping relationship between them is one-to-many. Previous works based on the pixel-wise reconstruction losses and deterministic processes fail to capture the complex conditional distribution of normally exposed images, which results in improper brightness, residual noise, and artifacts. In this paper, we investigate to model this one-to-many relationship via a proposed normalizing flow model. An invertible network that takes the low-light images/features as the condition and learns to map the distribution of normally exposed images into a Gaussian distribution. In this way, the conditional distribution of the normally exposed images can be well modeled, and the enhancement process, i.e., the other inference direction of the invertible network, is equivalent to being constrained by a loss function that better describes the manifold structure of natural images during the training. The experimental results on the existing benchmark datasets show our method achieves better quantitative and qualitative results, obtaining better-exposed illumination, less noise and artifact, and richer colors. Yufei Wang 0006, Renjie Wan, Wenhan Yang, Haoliang Li, Lap-Pui Chau, Alex Chichung Kot |
AAAI | 2 |
| 2022 | Neural Transmitted Radiance FieldsabstractNeural radiance fields (NeRF) have brought tremendous progress to novel view synthesis. Though NeRF enables the rendering of subtle details in a scene by learning from a dense set of images, it also reconstructs the undesired reflections when we capture images through glass. As a commonly observed interference, the reflection would undermine the visibility of the desired transmitted scene behind glass by occluding the transmitted light rays. In this paper, we aim at addressing the problem of rendering novel transmitted views given a set of reflection-corrupted images. By introducing the transmission encoder and recurring edge constraints as guidance, our neural transmitted radiance fields can resist such reflection interference during rendering and reconstruct high-fidelity results even under sparse views. The proposed method achieves superior performance from the experiments on a newly collected dataset compared with state-of-the-art methods. Chengxuan Zhu, Renjie Wan, Boxin Shi |
NeurIPS | 2 |
| 2022 | GMFAD: Towards Generalized Visual Recognition via Multilayer Feature Alignment and DisentanglementabstractThe deep learning based approaches which have been repeatedly proven to bring benefits to visual recognition tasks usually make a strong assumption that the training and test data are drawn from similar feature spaces and distributions. However, such an assumption may not always hold in various practical application scenarios on visual recognition tasks. Inspired by the hierarchical organization of deep feature representation that progressively leads to more abstract features at higher layers of representations, we propose to tackle this problem with a novel feature learning framework, which is called GMFAD, with better generalization capability in a multilayer perceptron manner. We first learn feature representations at the shallow layer where shareable underlying factors among domains (e.g., a subset of which could be relevant for each particular domain) can be explored. In particular, we propose to align the domain divergence between domain pair(s) by considering both inter-dimension and inter-sample correlations, which have been largely ignored by many cross-domain visual recognition methods. Subsequently, to learn more abstract information which could further benefit transferability, we propose to conduct feature disentanglement at the deep feature layer. Extensive experiments based on different visual recognition tasks demonstrate that our proposed framework can learn better transferable feature representation compared with state-of-the-art baselines. Haoliang Li, Shiqi Wang 0001, Renjie Wan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Learning Meta Pattern for Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is essential to secure face recognition systems and has been extensively studied in recent years. Although deep neural networks (DNNs) for the FAS task have achieved promising results in intra-dataset experiments with similar distributions of training and testing data, the DNNs’ generalization ability is limited under the cross-domain scenarios with different distributions of training and testing data. To improve the generalization ability, recent hybrid methods have been explored to extract task-aware handcrafted features (e.g., Local Binary Pattern) as discriminative information for the input of DNNs. However, the handcrafted feature extraction relies on experts’ domain knowledge, and how to choose appropriate handcrafted features is underexplored. To this end, we propose a learnable network to extract Meta Pattern (MP) in our learning-to-learn framework. By replacing handcrafted features with the MP, the discriminative information from MP is capable of learning a more generalized model. Moreover, we devise a two-stream network to hierarchically fuse the input RGB image and the extracted MP by using our proposed Hierarchical Fusion Module (HFM). We conduct comprehensive experiments and show that our MP outperforms the compared handcrafted features. Also, our proposed method with HFM and the MP can achieve state-of-the-art performance on two different domain generalization evaluation benchmarks. Rizhao Cai, Zhi Li 0054, Renjie Wan, Haoliang Li, Yongjian Hu, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Multi-Scale Feature Guided Low-Light Image EnhancementabstractLow-light image enhancement aims at enlarging the intensity of image pixels to better match human perception and to improve the performance of subsequent vision tasks. While it is relatively easy to enlighten a globally low-light image, the lighting condition of realistic scenes is usually non-uniform and complex, e.g., some images may contain both bright and extremely dark regions, with or without rich features and information. Existing methods often generate abnormal light-enhancement results with over-exposure artifacts without proper guidance. To tackle this challenge, we propose a multi-scale feature guided attention mechanism in the deep generator, which can effectively perform a spatially-varying light enhancement. The attention map is fused by both the gray map and extracted feature map of the input image, to focus more on those dark and informative regions. Our baseline is an unsupervised generative adversarial network, which can be trained without any low/normal light image pair. Experimental results demonstrate the superiority in visual quality and performance of subsequent object detection over state-of-the-art alternatives. Lanqing Guo, Renjie Wan, Guan-Ming Su, Alex Chichung Kot, Bihan Wen |
ICIP | 2 |
| 2021 | Unsupervised Domain Adaptation in the Wild via Disentangling Representation Learning
Haoliang Li, Renjie Wan, Shiqi Wang 0001, Alex Chichung Kot |
Int. J. Comput. Vis. | 2 |
| 2021 | Face Image Reflection Removal
Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot |
Int. J. Comput. Vis. | 1 |
| 2020 | Reflection Scene Separation From a Single ImageabstractFor images taken through glass, existing methods focus on the restoration of the background scene by regarding the reflection components as noise. However, the scene reflected by glass surface also contains important information to be recovered, especially for the surveillance or criminal investigations. In this paper, instead of removing reflection components from the mixture image, we aim at recovering reflection scenes from the mixture image. We first propose a strategy to obtain such ground truth and its corresponding input images. Then, we propose a two-stage framework to obtain the visible reflection scene from the mixture image. Specifically, we train the network with a shift-invariant loss which is robust to misalignment between the input and output images. The experimental results show that our proposed method achieves promising results. Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 1 |
| 2020 | Domain Generalization for Medical Imaging Classification with Linear-Dependency RegularizationabstractRecently, we have witnessed great progress in the field of medical imaging classification by adopting deep neural networks. However, the recent advanced models still require accessing sufficiently large and representative datasets for training, which is often unfeasible in clinically realistic environments. When trained on limited datasets, the deep neural network is lack of generalization capability, as the trained deep neural network on data within a certain distribution (e.g. the data captured by a certain device vendor or patient population) may not be able to generalize to the data with another distribution. In this paper, we introduce a simple but effective approach to improve the generalization capability of deep neural networks in the field of medical imaging classification. Motivated by the observation that the domain variability of the medical images is to some extent compact, we propose to learn a representative feature space through variational encoding with a novel linear-dependency regularization term to capture the shareable information among medical data collected from different domains. As a result, the trained neural network is expected to equip with better generalization capability to the ``unseen" medical data. Experimental results on two challenging medical imaging classification tasks indicate that our method can achieve better cross-domain generalization capability compared with state-of-the-art baselines. Haoliang Li, Yufei Wang 0006, Renjie Wan, Shiqi Wang 0001, Tie-Qiang Li, Alex Chichung Kot |
NeurIPS | 3 |
| 2020 | Improving Robustness of DNNs against Common Corruptions via Gaussian Adversarial TrainingabstractDeep neural networks have demonstrated tremendous success in image classification, but their performance sharply degrades when evaluated on slightly different test data (e.g., data with corruptions). To address these issues, we propose a minimax approach to improve common corruption robustness of deep neural networks via Gaussian Adversarial Training. To be specific, we propose to train neural networks with adversarial examples where the perturbations are Gaussian-distributed. Our experiments show that our proposed GAT can improve neural networks' robustness to noise corruptions more than other baseline methods. It also outperforms the state-of-the-art method in improving the overall robustness to common corruptions. Chenyu Yi, Haoliang Li, Renjie Wan, Alex Chichung Kot |
VCIP | 3 |
| 2020 | The Enhancement of Underexposed Images with Blurred ReflectanceabstractThe images captured in the low-light conditions always suffer from low visibility. Enhancing the visibility of the low-light image is of broad application to various computer vision tasks. Based on the classical Retinex model, previous methods assume the reflectance components as a well-exposed image. In this paper, we introduce the blurring distortion into the Retinex model to cover more general and challenging scenarios. We further propose a two-stage framework to extract the reflectance images and remove the blurring distortion separately. Specifically, we optimize the whole network by embedding a mechanism robust to the pixel misalignment in the training dataset. The experimental results show that our proposed method achieves promising results. Jinchao Zhou, Renjie Wan, Haoliang Li, Alex Chichung Kot |
VCIP | 2 |
| 2020 | CoRRN: Cooperative Reflection Removal NetworkabstractRemoving the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their limited description capability to the properties of real-world reflections. In this paper, we propose a network with the feature-sharing strategy to tackle this problem in a cooperative and unified framework, by integrating image context information and the multi-scale gradient information. To remove the strong reflections existed in some local regions, we propose a statistic loss by considering the gradient level statistics between the background and reflections. Our network is trained on a new dataset with 3250 reflection images taken under diverse real-world scenes. Experiments on a public benchmark dataset show that the proposed method performs favorably against state-of-the-art methods. Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Heterogeneous Transfer Learning via Deep Matrix Completion with Adversarial Kernel EmbeddingabstractHeterogeneous Transfer Learning (HTL) aims to solve transfer learning problems where a source domain and a target domain are of heterogeneous types of features. Most existing HTL approaches either explicitly learn feature mappings between the heterogeneous domains or implicitly reconstruct heterogeneous cross-domain features based on matrix completion techniques. In this paper, we propose a new HTL method based on a deep matrix completion framework, where kernel embedding of distributions is trained in an adversarial manner for learning heterogeneous features across domains. We conduct extensive experiments on two different vision tasks to demonstrate the effectiveness of our proposed method compared with a number of baseline methods. Haoliang Li, Sinno Jialin Pan, Renjie Wan, Alex Chichung Kot |
AAAI | 3 |
| 2019 | Learning to Jointly Generate and Separate ReflectionsabstractExisting learning-based single image reflection removal methods using paired training data have fundamental limitations about the generalization capability on real-world reflections due to the limited variations in training pairs. In this work, we propose to jointly generate and separate reflections within a weakly-supervised learning framework, aiming to model the reflection image formation more comprehensively with abundant unpaired supervision. By imposing the adversarial losses and combinable mapping mechanism in a multi-task structure, the proposed framework elegantly integrates the two separate stages of reflection generation and separation into a unified model. The gradient constraint is incorporated into the concurrent training process of the multi-task learning as well. In particular, we built up an unpaired reflection dataset with 4,027 images, which is useful for facilitating the weakly-supervised learning of reflection removal model. Extensive experiments on a public benchmark dataset show that our framework performs favorably against state-of-the-art methods and consistently produces visually appealing results. Daiqian Ma, Renjie Wan, Boxin Shi, Alex Chichung Kot, Ling-Yu Duan |
ICCV | 2 |
| 2019 | Learning to Remove Reflections for Text ImagesabstractText images taken behind a piece of glass in the wild are largely contaminated by reflections. Directly applying existing reflection removal methods on text images with reflections cannot recover clear and correct text contents due to the ignorance of special characteristics of texts. This paper proposes a stacked framework to solve the text image reflection removal problem by specifically considering the regional properties of reflection and embedding the specific text priors into the estimation process in a unified manner. Experiment results on a newly collected dataset demonstrate that the proposed method outperforms state-of-the-art methods in recovering visually pleasant reflection-free images and recognizable text features. Ce Wang 0007, Renjie Wan, Feng Gao 0014, Boxin Shi, Ling-Yu Duan |
ICME | 2 |
| 2019 | See Through the Windshield from Surveillance CameraabstractThis paper attempts to address the challenging task of seeing through the windshield images captured by surveillance cameras in the wild. Such images usually have very low visibility due to heterogeneous degradations caused by blur, haze, reflection, noise etc., which makes existing image enhancing methods inapplicable. We propose a windshield image restoration generative adversarial network (WIRE-GAN) to restore and enhance the visibility of windshield images. We adopt the weakly supervised framework based on the generative model, which has effectively released the request of paired training data for a specific type of degradation. To generate more semantically consistent results even in extreme lighting conditions, we introduce a novel content-preserving strategy into the proposed weakly-supervised framework. To make the image restoration more reliable, the WIRE-GAN network constructs a sort of content-aware embedding space and enforces the constraint of the restored windshield images being closer to the original input in the embedding space. Moreover, we collect a large-scale windshield image dataset (WIRE dataset) to validate the advantage of our method in improving the image quality, and further evaluate the impact of windshield restoration on the vehicle ReID performance. Daiqian Ma, Renjie Wan, Ce Wang 0007, Boxin Shi, Ling-Yu Duan |
ACM Multimedia | 3 |
| 2018 | CRRN: Multi-Scale Guided Concurrent Reflection Removal NetworkabstractRemoving the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their limited description capability to the properties of real-world reflections. In this paper, we propose the Concurrent Reflection Removal Network (CRRN) to tackle this problem in a unified framework. Our proposed network integrates image appearance information and multi-scale gradient information with human perception inspired loss function, and is trained on a new dataset with 3250 reflection images taken under diverse real-world scenes. Extensive experiments on a public benchmark dataset show that the proposed method performs favorably against state-of-the-art methods. Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
CVPR | 1 |
| 2018 | Region-Aware Reflection Removal With Unified Content and Gradient PriorsabstractRemoving the undesired reflections in images taken through the glass is of broad application to various image processing and computer vision tasks. Existing single image based solutions heavily rely on scene priors such as separable sparse gradients caused by different levels of blur, and they are fragile when such priors are not observed. In this paper, we notice that strong reflections usually dominant a limited region in the whole image, and propose a Region-aware Reflection Removal (R3) approach by automatically detecting and heterogeneously processing regions with and without reflections. We integrate content and gradient priors to jointly achieve missing contents restoration as well as background and reflection separation in a unified optimization framework. Extensive validation using 50 sets of real data shows that the proposed method outperforms state-of-the-art on both quantitative metrics and visual qualities. Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Wen Gao 0001, Alex Chichung Kot |
IEEE Trans. Image Process. | 1 |
| 2017 | Benchmarking Single-Image Reflection Removal AlgorithmsabstractRemoving undesired reflections from a photo taken in front of a glass is of great importance for enhancing the efficiency of visual computing systems. Various approaches have been proposed and shown to be visually plausible on small datasets collected by their authors. A quantitative comparison of existing approaches using the same dataset has never been conducted due to the lack of suitable benchmark data with ground truth. This paper presents the first captured Single-image Reflection Removal dataset ‘SIR2’ with 40 controlled and 100 wild scenes, ground truth of background and reflection. For each controlled scene, we further provide ten sets of images under varying aperture settings and glass thicknesses. We perform quantitative and visual quality comparisons for four state-of-the-art single-image reflection removal algorithms using four error metrics. Open problems for improving reflection removal algorithms are discussed at the end. Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
ICCV | 1 |
| 2017 | Sparsity based reflection removal using external patch searchabstractReflection removal aims at separating the mixture of the desired background scenes and the undesired reflections, when the photos are taken through the glass. It has both aesthetic and practical applications which can largely improve the performance of many multimedia tasks. Existing reflection removal approaches heavily rely on scene priors such as separable sparse gradients brought by different levels of blur, and they easily fail when such priors are not observed in many real scenes. Sparse representation models and nonlocal image priors have shown their effectiveness in image restoration with self similarity. In this work, we propose a reflection removal method benefited from the sparsity and nonlocal image prior as a unified optimization framework. We leverage the retrieved image patch from an external database to overcome the limited prior information in the input mixture image and self similarity search. The experimental results show that our proposed model performs better than the existing state-of-the-art reflection removal method for both objective and subjective image qualities. Renjie Wan, Boxin Shi, Ah-Hwee Tan, Alex Chichung Kot |
ICME | 1 |
| 2016 | Depth of field guided reflection removalabstractReflection removal aims at separating the mixture of the desired scene and the undesired reflections. Locating reflection and background edges is a key step for reflection removal. In this paper, we present a visual depth guided method to remove reflections. Our idea is to use Depth of Field (DoF) to label the background and reflection edges. We propose a DoF confidence map where pixels with higher DoF values are assumed to belong to the desired background components. Moreover, we observe that images with different resolutions show different properties in the DoF map. Thus, we introduce a multi-scale DoF computing strategy to classify edge pixels more efficiently. Based on the results of edge classification, the background and reflection layers can be separated. Experimental results validate the effectiveness of our method using real-world photos. Renjie Wan, Boxin Shi, Ah-Hwee Tan, Alex Chichung Kot |
ICIP | 1 |