VLDB 2026 Research / reviewers in the wild / expert
Yuenan Li 0001
dblp:46/2153-1 · also Yue-Nan Li 0001
· DBLP profile ↗
17ranked-venue papers
12as first author
8since 2021 · last 2026
0000-0001-8932-851XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 12 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optical-Cue-Guided Diffusion Probabilistic Model for Reflection RemovalabstractThis article presents a novel reflection removal algorithm that integrates a flash-based optical cue into a diffusion model to control the recovery of the transmission image. The algorithm accepts a pair of ambient and flash images as inputs, and a flash-only image, which corresponds to the one captured with flash as the sole illumination source, is derived from the inputs. In light of the reflection-free nature of the flash-only image, we use it to guide the diffusion model to reconstruct the structures of the transmission image. A feature distillation scheme is designed to infer the chromatic attributes of the transmission image from the ambient image, and the features are used to modulate the generative priors learned by the diffusion model. We use time-aware strategies to ensure the synchronization between feature distillation and the dynamic image generation process of the diffusion model. The performance of the proposed algorithm is sequentially optimized in latent and pixel spaces. We also develop a plug-and-play fidelity-enhancing module (FEM) and integrate it into the proposed model to enable the faithful reconstruction of fine-granular visual characteristics of the target scene and reduce artifacts. Comparative experiments demonstrate that the proposed algorithm shows superior quantitative and qualitative performance over state-of-the-art methods in real-world scenarios. By leveraging the optical cue and the generative capability of the diffusion model, the algorithm can accurately restore the visual details of the transmission image even in the presence of strong reflections, and it also exhibits satisfactory robustness against nonlinear image representation and misalignment. Yuenan Li 0001, Xiaoliang Chang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Robust Image Fingerprinting Algorithm Based on Local and Global Content DependenciesabstractImage fingerprinting summarizes the unique visual characteristics of an image into a robust and compact ID for content identification. This technique is widely adopted by social networks to identify unauthorized uploads of copyrighted content. In this letter, we propose a deep neural network based image fingerprinting algorithm, where a neural network is designed to capture the short and long-range structural dependencies of an image and compress the representative features into fingerprints. The training algorithm optimizes the content identification accuracy of the fingerprinting model from a hypothesis-testing perspective. We propose a differentiable training objective for minimizing the error rate of the hypothesis-testing problem. Since real applications prefer binary fingerprints, we also develop an adversarial training scheme to progressively force the outputs of the neural network to approach binary states, aiming to minimize the performance loss caused by fingerprint binarization. The experimental results show that the proposed algorithm achieves more accurate content identification than state-of-theart methods and is insensitive to fingerprint binarization. Yuenan Li 0001, Wei Zhang 0332 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Lightweight Neural Network for Enhancing Imaging Performance of Under-Display CameraabstractUnder-Display Camera (UDC) is an emerging feature of cellphone. This technology makes full-screen cellphones possible by hiding the front-facing camera below the display panel, which is in contrast to the conventional designs that place the camera in a bezel or punch-hole on the screen border. However, this novel imaging paradigm also causes degradation. The display panel attenuates and diffracts incoming light, so the images captured by UDC contain multiple artifacts, such as blurring, color shift, and low intensities. This paper proposes a lightweight deep learning approach to restore UDC images in a blind setting. The restoration network uses cross-scale modulation to exploit complementary information from multi-scale representations and capture the self-similarity across scales, aiming to find the cues for recovering distortion-free images. To facilitate the deployment of this scheme across mobile devices, especially on those with limited memory space and computing power, we compress the restoration network by reducing architectural redundancy. An adaptive distillation algorithm is designed to exploit knowledge from a pre-trained full-size model. The proposed work also interprets the behavior of the neural network in utilizing local and non-local information to restore UDC images. The proposed algorithm is evaluated over three datasets of the images captured by the cameras below different types of display panels. The results of comparative experiments demonstrate that our algorithm shows comparable or superior performance to the competing ones that are much heavier in parameter amount and computational complexities. Yuenan Li 0001, Zetao Shi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Reflection Removal via Recurrent Learning Guided by Physics Prior and Focal Perceptual LossabstractRemoving reflection from a single image is an ill-posed problem, while exploiting physics priors can ease this inverse problem. In this paper, we integrate a physics prior of reflection-free images derived from flash illumination into deep learning. The algorithm first estimates an approximation of the transmission scene (i.e., flash-only image) from a pair of images captured with and without flash illumination. We design two collaborative neural networks to make recurrent recovery of the transmission and reflection scenes at increasing resolutions under the guidance of the physics prior. The neural networks learn the cues for scene separation and reconstruction by embedding multi-scale feature extraction components into a nested topology. We also propose a focal perceptual loss for penalizing the artifacts in output images, where the perceptual distances computed in different feature spaces are weighted adaptively to emphasize the hard-to-restore visual attributes. The comparative experiments demonstrate that the proposed algorithm performs better than the state-of-the-art methods in real-world scenarios, and it outperforms the currently best-performing algorithm by 1.5 dB in PSNR. To understand the mechanism behind the performance enhancement brought by the physics prior, we use the attribution-based model interpretation approach to quantify the pixel-to-pixel influence of the flash-only image on the results of reflection removal. The results of model interpretation reveal that the physics prior plays a significant role in dealing with non-uniform and strong reflections. Zetao Shi, Yuenan Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Cross-Domain Joint Dictionary Learning for ECG Inference From PPGabstractThe inverse problem of inferring clinical gold-standard electrocardiogram (ECG) from photoplethysmogram (PPG) that can be measured by affordable wearable Internet of Healthcare Things (IoHT) devices is a research direction receiving growing attention. It combines the easy measurability of PPG and the rich clinical knowledge of ECG for long-term continuous cardiac monitoring. The prior art for reconstruction using a universal basis, such as discrete cosine transform (DCT), has limited fidelity for uncommon ECG shapes due to the lack of representative power. To better utilize the data and improve data representation, we design two dictionary learning frameworks, the cross-domain joint dictionary learning (XDJDL), and the label-consistent XDJDL (LC-XDJDL), to further improve the ECG inference quality and enrich the PPG-based diagnosis knowledge. Building on the K-SVD technique, the proposed joint dictionary learning frameworks extend the expressive power by optimizing simultaneously a pair of signal dictionaries for PPG and ECG with the transforms to relate their sparse codes and disease information. The proposed models are evaluated with a variety of PPG and ECG morphologies from two benchmark datasets that cover various age groups and disease types. The results show the proposed frameworks achieve better inference performance than previous methods with average Pearson coefficients being 0.88 using XDJDL and 0.92 using LC-XDJDL, suggesting an encouraging potential for ECG screening using PPG based on the proactively learned PPG-ECG relationship. By enabling the dynamic monitoring and analysis of the health status of an individual, the proposed frameworks contribute to the emerging digital twins paradigm for personalized healthcare. Xin Tian 0018, Qiang Zhu 0015, Yuenan Li 0001, Min Wu 0001 |
IEEE Internet Things J. | 3 |
| 2022 | Image Reflection Removal via Contextual Feature Fusion Pyramid and Task-Driven RegularizationabstractIn this paper, we propose a deep neural network for single image reflection removal. More specifically, we design a convolutional-grid module and take it as the building block of a feature fusion pyramid. The module leverages the combination effect of the grid topology to create a rich ensemble of receptive fields. Embedding the modules into a pyramidal architecture further expands the coverage of receptive fields. Another benefit of the pyramid is to fuse the multi-scale features learned by the modules locating at the ascending and descending pathways. The rich diversity of features helps the neural network analyze the contexts around overlapping objects at various spatial ranges and harvest the cues for layer separation. The proposed work also exploits useful semantic cues from the hyper-column descriptors generated by a pre-trained VGG-19 model to reduce the ambiguity of layer separation. In light of the low correlation between background and reflection layers, we design a channel-correlation based conditional discriminator to penalize residual reflection. The discriminator uses channel-wise attention to screen the features that can distinguish real background images from estimated ones. This paper also presents a task-driven regularization strategy. The high sensitivity of semantic segmentation to reflection is exploited for assessing the completeness of reflection removal. Training with this regularization strategy can boost the performance of both reflection removal and high-level task. The comparison against state-of-the-art algorithms on four public benchmark datasets demonstrates that this work exhibits superior performance in handling the complex reflections in wild scenarios. The proposed network architecture is also applicable to haze removal, which is another ill-posed layer separation problem, and has shown encouraging performance. Yuenan Li 0001, Qixin Yan, Kuangshi Zhang, Haoyu Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Single Image Dehazing via Semi-Supervised Domain Translation and Architecture SearchabstractData-driven methods have demonstrated great potential in single image dehazing, most of which are trained using the hazy images synthesized by the atmospheric scattering model. However, the model cannot accurately describe the complicated degradation caused by haze. At the same time, it is impractical to obtain the pixel-to-pixel aligned clear-scene counterparts of real-world hazy images for supervised training. In this letter, we formulate dehazing as a semi-supervised domain translation problem. For better generalization, two auxiliary domain translation tasks are designed to capture the properties of real-world haze and align synthetic hazy images to real-world ones to reduce the domain gap. Dehazing and the auxiliary tasks are conducted in shared latent spaces by a unified framework, and we use differential optimization to search the architectures of the framework. We evaluate the efficacy of the proposed work using one synthetic and three real-world benchmarks that cover the challenging cases in wild scenarios, and it outperforms state-of-the-art algorithms on these benchmarks. The benefits brought by auxiliary domain translation tasks and architecture search are also verified by ablation experiments. Kuangshi Zhang, Yuenan Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Deep Dehazing Network With Latent Ensembling Architecture and Adversarial LearningabstractMost existing dehazing algorithms recover haze-free image by solving the hazy imaging model using estimated transmission map and global atmospheric light. However, inaccurate estimation of these variables and the strong assumptions of imaging model result in unrealistic dehazing results. In this paper, we use the adversarial game between a pair of neural networks to accomplish end-to-end photo-realistic dehazing. To avoid uniform contrast enhancement, the generator learns to simultaneously restore haze-free image and capture the non-uniformity of haze. The modules for the two tasks are assembled in sequential and parallel manners to enable information sharing at different levels, and the architecture of the generator implicitly forms an ensemble of dehazing models that allows for feature selection. A multi-scale discriminator competes with the generator by learning to detect dehazing artifacts and the inconsistency between dehazed image and the spatial variation of haze. Unlike existing works that penalize dehazing artifacts via hand-crafted loss, the proposed algorithm uses the identity mapping in the space of clear-scene images to regularize data-driven dehazing. The proposed work also addresses the adaptability of data-driven dehazing to high-level computer vision task. We propose a task-driven training strategy that can optimize the object detection performance on dehazed images without updating the parameters of object detector. Performance of the proposed algorithm is assessed on the RESIDE, I-Haze, and O-Haze benchmarks. The comparison with ten state-of-the-art algorithms shows that the proposed work is the best performer in most competitions. Yuenan Li 0001, Qixin Yan, Kuangshi Zhang |
IEEE Trans. Image Process. | 1 |
| 2020 | Cross-Domain Joint Dictionary Learning for ECG Reconstruction from PPGabstractAn emerging research direction considers the inverse problem of inferring electrocardiogram (ECG) from photoplethysmogram (PPG) to bring about the synergy between the easy measurability of PPG and the rich clinical knowledge of ECG to facilitate preventive healthcare. Previous reconstruction using a universal basis has limited accuracy due to the lack of rich representative power. This paper proposes a cross-domain joint dictionary learning (XDJDL) framework to maximize the expressive power for the two cross-domain signals. Building on K-SVD technique, XDJDL optimizes simultaneously the PPG and ECG signal representations and the transform between them, enabling the joint learning of a pair of signal dictionaries with a transform to characterize the relation between their sparse codes. The proposed model is evaluated with 34,000+ ECG/PPG cycle pairs containing a variety of ECG morphologies and cardiovascular diseases. Experimental results validate the accuracy and the generality of the proposed algorithm, suggesting an encouraging potential for disease screening using PPG measurement based on the proactive learned PPG-ECG relationship. Xin Tian 0018, Qiang Zhu 0015, Yuenan Li 0001, Min Wu 0001 |
ICASSP | 3 |
| 2020 | Robust and Secure Image Fingerprinting Learned by Neural NetworkabstractImage fingerprinting is a technique that summarizes the perceptual characteristics of a digital image into an invariant digest, and it is one of the most effective solutions for digital rights management. Most conventional fingerprinting algorithms were developed by assembling manually designed feature extractor and quantizer, which requires extensive expert knowledge and may not capture the intrinsic or abstract visual characteristics of the digital image. Focusing on content identification related applications, we propose a data-driven image fingerprinting algorithm in this paper, where neural network is trained to automatically discover the optimal mapping from image to fingerprint. To ameliorate the difficulty of training, we start by training the fingerprint-computation network in a layer-wise manner to progressively improve its robustness against content-preserving distortions. Initialized by the states learned by layer-wise training, the network is then re-trained as a holistic unit, with the objective of maximizing its content identification accuracy. Moreover, we also develop a key-dependent version of the neural network-based fingerprinting algorithm. By quantifying its security using information-theoretic metrics, we have proved that the hierarchical architecture of neural network is beneficial to the security of fingerprinting algorithm. The experimental results on a large testing database show that the proposed work exhibits much higher content identification accuracy than state-of-the-art algorithms, and its execution speed is in the millisecond time scale. Yuenan Li 0001, Dongdong Wang 0005, Linlin Tang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Robust Image Fingerprinting via Distortion-Resistant Sparse CodingabstractContent fingerprinting recently emerges as an effective nonintrusive solution for copyright protection. Fingerprinting algorithm maps the perceptual contents of media file to an invariant digest, so that unauthorized copies can be identified via fingerprint comparison. This letter presents a distortion-resistant sparse coding strategy for image fingerprinting that simulates the hierarchical information processing flow of visual system. Sparse coding, which seeks a small set of atoms that can best represent input signal, helps fingerprinting algorithm detect the intrinsic visual features of image. However, the high freedom of atom selection makes sparse coding sensitive to distortion. In this letter, several measures are applied on sparse coding and dictionary learning to jointly ensure the invariance of fingerprint, such as imposing the neighborhood-priority principle on atom selection, regulating the layout of atoms, and forcing sparse codes to preserve the distance in the image space. Content identification performance of the proposed work was tested on a database of 219 000 images. The error rate of the proposed algorithm is at least ten times lower than state-of-the-arts, and satisfactory performance was observed even under extremely low bit budget. Yuenan Li 0001, Linlin Guo |
IEEE Signal Process. Lett. | 1 |
| 2017 | Robust and compact video descriptor learned by deep neural networkabstractIn this paper, we propose to extract robust video descriptor by training deep neural network to automatically capture the intrinsic visual characteristics of digital video. More specifically, we first train a conditional generative model to capture the spatio-temporal correlations among visual contents and represent them as an intermediate descriptor. A nonlinear encoder, with the functions of dimension reduction and error correcting, is then trained to learn a compressed yet more robust representation of the intermediate descriptor. The cascade of the conditional generative model and the encoder constitutes the building block of the deep network for learning video descriptor. As a post-processing component, the top layers of the network are trained to optimize the robustness and discriminative capability of the output descriptor. Experimental results on benchmark databases confirm that the descriptor learned by deep neural network shows excellent robustness against photometric, geometric, temporal and combined distortions, and it can attain an F1score of 0.982 in content identification, which is much higher than hand-engineered descriptors. Yuenan Li 0001, Xue Piao Chen |
ICASSP | 1 |
| 2016 | Robust image hashing based on low-rank and sparse decompositionabstractWe propose in this paper a low-rank and sparse decomposition based image hashing algorithm, aiming to summarize the structural information and sparse salient components of digital image to compact digest. More specifically, we leverage compressive sampling and random projection to separately aggregate the low-rank approximation of input image and the spatial layout of salient components into binary hash. Owing to its capability of capturing and fusing intrinsic visual characteristics, the proposed work demonstrates high robustness and discriminability. As observed in content identification experiments, it shows much higher accuracy than state-of-the-art algorithms. Furthermore, we also analytically evaluate the security of the proposed hashing algorithm using the entropy based metric, and its performance in content identification is analyzed using the channel coding theorem. Yuenan Li 0001 |
ICASSP | 1 |
| 2015 | Robust Content Fingerprinting Algorithm Based on Sparse CodingabstractContent fingerprinting is a powerful solution for media indexing, searching and digital right management, in which the perceptual content of digital media is summarized to a robust and discriminative digest. In this letter, we develop a general paradigm for image fingerprinting by exploiting the capability of sparse coding in capturing the visual characteristics of digital image. Furthermore, the impact of the dictionary for sparse coding on the performance of fingerprinting algorithm is analyzed. Accordingly, the problem of dictionary learning is studied in the context of content fingerprinting by incorporating the robustness and discriminability requirements. Comparative experiments indicate that the proposed work exhibits much higher content identification accuracy than the state-of-the-art ones, and the dictionary learned by the proposed work can substantially improve the performance of fingerprinting algorithm. In addition, our algorithm is highly efficient, and its average fingerprint computation time is less than 0.024s. Yuenan Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2015 | Robust Image Hashing Based on Selective Quaternion InvarianceabstractRobust image hashing, which maps the perceptual contents of image to a short digest, is a key technique for tackling the challenges of content-based indexing, searching and copyright protection. In this letter, we propose a quaternion invariance based hashing algorithm that can fuse complementary visual features to compact hash. The proposed algorithm leverages quaternion polar cosine transform to holistically capture the spatial and chromatic characteristics of digital image, and rotation-invariant features are derived from the phase information in the quaternion frequency domain. In addition, we also propose an information theoretic based metric for feature quality assessment and a greedy strategy to select the optimal subset of features for hash computation. Extensive experiments over a large database demonstrate that the proposed work shows higher content identification accuracy than most competing algorithms, and the resulting hash is compact and easy to compute. Yuenan Li 0001, Yuting Su 0001 |
IEEE Signal Process. Lett. | 1 |
| 2013 | Quaternion Polar Harmonic Transforms for Color ImagesabstractRobust and compact content representation is a fundamental problem in image processing. The recently proposed polar harmonic transforms (PHTs) have provided a set of powerful tools for image representation. However, two-dimensional transforms cannot handle color image in a holistic manner. To extend the nice properties of PHTs to color image processing, we generalize PHTs from the complex field to hypercomplex field in this letter, and quaternion polar harmonic transforms (QPHTs) are developed based on quaternion algebra. Furthermore, the properties of QPHTs are studied via quaternion computation, including the orthogonality of quaternion kernels, the relationships between different transforms and their rotation invariance. Experimental results reveal that compared with complex PHTs, the quaternion transforms can make a more compact and discriminative representation of color image. Moreover, QPHTs can well capture the chromatic features and exploit the inter-channel redundancies of color image. Yuenan Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2012 | Robust Image Hashing Based on Random Gabor Filtering and Dithered Lattice Vector QuantizationabstractIn this paper, we propose a robust-hash function based on random Gabor filtering and dithered lattice vector quantization (LVQ). In order to enhance the robustness against rotation manipulations, the conventional Gabor filter is adapted to be rotation invariant, and the rotation-invariant filter is randomized to facilitate secure feature extraction. Particularly, a novel dithered-LVQ-based quantization scheme is proposed for robust hashing. The dithered-LVQ-based quantization scheme is well suited for robust hashing with several desirable features, including better tradeoff between robustness and discrimination, higher randomness, and secrecy, which are validated by analytical and experimental results. The performance of the proposed hashing algorithm is evaluated over a test image database under various content-preserving manipulations. The proposed hashing algorithm shows superior robustness and discrimination performance compared with other state-of-the-art algorithms, particularly in the robustness against rotations (of large degrees). Yuenan Li 0001, Zheming Lu 0001, Ce Zhu, Xiamu Niu |
IEEE Trans. Image Process. | 1 |