VLDB 2026 Research / reviewers in the wild / expert
Di Li 0006
dblp:96/1434-6
· DBLP profile ↗
11ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0001-8059-8783ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Image and video processing · 82% Computational photography and imaging · 18% | |
| Artificial intelligence
1 paper |
Generative modeling · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing
image enhancement |
1.6 | 2 | 2025 | Rethinking Affine Transform for Efficient Image Enhancement: A Color Space Perspective · IEEE Trans. Multim. 2025 Learning Deep Representations for Photo Retouching · IEEE Trans. Multim. 2024 |
Image and video processing › color image processing
color space transformation |
0.9 | 1 | 2025 | Rethinking Affine Transform for Efficient Image Enhancement: A Color Space Perspective · IEEE Trans. Multim. 2025 |
Image and video processing › image enhancement
real-time image enhancement |
0.9 | 1 | 2025 | Rethinking Affine Transform for Efficient Image Enhancement: A Color Space Perspective · IEEE Trans. Multim. 2025 |
Computational photography and imaging › image aesthetics › aesthetic image enhancement
image retouching |
0.8 | 1 | 2024 | Learning Deep Representations for Photo Retouching · IEEE Trans. Multim. 2024 |
Machine learning › Generative modeling
generative adversarial network |
0.2 | 1 | 2024 | Learning Deep Representations for Photo Retouching · IEEE Trans. Multim. 2024 |
Methods — techniques the papers use, named apart from their topics
representation learning · 1.5one-way loss · 1.5multi-scale GAN · 1.5hierarchical transformer · 0.9bilateral grid learning · 0.9affine transform · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Real-Valued Discrete Fractional Hadamard Transform: Fast Algorithms and ImplementationsabstractThis paper introduces a new real-valued discrete fractional Hadamard transform (RFHT) designed to address the issues of high computational complexity and large storage demands found in the traditional discrete fractional Hadamard transform (FHT) and its modifications. Additionally, a fast algorithm for the RFHT is developed, and the corresponding computational complexity analysis demonstrates that for sizes ranging from$N = 2$to 1024, the proposed fast algorithm can reduce the number of multiplications and additions by up to 50.0% and 83.3%, respectively, compared to state-of-the-art fast algorithms. The comparison results show that the RFHT also has lower execution time and power consumption. Furthermore, due to its real-valued property, the RFHT has been applied in an image watermarking system and implemented on real iOS devices, demonstrating enhanced information security. Lower execution and power consumption, reduced storage and transmission requirements, and superior information protection make the RFHT a superior candidate compared to WHT, FHT, and its modifications. Zi-Chen Fan, Di Li 0006, Susanto Rahardja |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | Rethinking Affine Transform for Efficient Image Enhancement: A Color Space PerspectiveabstractIn recent years, we have observed significant advancements in learning-based techniques for image-enhancement tasks. However, most of the existing methods are either purely based on image-to-image convolutional neural networks, which cannot handle high-resolution images in real-time, or resort to 3D Lookup Tables, which fall short of local tone adjustments. In this paper, we rethink affine transform through a color space perspective, and then propose AttnBL (Attentional Bilateral Grid Learning), a novel hybrid image enhancement algorithm to process ultra-high-definition images in real-time. Our algorithm consists of two paths, the low-resolution chroma prediction path that aims to learn the chroma coefficients and the full-resolution luma adaptation path that aims to preserve brightness details. Specifically, we propose a carefully designed hierarchical transformer to capture the global information in an efficient way and introduce a feature extraction module to adaptively learn a luma guidance for bilateral upsampling. Our algorithm can process a 4K-resolution image in 20 milliseconds. This efficiency provides a practical solution for high resolution real-time preview. Without bells and whistles, our model outperforms previous state-of-the-art methods on two well-known datasets in image enhancement tasks both quantitatively and qualitatively. Our analysis also provides some interesting findings that may enlighten further studies. Di Li 0006, Susanto Rahardja |
IEEE Trans. Multim. | 1 |
| 2024 | A Novel Discrete Fractional Complex Hadamard Transform for Medical Image EncryptionabstractThis paper introduces a new discrete fractional complex Hadamard transform (FCHT) and its generalized form, the multiple-parameter FCHT (MFCHT). The MFCHT is applied to the medical image encryption. Both subjective observations and objective evaluations are conducted to validate the effectiveness of the proposed algorithm, and the simulation results demonstrate that the proposed MFCHT outperforms previous transform-based algorithms, including the discrete fractional Fourier transform and the discrete fractional Hadamard transform, in terms of robust image preservation against blind attacks in the medical image encryption. Zi-Chen Fan, Di Li 0006, Susanto Rahardja |
ICASSP | 2 |
| 2024 | Unsupervised Image Enhancement via Contrastive LearningabstractRecent years have witnessed significant achievements for image enhancement tasks. However, many advanced algorithms are trained in a supervised manner and thus rely on a huge collection of paired data, for which the collection is itself a challenge especially for real-world scenarios. We address this issue by proposing a novel GAN framework designed for unsupervised training. To be specific, our approach introduces a contrastive loss to ensure that the content remains consistent across multiple scales in both input and output representations. In addition, we propose a multi-scale discriminator to strengthen the adversarial learning. Extensive experiments conducted in this paper showed that our algorithm achieved state-of-the-art performance on MIT-Adobe-FiveK dataset both quantitively and qualitatively. Di Li 0006, Susanto Rahardja |
ISCAS | 1 |
| 2024 | Contrastive learning for deep tone mapping operator
Di Li 0006, Mou Wang, Susanto Rahardja |
Signal Process. Image Commun. | 1 |
| 2024 | Efficient Computation for Discrete Fractional Hadamard TransformabstractThis paper introduces a new fast algorithm for the discrete fractional Hadamard transform (FHT). The proposed algorithm demonstrates superior computational efficiency. For data lengths ranging from$2 \leq N \leq 1024$, our algorithm achieves a reduction in the number of multiplications by up to 96.53%, 81.82%, 33.33%, and 90% compared to four existing fast algorithms for the FHT. Additionally, we compare the execution times with those of existing fast algorithms, and the results show that the proposed algorithm has better performance. The reduced computational complexity makes the proposed algorithm a potential candidate for calculating the FHT. Zi-Chen Fan, Di Li 0006, Susanto Rahardja |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Learning Deep Representations for Photo RetouchingabstractPhoto enhancement is a long-standing and challenging problem in image processing community. Despite having witnessed significant achievements in recent years, many of them are built upon supervised learning theories and thus required expertise in constructing a huge collection of paired data, which is well-known to be a problem as the acquisition of such data in real life can be impractical. We address this issue by proposing a multi-scale GAN framework that can be trained in an unsupervised fashion. Notably, we unify the design principle of the generator and discriminator in our framework so as to maximize the ability to learn deep latent representations. Specifically, rather than maintaining the content consistency through complicated two-way loss, we present a one-way loss that measures the content distance between multi-scale latent representations of inputs and outputs to speed up the training by$\text{1.7}\times$. Furthermore, we redesign the discriminator into a multi-scale-multi-stage manner to strengthen the adversarial learning, where the multiple latent features with varying scales are produced by the main discriminator and these features are then sent to auxiliary discriminators for final recognition. Extensive experiments have been conducted in the well-known MIT-Adobe-fivek and HDR+ datasets, and the results demonstrated that the proposed multi-scale representation learning framework shows outstanding performance in photo enhancement task. Di Li 0006, Susanto Rahardja |
IEEE Trans. Multim. | 1 |
| 2023 | Pure Number Discrete Fractional Complex Hadamard TransformabstractThis paper introduces a novel discrete fractional transform termed as pure number discrete fractional complex Hadamard transform (PN-FCHT). The proposed PN-FCHT offers three advantages over the traditional discrete fractional Hadamard transform (FHT). Firstly, the higher-order PN-FCHT matrix exhibits the Self-Kronecker product structure, which allows for the recursive generation from the$2\times 2$core PN-FCHT matrix. Secondly, it possesses two important properties for computation, i.e. pure number property. Lastly, compared to existing state-of-the-art fast FHT algorithms, the PN-FCHT can reduce the transform multiplication computational complexity by up to 80% and this results in a more efficient hardware implementation. Zi-Chen Fan, Di Li 0006, Susanto Rahardja |
IEEE Signal Process. Lett. | 2 |
| 2021 | DecomVQANet: Decomposing visual question answering deep network via tensor decomposition and regression
Zongwen Bai, Ying Li 0017, Marcin Wozniak, Meili Zhou, Di Li 0006 |
Pattern Recognit. | 5 |
| 2020 | Bilinear Semi-Tensor Product Attention (BSTPA) model for visual question answeringabstractWe propose a semi-tensor product attention network model as a visual question answering tool for complex interaction over image features. Proposed model performs matrix multiplication of two arbitrary dimensions, which is used to overcome possible dimensional limitations and improve recognition flexibility. In used block-wise operation we preserve spatial and temporal information but reduce the number of parameters by using low-rank pooling scheme. Applied BERT pre-train model is tuned to recognize question features. The proposed model is evaluated on the VQA2.0 dataset. Research results show that our model has good accuracy and easy reconfiguration for future research. Zongwen Bai, Ying Li 0017, Meili Zhou, Di Li 0006, Dong Wang 0022, Dawid Polap, Marcin Wozniak |
IJCNN | 4 |
| 2019 | Residual U-Net for Retinal Vessel SegmentationabstractIn recent years, the influence of deep learning on retinal vessel segmentation has grown rapidly. Most of the available deep learning based methods use relatively shallow structures. However, due to the limited representative capacity, shallow networks will restrain deep learning models to segment both vessel and non-vessel pixels accurately. In this paper, we propose a residual U-Net for retinal vessel segmentation. Our network has several advantages. First, the network uses a new residual block structure. In the new structure, batch normalization layers are placed before the activation unit to achieve better performance and accelerate the convergence. Also, a dropout layer is utilized in the structure to alleviate over-fitting problems. Second, the depth of the network is increased by adding more residual blocks and strong dropouts which then allow the network to extract features better. Fundus images from the publicly available DRIVE and STARE datasets are used to evaluate the proposed network. Experimental result shows that the proposed modified residual U-Net has better performance than existing state-of-the-art algorithms. Di Li 0006, Dhimas Arief Dharmawan, Boon Poh Ng, Susanto Rahardja |
ICIP | 1 |