EDBT 2026 Demo / reviewers in the wild / expert
Zhaoyu Zhang 0001
dblp:75/1682-1
· DBLP profile ↗
17ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0003-4303-2806ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LightHD: A Lightweight and High-Performance Hardware Accelerator of CRYSTALS-DilithiumabstractCRYSTALS-Dilithium serves as the foundation of the NIST-standardised PQC digital signature scheme, and has been declared as the first recommended digital signature algorithm. However, due to the computational complexity and intricate processing flow of CRYSTALS-Dilithium, two limitations are shown in existing methods: its applicability on resource-constrained devices is limited and the performance reported so far remains relatively low. This paper presents a lightweight yet high-performance hardware architecture that optimises the core computational units of CRYSTALS-Dilithium. First, an iterative dual-Keccak SHA-3 module is proposed, where two cores operate with a 26-cycle offset to accelerate processing without compromising frequency. In addition, the rejection sampler is streamlined by two compact registers for intermediate values and counters, improving efficiency when consuming interleaved SHA-3 outputs. Second, for small bit-width polynomials, we eliminate the first NTT stage via lookup tables and data regrouping, reducing NTT cycles by 11.7% with little hardware overhead. Further hardware savings are achieved by maximising IP core utilisation and simplifying input multiplexers. Furthermore, a compact scheduling strategy ensures that all intermediate storage fits within a single polynomial-sized memory block. On Xilinx Artix-7 FPGAs, the design reduces hardware overhead by 14.2% compared with state-of-the-art lightweight implementations. Across three security levels, KeyGen and Verify are 27.3% and 13.5% faster, respectively, than high-performance prior designs. At level 5, the best-case Sign latency is only 120 µs. Ziying Ni, Ayesha Khalid, Zhaoyu Zhang 0001, Yijun Cui, Weiqiang Liu 0001, Máire O'Neill |
IEEE Trans. Computers | 3 |
| 2025 | Improving the Training of Data-Efficient GANs via Quality Aware Dynamic Discriminator Rejection SamplingabstractData-Efficient Generative Adversarial Nets (DE-GANs) have become more and more popular in recent years. Existing methods apply data augmentation, noise injection and pre-trained models to maximumly increase the number of training samples thus improving the training of DE-GANs. However, none of these methods considers the sample quality during training, which can also significantly influence the training of DE-GANs. Focusing on sample quality during training, in this paper, we are the first to incorporate discriminator rejection sampling (DRS) into the training process and introduce a novel method, called quality aware dynamic discriminator rejection sampling (QADDRS). Specifically, QADDRS consists of two steps: (1) the sample quality aware step, which aims to obtain the sorted critic scores, i.e., the ordered discriminator outputs, on real/fake samples in the current training stage; (2) the dynamic rejection step that obtains dynamic rejection number N, where N is controlled by the overfitting degree of discriminator (D) during training. When updating the parameters of D, the N high critic score real samples and the N low critic score fake samples in the minibatch are rejected dynamically based on the overfitting degree of D. As a result, QAD-DRS can avoid D becoming overly confident in distinguishing both real and fake samples, thereby alleviating the over-fitting of D issue during training. Extensive experiments on several datasets demonstrate that integrating QADDRS into different DE-GANs can achieve better performance and deliver state-of-the-art results. Codes are available at https://github.com/zzhang05/QADDRS. Zhaoyu Zhang 0001, Yang Hua 0001, Guanxiong Sun, Hui Wang 0001, Seán F. McLoone |
CVPR | 1 |
| 2025 | Training Diffusion-based Generative Models with Limited DataabstractDiffusion-based generative models (diffusion models) often require a large amount of data to train a score-based model that learns the score function of the data distribution through denoising score matching. However, collecting and cleaning such data can be expensive, time-consuming, and even infeasible. In this paper, we present a novel theoretical insight for diffusion models that two factors, i.e., the denoiser function hypothesis space and the number of training samples, can affect the denoising score matching error of all training samples. Based on this theoretical insight, it is evident that minimizing the total denoising score matching error is challenging within the denoiser function hypothesis space in existing methods, when training diffusion models with limited data. To address this, we propose a new diffusion model called Limited Data Diffusion (LD-Diffusion), which consists of two main components: a compressing model and a novel mixed augmentation with fixed probability (MAFP) strategy. Specifically, the compressing model can constrain the complexity of the denoiser function hypothesis space and MAFP can effectively increase the training samples by providing more informative guidance than existing data augmentation methods in the compressed hypothesis space. Extensive experiments on several datasets demonstrate that LD-Diffusion can achieve better performance compared to other diffusion models. Codes are available at https://github.com/zzhang05/LD-Diffusion. Zhaoyu Zhang 0001, Yang Hua 0001, Guanxiong Sun, Hui Wang 0001, Seán F. McLoone |
ICML | 1 |
| 2025 | OBELLA: Open the Book for Evaluating Long-Form Large Language Model Answers in Open-Domain Question AnsweringabstractReliable factuality evaluation is critical for the iterative development of open-domain question answering (ODQA) systems, especially given the rise of large language models (LLMs) and their propensity for hallucination. However, state-of-the-art (SOTA) automatic metrics, which are mostly supervised, remain notably less reliable than humans. In this paper, we find two key challenges behind this gap: (1) length distribution mismatch between lengthy LLM answers and shorter training answers used by current metrics; and (2) reference incompleteness, where current metrics often misjudge valid system answers absent from given references-a challenge worsened by the diversity of LLM outputs. To address these issues, we present a new ODQA factuality evaluation dataset called OBELLA (Open-Book Evaluation for Long-form LLM Answers). OBELLA narrows the length distribution mismatch by significantly increasing the candidate answer length to align with LLM outputs. Moreover, it introduces a neutral class for plausible yet under-supported candidate answers to differentiate reference incompleteness from outright incorrectness, thus enabling flexible reevaluation by consulting external knowledge for more references. Based on OBELLA, we propose a novel metric named OBELLAM (OBELLA Metric). OBELLAM integrates a cross-attention mechanism to enhance long-form candidate answer representations and employs a dynamic closed-open book evaluation strategy to tackle reference incompleteness. Our OBELLAM sets a new SOTA in aligning with human judgments across two ODQA evaluation benchmarks, marking a promising step toward more robust ODQA factuality evaluation. Zhaoyu Zhang 0001, Hui Wang 0001, Karen Rafferty |
SIGIR | 2 |
| 2024 | Improving the Training of the GANs with Limited Data via Dual Adaptive Noise InjectionabstractRecently, many studies have highlighted that training Generative Adversarial Networks (GANs) with limited data suffers from the overfitting of the discriminator (D). Existing studies mitigate the overfitting of D by employing data augmentation, model regularization, or pre-trained models. Despite the success of existing methods in training GANs with limited data, noise injection is another plausible, complementary, yet not well-explored approach to alleviate the overfitting of D issue. In this paper, we propose a simple yet effective method called Dual Adaptive Noise Injection (DANI), to further improve the training of GANs with limited data. Specifically, DANI consists of two adaptive strategies: adaptive injection probability and adaptive noise strength. For the adaptive injection probability, Gaussian noise is injected into both real and fake images for generator (G) and D with a probability p, respectively, where the probability p is controlled by the overfitting degree of D. For the adaptive noise strength, the Gaussian noise is produced by applying the adaptive forward diffusion process to both real and fake images, respectively. As a result, DANI can effectively increase the overlap between the distributions of real and fake data during training, thus alleviating the overfitting of D issue. Extensive experiments on several commonly-used datasets with both StyleGAN2 and FastGAN backbones demonstrate that DANI can further improve the training of GANs with limited data and achieve state-of-the-art results compared with other methods. Codes are available at https://github.com/zzhang05/DANI. Zhaoyu Zhang 0001, Yang Hua 0001, Guanxiong Sun, Hui Wang 0001, Seán F. McLoone |
ACM Multimedia | 1 |
| 2024 | Improving the Leaking of Augmentations in Data-Efficient GANs via Adaptive Negative Data AugmentationabstractData augmentation (DA) has shown its effectiveness in training Data-Efficient GANs (DE-GANs). However, applying DA in DE-GANs results in transforming the distributions of generated data and real data to augmented distributions of generated data and real data. This augmentation process could produce some out-of-distribution samples, known as the leaking of augmentations problem, which is highly undesirable in DE-GANs training. Although some methods propose "leaking-free" DAs for DE-GANs, we theoretically and practically argue that the leaking of augmentations problem still exists in these methods. To alleviate the leaking of augmentations in DE-GANs, in this paper, we propose a simple yet effective method called adaptive negative data augmentation (ANDA) for DE-GANs, with a negligible computational cost increase. Specifically, ANDA adaptively augments the augmented distribution of generated data using the augmented distribution of negative real data, where the negative real data is produced by applying negative data augmentation (NDA) on the real data. In this case, potential leaking samples can be presented as "fake" instances to the discriminator adaptively, which avoids the generator (G) learning such samples, thus resulting in better performance. Extensive experiments on several datasets with different DE-GANs demonstrate that ANDA can effectively alleviate the leaking of augmentations problem during training and achieve better performance. Codes are available at https://github.com/zzhang05/ANDA Zhaoyu Zhang 0001, Yang Hua 0001, Guanxiong Sun, Hui Wang 0001, Seán F. McLoone |
WACV | 1 |
| 2024 | Improving the Fairness of the Min-Max Game in GANs TrainingabstractGenerative adversarial networks (GANs) have achieved great success and become more and more popular in recent years. However, understanding of the min-max game in GANs training is still limited. In this paper, we first utilize information game theory to analyze the min-max game in GANs and introduce a new viewpoint on the GANs training that the min-max game in existing GANs is unfair during training, leading to sub-optimal convergence. To tackle this, we propose a novel GAN called Information Gap GAN (IGGAN), which consists of one generator (G) and two discriminators (D1and D2). Specifically, we apply different data augmentation methods to D1and D2, respectively. The information gap between different data augmentation methods can change the information received by each player in the min-max game and lead to all three players G, D1and D2in IGGAN obtaining incomplete information, which improves the fairness of the min-max game, yielding better convergence. We conduct extensive experiments for large-scale and limited data settings on several common datasets with two backbones, i.e., BigGAN and StyleGAN2. The results demonstrate that IGGAN can achieve a higher Inception Score (IS) and a lower Fréchet Inception Distance (FID) compared with other GANs. Codes are available at https://github.com/zzhang05/IGGAN Zhaoyu Zhang 0001, Yang Hua 0001, Hui Wang 0001, Seán F. McLoone |
WACV | 1 |
| 2023 | Spatio-temporal Prompting Network for Robust Video Feature ExtractionabstractFrame quality deterioration is one of the main challenges in the field of video understanding. To compensate for the information loss caused by deteriorated frames, recent approaches exploit transformer-based integration modules to obtain spatio-temporal information. However, these integration modules are heavy and complex. Furthermore, each integration module is specifically tailored for its target task, making it difficult to generalise to multiple tasks. In this paper, we present a neat and unified framework, called Spatio-Temporal Prompting Network (STPN). It can efficiently extract robust and accurate video features by dynamically adjusting the input features in the backbone network. Specifically, STPN predicts several video prompts containing spatio-temporal information of neighbour frames. Then, these video prompts are prepended to the patch embeddings of the current frame as the updated input for video feature extraction. Moreover, STPN is easy to generalise to various video tasks because it does not contain task-specific modules. Without bells and whistles, STPN achieves state-of-the-art performance on three widely-used datasets for different video understanding tasks, i.e., ImageNetVID for video object detection, YouTubeVIS for video instance segmentation, and GOT-10k for visual object tracking. Codes are available at https://github.com/guanxiongsun/STPN Guanxiong Sun, Zhaoyu Zhang 0001, Jiankang Deng, Stefanos Zafeiriou, Yang Hua 0001 |
ICCV | 3 |
| 2022 | TWGAN: Twin Discriminator Generative Adversarial NetworksabstractGenerative Adversarial Networks (GAN) has become more and more popular these years. However, it is difficult to train and suffers from the training instability problem. To tackle this difficulty, this paper proposes a novel approach. Our idea is intuitive but proven to be very useful. In essence, it combines saturating loss and non-saturating loss into the loss function. Thus it will exploit the complementary statistical properties from two kinds of loss functions to effectively improve the training stability. We term our method twin discriminator Generative Adversarial Networks (TWGAN), which, unlike GAN, has a generator and a twin discriminator. The twin discriminator consists of two discriminators with identical architecture and both of them aim to distinguish whether the samples are from real data or fake data. We develop theoretical analysis to show that, given the optimal discriminators, optimizing the generator of TWGAN reduces to minimizing the Kullback-Leibler (KL) divergence between the distribution of generated data ($P_g$) and the distribution of real data ($P_data$), hence effectively addressing the training instability problem. Extensive experiments on MNIST, Fashion MNIST, CIFAR-10/100 and STL-10 datasets demonstrate that the competitive performance of our TWGAN in generating good quality and diverse samples over baselines. The obtained highest inception score (IS) and lowest Fr$\acute{e}$chet Inception Distance (FID), compared with other state-of-the-art GANs, show the superiority of our TWGAN. Zhaoyu Zhang 0001, Haonian Xie, Jun Yu 0001, Tongliang Liu, Chang Wen Chen |
IEEE Trans. Multim. | 1 |
| 2021 | Learning Face Image Super-Resolution Through Facial Semantic Attribute Transformation and Self-Attentive Structure EnhancementabstractFace super-resolution is a domain-specific super-resolution (SR) problem of generating high-resolution (HR) face images from low-resolution (LR) inputs. Even though existing face SR methods have achieved great performance on the global region evaluation, most of them cannot restore local attributes and structure reasonably, especially to ultra-resolve tiny LR face images (16 × 16 pixels) to its larger version (8 × upscaling factor). In this paper, we propose an open source face SR framework based on facial semantic attribute transformation and self-attentive structure enhancement. Specifically, the proposed framework introduces face semantic information (i.e., face attributes) and face structure information (i.e., face boundaries) in a successive two-stage fashion. In the first stage, an Attribute Transformation Network (AT-Net) is established. It upsamples LR face images to HR feature maps and then combines facial attributes with these features to generate the intermediate HR results with rational attributes. In the second stage, a Structure Enhancement Network (SE-Net) is built. It simultaneously extracts face features and estimates facial boundary heatmaps from the inputs, and then fuses them to output the final HR face images. Extensive experiments demonstrate that our method achieves superior super-resolved results and outperforms the state-of-the-art methods. Zhaoyu Zhang 0001, Jun Yu 0001, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2020 | A Deep Learning Approach for Face Hallucination Guided by Facial Boundary ResponsesabstractFace hallucination is a domain-specific super-resolution (SR) problem of learning a mapping between a low-resolution (LR) face image and its corresponding high-resolution (HR) image. Tremendous progress on deep learning has shown exciting potential for a variety of face hallucination tasks. However, most deep-learning–based methods are limited to handle facial appearance information without paying attention to facial structure priors. In this article, we propose an open source 1 Boundary-aware Dual-branch Network (BDN) for face hallucination, which simultaneously extracts face features and estimates facial boundary responses from LR inputs, ultimately fusing them to reconstruct HR results. Specifically, we first upsample LR face images to HR feature maps, and then feed the upsampled HR features into a memory unit and an attention unit synchronously to obtain the refined features and predict facial boundary responses. Next, they are fed into a feature map fusion unit to combine facial appearance and structure information by a spatial attention mechanism. Moreover, we employ a series of stacked units to boost performance before recovering HR face images. Finally, a discriminative network is developed to improve visual quality by introducing adversarial learning strategy. Extensive experiments show that the proposed approach achieves superior face hallucination results against the state-of-the-art ones. Zhaoyu Zhang 0001, Guochen Xie, Jun Yu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Towards the Gradient Vanishing, Divergence Mismatching and Mode Collapse of Generative Adversarial NetsabstractGenerative adversarial network (GAN) is a powerful generative model. However, it suffers from gradient vanishing, divergence mismatching and mode collapse. To overcome these problems, we propose a novel GAN, which consists of one generator G and two discriminators (D1, D2). Focusing on the gradient vanishing, Spectral Normalization (SN) and ResBlock are first adopted in D1 and D2. Then, Scaled Exponential Linear Units (SELU) is adopted at last half layers of D2 to further address the problem. To divergence mismatching, relativistic discriminator is adopted in our GAN to make the loss function minimization in the training of generator equal to the theoretical divergence minimization. Concentrating on the mode collapse, D1 rewards high scores for the samples from the data distribution, while D2 favors the samples from the generator conversely. In addition, the minibatch discrimination is adopted in D1 to further address the problem. Extensive experiments on CIFAR-10/100 and ImageNet datasets demonstrate that our GAN can obtain the highest inception score (IS) and lowest Frechet Inception Distance (FID) compared with other state-of-the-art GANs. Zhaoyu Zhang 0001, Changwei Luo, Jun Yu 0001 |
CIKM | 1 |
| 2019 | D2PGGAN: Two Discriminators Used in Progressive Growing of GANSabstractGenerative adversarial network (GAN) is a powerful generative model. However, it suffers from two key problems which are convergence instability and mode collapse. Recently, progressive growing of GANs for improving quality, stability and variation (PGGAN) is proposed to better solve these two problems. Although the performance of PGGAN is good on these two problems, it is still not satisfied on mode collapse problem. In this paper, we propose a new architecture based on PGGAN called D2PGGAN to better solve the mode collapse problem. The key idea consists of one generator and two different discriminators in PGGAN. With the fact that GAN is the analogy of a minimax game, the proposed architecture is as follows. The generator (G) aims to produce realistic-looking samples to fool both of two discriminators. The first discriminator (D1) rewards high scores for samples from the data distribution, while the second one (D2) favors samples from the generator conversely. Specifically, a novel loss function is designed to optimize the proposed D2PGGAN. Extensive experiments on CIFAR-10 and CIFAR-100 datasets demonstrate that the proposed method is effective and obtains the highest inception scores compared with others state-of-the-art GANs. Zhaoyu Zhang 0001, Jun Yu 0001 |
ICASSP | 1 |
| 2019 | Deep Learning Face Hallucination via Attributes Transfer and EnhancementabstractFace hallucination technique aims to generate high-resolution (HR) face images from low-resolution (LR) inputs. Even though existing face hallucination methods have achieved great performance on the global region evaluation, most of them cannot reasonably restore local attributes, especially when ultra-resolving tiny LR face image (16 × 16 pixels) to its larger version (8× upscaling factor). In this paper, we propose a novel attribute-guided face transfer and enhancement network for face hallucination. Specifically, we first construct a face transfer network, which upsamples LR face images to HR feature maps, and then fuses facial attributes and the upsampled features to generate HR face images with rational attributes. Finally, a face enhancement network is developed based on generative adversarial network (GAN) to improve visual quality by exploiting a composite loss that combines image color, texture and content. Extensive experiments demonstrate that our method achieves superior face hallucination results and outperforms the state-of-the-art. Yuechuan Sun, Zhaoyu Zhang 0001, Haonian Xie, Jun Yu 0001 |
ICME | 3 |
| 2019 | STDGAN: ResBlock Based Generative Adversarial Nets Using Spectral Normalization and Two Different DiscriminatorsabstractGenerative adversarial network (GAN) is a powerful generative model. However, it suffers from two key problems, which are convergence and mode collapse. To overcome these drawbacks, this paper presents a novel architecture of GAN, called STDGAN, which consists of one generator and two different discriminators. With the fact that GAN is the analogy of a minimax game, the proposed architecture is as follows. The generator G aims to produce realistic-looking samples to fool both of two discriminators. The first discriminator D1 rewards high scores for the samples from the data distribution, while the second one D2 favors the samples from the generator conversely. Specifically, the minibatch discrimination and Spectral Normalization (SN) are first adopted in D1. Then, based on the ResBlock architecture, Spectral Normalization (SN) and Scaled Exponential Linear Units (SELU) are adopted in the first and last half layers of D2 respectively. In particular, a novel loss function is designed to optimize the STDGAN by minimizing the KL divergence. Extensive experiments on CIFAR-10/100 and ImageNet datasets demonstrate that the proposed STDGAN can effectively solve the problems of convergence and mode collapse and obtain the higher inception score (IS) and lower Frechet Inception Distance (FID) compared with other state-of-the-art GANs. Zhaoyu Zhang 0001, Jun Yu 0001 |
ACM Multimedia | 1 |
| 2018 | A Coarse-to-Fine Face Hallucination Method by Exploiting Facial Prior KnowledgeabstractFace hallucination technique generates high-resolution (HR) face images from low-resolution (LR) ones. In this paper, we propose to use a coarse-to-fine method for face hallucination by constructing a two-branch network, which makes full use of the specific prior knowledge of face images and the advantages of generic image super-resolution (SR) methods. Specifically, we jointly build a deep neural network (DNN) with a face image SR branch and a semantic face parsing branch. The former branch implements the image upsampling and feature extraction using a cascade of convolutional layers. The latter branch extracts facial semantic parsing as prior knowledge. Then, we combine the image features and the prior know ledge to reconstruct HR face images. Finally, we optimize the DNN, by using adversarial training and a perceptual loss, in order to obtain high realism. Extensive experiments show that the proposed method outperforms the state-of-the-art alternatives in terms of accuracy and realism. Yuechuan Sun, Zhaoyu Zhang 0001, Jun Yu 0001 |
ICIP | 3 |
| 2018 | A Cross-Layer Based Network for Faster Image GenerationabstractOwing to the great success of generative adversarial networks (GANs), unsupervised learning based image generation is popular currently. This paper presents a cross-layer architecture for the generators in GANs, which passes inputs to every subsequent layer and encourages the flow of information and gradients throughout the network. While traditional networks with L layers have L connections, the proposed network with cross-layer architecture has 2L-1 direct connections. For each layer, the network input and the output of its previous layer are used as inputs. Extensive experiments demonstrate that our network can generate images at a higher speed without introducing extra parameters by comparing with two state-of-the-art GANs, namely deep convolutional GAN and Wasserstein GAN-GP, on two datasets: Fashion-MNIST and CelebA. Zhaoyu Zhang 0001, Yuechuan Sun, Jun Yu 0001 |
ICIP | 1 |