EDBT 2026 Demo / reviewers in the wild / expert
Weilun Wang
dblp:254/5358
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ZergPPU: an evolutionary multi-mode post processor for vision neural networksabstractWith the widespread adoption of AI in various industries, there is increasing demand for edge devices to efficiently utilize AI, particularly in object detection and related applications. Current studies aim to accelerate specific computations but struggle to keep up with rapid advancements in neural network optimization techniques like quantization and pruning. So the goal of this work is to propose a microarchitecture approach that can flexibly support multiple generations of visual neural networks and dynamic data shapes.This work presents the Zerg Post Processing Unit (ZergPPU), which determines its computational architecture based on data information, partitioning into regions for efficient processing. With software support, ZergPPU adapts to evolving AI algorithms through multi-mode processing (e.g., from YOLOv3 to YOLOv9). By incorporating a path prediction unit and dynamic address generation, it supports a wide range of neural networks and optimizations. To minimize area, ZergPPU uses selective reuse executors with automatic software optimizations, resulting in an area of 0.064mm2. Experiments show that this design achieves 49.3% lower computational latency compared to the baseline, making it well-suited for edge AI vision applications. Weilun Wang, Chen Chao, Zheng Wang 0027 |
ISCAS | 1 |
| 2025 | SinDiffusion: Learning a Diffusion Model From a Single Natural ImageabstractWe present SinDiffusion, leveraging denoising diffusion models to capture internal distribution of patches from a single natural image. The default approach of previous GAN-based methods on this problem is to train multiple models at progressive growing scales, which leads to the accumulation of errors and causes characteristic artifacts in generated results. In this paper, we uncover that multiple models at progressive growing scales are not essential for learning from a single image and propose SinDiffusion, a single diffusion-based model trained on a single scale, which is better-suited for this task. Furthermore, we identify that a patch-level receptive field is crucial and effective for diffusion models to capture the image's patch statistics, therefore we redesign an patch-wise denoising network for SinDiffusion. Coupling these two designs enables SinDiffusion to generate more photorealistic and diverse images from a single image compared with GAN-based approaches. SinDiffusion can also be applied to various applications, i.e., text-guided image generation, and image outpainting beyond the capability of SinGAN. Extensive experiments on a wide range of images demonstrate the superiority of SinDiffusion for modeling the patch distribution. Weilun Wang, Jianmin Bao, Wengang Zhou 0001, Dongdong Chen 0001, Dong Chen 0003, Lu Yuan 0001, Houqiang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | DISA: Disentangled Dual-Branch Framework for Affordance-Aware Human InsertionabstractAffordance-aware human insertion is a controllable human synthesis task aimed at seamlessly integrating a person into a scene while aligning human pose with contextual scene affordance and preserving human visual identity. Previous methods, typically reliant on a general framework of inpainting that injects all conditional information into a single branch, often struggle with the complexities of real-world contexts and the nuanced attributes of human figures. To this end, we present a novel Disentangled dual-branch framework for Affordance-aware human insertion task (DISA) , which focuses on both scene context comprehension and precise person attribute extraction. Specifically, our dual-branch design facilitates diffusion models to ensure disentangled and precise manipulations: one branch utilizes an additional network for deep scene context comprehension and control, while the other branch employs a parallel encoder to extract the feature of the reference person and injects this information through cross-attention mechanism. Furthermore, to comprehensively evaluate affordance-aware human insertion task, we introduce a new metric to assess the preservation of visual identity. We conduct a broad variety of evaluation experiments and validate the diversity and robustness of our method in different settings and downstream applications. Both qualitative and quantitative experimental analysis demonstrates that our approach outperforms previous methods in terms of image quality, pose accuracy, and visual identity preservation. Xuanqing Cao, Wengang Zhou 0001, Qi Sun 0005, Weilun Wang, Li Li 0040, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Unmanned Aerial Vehicle Defense Penetration on a Digital-Twin-City Based on an Improved MASAC Reinforcement LearningabstractIn this paper, we investigate the use of multi-UAVs (Unmanned Aerial Vehicles) for defense penetration in an urban environment with obstacles. To ensure a high-quality scenario of the urban environment, a digital-twin-city simulation platform is proposed featuring a powerful physics engine and real-time interaction technology. Subsequently, an improved multi-agent soft actor-critic (MASAC) method with a target prediction network is introduced to enhance the effectiveness of the cooperative multi-UAV system in defense penetration missions. Finally, the performance of the proposed improved MASAC method is evaluated through comparative experiments. Ruilong Zhang 0002, Xiuyun Zhang, Weilun Wang, Qun Zong |
ICARCV | 5 |
| 2024 | Dual-Assessment Driven Pruning: Iterative Optimizing Layer-wise Sparsity for Large Language ModelabstractLarge Language Models (LLMs) have demonstrated efficacy in various domains, but deploying these models is economically challenging due to extensive parameter counts. Numerous efforts have been dedicated to reducing the parameter count of these models without compromising performance, employing a technique known as model pruning. Conventional pruning methods assess the significance of weights within individual layers and typically apply uniform sparsity levels across all layers, potentially neglecting the varying significance of each layer. To address this oversight, we first propose a dual-assessment driven pruning strategy that employs both intra-layer metric and global performance metric to comprehensively evaluate the impact of pruning. Then our method leverages an iterative optimization algorithm to find the optimal layer-wise sparsity distribution, thereby minimally impacting model performance. Extensive benchmark evaluations on state-of-the-art LLM architectures such as LLaMAv2 and OPT across a variety of NLP tasks demonstrate the effectiveness of our approach. When applied to the LLaMaV2-7B model with an overall pruning sparsity of 80%, our method achieves a 50% reduction in perplexity compared to the benchmark. The results indicate that our method significantly outperforms existing state-of-the-art methods in preserving performance after pruning. Qinghui Sun, Weilun Wang, Yanni Zhu, Shenghuan He, Zehua Cai |
KDD | 2 |
| 2024 | CLIP2GAN: Toward Bridging Text With the Latent Space of GANsabstractIn this work, we are dedicated to text-guided image generation and propose a novel framework,i.e., CLIP2GAN, by leveraging CLIP model and StyleGAN. The key idea of our CLIP2GAN is to bridge the output feature embedding space of CLIP and the input latent space of StyleGAN, which is realized by introducing a mapping network. In the training stage, we encode an image with CLIP and map the output feature to a latent code, which is further used to reconstruct the image. In this way, the mapping network is optimized in a self-supervised learning way. In the inference stage, since CLIP can embed both image and text into a shared feature embedding space, we replace CLIP image encoder in the training architecture with CLIP text encoder, while keeping the following mapping network as well as StyleGAN model. As a result, we can flexibly input a text description to generate an image. Moreover, by simply adding mapped text features of an attribute to a mapped CLIP image feature, we can effectively edit the attribute to the image. Extensive experiments demonstrate the superior performance of our proposed CLIP2GAN compared to previous methods. Wengang Zhou 0001, Jianmin Bao, Weilun Wang, Li Li 0040, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | AltFreezing for More General Video Face Forgery DetectionabstractExisting face forgery detection models try to discriminate fake images by detecting only spatial artifacts (e.g., generative artifacts, blending) or mainly temporal artifacts (e.g., flickering, discontinuity). They may experience significant performance degradation when facing out-domain artifacts. In this paper, we propose to capture both spatial and temporal artifacts in one model for face forgery detection. A simple idea is to leverage a spatiotemporal model (3D ConvNet). However, we find that it may easily rely on one type of artifact and ignore the other. To address this issue, we present a novel training strategy called AltFreezing for more general face forgery detection. The AltFreezing aims to encourage the model to detect both spatial and temporal artifacts. It divides the weights of a spatiotemporal network into two groups: spatial-related and temporal-related. Then the two groups of weights are alternately frozen during the training process so that the model can learn spatial and temporal features to distinguish real or fake videos. Furthermore, we introduce various video-level data augmentation methods to improve the generalization capability of the forgery detection model. Extensive experiments show that our framework outperforms existing methods in terms of generalization to unseen manipulations and datasets. Jianmin Bao, Wengang Zhou 0001, Weilun Wang, Houqiang Li |
CVPR | 4 |
| 2023 | DIRE for Diffusion-Generated Image DetectionabstractDiffusion models have shown remarkable success in visual synthesis, but have also raised concerns about potential abuse for malicious purposes. In this paper, we seek to build a detector for telling apart real images from diffusion-generated images. We find that existing detectors struggle to detect images generated by diffusion models, even if we include generated images from a specific diffusion model in their training data. To address this issue, we propose a novel image representation called DIffusion Reconstruction Error (DIRE), which measures the error between an input image and its reconstruction counterpart by a pre-trained diffusion model. We observe that diffusion-generated images can be approximately reconstructed by a diffusion model while real images cannot. It provides a hint that DIRE can serve as a bridge to distinguish generated and real images. DIRE provides an effective way to detect images generated by most diffusion models, and it is general for detecting generated images from unseen diffusion models and robust to various perturbations. Furthermore, we establish a comprehensive diffusion-generated benchmark including images generated by various diffusion models to evaluate the performance of diffusion-generated image detectors. Extensive experiments on our collected benchmark demonstrate that DIRE exhibits superiority over previous generated-image detectors. The code, models, and dataset are available at https://github.com/ZhendongWang6/DIRE. Jianmin Bao, Wengang Zhou 0001, Weilun Wang, Hezhen Hu, Houqiang Li |
ICCV | 4 |
| 2023 | Grading of HCC Biopsy Images Using Nucleus and Texture FeaturesabstractHepatocellular carcinoma (HCC) is one of the most critical health problems in the world. For proper treatment, it is important to identify the grade of cancer morbidity from HCC biopsy image. The diagnostic work is not only time-consuming but also subjective. The same biopsy image may be diagnosed as of different grades by different doctors, due to lack of experience or difference in opinion. In this work, we proposed an automatic grading system with classification accuracy matching to an experienced doctor, to help augment the diagnosis process. First, we proposed a segmentation method to isolate all nucleus-like objects present in a biopsy image. Non-target objects (here the target is a single HCC nucleus) present in the biopsy image are isolated too in the segmentation process. To eliminate such non-target objects, we proposed clustering of segmented images and a novel method to filter out target objects. Next, we proposed a two track neural network, where input consists of 2 different images. It combines a single segmented nucleus and a random cropped texture patch of the biopsy image to which the nucleus belongs. At this classifier output, we grade the single nucleus. Finally, a majority voting method is used to identify the grade of the whole biopsy image. We achieved an accuracy of 99.03% for nucleus image grading and 99.66% accuracy for grading biopsy images. Goutam Chakraborty, Weilun Wang, Basabi Chakraborty, Shao-Kuo Tai, Yi-Shun Lo |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Coherent Image Animation Using Spatial-Temporal CorrespondenceabstractRecent studies have achieved remarkable success using deep generative models for the image animation of an arbitrary object.However, previous methods synthesize animated results in a frame-by-frame manner, which is prone to producing flickering and temporally inconsistent results. In this paper, we propose a novel self-supervised framework leveraging temporal information for image animation. Our framework processes a video clip directly instead of processing each frame independently. To achieve coherence in the animated video, we design a spatial-temporal correspondence network (STCN) to maintain the consistency of the keypoints. Specifically, the STCN takes full advantage of temporal information to propagate the keypoints between adjacent frames, and it can be trained with consistent keypoints during the forward and backward process. Furthermore, we apply a 3D-CNN-based generator and discriminator in our framework to ensure coherence in the final output video. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our method. Weilun Wang, Wengang Zhou 0001, Jianmin Bao, Houqiang Li |
IEEE Trans. Multim. | 1 |
| 2022 | Hand-Object Interaction Image GenerationabstractIn this work, we are dedicated to a new task, i.e., hand-object interaction image generation, which aims to conditionally generate the hand-object image under the given hand, object and their interaction status. This task is challenging and research-worthy in many potential application scenarios, such as AR/VR games and online shopping, etc. To address this problem, we propose a novel HOGAN framework, which utilizes the expressive model-aware hand-object representation and leverages its inherent topology to build the unified surface space. In this space, we explicitly consider the complex self- and mutual occlusion during interaction. During final image synthesis, we consider different characteristics of hand and object and generate the target image in a split-and-combine manner. For evaluation, we build a comprehensive protocol to access both the fidelity and structure preservation of the generated image. Extensive experiments on two large-scale datasets, i.e., HO3Dv3 and DexYCB, demonstrate the effectiveness and superiority of our framework both quantitatively and qualitatively. The code will be available at https://github.com/play-with-HOI-generation/HOIG. Hezhen Hu, Weilun Wang, Wengang Zhou 0001, Houqiang Li |
NeurIPS | 2 |
| 2021 | Model-Aware Gesture-to-Gesture TranslationabstractHand gesture-to-gesture translation is a significant and interesting problem, which serves as a key role in many applications, such as sign language production. This task involves fine-grained structure understanding of the mapping between the source and target gestures. Current works follow a data-driven paradigm based on sparse 2D joint representation. However, given the insufficient representation capability of 2D joints, this paradigm easily leads to blurry generation results with incorrect structure. In this paper, we propose a novel model-aware gesture-to-gesture translation framework, which introduces hand prior with hand meshes as the intermediate representation. To take full advantage of the structured hand model, we first build a dense topology map aligning the image plane with the encoded embedding of the visible hand mesh. Then, a transformation flow is calculated based on the correspondence of the source and target topology map. During the generation stage, we inject the topology information into generation streams by modulating the activations in a spatially-adaptive manner. Further, we incorporate the source local characteristic to enhance the translated gesture image according to the transformation flow. Extensive experiments on two benchmark datasets have demonstrated that our method achieves new state-of-the-art performance. Hezhen Hu, Weilun Wang, Wengang Zhou 0001, Weichao Zhao, Houqiang Li |
CVPR | 2 |
| 2021 | Instance-wise Hard Negative Example Generation for Contrastive Learning in Unpaired Image-to-Image TranslationabstractContrastive learning shows great potential in unpaired image-to-image translation, but sometimes the translated results are in poor quality and the contents are not preserved consistently. In this paper, we uncover that the negative examples play a critical role in the performance of contrastive learning for image translation. The negative examples in previous methods are randomly sampled from the patches of different positions in the source image, which are not effective to push the positive examples close to the query examples. To address this issue, we present instance-wise hard Negative Example Generation for Contrastive learning in Unpaired image-to-image Translation (NEGCUT). Specifically, we train a generator to produce negative examples online. The generator is novel from two perspectives: 1) it is instance-wise which means that the generated examples are based on the input image, and 2) it can generate hard negative examples since it is trained with an adversarial loss. With the generator, the performance of unpaired image-to-image translation is significantly improved. Experiments on three benchmark datasets demonstrate that the proposed NEGCUT framework achieves state-of-the-art performance compared to previous methods. Weilun Wang, Wengang Zhou 0001, Jianmin Bao, Dong Chen 0003, Houqiang Li |
ICCV | 1 |
| 2021 | Automatic prognosis of lung cancer using heterogeneous deep learning models for nodule detection and eliciting its morphological features
Weilun Wang, Goutam Chakraborty |
Appl. Intell. | 1 |
| 2019 | Evaluation of Malignancy of Lung Nodules from CT Image Using Recurrent Neural NetworkabstractThe efficacy of treatment of cancer depends largely on early detection and correct prognosis. It is more important in case of pulmonary cancer, where the detection is based on identifying malignant nodules in the Computed Tomography (CT) scans of the lung. There are two problems for making correct decision about malignancy: (1) At early stage, the nodule size is small (length 5 to 10 mm). As the CT scan covers a volume of 30cm. × 30cm. × 40cm., manually searching for nodules takes a very long time (approximately 10 minutes for an expert). (2) There are benign nodules and nodules due to other ailments like bronchitis, pneumonia, tuberculosis. To identify whether the nodule is carcinogenic needs long experience and expertise. In recent years, several works have been reported to classify lung cancer using not only the CT scan image, but also other features causing or related to cancer. In all recent works, for CT image analysis, 3-D Convolution Neural Network (CNN) is used to identify cancerous nodules. In spite of various preprocessing used to improve training efficiency, 3-D CNN is extremely slow. The aim of this work is to improve training efficiency by proposing a new deep NN model. It consists of a hierarchical (sliced) structure of recurrent neural network (RNN), where different layers of the hierarchy can be trained simultaneously, decreasing training time. In addition, selective attention (alignment) during training improves convergence rate. The result shows a 3-fold increase in training efficiency, compared to recent state-of-the-art work using 3-D CNN. Weilun Wang, Goutam Chakraborty |
SMC | 1 |