Huaizhong Lin

dblp:96/2371 · DBLP profile ↗
← Back
33ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0001-6313-5349ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 3 first-authorArtificial intelligence and machine learning · 11 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorSystems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Fast and Robust Deformable 3D Gaussian Splatting
abstract
3D Gaussian Splatting has demonstrated remarkable real-time rendering capabilities and superior visual quality in novel view synthesis for static scenes. Building upon these advantages, researchers have progressively extended 3D Gaussians to dynamic scene reconstruction. Deformation field-based methods have emerged as a promising approach among various techniques. These methods maintain 3D Gaussian attributes in a canonical field and employ the deformation field to transform this field across temporal sequences. Nevertheless, these approaches frequently encounter challenges such as suboptimal rendering speeds, significant dependence on initial point clouds, and vulnerability to local optima in dim scenes. To overcome these limitations, we present FRoG, an efficient and robust framework for high-quality dynamic scene reconstruction. FRoG integrates per-Gaussian embedding with a coarse-to-fine temporal embedding strategy, accelerating rendering through the early fusion of temporal embeddings. Moreover, to enhance robustness against sparse initializations, we introduce a novel depth- and error-guided sampling strategy. This strategy populates the canonical field with new 3D Gaussians at low-deviation initial positions, significantly reducing the optimization burden on the deformation field and improving detail reconstruction in both static and dynamic regions. Furthermore, by modulating opacity variations, we mitigate the local optima problem in dim scenes, improving color fidelity. Comprehensive experimental results validate that our method achieves accelerated rendering speeds while maintaining state-of-the-art visual quality.
Han Jiao 0005, Jiakai Sun, Lei Zhao 0011, Zhanjie Zhang, Wei Xing 0001, Huaizhong Lin
IEEE Trans. Vis. Comput. Graph.6
2025 Cascaded Diffusion Models for Virtual Try-On: Improving Control and Resolution
abstract
Previous virtual try-on methods have employed ControlNet architecture in exemplar-based inpainting diffusion models to guide the generation of try-on images, preserving the garment's features and enhancing the realism of the generated images. While these methods have maintained the identity of the garment and improved the naturalness of the generated images, they still face the following limitations: (1) For garments with complex features, such as intricate text, patterns, and uncommon styles, they struggle to retain these detailed features in the generated try-on images. (2) They are limited to generating try-on images at a maximum resolution of 1K, which may not meet the demands of real-world scenarios, where higher resolutions might be required. To address the aforementioned issues, in this paper, we propose a Cascaded Diffusion Model for virtual try-on to enhance both image controllability and resolution. We call it CDM-VTON. Specifically, we design two diffusion models: the Multi-Conditioned Diffusion Model (MC-DM) and the Super-Resolution Diffusion Model (SR-DM). The former generates low-resolution try-on images while preserving the garment's complex features, and the latter enhances the resolution of these images. Additionally, we incorporate a multi-control integration module in the MC-DM, which injects multiple control conditions into a frozen denoising U-Net to ensure that the generated try-on images retain complex garment features. Our experimental results demonstrate that our method outperforms previous approaches in preserving garment details and generating authentic virtual try-on images, both qualitatively and quantitatively.
Junsheng Luan, Lei Zhao 0011, Wei Xing 0001, Huaizhong Lin, Binkai Ou
AAAI6
2024 PNeSM: Arbitrary 3D Scene Stylization via Prompt-Based Neural Style Mapping
abstract
3D scene stylization refers to transform the appearance of a 3D scene to match a given style image, ensuring that images rendered from different viewpoints exhibit the same style as the given style image, while maintaining the 3D consistency of the stylized scene. Several existing methods have obtained impressive results in stylizing 3D scenes. However, the mod- els proposed by these methods need to be re-trained when applied to a new scene. In other words, their models are cou- pled with a specific scene and cannot adapt to arbitrary other scenes. To address this issue, we propose a novel 3D scene stylization framework to transfer an arbitrary style to an ar- bitrary scene, without any style-related or scene-related re- training. Concretely, we first map the appearance of the 3D scene into a 2D style pattern space, which realizes complete disentanglement of the geometry and appearance of the 3D scene and makes our model be generalized to arbitrary 3D scenes. Then we stylize the appearance of the 3D scene in the 2D style pattern space via a prompt-based 2D stylization al- gorithm. Experimental results demonstrate that our proposed framework is superior to SOTA methods in both visual qual- ity and generalization.
Jiafu Chen, Wei Xing 0001, Jiakai Sun, Tianyi Chu, Boyan Ji, Lei Zhao 0011, Huaizhong Lin, Haibo Chen 0006, Zhizhong Wang
AAAI8
2024 Attack Deterministic Conditional Image Generative Models for Diverse and Controllable Generation
abstract
Existing generative adversarial network (GAN) based conditional image generative models typically produce fixed output for the same conditional input, which is unreasonable for highly subjective tasks, such as large-mask image inpainting or style transfer. On the other hand, GAN-based diverse image generative methods require retraining/fine-tuning the network or designing complex noise injection functions, which is computationally expensive, task-specific, or struggle to generate high-quality results. Given that many deterministic conditional image generative models have been able to produce high-quality yet fixed results, we raise an intriguing question: is it possible for pre-trained deterministic conditional image generative models to generate diverse results without changing network structures or parameters? To answer this question, we re-examine the conditional image generation tasks from the perspective of adversarial attack and propose a simple and efficient plug-in projected gradient descent (PGD) like method for diverse and controllable image generation. The key idea is attacking the pre-trained deterministic generative models by adding a micro perturbation to the input condition. In this way, diverse results can be generated without any adjustment of network structures or fine-tuning of the pre-trained models. In addition, we can also control the diverse results to be generated by specifying the attack direction according to a reference text or image. Our work opens the door to applying adversarial attack to low-level vision tasks, and experiments on various conditional image generation tasks demonstrate the effectiveness and superiority of the proposed method.
Tianyi Chu, Wei Xing 0001, Jiafu Chen, Zhizhong Wang, Jiakai Sun, Lei Zhao 0011, Haibo Chen 0006, Huaizhong Lin
AAAI8
2024 ArtBank: Artistic Style Transfer with Pre-trained Diffusion Model and Implicit Style Prompt Bank
abstract
Artistic style transfer aims to repaint the content image with the learned artistic style. Existing artistic style transfer methods can be divided into two categories: small model-based approaches and pre-trained large-scale model-based approaches. Small model-based approaches can preserve the content strucuture, but fail to produce highly realistic stylized images and introduce artifacts and disharmonious patterns; Pre-trained large-scale model-based approaches can generate highly realistic stylized images but struggle with preserving the content structure. To address the above issues, we propose ArtBank, a novel artistic style transfer framework, to generate highly realistic stylized images while preserving the content structure of the content images. Specifically, to sufficiently dig out the knowledge embedded in pre-trained large-scale models, an Implicit Style Prompt Bank (ISPB), a set of trainable parameter matrices, is designed to learn and store knowledge from the collection of artworks and behave as a visual prompt to guide pre-trained large-scale models to generate highly realistic stylized images while preserving content structure. Besides, to accelerate training the above ISPB, we propose a novel Spatial-Statistical-based self-Attention Module (SSAM). The qualitative and quantitative experiments demonstrate the superiority of our proposed method over state-of-the-art artistic style transfer methods. Code is available at https://github.com/Jamie-Cheung/ArtBank.
Zhanjie Zhang, Quanwei Zhang, Wei Xing 0001, Lei Zhao 0011, Jiakai Sun, Zehua Lan, Junsheng Luan, Huaizhong Lin
AAAI10
2024 Rethinking Video Deblurring with Wavelet-Aware Dynamic Transformer and Diffusion Model
Chen Rao, Zehua Lan, Jiakai Sun, Junsheng Luan, Wei Xing 0001, Lei Zhao 0011, Huaizhong Lin, Jianfeng Dong, Dalong Zhang
ECCV (45)8
2024 Towards Highly Realistic Artistic Style Transfer via Stable Diffusion with Step-aware and Layer-aware Prompt
Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin, Wei Xing 0001, Juncheng Mo, Shuaicheng Huang, Jinheng Xie, Junsheng Luan, Lei Zhao 0011, Dalong Zhang, Lixia Chen
IJCAI3
2024 PKD-Net: Distillation of prior knowledge for image completion by multi-level semantic attention
abstract
Summary Prior knowledge plays a crucial role in image completion. Although almost all of the existing image completion methods use prior knowledge to complete the image to be repaired from different perspectives, the learning and modeling of the prior knowledge is still a challenging problem. In order to address this issue, we propose a novel prior knowledge distillation framework (PKD‐Net) which could distill prior knowledge of structure and style from multiple semantic space and generates not only plausible content but also consistent style with surrounding image area. Our PKD‐Net replaces the skip connection in the vanilla U‐Net with a semantic shift attention module. The semantic shift attention module takes features from encoder layer and those from decoder layer as input pairs to output shifted features which take into account the long‐range dependency of encoder layer features and corresponding decoder layer features from the perspective of local structure and style. Semantic shift attention module models the global interdependencies in local spatial structures (patches centered at each position) and style (appearance texture) dimensions respectively, which could implement distillation of prior knowledge from two aspects: structure and style. Experiments on multiple datasets including faces (CelebA, CelebA‐HQ) and natural images (ImageNet, Places2, Paris Street View) demonstrate that our proposed approach generates higher quality completion results than existing ones.
Qiong Lu, Huaizhong Lin, Wei Xing 0001, Lei Zhao 0011, Jingjing Chen 0002
Concurr. Comput. Pract. Exp.2
2024 Rethink arbitrary style transfer with transformer and contrastive learning
Zhanjie Zhang, Jiakai Sun, Lei Zhao 0011, Quanwei Zhang, Zehua Lan, Haolin Yin, Huaizhong Lin, Zhiwen Zuo
Comput. Vis. Image Underst.9
2023 Rethinking Multi-Contrast MRI Super-Resolution: Rectangle-Window Cross-Attention Transformer and Arbitrary-Scale Upsampling
abstract
Recently, several methods have explored the potential of multi-contrast magnetic resonance imaging (MRI) super-resolution (SR) and obtain results superior to single-contrast SR methods. However, existing approaches still have two shortcomings: (1) They can only address fixed integer upsampling scales, such as 2×, 3×, and 4×, which require training and storing the corresponding model separately for each upsampling scale in clinic. (2) They lack direct interaction among different windows as they adopt the square window (e.g., 8×8) transformer network architecture, which results in inadequate modelling of longer-range dependencies. Moreover, the relationship between reference images and target images is not fully mined. To address these issues, we develop a novel network for multi-contrast MRI arbitrary-scale SR, dubbed as McASSR. Specifically, we design a rectangle-window cross-attention transformer to establish longer-range dependencies in MR images without increasing computational complexity and fully use reference information. Besides, we propose the reference-aware implicit attention as an upsampling module, achieving arbitrary-scale super-resolution via implicit neural representation, further fusing supplementary information of the reference image. Extensive and comprehensive experiments on both public and clinical datasets show that our McASSR yields superior performance over SOTA methods, demonstrating its great potential to be applied in clinical practice. Code will be available at https://github.com/GuangYuanKK/McASSR.
Lei Zhao 0011, Jiakai Sun, Zehua Lan, Zhanjie Zhang, Jiafu Chen, Huaizhong Lin, Wei Xing 0001
ICCV8
2023 Self-Reference Image Super-Resolution via Pre-trained Diffusion Large Model and Window Adjustable Transformer
abstract
Currently, reference-based super-resolution (RefSR) techniques leverage high-resolution (HR) reference images to provide useful content and texture information for low-resolution (LR) images during the super-resolution (SR) process. Nevertheless, it is time-consuming, laborious, and even impossible in some cases to find high-quality reference images. To tackle this problem, we propose a brand-new self-reference image super-resolution approach using a pre-trained diffusion large model and a window adjustable transformer, termed DWTrans. Our proposed method does not require explicitly inputting manually acquired reference images during training and inference. Specifically, we feed the degraded LR images into a pre-trained stable diffusion large model to automatically generate corresponding high-quality self-reference (SRef) images that provide valuable high-frequency details for the LR images in the process of SR. To extract valuable high-frequency information in SRef images, we design a window adjustable transformer with both non-adjustable window layer (NWL) and adjustable window layer (AWL). The NWL learns local features from LR images using a dense window, while the AWL acquires global features from the SRef images using a random sparse window. Furthermore, to fully utilize the high-frequency features in the SRef image, we introduce the adaptive deformable fusion module to adaptively fuse the features of the LR and SRef images. Experimental results validate that our proposed DWTrans outperforms state-of-the-art methods on various benchmark datasets both quantitatively and visually.
Wei Xing 0001, Lei Zhao 0011, Zehua Lan, Jiakai Sun, Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin
ACM Multimedia8
2023 DuDoINet: Dual-Domain Implicit Network for Multi-Modality MR Image Arbitrary-scale Super-Resolution
abstract
Compared to single-modality magnetic resonance (MR) image super-resolution (SR) methods, multi-modality MR image methods can utilize high-resolution reference modality (e.g., T1 modality) to provide valuable complementary information for low-resolution target modality (e.g., T2 modality) in SR reconstruction, which can further improve the quality of the SR images. Although they have achieved impressive results, these methods still suffer from the following drawbacks: (1) They can only handle fixed integer upsampling factors, such as 2X, 3X, and 4X, and require training and storing corresponding models for each upsampling factor, which is infeasible in clinical practice; (2) They only perform feature extraction and reconstruction in the image domain. However, the aliasing artifacts produced in the image domain are structural and non-local. Therefore, using only the image domain cannot effectively reconstruct high-quality aliasing-free SR images. To address these issues, we develop a brand-new Dual-Domain Implicit Network (DuDoINet) for multi-modality MR image arbitrary-scale SR. Specifically, we propose a dual-domain learning scheme for multi-modality MR image SR, which allows the network to sufficiently exploit the frequency and image domain information in MR images. In addition, we design implicit attention to achieve arbitrary-scale upsampling of MR images, which utilizes a continuously differentiable function that generates pixel values from pixel coordinates. Furthermore, we designed a deformable cross-modality attention mechanism that can adaptively transfer high-frequency details from the T1 to the T2 modality, better integrating valuable complementary information from the T1 modality. Extensive and comprehensive experiments on healthy subjects and patient datasets demonstrate that our DuDoINet outperforms SOTA methods, demonstrating its great potential for clinical practice.
Wei Xing 0001, Lei Zhao 0011, Zehua Lan, Zhanjie Zhang, Jiakai Sun, Haolin Yin, Huaizhong Lin
ACM Multimedia8
2020 SpatialGAN: Progressive Image Generation Based on Spatial Recursive Adversarial Expansion
abstract
The image generation model based on generative adversarial networks has recently received significant attention and can produce diverse, sharp, and realistic images. However, generating high-resolution images has long been a challenge. In this paper, we propose a progressive spatial recursive adversarial expansion model(called SpatialGAN) capable of producing high-quality samples of the natural image. Our approach uses a cascade of convolutional networks to progressively generate images in a part-to-whole fashion. At each level of spatial expansion, a separate image-to-image spatial adversarial expansion network (conditional GAN) is recursively trained based on context image generated by previous GAN or CGAN. Unlike other coarse-to-fine generative methods that constraint on generative process either by multi-scale resolution or by hierarchical feature, the SpatialGAN decomposes image space into multiple subspaces and gradually resolves uncertainties in the local-to-whole generative process. The SpatialGAN greatly stabilizes and speeds up the training, which allows us to produce images of high quality. Based on visual Inception Score and Fréchet Inception Distance, we demonstrate that the quality of images generated by SpatialGAN on several typical datasets is better than that of images generated by GANs without cascading and comparative with the state of art methods with cascading.
Lei Zhao 0011, Sihuan Lin, Ailin Li, Huaizhong Lin, Wei Xing 0001, Dongming Lu
ACM Multimedia4
2020 LERI: Local Exploration for Rare-Category Identification
abstract
To identify the data examples of rare categories that form small compact clusters in large data sets, existing approaches mostly require enough labeled data examples as a training set to learn a classifier, assuming that the rare-category clusters are spherical or nearly spherical. Nonetheless, a large enough training set is usually difficult to obtain in practice, and rare categories in many real-world applications often form small compact clusters with arbitrary shapes. In this paper, we investigate how to identify all data examples of a rare category with an arbitrary shape based on only one seed (i.e., a labeled rare-category data example). Instead of finding a compact and spherical local region around the seed, we locally explore the data set from the seed by continuously searching and visiting the k-nearest neighbors of each newly visited data example. The local exploration connects the data examples in the objective rare category by the relationship of k-nearest neighbors, and meanwhile, suspected external data examples are filtered out if they are not close enough to any visited data example. Experimental results on both synthetic and real-world data sets are conducted, and the results verify the effectiveness and efficiency of our approach.
Hao Huang 0001, Qian Yan 0001, Wei Lu 0015, Huaizhong Lin, Yunjun Gao, Lei Chen 0002
IEEE Trans. Knowl. Data Eng.4
2019 GOAL: a clustering-based method for the group optimal location problem
Fangshu Chen, Jianzhong Qi 0001, Huaizhong Lin, Yunjun Gao, Dongming Lu
Knowl. Inf. Syst.3
2018 Aggregate keyword nearest neighbor queries on road networks
Pengfei Zhang 0004, Huaizhong Lin, Yunjun Gao, Dongming Lu
GeoInformatica2
2018 Finding the hottest item in data streams
Huaizhong Lin, Leong Hou U, Ngai Meng Kou, Yunjun Gao, Dongming Lu
Inf. Sci.1
2017 Collective-k Optimal Location Selection
Fangshu Chen, Huaizhong Lin, Jianzhong Qi 0001, Yunjun Gao
SSTD2
2017 Level-aware collective spatial keyword queries
Pengfei Zhang 0004, Huaizhong Lin, Bin Yao 0002, Dongming Lu
Inf. Sci.2
2017 RLC: ranking lag correlations with flexible sliding windows in data streams
Huaizhong Lin, Wenxiang Wang, Dongming Lu, Leong Hou U, Yunjun Gao
Pattern Anal. Appl.2
2017 Novel structures for counting frequent items in time decayed streams
Huaizhong Lin, Leong Hou U, Yunjun Gao, Dongming Lu
World Wide Web2
2016 Finding Frequent Items in Time Decayed Data Streams
Huaizhong Lin, Leong Hou U, Yunjun Gao, Dongming Lu
APWeb (2)2
2016 Capacity constrained maximizing bichromatic reverse nearest neighbor search
Fangshu Chen, Huaizhong Lin, Yunjun Gao, Dongming Lu
Expert Syst. Appl.2
2016 Finding optimal region for bichromatic reverse nearest neighbor in two- and three-dimensional spaces
Huaizhong Lin, Fangshu Chen, Yunjun Gao, Dongming Lu
GeoInformatica1
2013 OptRegion: Finding Optimal Region for Bichromatic Reverse Nearest Neighbors
Huaizhong Lin, Fangshu Chen, Yunjun Gao, Dongming Lu
DASFAA (1)1
2013 Improving Semi-supervised Text Classification by Using Wikipedia Knowledge
Huaizhong Lin, Huazhong Wang, Dongming Lu
WAIM2
2006 Personalized Web Recommendation Based on Path Clustering
Yijun Yu 0005, Huaizhong Lin, Yimin Yu, Chun Chen 0001
FQAS2
2006 Mining Interest Navigation Patterns Based on Hybrid Markov Model
Yijun Yu 0005, Huaizhong Lin, Yimin Yu, Chun Chen 0001
FQAS2
2005 Transaction Reordering for Epidemic Quorum in Replicated Databases
Huaizhong Lin, Zengwei Zheng, Chun Chen 0001
ICCSA (4)1
2005 Dynamic self-adaptive multicast in wireless networks
abstract
Data multicast technology has received much attention and many achievements have been attained for it. However, most existing multicast techniques rely heavily on client input for data dissemination. In this paper, we investigated an efficient way of automatically identifying the changes of clients' preferences. Our scheme utilizes clustering technique to realize the group of similar clients, and furthermore, dynamically adjusts these clusters to cope with the changes of clients' demands. The rationale underneath is that clients within the same cluster possess similar preferences. The goal of our method is to dynamically present information tailored to specific groups of clients and ensure that the multicast information clients received can satisfy their interests. To conclude this paper, we experimented the recommended method on three different modes of clients' preferences. The simulation results prove it an effective scheme to reduce the amount of superfluous data clients received for certain preference distributions.
Keke Cai, Huaizhong Lin, Guang Qiu
SMC2
2005 Two-Phase Exclusion Based Broadcast Adaptation in Wireless Networks
Keke Cai, Huaizhong Lin, Chun Chen 0001
WAIM2
2004 An event-driven clustering routing algorithm for wireless sensor networks
abstract
Wireless sensor networks (WSNs) are an area of emerging networking research. One obstacle is its limited supply of energy. Therefore, minimizing energy consumption and maximizing system lifetime have been a major design goal for WSNs. This paper presents an energy-efficient event-driven clustering routing algorithm (EDC algorithm) based on unique features of the event-driven data model of WSNs. The algorithm can decide which nodes would become cluster-head nodes according to the maximum remainder energy of nodes, which are sensing an event occurred and are firstly switched to the active state if their several components are in the sleeping state. This strategy can keep sensor nodes with lower remainder energy out of being used up quickly. Detailed simulations of sensor network environments demonstrate that EDC algorithm saves node energy, prolongs system lifetime, and improves evenness of dissipated network energy.
Zengwei Zheng, Zhaohui Wu 0001, Huaizhong Lin
IROS3
2002 Optimistic Voting for Managing Replicated Data
Huaizhong Lin, Chun Chen 0001
J. Comput. Sci. Technol.1