Weiling Cai

dblp:02/1548 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Semantic-Guided Coarse-to-Fine Diffusion Model for Self-Supervised Image Shadow Removal
Ziqi Zeng, Weiling Cai
CVM (2)3
2025 O-Mamba: O-Shape State-Space Model for Underwater Image Enhancement
Chenyu Dong, Weiling Cai
PRCV (9)3
2025 Multi-cropping contrastive learning and domain consistency for unsupervised image-to-image translation
abstract
Abstract Recently, unsupervised image‐to‐image (i2i) translation methods based on contrastive learning have achieved state‐of‐the‐art results. However, in previous works, the negatives are sampled from the input image itself, which inspires us to design a data augmentation method to improve the quality of the selected negatives. Moreover, the previous methods only preserve the content consistency via patch‐wise contrastive learning, which ignores the domain consistency between the generated images and the real images of the target domain. This paper proposes a novel unsupervised i2i translation framework based on multi‐cropping contrastive learning and domain consistency, called MCDUT. Specifically, the multi‐cropping views are obtained with the aim of further generating high‐quality negative examples. To constrain the embeddings in the deep feature space, a new domain consistency loss is formulated, which encourages the generated images to be close to the real images. In many i2i translation tasks, this method achieves state‐of‐the‐art results, and the advantages of this method have been proven through extensive comparison experiments and ablation research. The code of MCDUT is available at https://github.com/zhihefang/MCDUT .
Weiling Cai, Chengwei Hu
IET Image Process.2
2025 Spectral normalization and dual contrastive regularization for image-to-image translation
Weiling Cai
Vis. Comput.2
2024 Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image Restoration
abstract
Underwater images are subject to intricate and diverse degradation, inevitably affecting the effectiveness of underwater visual tasks. However, most approaches primarily operate in the raw pixel space of images, which limits the exploration of the frequency characteristics of underwater images, leading to an inadequate utilization of deep models' representational capabilities in producing high-quality images. In this paper, we introduce a novel Underwater Image Enhancement (UIE) framework, named WF-Diff, designed to fully leverage the character-istics of frequency domain information and diffusion models. WF-Diff consists of two detachable networks: Wavelet-based Fourier information interaction network (WFI2-net) and Frequency Residual Diffusion Adjustment Module (FR-DAM). With our full exploration of the frequency domain in-formation, WFI2-net aims to achieve preliminary enhancement of frequency information in the wavelet space. Our proposed FRDAM can further refine the highand low-frequency information of the initial enhanced images, which can be viewed as a plug-and-play universal module to adjust the detail of the underwater images. With the above techniques, our algorithm can show SOTA performance on real-world underwater image datasets, and achieves competitive performance in visual quality. The code is available at https://github.com/zhihefang/WF-Diff.
Weiling Cai, Chenyu Dong, Chengwei Hu
CVPR2
2024 Toward Sufficient Spatial-Frequency Interaction for Gradient-Aware Underwater Image Enhancement
abstract
Underwater images suffer from complex and diverse degradation, which inevitably affects the performance of underwater visual tasks. However, most existing learning-based underwater image enhancement (UIE) methods mainly restore such degradations in the spatial domain, and rarely pay attention to the fourier frequency information. In this paper, we develop a novel UIE framework based on spatial-frequency interaction and gradient maps, namely SFGNet, which consists of two stages. Specifically, in the first stage, we propose a dense spatial-frequency fusion network (DSFFNet), mainly including our designed dense fourier fusion block and dense spatial fusion block, achieving sufficient spatial-frequency interaction by cross connections between these two blocks. In the second stage, we propose a gradient-aware corrector (GAC) to further enhance perceptual details and geometric structures of images by gradient map. Experimental results on two real-world underwater image datasets show that our approach can successfully enhance underwater images, and achieves competitive performance in visual quality improvement. The code is available at https://github.com/zhihefang/SFGNet.
Weiling Cai, Chenyu Dong, Ziqi Zeng
ICASSP2
2024 ECT: Efficient Cross Transformer for Image Deblurring
abstract
Image deblurring presents a complex challenge intending to renew visual clarity in images affected by camera shake or object motion. However, traditional deblurring methodologies tend to emphasize local features, ignoring critical contextual information, which consequently limits their efficacy in addressing common blurry image issues. This paper proposes Efficient Cross Transformer (ECT) to overcome the limitations of inadequate global features in image restoration across different scales. ECT designs the cross-attention layers to achieve interaction between input tokens and network tokens, aiming to efficiently capture image details in the input space and feature information in the latent space of each layer. The cross-attention mechanism works in latent layers, integrating image features from different scales while reducing the computational burden compared to traditional Transformers. Furthermore, ECT employs windowing techniques in a strategic manner to capture local hazy components, duly amplifying its image restoration abilities. Empirical evidence shows that ECT has achieved state-of-the-art results in deblurring images, demonstrating excellent performance without requiring a large training dataset. Consequently, ECT emerges as a promising solution for rectifying blurred images across artificial and real-world environments.
Chengwei Hu, Weiling Cai
IJCNN2
2024 Cycle contrastive adversarial learning with structural consistency for unsupervised high-quality image deraining transformer
Weiling Cai, Chengwei Hu
Neural Networks2
2023 Creative Geotechnical Engineering Education Module Based on an Educational Game Using Multiphysics Enriched Mixed Reality
abstract
This work-in-progress paper discusses the development of an educational game to provide integrated geotechnical engineering education modules that connect theoretical concepts, laboratory testing, field investigation, and engineering design. The game, Earth Trek, is developed based on the design of geothermal piles, which are an innovative and sustainable geotechnical engineering approach to combat climate change. Virtual reality is applied to visualize the field environments, laboratory conditions, and design components for structural simulation. The game uses a combination of storytelling and tasks to engage students with geotechnical concepts in an enjoyable way. With the newly developed game, geotechnical engineering instructors can provide students with exposure to laboratory testing and field environments, improving the quality of geotechnical engineering education. The use of multiphysics enriched mixed reality gaming allows for a visual representation of the connections between theoretical concepts, laboratory testing, field investigation, and engineering design. Additionally, this study discusses the challenges that geotechnical students face when dealing with worldwide concerns such as energy demand, environmental protection, infrastructure sustainability, and hazard reduction. Earth Trek allows students to apply geotechnical engineering knowledge to explore the underground space and the associated geothermal energy to tackle the engineering problems using only their smartphones. Through exploring the virtual environment and completing game tasks, students can obtain different testing tools used for geotechnical experiments, including thermal conductivity measurement and direct shear test. They are also trained to conduct parametric study to explore the influence of boundary conditions on thermal transfer efficiency of the geothermal pile. The key contribution of this work is to illustrate an educational paradigm based on mixed reality, moving towards creative engineering education in geotechnical engineering. The newly developed educational game and Earth Trek are expected to enhance geotechnical engineering education and provide students with an interdisciplinary knowledge to tackle worldwide concerns.
Luobin Cui, Weiling Cai, Ryan Hare, ChenChen Huang, Ying Tang 0001
FIE2
2023 Context-FPN and Memory Contrastive Learning for Partially Supervised Instance Segmentation
Weiling Cai
PRCV (12)2
2022 Unsupervised Photo-to-Caricature Generation with Adaptive Select Layer-Instance Normalization and Semi-cycle Consistency
Weiling Cai, Cairun Wang
ACML2
2022 Makeup Transfer Based on Generative Adversarial Network for Large Angle Spatial Misalignment
Cairun Wang, Weiling Cai
ICANN (3)2
2022 Semantic Diversity Image Translation Based on Deep Feature Difference and Attention Mechanism
Weiling Cai, Cairun Wang
ICANN (3)2
2022 C-Lop: Accurate contention-based modeling of MPI concurrent communication
Ziheng Wang 0002, Heng Chen 0002, Weiling Cai, Xiaoshe Dong, Xingjun Zhang
Parallel Comput.3
2021 Multi-view Latent Subspace Clustering based on both Global and Local Structure
abstract
Most existing multi-view clustering methods focus on the global structure or local structure among samples, and few methods focus on the two structures at the same time. In this paper, we propose a Multi-view Latent subspace Clustering based on both Global and Local structure (MLCGL). In this method, a latent embedding representation is learned by exploring the complementary information from different views. In the latent space, not only the global reconstruction relationship but also the local geometric structure among the latent variables are discovered. In this way, a unified affinity graph matrix is constructed in the latent space for different views, which indicates a clear between-class relationship. Meanwhile, a rank constraint is introduced on the Laplacian graph to facilitate the division of samples into the required clusters. In MLCGL, the affinity graph also provides positive feedback to optimize the learned latent representation and contribute to divided it into reasonable clusters. Moreover, we present an alternating iterative optimization scheme to optimize objective functions. Compared with the state-of-art algorithms, MLCGL has achieved excellent experimental performance on several real-world datasets.
Honghan Zhou, Weiling Cai
ACML2
2021 Semi-paired Image-to-Image Translation using Neighbor-based Generative Adversarial Networks
abstract
Image-to-image translation aims at learning the mapping between an input image and an output image using a training set of aligned image pairs. In reality, obtaining paired images is difficult and expensive. Generally, the data often exist in the form of partial pairing, that is, a small number of images are paired and most of the images are not paired. In this paper, we present a semi-paired image-to-image translation approach using neighbor-based generative adversarial networks. Our goal is to break the restriction that training images must be paired, and meanwhile guarantee the quality of image translation. For the unpaired images, we introduce an inverse mapping and cycle consistency loss to enforce the image reconstruction; for the paired images, we make full use of the one-to-one strong correlation to guide the image translation. To further take advantage of the paired images, our approach employs neighbor images to further expand the paired information and establishes the neighbor-based cycle consistency. Our method is characterized by flexibility and adaptability under various scenarios, such as target deformation, day-night transformation, etc. Compared with the previous methods, the experimental results prove the superiority of our method.
Weiling Cai, Honghan Zhou
IJCNN2
2019 Multi-view Locality Preserving Embedding with View Consistent Constraint for Dimension Reduction
Weiling Cai, Ming Yang 0014, Fengyi Song
KSEM (1)2
2018 Attributes Consistent Faces Generation Under Arbitrary Poses
Fengyi Song, Jinhui Tang 0001, Ming Yang 0014, Weiling Cai, Wanqi Yang
ACCV (2)4
2018 Image filtering method using trimmed statistics and edge preserving
abstract
Image filtering is to retain the details of the image as much as possible and meanwhile suppress the noise pollution to great extent. This study presents an image filtering using the truncated statistics and edge preserving. In the first step of our method, the alpha‐trimmed filter is utilized to remove a variety of types of noises; in the second step, taking the image after alpha‐trimmed filtering as a guide image, the local linear model between the guide image and the target image is established; in the third step, the obtained local linear model is further simplified to reduce the time complexity; and finally, using the relationship between image local variance and the global variance, the local linear model is modified to enhance the details of the image and meanwhile remove halo phenomenon. This method has three advantages: (i) it is flexible to deal with the images stained by various types of high‐intensity noise; (ii) it is effective to keep the image details and profile information, and remove the halo phenomenon; and (iii) it runs in time linear in the image size, thus its computation complexity is low. Experimental results show that the proposed filter is robust and efficient.
Weiling Cai, Ming Yang 0014, Fengyi Song
IET Image Process.1
2017 A dimension reduction algorithm preserving both global and local clustering structure
Weiling Cai
Knowl. Based Syst.1
2015 A manifold learning framework for both clustering and classification
Weiling Cai
Knowl. Based Syst.1
2012 Simultaneous clustering and classification over cluster structure representation
Songcan Chen, Weiling Cai
Pattern Recognit.3
2010 A multiobjective simultaneous learning framework for clustering and classification
abstract
Traditional pattern recognition involves two tasks: clustering learning and classification learning. Clustering result can enhance the generalization ability of classification learning, while the class information can improve the accuracy of clustering learning. Hence, both learning methods can complement each other. To fuse the advantages of both learning methods together, many existing algorithms have been developed in a sequential fusing way by first optimizing the clustering criterion and then the classification criterion associated with the obtained clustering results. However, such kind of algorithms naturally fails to achieve the simultaneous optimality for two criteria, and thus has to sacrifice either the clustering performance or the classification performance. To overcome that problem, in this paper, we present a multiobjective simultaneous learning framework (MSCC) for both clustering and classification learning. MSCC utilizes multiple objective functions to formulate the clustering and classification problems, respectively, and more importantly, it employs the Bayesian theory to make these functions all only dependent on a set of the same parameters, i.e., clustering centers which play a role of the bridge connecting the clustering and classification learning. By simultaneously optimizing the clustering centers embedded in these functions, not only the effective clustering performance but also the promising classification performance can be simultaneously attained. Furthermore, from the multiple Pareto-optimality solutions obtained in MSCC, we can get an interesting observation that there is complementarity to great extent between clustering and classification learning processes. Empirical results on both synthetic and real data sets demonstrate the effectiveness and potential of MSCC.
Weiling Cai, Songcan Chen, Daoqiang Zhang
IEEE Trans. Neural Networks1
2009 A simultaneous learning framework for clustering and classification
Weiling Cai, Songcan Chen, Daoqiang Zhang
Pattern Recognit.1
2007 Fast and robust fuzzy c-means clustering algorithms incorporating local information for image segmentation
Weiling Cai, Songcan Chen, Daoqiang Zhang
Pattern Recognit.1
2007 Robust fuzzy relational classifier incorporating the soft class labels
Weiling Cai, Songcan Chen, Daoqiang Zhang
Pattern Recognit. Lett.1