Jiali Cui

dblp:17/2469 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
7since 2021 · last 2026
0000-0002-7834-4562ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 ShaLa: Multimodal Shared Latent Generative Modelling
abstract
This paper presents a novel generative framework for learning shared latent representations across multimodal data. Many advanced multimodal methods focus on capturing all combinations of modality-specific details across inputs, which can inadvertently obscure the high-level semantic concepts that are shared across modalities. Notably, Multimodal VAEs with low-dimensional latent variables are designed to capture shared representations, enabling various tasks such as joint multimodal synthesis and cross-modal inference. However, multimodal VAEs often struggle to design expressive joint variational posteriors and suffer from low-quality synthesis. In this work, ShaLa addresses these challenges by integrating a novel architectural inference model and a second-stage expressive diffusion prior, which not only facilitates effective inference of shared latent representation but also significantly improves the quality of downstream multimodal synthesis. We validate ShaLa extensively across multiple benchmarks, demonstrating superior coherence and synthesis quality compared to state-of-the-art multimodal VAEs. Furthermore, ShaLa scales to many more modalities while prior multimodal VAEs have fallen short in capturing the increasing complexity of the shared latent space.
Jiali Cui, Yan-Ying Chen, Matthew Klenk 0001
AAAI1
2025 Secure transmission cryptographic approach for remote-sensing image based on discrete memristor-coupled Rulkov neuron map and TIMG
Jiali Cui, Yinghong Cao, Hadi Jahanshahi, Jun Mou
Multim. Tools Appl.1
2024 Learning Multimodal Latent Generative Models with Energy-Based Prior
Shiyu Yuan, Jiali Cui, Hanao Li, Tian Han 0001
ECCV (89)2
2024 Learning Latent Space Hierarchical EBM Diffusion Models
abstract
This work studies the learning problem of the energy-based prior model and the multi-layer generator model. The multi-layer generator model, which contains multiple layers of latent variables organized in a top-down hierarchical structure, typically assumes the Gaussian prior model. Such a prior model can be limited in modelling expressivity, which results in a gap between the generator posterior and the prior model, known as the prior hole problem. Recent works have explored learning the energy-based (EBM) prior model as a second-stage, complementary model to bridge the gap. However, the EBM defined on a multi-layer latent space can be highly multi-modal, which makes sampling from such marginal EBM prior challenging in practice, resulting in ineffectively learned EBM. To tackle the challenge, we propose to leverage the diffusion probabilistic scheme to mitigate the burden of EBM sampling and thus facilitate EBM learning. Our extensive experiments demonstrate a superior performance of our diffusion-learned EBM prior on various challenging tasks.
Jiali Cui, Tian Han 0001
ICML1
2023 Learning Joint Latent Space EBM Prior Model for Multi-layer Generator
abstract
This paper studies the fundamental problem of learning multi-layer generator models. The multi-layer generator model builds multiple layers of latent variables as a prior model on top of the generator, which benefits learning complex data distribution and hierarchical representations. However, such a prior model usually focuses on modeling inter-layer relations between latent variables by assuming non-informative (conditional) Gaussian distributions, which can be limited in model expressivity. To tackle this issue and learn more expressive prior models, we propose an energy-based model (EBM) on the joint latent space over all layers of latent variables with the multi-layer generator as its backbone. Such joint latent space EBM prior model captures the intra-layer contextual relations at each layer through layer-wise energy terms, and latent variables across different layers are jointly corrected. We develop a joint training scheme via maximum likelihood estimation (MLE), which involves Markov Chain Monte Carlo (MCMC) sampling for both prior and posterior distributions of the latent variables from different layers. To ensure efficient inference and learning, we further propose a variational training scheme where an inference model is used to amortize the costly posterior MCMC sampling. Our experiments demonstrate that the learned model can be expressive in generating high-quality images and capturing hierarchical features for better outlier detection.
Jiali Cui, Ying Nian Wu, Tian Han 0001
CVPR1
2023 Learning Hierarchical Features with Joint Latent Space Energy-Based Prior
abstract
This paper studies the fundamental problem of multilayer generator models in learning hierarchical representations. The multi-layer generator model that consists of multiple layers of latent variables organized in a top-down architecture tends to learn multiple levels of data abstraction. However, such multi-layer latent variables are typically parameterized to be Gaussian, which can be less informative in capturing complex abstractions, resulting in limited success in hierarchical representation learning. On the other hand, the energy-based (EBM) prior is known to be expressive in capturing the data regularities, but it often lacks the hierarchical structure to capture different levels of hierarchical representations. In this paper, we propose a joint latent space EBM prior model with multi-layer latent variables for effective hierarchical representation learning. We develop a variational joint learning scheme that seamlessly integrates an inference model for efficient inference. Our experiments demonstrate that the proposed joint EBM prior is effective and expressive in capturing hierarchical representations and modelling data distribution.
Jiali Cui, Ying Nian Wu, Tian Han 0001
ICCV1
2023 Learning Energy-based Model via Dual-MCMC Teaching
abstract
This paper studies the fundamental learning problem of the energy-based model (EBM). Learning the EBM can be achieved using the maximum likelihood estimation (MLE), which typically involves the Markov Chain Monte Carlo (MCMC) sampling, such as the Langevin dynamics. However, the noise-initialized Langevin dynamics can be challenging in practice and hard to mix. This motivates the exploration of joint training with the generator model where the generator model serves as a complementary model to bypass MCMC sampling. However, such a method can be less accurate than the MCMC and result in biased EBM learning. While the generator can also serve as an initializer model for better MCMC sampling, its learning can be biased since it only matches the EBM and has no access to empirical training examples. Such biased generator learning may limit the potential of learning the EBM. To address this issue, we present a joint learning framework that interweaves the maximum likelihood learning algorithm for both the EBM and the complementary generator model. In particular, the generator model is learned by MLE to match both the EBM and the empirical data distribution, making it a more informative initializer for MCMC sampling of EBM. Learning generator with observed examples typically requires inference of the generator posterior. To ensure accurate and efficient inference, we adopt the MCMC posterior sampling and introduce a complementary inference model to initialize such latent MCMC sampling. We show that three separate models can be seamlessly integrated into our joint framework through two (dual-) MCMC teaching, enabling effective and efficient EBM learning.
Jiali Cui, Tian Han 0001
NeurIPS1
2010 DHV Image Registration Using Boundary Optimization
Jiali Cui
ICIC (1)1
2010 Study of Hand-Dorsa Vein Recognition
Jiali Cui, Lik-Kwan Shark, Martin R. Varley
ICIC (1)3
2005 Improving iris recognition accuracy via cascaded classifiers
abstract
As a reliable approach to human identification, iris recognition has received increasing attention in recent years. The most distinguishing feature of an iris image comes from the fine spatial changes of the image structure. So iris pattern representation must characterize the local intensity variations in iris signals. However, the measurements from minutiae are easily affected by noise, such as occlusions by eyelids and eyelashes, iris localization error, nonlinear iris deformations, etc. This greatly limits the accuracy of iris recognition systems. In this paper, an elastic iris blob matching algorithm is proposed to overcome the limitations of local feature based classifiers (LFC). In addition, in order to recognize various iris images efficiently a novel cascading scheme is proposed to combine the LFC and an iris blob matcher. When the LFC is uncertain of its decision, poor quality iris images are usually involved in intra-class comparison. Then the iris blob matcher is resorted to determine the input iris' identity because it is capable of recognizing noisy images. Extensive experimental results demonstrate that the cascaded classifiers significantly improve the system's accuracy with negligible extra computational cost.
Zhenan Sun, Yunhong Wang 0001, Tieniu Tan, Jiali Cui
IEEE Trans. Syst. Man Cybern. Part C4
2004 Fast recursive mathematical morphological transforms
abstract
Since many mathematical morphology operations are recursive transforms of dilation and erosion, this paper proposes fast recursive transforms to reduce computational complexity. The basic idea of the method is to compute the temporary results within a series of adaptive windows and the computing is performed on specific pixels. Each step of the recursive process consists of two parts: 1) computation is limited to the specific pixels (foreground or background pixels) within a window; 2) update the window adoptively and delete those varied pixels. Extensive results show that the time complexity of the method is proportional to the number of the specific pixels.
Jiali Cui, Yunhong Wang 0001, Tieniu Tan, Zhenan Sun
ICIG1
2004 Noise removal and impainting model for IRIS image
abstract
Noise removal is an important problem for iris recognition. If the iris regions were not correctly segmented in iris images, segmented iris regions possibly include noises, namely eyelashes, eyelids, reflections and pupil. Noises influence the features of both noise regions and their neighboring regions, which will result in poor recognition performance. To solve this problem, this paper proposes a method for removing noises and impainting iris images. The whole procedure includes three steps: 1) localization and normalization, 2) noise removal based on phase congruency and 3) iris image impainting. A series of experiments show that the proposed method has encouraging performance for improving the recognition accuracy.
Junzhou Huang, Yunhong Wang 0001, Jiali Cui, Tieniu Tan
ICIP3
2004 Cascading statistical and structural classifiers for iris recognition
abstract
Reliable human identification using iris pattern has recently gained growing interests from pattern recognition researchers. In literature of iris recognition, almost all algorithms are based on statistical information. In this paper, a structural iris image analysis method is proposed, which provides complementary information to statistical classifier. In order to save computational cost, the structural matcher is not consulted unless the statistical classifier is uncertain of its decision. At the second stage, the structural classifier may be combined with statistical classifier with different fusion strategies. The experimental results of decision-level classifiers combination are reported, which demonstrate that the cascaded classification system significantly outperforms single classifier.
Zhenan Sun, Yunhong Wang 0001, Tieniu Tan, Jiali Cui
ICIP4