Rongzhen Zhao

dblp:137/9968 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Predicting Video Slot Attention Queries from Random Slot-Feature Pairs
abstract
Unsupervised video Object-Centric Learning (OCL) is promising as it enables object-level scene representation and understanding as we humans do. Mainstream video OCL methods adopt a recurrent architecture: An aggregator aggregates current video frame into object features, termed slots, under some queries; A transitioner transits current slots to queries for the next frame. This is an effective architecture but all existing implementations both (i1) neglect to incorporate next frame features, the most informative source for query prediction, and (i2) fail to learn transition dynamics, the knowledge essential for query prediction. To address these issues, we propose Random Slot-Feature pair for learning Query prediction (RandSF.Q): (t1) We design a new transitioner to incorporate both slots and features, which provides more information for query prediction; (t2) We train the transitioner to predict queries from slot-feature pairs randomly sampled from available recurrences, which drives it to learn transition dynamics. Experiments on scene representation demonstrate that our method surpass existing video OCL methods significantly, e.g., up to 10 points on object discovery, setting new state-of-the-art. Such superiority also benefits downstream tasks like scene understanding.
Rongzhen Zhao, Juho Kannala, Joni Pajarinen
AAAI1
2026 Joint embedding for multi-structural hypergraph based dimensionality reduction in rotor fault diagnosis
Yongfei Zhang, Qibo Liang, Yuqiao Zheng, Rongzhen Zhao, Linfeng Deng, Mingkuan Shi, Kongyuan Wei
Eng. Appl. Artif. Intell.4
2025 Multi-Scale Fusion for Object Representation
abstract
Representing images or videos as object-level feature vectors, rather than pixel-level feature maps, facilitates advanced visual tasks. Object-Centric Learning (OCL) primarily achieves this by reconstructing the input under the guidance of Variational Autoencoder (VAE) intermediate representation to drive so-called slots to aggregate as much object information as possible. However, existing VAE guidance does not explicitly address that objects can vary in pixel sizes while models typically excel at specific pattern scales. We propose Multi-Scale Fusion (MSF) to enhance VAE guidance for OCL training. To ensure objects of all sizes fall within VAE's comfort zone, we adopt the image pyramid, which produces intermediate representations at multiple scales; To foster scale-invariance/variance in object super-pixels, we devise inter/intra-scale fusion, which augments low-quality object super-pixels of one scale with corresponding high-quality super-pixels from another scale. On standard OCL benchmarks, our technique improves mainstream methods, including state-of-the-art diffusion-based ones. The source code is available on https://github.com/Genera1Z/MultiScaleFusion.
Rongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni Pajarinen
ICLR1
2025 Slot Attention with Re-Initialization and Self-Distillation
abstract
Unlike popular solutions based on dense feature maps, Object-Centric Learning (OCL) represents visual scenes as sub-symbolic object-level feature vectors, termed slots, which are highly versatile for tasks involving visual modalities. OCL typically aggregates object superpixels into slots by iteratively applying competitive cross attention, known as Slot Attention, with the slots as the query. However, once initialized, these slots are reused naively, causing redundant slots to compete with informative ones for representing objects. This often results in objects being erroneously segmented into parts. Additionally, mainstream methods derive supervision signals solely from decoding slots into the input's reconstruction, overlooking potential supervision based on internal information. To address these issues, we propose Slot Attention with re-Initialization and self-Distillation (DIAS): i) We reduce redundancy in the aggregated slots and re-initialize extra aggregation to update the remaining slots; ii) We drive the bad attention map at the first aggregation iteration to approximate the good at the last iteration to enable self-distillation. Experiments demonstrate that DIAS achieves state-of-the-art on OCL tasks like object discovery and recognition, while also improving advanced visual prediction and reasoning. Our source code and model checkpoints are available on https://github.com/Genera1Z/DIAS.
Rongzhen Zhao, Yi Zhao 0014, Juho Kannala, Joni Pajarinen
ACM Multimedia1
2025 Vector-Quantized Vision Foundation Models for Object-Centric Learning
abstract
Object-Centric Learning (OCL) aggregates image or video feature maps into object-level feature vectors, termed slots. It's self-supervision of reconstructing the input from slots struggles with complex object textures, thus Vision Foundation Model (VFM) representations are used as the aggregation input and reconstruction target. Existing methods leverage VFM representations in diverse ways yet fail to fully exploit their potential. In response, we propose a unified architecture, Vector-Quantized VFMs for OCL (VQ-VFM-OCL, or VVO). The key to our unification is simply shared quantizing VFM representations in OCL aggregation and decoding. Experiments show that across different VFMs, aggregators and decoders, our VVO consistently outperforms baselines in object discovery and recognition, as well as downstream visual prediction and reasoning. We also mathematically analyze why VFM representations facilitate OCL aggregation and why their shared quantization as reconstruction targets strengthens OCL supervision. Our source code and model checkpoints are available on https://github.com/Genera1Z/VQ-VFM-OCL.
Rongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni Pajarinen
ACM Multimedia1
2025 MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
abstract
Learning object-level, structured representations is widely regarded as a key to better generalization in vision and underpins the design of next-generation Pre-trained Vision Models (PVMs). Mainstream Object-Centric Learning (OCL) methods adopt Slot Attention or its variants to iteratively aggregate objects' super-pixels into a fixed set of query feature vectors, termed slots. However, their reliance on a static slot count leads to an object being represented as multiple parts when the number of objects varies. We introduce MetaSlot, a plug-and-play Slot Attention variant that adapts to variable object counts. MetaSlot (i) maintains a codebook that holds prototypes of objects in a dataset by vector-quantizing the resulting slot representations; (ii) removes duplicate slots from the traditionally aggregated slots by quantizing them with the codebook; and (iii) injects progressively weaker noise into the Slot Attention iterations to accelerate and stabilize the aggregation. MetaSlot is a general Slot Attention variant that can be seamlessly integrated into existing OCL architectures. Across multiple public datasets and tasks--including object discovery and recognition--models equipped with MetaSlot achieve significant performance gains and markedly interpretable slot representations, compared with existing Slot Attention variants. The code is available at https://github.com/lhj-lhj/MetaSlot.
Hongjia Liu, Rongzhen Zhao, Haohan Chen, Joni Pajarinen
NeurIPS2
2025 Grouped Discrete Representation for Object-Centric Learning
Rongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni Pajarinen
ECML/PKDD (6)1
2025 Multimetric hypergraph embedding for dimensionality reduction in rotor fault diagnosis
Yongfei Zhang, Yuqiao Zheng, Rongzhen Zhao, Linfeng Deng, Mingkuan Shi, Kongyuan Wei
Adv. Eng. Informatics3
2025 Multi-source contrastive cluster center method for cross-domain bearing fault identification
Lizhen Wu, Rongzhen Zhao, Kongyuan Wei, Yuqiao Zheng, Linfeng Deng, Yongfei Zhang, Mingkuan Shi
Eng. Appl. Artif. Intell.3
2025 Dimensionality reduction of rolling bearing fault data based on graph-embedded semi-supervised deep auto-encoders
Wei Kongyuan, Rongzhen Zhao, Haixia Kou, Yongyong Cao, Yuqiao Zheng, Linfeng Deng
Eng. Appl. Artif. Intell.2
2023 Unsupervised structure subdomain adaptation based the Contrastive Cluster Center for bearing fault diagnosis
Rongzhen Zhao, Tianjing He, Kongyuan Wei, Jianhui Yuan
Eng. Appl. Artif. Intell.2
2022 Convolution of Convolution: Let Kernels Spatially Collaborate
abstract
In the biological visual pathway especially the retina, neurons are tiled along spatial dimensions with the electrical coupling as their local association, while in a convolution layer, kernels are placed along the channel dimension singly. We propose convolution of convolution, associating kernels in a layer and letting them collaborate spatially. With this method, a layer can provide feature maps with extra transformations and learn its kernels together instead of isolatedly. It is only used during training, bringing in negligible extra costs; then it can be re-parameterized to common convolution before testing, boosting performance gratuitously in tasks like classification, detection and segmentation. Our method works even better when larger receptive fields are demanded. The code is available on site: https://github.com/Genera1Z/ConvolutionOfConvolution.
Rongzhen Zhao, Zhenzhi Wu
CVPR1
2022 Intelligent fault diagnosis of rolling bearing based on novel CNN model considering data imbalance
Ziyang Xing, Rongzhen Zhao, Yaochun Wu, Tianjing He
Appl. Intell.2
2022 Modeling learnable electrical synapse for high precision spatio-temporal recognition
Zhenzhi Wu, Zhihong Zhang 0008, Huanhuan Gao, Rongzhen Zhao, Guang-She Zhao, Guoqi Li 0002
Neural Networks5
2021 Intelligent fault diagnosis of rolling bearings using a semi-supervised convolutional neural network
Yaochun Wu, Rongzhen Zhao, Wuyin Jin, Tianjing He, Sencai Ma, Mingkuan Shi
Appl. Intell.2
2021 Learnable Heterogeneous Convolution: Learning both topology and strength
Rongzhen Zhao, Zhenzhi Wu, Qikun Zhang
Neural Networks1