EDBT 2026 Demo / reviewers in the wild / expert
Yan Zhong 0001
dblp:81/5094-1
· DBLP profile ↗
22ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0003-0005-2620ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Category-level Articulated Object Pose Tracking on SE(3) ManifoldsabstractArticulated objects are prevalent in daily life and robotic manipulation tasks. However, compared to rigid objects, pose tracking for articulated objects remains an underexplored problem due to their inherent kinematic constraints. To address these challenges, this work proposes a novel point-pair-based pose tracking framework, termed PPF-Tracker. The proposed framework first performs quasi-canonicalization of point clouds in the SE(3) Lie group space, and then models articulated objects using Point Pair Features (PPF) to predict pose voting parameters by leveraging the invariance properties of SE(3). Finally, semantic information of joint axes is incorporated to impose unified kinematic constraints across all parts of the articulated object. PPF-Tracker is systematically evaluated on both synthetic datasets and real-world scenarios, demonstrating strong generalization across diverse and challenging environments. Experimental results highlight the effectiveness and robustness of PPF-Tracker in multi-frame pose tracking of articulated objects. We believe this work can foster advances in robotics, embodied intelligence, and augmented reality. Xianhui Meng, Yukang Huo, Li Zhang 0104, Liu Liu 0012, Yan Zhong 0001, Pingrui Zhang, Cewu Lu, Jun Liu 0004 |
AAAI | 6 |
| 2026 | MusicRec: Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative RecommendationabstractGenerative recommendation as a new paradigm is influencing the current development of recommender systems. It aims to assign identifiers that capture richer semantic and collaborative information to items, and subsequently predict item identifiers via autoregressive generation using Large Language Models (LLMs). Existing approaches primarily tokenize item text into codebooks with preserved semantic IDs through RQ-VAE, or separately tokenize different modality features of items. However, existing tokenization methods face two major challenges: (1) Learning decoupled multi-modal features limits the quality of the semantic representation. (2) Ignoring collaborative signals from interaction history limits the comprehensiveness of identifiers. To address these limitations, we propose a multi-modal semantic-enhanced identifier with collaborative signals for generative recommendation, named MusicRec. In MusicRec, we propose a tokenization approach based on shared-specific modal fusion, enabling the generated identifiers to preserve semantic information more comprehensively from all modalities. In addition, we incorporate collaborative signals from user interactions to guide identifier generation, preserving collaborative patterns in the semantic representation space. Extensive experiments on three public datasets demonstrate that MusicRec achieves state-of-the-art performance compared to existing baseline methods. Yuqiu Zhao, Lei Shi 0030, Yan Zhong 0001, Feifei Kou, Pengfei Zhang 0010, Jiwei Zhang 0007, Mingying Xu |
AAAI | 3 |
| 2026 | Semi-supervised multi-label feature selection with consistent sparse graph learning
Yan Zhong 0001, Xinping Zhao, Li Zhang 0104, Xinyuan Song 0002, Lei Shi 0030, Bingbing Jiang 0001 |
Neural Networks | 1 |
| 2026 | Pre-Defined Keypoints Worth It: Multi-Modal Learning for Category-Level Articulated Objects Pose EstimationabstractArticulated objects play a vital role in daily interactions, but traditional RGB-based pose estimation methods often face challenges such as lighting variations and shadows. To address these limitations, we introduce PAGE, a novel Pre-defined keypoint-based framework for category-level articulation pose estimation via multi-modal AliGnmEnt. Our approach is motivated by the observation that the distance distribution between heuristically generated keypoints and visible points exhibits a divergent pattern, a phenomenon previously overlooked. To tackle this, we propose a customized unsupervised keypoint estimation method that enhances the stability and robustness of model predictions. Furthermore, to minimize mutual information redundancy between point clouds and RGB images, we design a geometry-color alignment module that fuses features after aligning the two modalities. This is followed by decoding the radius for each visible point and applying our proposal integration scoring strategy to predict keypoints. The framework ultimately outputs the per-part 6D pose of the articulated object.We conduct extensive experiments across diverse datasets, ranging from synthetic to real-world scenarios, demonstrating the robustness and superior performance of PAGE. This work holds significant promise for applications in robotics, embodied intelligence, and augmented reality. Codes and datasets are available at the website: https://sites.google.com/view/pageforart. Li Zhang 0104, Liu Liu 0012, Rujing Wang, Yan Zhong 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2026 | I2EKD: Efficient and Versatile Image-to-Event Knowledge DistillationabstractRecently, general-purpose features for event camera data have become increasingly important in advancing event-based vision applications. Current methods typically adopt pre-training paradigms, yielding promising performance. However, the limited data and sparse spatial information of events hinder effective use of pretraining for rich semantic learning. In this paper, we tackle semantic scarcity by transferring knowledge from large pre-trained image models, without increasing event training data. Concretely, we propose a novel image-to-event knowledge distillation method named I2EKD. Acknowledging that different backbones suit different applications, we fix the teacher and keep the student architecture flexible. To improve versatility, we equip I2EKD with two model-agnostic objectives at the logit and feature levels. Additionally, without task-specific objectives or labels, I2EKD avoids re-distillation and transfers well to downstream applications. Furthermore, leveraging DINOv2 as the teacher, whose feature distribution is built from billions of data, the student can swiftly mimic the superior distribution in a data-efficient manner. Compared with the SOTA pre-training method, I2EKD generates outperforming or comparable features with 1/15 training cost (1/10 data × 2/3 epochs). Extensive experiments on different vision tasks (object recognition, semantic segmentation, and monocular depth) verify the effectiveness of our method. Notably, I2EKD achieves top-1 object recognition accuracy of 70.72%, leading the pre-training SOTA by 5.89%. Hu Cao, Sanqing Qu, Fan Lu 0001, Yan Zhong 0001, Zhichao Lu, Luziwei Leng, Guang Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space ModelsabstractKnown as low energy consumption networks, spiking neural networks (SNNs) have gained a lot of attention within the past decades. While SNNs are increasing competitive with artificial neural networks (ANNs) for vision tasks, they are rarely used for long sequence tasks, despite their intrinsic temporal dynamics. In this work, we develop spiking state space models (SpikingSSMs) for long sequence learning by leveraging on the sequence learning abilities of state space models (SSMs). Inspired by dendritic neuron structure, we hierarchically integrate neuronal dynamics with the original SSM block, meanwhile realizing sparse synaptic computation. Furthermore, to solve the conflict of event-driven neuronal dynamics with parallel computing, we propose a light-weight surrogate dynamic network which accurately predicts the after-reset membrane potential and compatible to learnable thresholds, enabling orders of acceleration in training speed compared with conventional iterative methods. On the long range arena benchmark task, SpikingSSM achieves competitive performance to state-of-the-art SSMs meanwhile realizing on average 90% of network sparsity. On language modeling, our network significantly surpasses existing spiking large language models (spikingLLMs) on the WikiText-103 dataset with only a third of the model size, demonstrating its potential as backbone architecture for low computation cost LLMs. Shuaijie Shen, Renzhuo Huang, Yan Zhong 0001, Qinghai Guo, Zhichao Lu, Jianguo Zhang 0001, Luziwei Leng |
AAAI | 4 |
| 2025 | R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render StrategyabstractHuman life is filled with articulated objects. Previous works for estimating the pose of category-level articulated objects rely on costly 3D point clouds or RGB-D images. In this paper, our goal is to estimate category-level articulation poses from a single RGB image, where we propose R2-Art, a novel category-level Articulation pose estimation framework from a single RGB image and a cascade Render strategy. Given an RGB image as input, R2-Art estimates per-part 6D pose for the articulation. Specifically, we design parallel regression branches tailored to generate camera-to-root translation and rotation. Using the predicted joint states, we perform PC prior transformation and deformation with a joint-centric modeling approach. For further refinement, a cascade render strategy is proposed for projecting the 3D deformed prior onto the 2D mask. Extensive experiments are provided to validate our R2-Art on various datasets ranging from synthetic datasets to real-world scenarios, demonstrating the superior performance and robustness of the R2-Art. We believe that this work has the potential to be applied in many fields including robotics, embodied intelligence, and augmented reality. Li Zhang 0104, Yukang Huo, Yan Zhong 0001, Rujing Wang, Liu Liu 0012 |
AAAI | 4 |
| 2025 | Diff-Art: Category-level Articulation Pose Estimation via Conditional DiffusionabstractArticulated objects are prevalent in people’s daily lives, yet their diverse motion structures pose significant challenges for category-level pose estimation. To address this, in this work, we introduce Diff-Art, a method tailored for category-level articulation pose estimation using conditional diffusion. Given a partial point cloud as input, Diff-Art predicts the per-part 6D pose of the articulated object. Our approach incorporates a novel modeling strategy that exploits the unique kinematic constraints of articulated objects, effectively handling self-occlusion scenarios. Furthermore, we propose a re-scoring strategy to enhance the accuracy of 6D pose estimation within the conditional diffusion framework. Extensive experiments validate the effectiveness of Diff-Art, demonstrating its strong performance on both synthetic datasets and its ability to generalize to real-world scenarios. We believe this work holds significant potential for applications in embodied AI and robotics. Yukang Huo, Xianhui Meng, Li Zhang 0104, Yan Zhong 0001, Mingyuan Yao |
ICME | 5 |
| 2025 | Semi-Supervised Blind Quality Assessment with Confidence-quantifiable Pseudo-label Learning for Authentic ImagesabstractThis paper presents CPL-IQA, a novel semi-supervised blind image quality assessment (BIQA) framework for authentic distortion scenarios. To address the challenge of limited labeled data in IQA area, our approach leverages confidence-quantifiable pseudo-label learning to effectively utilize unlabeled authentically distorted images. The framework operates through a preprocessing stage and two training phases: first converting MOS labels to vector labels via entropy minimization, followed by an iterative process that alternates between model training and label optimization. The key innovations of CPL-IQA include a manifold assumption-based label optimization strategy and a confidence learning method for pseudo-labels, which enhance reliability and mitigate outlier effects. Experimental results demonstrate the framework’s superior performance on real-world distorted image datasets, offering a more standardized semi-supervised learning paradigm without requiring additional supervision or network complexity. Yan Zhong 0001, Chenxi Yang 0004, Suyuan Zhao, Tingting Jiang 0001 |
ICML | 1 |
| 2025 | Pre-defined Keypoints Promote Category-level Articulation Pose Estimation via Multi-Modal AlignmentabstractArticulations are essential in everyday interactions, yet traditional RGB-based pose estimation methods often struggle with issues such as lighting variations and shadows. To overcome these challenges, we propose a novel Pre-defined keypoint based framework for category-level articulation pose estimation via multi-modal Alignment, coined PAGE. Specifically, we first propose a customized keypoint estimation method, aiming to avoid the divergent distance pattern between heuristically generated keypoints and visible points. In addition, to reduce the mutual information redundancy between point clouds and RGB images, we design the geometry-color alignment, which fuses the features after aligning two modalities. This is followed by decoding the radius for each visible point, and applying our proposal integration scoring strategy to predict keypoints. Ultimately, the framework outputs the per-part 6D pose of the articulation. We conduct extensive experiments to evaluate PAGE across a variety of datasets, from synthetic to real-world scenarios, demonstrating its robustness and superior performance. Li Zhang 0104, Liu Liu 0012, Yan Zhong 0001, Rujing Wang |
IJCAI | 4 |
| 2025 | Twin Co-Adaptive Dialogue for Progressive Image GenerationabstractModern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in user prompts. In this work, we present Twin-Co, a framework that leverages synchronized, co-adaptive dialogue to progressively refine image generation. Instead of a static generation process, Twin-Co employs a dynamic, iterative workflow where an intelligent dialogue agent continuously interacts with the user. Initially, a base image is generated from the user's prompt. Then, through a series of synchronized dialogue exchanges, the system adapts and optimizes the image according to evolving user feedback. The co-adaptive process allows the system to progressively narrow down ambiguities and better align with user intent. Experiments demonstrate that Twin-Co not only enhances user experience by reducing trial-and-error iterations but also improves the quality of the generated images, streamlining creative process across various applications. Jianhui Wang 0001, Yangfan He, Yan Zhong 0001, Xinyuan Song 0002, Jiayi Su, Yuheng Feng, Hongyang He, Wenyu Zhu, Xinhang Yuan, Miao Zhang 0010, Tianyu Shi 0003, Xueqian Wang 0001 |
ACM Multimedia | 3 |
| 2025 | Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets TrainingabstractDue to the high cost and small scale of Image Quality Assessment (IQA) datasets, achieving robust generalization remains challenging for prevalent Blind IQA (BIQA) methods. Traditional deep learning-based methods emphasize visual information to capture quality features, while recent developments in Vision-Language Models (VLMs) demonstrate strong potential in learning generalizable representations through textual information. However, applying VLMs to BIQA poses three major Challenges: (1) How to make full use of the multi-modal information. (2) The prompt engineering for appropriate quality description is extremely time-consuming. (3) How to use mixed data for joint training to enhance the generalization of VLM-based BIQA model. To this end, we propose a Multi-modal BIQA method with prompt learning, named MMP-IQA. For (1), we propose a conditional fusion module to better utilize the cross-modality information. By jointly adjusting visual and textual features, our model can capture quality information with a stronger representation ability. For (2), we model the quality prompt's context words with learnable vectors during the training process, which can be adaptively updated for superior performances. For (3), we jointly train a linearity-induced quality evaluator, a relative quality evaluator, and a dataset-specific absolute quality evaluator. In addition, we propose a dual automatic weight adjustment strategy to adaptively balance the loss weights between different datasets and among various losses within the same dataset. Extensive experiments illustrate the superior effectiveness of MMP-IQA. Yan Zhong 0001, Xinping Zhao, Li Zhang 0104, Xinyuan Song 0002, Tingting Jiang 0001 |
ACM Multimedia | 1 |
| 2024 | U-COPE: Taking a Further Step to Universal 9D Category-Level Object Pose Estimation
Li Zhang 0104, Weiqing Meng, Yan Zhong 0001, Jianming Du, Rujing Wang, Liu Liu 0012 |
ECCV (10) | 3 |
| 2024 | Causal-IQA: Towards the Generalization of Image Quality Assessment Based on Causal InferenceabstractDue to the high cost of Image Quality Assessment (IQA) datasets, achieving robust generalization remains challenging for prevalent deep learning-based IQA methods. To address this, this paper proposes a novel end-to-end blind IQA method: Causal-IQA. Specifically, we first analyze the causal mechanisms in IQA tasks and construct a causal graph to understand the interplay and confounding effects between distortion types, image contents, and subjective human ratings. Then, through shifting the focus from correlations to causality, Causal-IQA aims to improve the estimation accuracy of image quality scores by mitigating the confounding effects using a causality-based optimization strategy. This optimization strategy is implemented on the sample subsets constructed by a Counterfactual Division process based on the Backdoor Criterion. Extensive experiments illustrate the superiority of Causal-IQA. Yan Zhong 0001, Li Zhang 0104, Chenxi Yang 0004, Tingting Jiang 0001 |
ICML | 1 |
| 2024 | Large Language Model-Enhanced Algorithm Selection: Towards Comprehensive Algorithm Representation
Yan Zhong 0001, Jibin Wu, Bingbing Jiang 0001, Kay Chen Tan |
IJCAI | 2 |
| 2024 | Multi-View Semi-Supervised Feature Selection with Graph Convolutional NetworksabstractMulti-view semi-supervised feature selection, aims to simultaneously exploit both labeled and unlabeled samples to select a subset of features from multiple feature representations, has become an important task. However, the performance of existing methods is susceptible to the quality of graph due to the following reasons: 1) The samples from different classes located near the boundary are quite close and fail to be classified accurately, causing the unclear neighbor structures. 2) The graph is directly derived from the original space, such that the low-quality features will undermine the true relation between samples. To address above issues, we propose a novel multi-view semi-supervised feature selection method (MVFS), which exploits the regression losses of samples to correct the label information inaccurately propagated via the unreliable neighbor structures on the boundary samples, so as to enhance the discrimination of prediction labels. Moreover, the data representation generated by graph convolutional networks (GCN), which integrates the features, neighbor structures and label information, is incorporated to adaptively update similarity graph to better capture the neighbor structures of samples. Benefiting from these, the discriminative prediction labels and a reliable similarity graph are learned to facilitate the final feature selection. An efficient solution is designed to iteratively optimize MVFS, and comprehensive experiments demonstrate the effectiveness of MVFS. Zhaolong Ling, Peng Zhou 0006, Yan Zhong 0001, Li Li 0037, Weiguo Sheng 0001, Bingbing Jiang 0001 |
IJCNN | 5 |
| 2024 | VoCAPTER: Voting-based Pose Tracking for Category-level Articulated Object via Inter-frame PriorsabstractArticulated objects are common in our daily life. However, current category-level articulation pose works mostly focus on predicting 9D poses on statistical point cloud observations. In this paper, we deal with the problem of category-level online robust 9D pose tracking of articulated objects, where we propose VoCAPTER, a novel 3D Voting-based Category-level Articulated object Pose TrackER. Our VoCAPTER efficiently updates poses between adjacent frames by utilizing partial observations from the current frame and the estimated per-part 9D poses from the previous frame. Specifically, by incorporating prior knowledge of continuous motion relationships between frames, we begin by canonicalizing the input point cloud, casting the pose tracking task as an inter-frame pose increment estimation challenge. Subsequently, to obtain a robust pose-tracking algorithm, our main idea is to leverage SE(3)-invariant features during motion. This is achieved through a voting-based articulation tracking algorithm, which identifies keyframes as reference states for accurate pose updating throughout the entire video sequence. We evaluate the performance of VoCAPTER in the synthetic dataset and real-world scenarios, which demonstrates VoCAPTER's generalization ability to diverse and complicated scenes. Through these experiments, we provide evidence of VoCAPTER's superiority and robustness in multi-frame pose tracking of articulated objects. We believe that this work can facilitate the progress of various fields, including robotics, embodied intelligence, and augmented reality. All the codes will be made publicly available. Li Zhang 0104, Zean Han, Yan Zhong 0001, Qiaojun Yu, Rujing Wang |
ACM Multimedia | 3 |
| 2024 | Rethinking 3D Convolution in $\ell_p$-norm SpaceabstractConvolution is a fundamental operation in the 3D backbone. However, under certain conditions, the feature extraction ability of traditional convolution methods may be weakened. In this paper, we introduce a new convolution method based on $\ell_p$-norm.
For theoretical support, we prove the universal approximation theorem for $\ell_p$-norm based convolution, and analyze the robustness and feasibility of $\ell_p$-norms in 3D point cloud tasks. Concretely, $\ell_{\infty}$-norm based convolution is prone to feature loss. $\ell_2$-norm based convolution is essentially a linear transformation of the traditional convolution. $\ell_1$-norm based convolution is an economical and effective feature extractor. We propose customized optimization strategies to accelerate the training process of $\ell_1$-norm based Nets and enhance the performance. Besides, a theoretical guarantee is given for the convergence by \textit{regret} argument. We apply our methods to classic networks and conduct related experiments. Experimental results indicate that our approach exhibits competitive performance with traditional CNNs, with lower energy consumption and instruction latency. Li Zhang 0104, Yan Zhong 0001, Zhe Min, RujingWang, Liu Liu 0012 |
NeurIPS | 2 |
| 2024 | Nonlinear learning method for local causal structures
Yan Zhong 0001, Zhaolong Ling, Jie Yang 0052, Li Li 0037, Weiguo Sheng 0001, Bingbing Jiang 0001 |
Inf. Sci. | 2 |
| 2023 | Multi-Target Markov Boundary Discovery: Theory, Algorithm, and ApplicationabstractMarkov boundary (MB) has been widely studied in single-target scenarios. Relatively few works focus on the MB discovery for variable set due to the complex variable relationships, where an MB variable might contain predictive information about several targets. This paper investigates the multi-target MB discovery, aiming to distinguish the common MB variables (shared by multiple targets) and the target-specific MB variables (associated with single targets). Considering the multiplicity of MB, the relation between common MB variables and equivalent information is studied. We find that common MB variables are determined by equivalent information through different mechanisms, which is relevant to the existence of the target correlation. Based on the analysis of these mechanisms, we propose a multi-target MB discovery algorithm to identify these two types of variables, whose variant also achieves superiority and interpretability in feature selection tasks. Extensive experiments demonstrate the efficacy of these contributions. Bingbing Jiang 0001, Yan Zhong 0001, Huanhuan Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Multi-Label Local-to-Global Feature SelectionabstractRecent years have witnessed the proliferation of multi-label feature selection, which is an effective data pre-processing step for multi-label learning. And label correlations provide critical information for multi-label feature selection. Existing methods consider global label correlations roughly by assuming that the same label correlations are shared by all samples. Nevertheless, there may exist different local inner label correlations in different local sample sets. Although some methods try to explore them respectively in these local sample sets separated by clustering, they ignore the exterior correlation structure of different local inner correlations, leading to performance degradation. To address this problem, we novelly extract local sample sets for each label through mining Markov blankets of these labels, based on which the proposed multi-label local-to-global feature selection algorithm (ML2G) is employed to select predictive features on these sample subsets. ML2G simultaneously learns different local inner label correlations in each local sample set and the exterior structure of these local inner correlations. Moreover, different from existing methods, ML2G extra considers the asymmetric label correlations, which could describe label correlations more accurately and thus improve the performance of ML2G. Empirical studies validate the superiority of ML2G against state-of-the-art methods on realworld datasets. Yan Zhong 0001, Bingbing Jiang 0001, Huanhuan Chen 0001 |
IJCNN | 1 |
| 2020 | Tolerant Markov Boundary Discovery for Feature SelectionabstractDue to the interpretability and robustness, Markov boundary (MB) has received much attention and been widely applied to causal feature selection. However, enormous empirical studies show that, existing algorithms achieve outstanding performance only on the standard Bayesian network data. While on the real-world data, they could not identify some of the relevant features since the large conditioning set and the ignored multivariate dependence lead to performance degradation. In this paper, we propose a tolerant MB discovery algorithm (TLMB), which maps the feature space and target space to a reproducing kernel Hilbert space through the conditional covariance operator, to measure the causal information carried by a feature. Specifically, TLMB uses a score function to filter the redundant features first and then minimize the trace of the conditional covariance operator, where both of the score function and the optimization problem work in the reproducing kernel Hilbert space so that TLMB can select features with not only pairwise dependence but also multivariate dependence. Moreover, as a MB-based method, TLMB can automatically determine the number of selected features due to the property of MB. Bingbing Jiang 0001, Yan Zhong 0001, Huanhuan Chen 0001 |
CIKM | 3 |