EDBT 2026 Demo / reviewers in the wild / expert
Li Zhang 0104
dblp:89/5992-104
· DBLP profile ↗
29ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0003-1610-6056ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 19 · 7 first-author · 18 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SIAM: Towards Generalizable Articulated Object Modeling via Single Robot-Object InteractionabstractArticulated object modeling, which represents interconnected rigid bodies with their geometry, part segmentation, articulation tree, and physical properties, is crucial for robotic perception and manipulation. Recently existing methods like SAGCI leverage Interactive Perception (IP) to refine models through robot interaction. However, SAGCI suffers from prior-dependency (requiring initialization), neglects kinematic/dynamic constraints, and generates non-watertight meshes. To overcome these limitations, we propose SIAM, a novel framework for efficient and generalizable Single-Interaction Articulated Modeling. Given an initial point cloud, SIAM first enables minimal robot interaction to trigger object motion. It then precisely segments parts by analyzing point cloud differences pre- and post-interaction. For joint parameter estimation, we introduce an optimization incorporating novel kinematic energy constraints, enhancing physical consistency. Finally, we reconstruct a high-quality, topologically watertight mesh by learning 3D Gaussian Primitives from multi-view RGB-D observations under deformation. Extensive experiments on the PartNet-Mobility benchmark demonstrate state-of-the-art articulation modeling performance. Successful real-world deployment with an xArm robot further validates the framework's practicality and transferability. SIAM achieves accurate, prior-free modeling with significantly reduced interaction cost. Li Zhang 0104, Yan Zhang 0053, Anran Huang, Liu Liu 0012, Dan Guo 0001 |
AAAI | 2 |
| 2026 | Uncertainty-Guided View-Strength-Aware Feature Utilization for Multi-View ClassificationabstractIn multi-view classification tasks (MVC), each view provides an unique perspective on the data, offering complementary information that can improve classification performance when properly integrated. However, traditional methods typically adopt a uniform processing strategy for all views before fusion, overlooking the fact that different views may require different treatments due to variations in their quality and informativeness. To address this limitation, we propose a novel framework called Uncertainty-Guided View-Strength-Aware Feature Utilization (UVF) for multi-view classification. Our approach introduces a view uncertainty estimation module to quantify the discriminative strength of each view. Based on this estimation, a Differentiated Feature Selector (DFS) adaptively selects features, retaining informative dimensions in weak views while preserving original features in strong views. Furthermore, we employ an uncertainty-guided fusion strategy that assigns dynamic weights to each view's contribution based on its uncertainty score, enhancing the robustness and reliability of the final decision. Experimental results on benchmark datasets demonstrate that our method significantly outperforms conventional approaches, achieving better classification accuracy and interpretability through strength-aware feature processing and fusion. Qian Guo 0005, Li Zhang 0104, Liang Du 0003, Bingbing Jiang 0001, Lu Chen 0003, Xinyan Liang |
AAAI | 3 |
| 2026 | Exploring Category-level Articulated Object Pose Tracking on SE(3) ManifoldsabstractArticulated objects are prevalent in daily life and robotic manipulation tasks. However, compared to rigid objects, pose tracking for articulated objects remains an underexplored problem due to their inherent kinematic constraints. To address these challenges, this work proposes a novel point-pair-based pose tracking framework, termed PPF-Tracker. The proposed framework first performs quasi-canonicalization of point clouds in the SE(3) Lie group space, and then models articulated objects using Point Pair Features (PPF) to predict pose voting parameters by leveraging the invariance properties of SE(3). Finally, semantic information of joint axes is incorporated to impose unified kinematic constraints across all parts of the articulated object. PPF-Tracker is systematically evaluated on both synthetic datasets and real-world scenarios, demonstrating strong generalization across diverse and challenging environments. Experimental results highlight the effectiveness and robustness of PPF-Tracker in multi-frame pose tracking of articulated objects. We believe this work can foster advances in robotics, embodied intelligence, and augmented reality. Xianhui Meng, Yukang Huo, Li Zhang 0104, Liu Liu 0012, Yan Zhong 0001, Pingrui Zhang, Cewu Lu, Jun Liu 0004 |
AAAI | 3 |
| 2026 | EvoFMVC: Trusted Federated Multi-View Clustering with Evolutionary FusionabstractWith the growing demand for decentralized collaborative analysis of privacy-sensitive data, federated multi-view clustering (FMVC) has attracted widespread attention due to its ability to balance privacy protection and collaborative modeling. However, current methods still face the following challenges: (1) Clients need to frequently upload high-dimensional data such as model parameters or graph structures, resulting in high communication costs; (2) The structured data uploaded often contains semantic features and has a high risk of being inverted; (3) The server usually merges the data from all clients with the fixed fusion rule, which may result in a suboptimized clustering result when there exist low-quality clients. To address the issues, we propose a new trusted federated multi-view clustering framework (EvoFMVC) that introduces three key innovations: First, lightweight trusted evidence serves as a compact communication medium, significantly reducing overhead compared to conventional model parameters or graph structures. Second, trusted evidences express clustering results in the form of probability distribution, which avoids the risk of structured information being easily inverted. Lastly, we formalize the server-side aggregation process as a neural architecture search (NAS) task where the server flexibly uses different fusion operators to filter and fuse necessary views through evolutionary algorithms, which significantly improves the fusion effect and model performance. Experimental results on multiple datasets show that our method is superior to existing FMVC methods in terms of clustering accuracy and communication efficiency. Li Zhang 0104, Pinhan Fu, Qian Guo 0005, Liang Du 0003, Xinyan Liang |
AAAI | 1 |
| 2026 | Semi-supervised multi-label feature selection with consistent sparse graph learning
Yan Zhong 0001, Xinping Zhao, Li Zhang 0104, Xinyuan Song 0002, Lei Shi 0030, Bingbing Jiang 0001 |
Neural Networks | 4 |
| 2026 | Probing Effective and Efficient Category-Level Articulated Object Pose PerceptionabstractCategory-level articulated object pose perception-encompassing both static pose estimation and dynamic pose tracking-is critical for embodied AI systems interacting with complex environments. Due to the inherent complexity and diverse motion structures of articulated objects, existing methods often exhibit limitations in adequately modeling kinematic constraints, handling self-occlusions, and meeting optimization requirements. Building upon EfficientCAPER (Yu et al., 2024), this work introduces CAPER++, a unified framework addressing these limitations through three key innovations: first, a joint-centric hierarchical model decomposes objects into a root part and constrained parts linked by joints, explicitly embedding kinematic constraints for geometrically consistent pose recovery. Second, an SE(3) manifold formulation leverages Lie algebra in the tangent space for singularity-free rotation representation and stable optimization, replacing error-prone direct regression. Third, for tracking, a proxy canonicalization strategy reformulates pose updates as SE(3) increment predictions relative to keyframes, enhanced by a dynamic keyframe mechanism to suppress drift. Extensive experiments on synthetic (ArtImage, PM-Videos), semi-synthetic (ReArtMix, ReArt-Videos), and real-world (RobotArm, RobotArm-Videos) benchmarks demonstrate state-of-the-art accuracy and robustness. CAPER++ achieves real-time inference (50 FPS) without post-processing, significantly advancing category-level articulated perception for real-world applications. Li Zhang 0104, Xianhui Meng, Liu Liu 0012, Rujing Wang, Cewu Lu, Jun Liu 0004, Hong Zhang 0013 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | iClickSeg: Interactive click segmentation for zero-shot cross-category 3D part segmentation
Xing Yi, Liu Liu 0012, Qiupu Chen, Li Zhang 0104, Dan Guo 0001 |
Pattern Recognit. | 4 |
| 2026 | Pre-Defined Keypoints Worth It: Multi-Modal Learning for Category-Level Articulated Objects Pose EstimationabstractArticulated objects play a vital role in daily interactions, but traditional RGB-based pose estimation methods often face challenges such as lighting variations and shadows. To address these limitations, we introduce PAGE, a novel Pre-defined keypoint-based framework for category-level articulation pose estimation via multi-modal AliGnmEnt. Our approach is motivated by the observation that the distance distribution between heuristically generated keypoints and visible points exhibits a divergent pattern, a phenomenon previously overlooked. To tackle this, we propose a customized unsupervised keypoint estimation method that enhances the stability and robustness of model predictions. Furthermore, to minimize mutual information redundancy between point clouds and RGB images, we design a geometry-color alignment module that fuses features after aligning the two modalities. This is followed by decoding the radius for each visible point and applying our proposal integration scoring strategy to predict keypoints. The framework ultimately outputs the per-part 6D pose of the articulated object.We conduct extensive experiments across diverse datasets, ranging from synthetic to real-world scenarios, demonstrating the robustness and superior performance of PAGE. This work holds significant promise for applications in robotics, embodied intelligence, and augmented reality. Codes and datasets are available at the website: https://sites.google.com/view/pageforart. Li Zhang 0104, Liu Liu 0012, Rujing Wang, Yan Zhong 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Distilling Textual Priors From LLM to Efficient Image FusionabstractMulti-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but struggle to handle low-quality or complex inputs. Recent advances in text-guided methods leverage large model priors to overcome these limitations, but at the cost of significant computational overhead, both in memory and inference time. To address this challenge, we propose a novel framework for distilling large model priors, eliminating the need for text guidance during inference while dramatically reducing model size. Our framework utilizes a teacher-student architecture, where the teacher network incorporates large model priors and transfers this knowledge to a smaller student network via a tailored distillation process. Crucially, our experiments demonstrate that this knowledge transfer is the primary driver of performance gains, rather than mere architectural optimization. Additionally, we introduce a spatial-channel cross-fusion module to enhance the model’s ability to leverage textual priors across both spatial and channel dimensions. Our method achieves a favorable trade-off between computational efficiency and fusion quality. The distilled network, requiring only 10% of the parameters and inference time of the teacher network, retains 90% of its performance and outperforms existing SOTA methods. Extensive experiments demonstrate the effectiveness of our approach. Codes are available at https://github.com/Zirconium233/DTPF. Xuanhua He, Ke Cao 0001, Liu Liu 0012, Li Zhang 0104, Man Zhou 0003, Jie Zhang 0033, Dan Guo 0001, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | R^2-Art: Category-Level Articulation Pose Estimation from Single RGB Image via Cascade Render StrategyabstractHuman life is filled with articulated objects. Previous works for estimating the pose of category-level articulated objects rely on costly 3D point clouds or RGB-D images. In this paper, our goal is to estimate category-level articulation poses from a single RGB image, where we propose R2-Art, a novel category-level Articulation pose estimation framework from a single RGB image and a cascade Render strategy. Given an RGB image as input, R2-Art estimates per-part 6D pose for the articulation. Specifically, we design parallel regression branches tailored to generate camera-to-root translation and rotation. Using the predicted joint states, we perform PC prior transformation and deformation with a joint-centric modeling approach. For further refinement, a cascade render strategy is proposed for projecting the 3D deformed prior onto the 2D mask. Extensive experiments are provided to validate our R2-Art on various datasets ranging from synthetic datasets to real-world scenarios, demonstrating the superior performance and robustness of the R2-Art. We believe that this work has the potential to be applied in many fields including robotics, embodied intelligence, and augmented reality. Li Zhang 0104, Yukang Huo, Yan Zhong 0001, Rujing Wang, Liu Liu 0012 |
AAAI | 1 |
| 2025 | GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D ReconstructionabstractGarments are common in daily life and are important for embodied intelligence community. Current category-level garments pose tracking works focus on predicting point-wise canonical correspondence and learning a shape deformation in point cloud sequences. In this paper, motivated by the 2D warping space and shape prior, we propose GaPT-DAR, a novel category-level Garments Pose Tracking framework with integrated 2D Deformation And 3D Reconstruction function, which fully utilize 3D-2D projection and 2D-3D reconstruction to transform the 3D point-wise learning into 2D warping deformation learning. Specifically, GaPT-DAR firstly builds a Voting-based Project module that learns the optimal 3D-2D projection plane for maintaining the maximum orthogonal entropy during point projection. Next, a Garments Deformation module is designed in 2D space to explicitly model the garments warping procedure with deformation parameters. Finally, we build a Depth Reconstruction module to recover the 2D images into 3D warp field. We provide extensive experiments on VR-Folding dataset to evaluate our GaPT-DAR and the results show obvious improvements on most of the metrics compared to state-of-the-arts (i.e. Garment-Nets [8] and GarmentTracking [32]). More details are available at https://sites.google.com/view/gapt-dar. Li Zhang 0104, Qiaojun Yu, Lixin Yang 0001, Yong-Lu Li 0001, Cewu Lu, Rujing Wang, Liu Liu 0012 |
CVPR | 1 |
| 2025 | Generalizable Articulated Object Perception with SuperpointsabstractManipulating articulated objects with robotic arms is challenging due to the complex kinematic structure, which requires precise part segmentation for efficient manipulation. In this work, we introduce a novel superpoint-based perception method designed to improve part segmentation in 3D point clouds of articulated objects. We propose a learnable, part-aware superpoint generation technique that efficiently groups points based on their geometric and semantic similarities, resulting in clearer part boundaries. Furthermore, by leveraging the segmentation capabilities of the 2D foundation model SAM, we identify the centers of pixel regions and select corresponding superpoints as candidate query points. Integrating a query-based transformer decoder further enhances our method's ability to achieve precise part segmentation. Experimental results on the GAPartNet dataset show that our method outperforms existing state-of-the-art approaches in cross-category part segmentation, achieving AP50 scores of 77.9% for seen categories (4.4% improvement) and 39.3% for unseen categories (11.6% improvement), with superior results in 5 out of 9 part categories for seen objects and outperforming all previous methods across all part categories for unseen objects. Qiaojun Yu, Ce Hao, Xibin Yuan, Li Zhang 0104, Liu Liu 0012, Yukang Huo, Cewu Lu |
ICASSP | 4 |
| 2025 | Towards Robust Category-level Articulation Pose Estimation via Integrated Differentiable RenderingabstractAccurate object pose estimation is crucial for embodied intelligence tasks such as manipulation, grasping, and human-robot interaction. However, due to the inherent characteristics of articulated objects, such as kinematic constraints and self-occlusion, pose estimation for articulated objects has remained a significant challenge. To address these issues, this paper proposes CAPED, an end-to-end robust Category-level Articulated object Pose Estimator integrated differentiable rendering. Given partial point cloud as input, CAPED outputs the per-part 6D pose for articulation. Specifically, with the proposed joint-centric modeling manner, CAPED firstly estimates the pose for the free part. Afterward, we canonicalize the input point cloud to estimate constrained parts’ poses by predicting the joint parameters and states as replacements. For further refinement, we propose a differentiable rendering scheme for pose optimization. Evaluations of the ArtImage and RobotArm datasets demonstrate that CAPED exhibits outstanding effectiveness and generalization in tasks ranging from synthetic data to real-world scenarios. We will publicly release the code. Li Zhang 0104, Yukang Huo, Lin Wu 0001, Yanyan Wei, Harshal Suresh Shende, Liu Liu 0012, Linlin Ou |
ICASSP | 3 |
| 2025 | Diff-Art: Category-level Articulation Pose Estimation via Conditional DiffusionabstractArticulated objects are prevalent in people’s daily lives, yet their diverse motion structures pose significant challenges for category-level pose estimation. To address this, in this work, we introduce Diff-Art, a method tailored for category-level articulation pose estimation using conditional diffusion. Given a partial point cloud as input, Diff-Art predicts the per-part 6D pose of the articulated object. Our approach incorporates a novel modeling strategy that exploits the unique kinematic constraints of articulated objects, effectively handling self-occlusion scenarios. Furthermore, we propose a re-scoring strategy to enhance the accuracy of 6D pose estimation within the conditional diffusion framework. Extensive experiments validate the effectiveness of Diff-Art, demonstrating its strong performance on both synthetic datasets and its ability to generalize to real-world scenarios. We believe this work holds significant potential for applications in embodied AI and robotics. Yukang Huo, Xianhui Meng, Li Zhang 0104, Yan Zhong 0001, Mingyuan Yao |
ICME | 3 |
| 2025 | GASEM: Boosting Generalized and Actionable Parts Segmentation and Pose Estimation via Object Motion PerceptionabstractCategory-level object understanding has progressed, but generalized part perception remains underexplored. This paper introduces GASEM, a framework for Generalizable and Actionable Parts GAPart Segmentation and pose Estimation via object Motion perception. GASEM utilizes point-wise motion data from observed point clouds and cross-perspective alignment to learn object motion using a scene flow model. It features a segmentation proposal architecture for GAPart segmentation and an Normalized Object Coordinate Space(NPCS) branch for pose estimation. Additionally, a reinforcement learning agent is trained for robust GAPart manipulation in both simulations and real-world environments. Experiments on the GAPartNet dataset show GASEM outperforms state-of-the-art methods. This work promises advancements in embodied intelligence applications like robot-object interaction and generalizable manipulation. Codes are available at https://github.com/Zirconium233/GASEM. Liu Liu 0012, Li Zhang 0104, Yiming Tang 0001, Qi Wu 0007, Hao Wu 0040 |
ICME | 4 |
| 2025 | Pre-defined Keypoints Promote Category-level Articulation Pose Estimation via Multi-Modal AlignmentabstractArticulations are essential in everyday interactions, yet traditional RGB-based pose estimation methods often struggle with issues such as lighting variations and shadows. To overcome these challenges, we propose a novel Pre-defined keypoint based framework for category-level articulation pose estimation via multi-modal Alignment, coined PAGE. Specifically, we first propose a customized keypoint estimation method, aiming to avoid the divergent distance pattern between heuristically generated keypoints and visible points. In addition, to reduce the mutual information redundancy between point clouds and RGB images, we design the geometry-color alignment, which fuses the features after aligning two modalities. This is followed by decoding the radius for each visible point, and applying our proposal integration scoring strategy to predict keypoints. Ultimately, the framework outputs the per-part 6D pose of the articulation. We conduct extensive experiments to evaluate PAGE across a variety of datasets, from synthetic to real-world scenarios, demonstrating its robustness and superior performance. Li Zhang 0104, Liu Liu 0012, Yan Zhong 0001, Rujing Wang |
IJCAI | 2 |
| 2025 | Generalizable and Actionable Part Detection and Manipulation with SAM-rectified Segmentation and Iterative Pose RefinementabstractThe ability to perform cross-category object perception and manipulation is highly desirable in building intelligent robots. One promising approach is to define the concept of Generalizable and Actionable Parts (GAParts), such as buttons and handles, on both seen and unseen object categories. However, the accurate cross-category perception of GAParts is still challenging due to the large inter-category object shape variations. To address this issue, we introduce SAMIR, a novel framework using SAM-rectified segmentation and Iterative pose Refinement for GAPart detection and manipulation. Firstly, we introduce a Segment Anything (SAM) segmentation prior to rectify the unconfident, fragmented GAPart instance proposals. Secondly, in addition to the zero-shot generalization of the SAM foundation model, we further finetune it with a lightweight adaptor model on our task dataset. Finally, we propose an iterative pose refinement procedure that improves the accuracy of GAPart pose estimation. Our perception experiments on GAPartNet dataset show that SAMIR consistently outperforms the baseline method on instance segmentation and pose estimation tasks. Our manipulation experiments in Sapien simulator illustrate that SAMIR leads to an improved manipulation success rate. We also deploy our method to a real robot for real-world manipulation. Our code and video are available at sites.google.com/view/samir-gapart. Sucheng Qian, Li Zhang 0104, Yanyan Wei, Liu Liu 0012, Cewu Lu |
IROS | 2 |
| 2025 | DFGAP: Towards Depth-Free Cross-Category GAParts Perception via Uncertainty-Quantified ModelingabstractCross-category object perception is one of the essential upstream tasks for generelizable robot object interaction and manipulation. Recently, an increasing number of researchers are focusing on investigating visual Generalizable and Actionable Parts understanding at cross-category level perception. However, these works are built upon the RGB-D or point cloud input, that relies on the depth information capture. Under the circumstances of limited depth camera performance, e.g. transparent or light absorbing material, perception algorithms that do not require depth information are urgently needed. In this paper, we propose DFGAP, a novel depth-free framework for RGB-based GAParts segmentation and pose estimation. Specifically, we independently model the ill-pose problems from the absence of depth for GAPart segmentation and pose estimation, by clearly quantifying the pixel-wise segmentation probability and relative depth. We reduce the uncertainty and benefit learning in these two tasks. The experimental results demonstrate the superior performance and robustness of our DFGAP. Our work provides a new research paradigm in GAParts perception. We believe that our work has the enormous potential to be applied in many areas of embodied AI system. Xueyu Yuan, Jiangqi Song, Liu Liu 0012, Li Zhang 0104, Dan Guo 0001, Richang Hong, Meng Wang 0002 |
ACM Multimedia | 5 |
| 2025 | Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets TrainingabstractDue to the high cost and small scale of Image Quality Assessment (IQA) datasets, achieving robust generalization remains challenging for prevalent Blind IQA (BIQA) methods. Traditional deep learning-based methods emphasize visual information to capture quality features, while recent developments in Vision-Language Models (VLMs) demonstrate strong potential in learning generalizable representations through textual information. However, applying VLMs to BIQA poses three major Challenges: (1) How to make full use of the multi-modal information. (2) The prompt engineering for appropriate quality description is extremely time-consuming. (3) How to use mixed data for joint training to enhance the generalization of VLM-based BIQA model. To this end, we propose a Multi-modal BIQA method with prompt learning, named MMP-IQA. For (1), we propose a conditional fusion module to better utilize the cross-modality information. By jointly adjusting visual and textual features, our model can capture quality information with a stronger representation ability. For (2), we model the quality prompt's context words with learnable vectors during the training process, which can be adaptively updated for superior performances. For (3), we jointly train a linearity-induced quality evaluator, a relative quality evaluator, and a dataset-specific absolute quality evaluator. In addition, we propose a dual automatic weight adjustment strategy to adaptively balance the loss weights between different datasets and among various losses within the same dataset. Extensive experiments illustrate the superior effectiveness of MMP-IQA. Yan Zhong 0001, Xinping Zhao, Li Zhang 0104, Xinyuan Song 0002, Tingting Jiang 0001 |
ACM Multimedia | 3 |
| 2024 | CatmullRom Splines-Based Regression for Image Forgery LocalizationabstractIFL (Image Forgery Location) helps secure digital media forensics. However, many methods suffer from false detections (i.e., FPs) and inaccurate boundaries. In this paper, we proposed the CatmullRom Splines-based Regression Network (CSR-Net), which first rethinks the IFL task from the perspective of regression to deal with this problem. Specifically speaking, we propose an adaptive CutmullRom splines fitting scheme for coarse localization of the tampered regions. Then, for false positive cases, we first develop a novel re-scoring mechanism, which aims to filter out samples that cannot have responses on both the classification branch and the instance branch. Later on, to further restrict the boundaries, we design a learnable texture extraction module, which refines and enhances the contour representation by decoupling the horizontal and vertical forgery features to extract a more robust contour representation, thus suppressing FPs. Compared to segmentation-based methods, our method is simple but effective due to the unnecessity of post-processing. Extensive experiments show the superiority of CSR-Net to existing state-of-the-art methods, not only on standard natural image datasets but also on social media datasets. Li Zhang 0104, Dong Li 0055, Jianming Du, Rujing Wang |
AAAI | 1 |
| 2024 | ICAF-4: An Integrated Framework of Category-level Articulated Object Perception and Manipulation for Embodied Intelligence
Li Zhang 0104, Qiankun Li 0004, Qi Wu 0007, Lin Wu 0001, Liu Liu 0012 |
BMVC | 2 |
| 2024 | U-COPE: Taking a Further Step to Universal 9D Category-Level Object Pose Estimation
Li Zhang 0104, Weiqing Meng, Yan Zhong 0001, Jianming Du, Rujing Wang, Liu Liu 0012 |
ECCV (10) | 1 |
| 2024 | Causal-IQA: Towards the Generalization of Image Quality Assessment Based on Causal InferenceabstractDue to the high cost of Image Quality Assessment (IQA) datasets, achieving robust generalization remains challenging for prevalent deep learning-based IQA methods. To address this, this paper proposes a novel end-to-end blind IQA method: Causal-IQA. Specifically, we first analyze the causal mechanisms in IQA tasks and construct a causal graph to understand the interplay and confounding effects between distortion types, image contents, and subjective human ratings. Then, through shifting the focus from correlations to causality, Causal-IQA aims to improve the estimation accuracy of image quality scores by mitigating the confounding effects using a causality-based optimization strategy. This optimization strategy is implemented on the sample subsets constructed by a Counterfactual Division process based on the Backdoor Criterion. Extensive experiments illustrate the superiority of Causal-IQA. Yan Zhong 0001, Li Zhang 0104, Chenxi Yang 0004, Tingting Jiang 0001 |
ICML | 3 |
| 2024 | VoCAPTER: Voting-based Pose Tracking for Category-level Articulated Object via Inter-frame PriorsabstractArticulated objects are common in our daily life. However, current category-level articulation pose works mostly focus on predicting 9D poses on statistical point cloud observations. In this paper, we deal with the problem of category-level online robust 9D pose tracking of articulated objects, where we propose VoCAPTER, a novel 3D Voting-based Category-level Articulated object Pose TrackER. Our VoCAPTER efficiently updates poses between adjacent frames by utilizing partial observations from the current frame and the estimated per-part 9D poses from the previous frame. Specifically, by incorporating prior knowledge of continuous motion relationships between frames, we begin by canonicalizing the input point cloud, casting the pose tracking task as an inter-frame pose increment estimation challenge. Subsequently, to obtain a robust pose-tracking algorithm, our main idea is to leverage SE(3)-invariant features during motion. This is achieved through a voting-based articulation tracking algorithm, which identifies keyframes as reference states for accurate pose updating throughout the entire video sequence. We evaluate the performance of VoCAPTER in the synthetic dataset and real-world scenarios, which demonstrates VoCAPTER's generalization ability to diverse and complicated scenes. Through these experiments, we provide evidence of VoCAPTER's superiority and robustness in multi-frame pose tracking of articulated objects. We believe that this work can facilitate the progress of various fields, including robotics, embodied intelligence, and augmented reality. All the codes will be made publicly available. Li Zhang 0104, Zean Han, Yan Zhong 0001, Qiaojun Yu, Rujing Wang |
ACM Multimedia | 1 |
| 2024 | EfficientCAPER: An End-to-End Framework for Fast and Robust Category-Level Articulated Object Pose EstimationabstractHuman life is populated with articulated objects. Pose estimation for category-level articulated objects is a significant challenge due to their inherent complexity and diverse kinematic structures. Current methods for this task usually meet the problems of insufficient consideration of kinematic constraints, self-occlusion, and optimization requirements. In this paper, we propose EfficientCAPER, an end-to-end Category-level Articulated object Pose EstimatoR, eliminating the need for optimization functions as post-processing and utilizing the kinematic structure for joint-centric pose modeling, thus enhancing the efficiency and applicability. Given a partial point cloud as input, the EfficientCAPER firstly estimates the pose for the free part of an articulated object using decoupled rotation representation. Next, we canonicalize the input point cloud to estimate constrained parts' poses by predicting the joint parameters and states as replacements. Evaluations on three diverse datasets, ArtImage, ReArtMix, and RobotArm, show EfficientCAPER's effectiveness and generalization ability to real-world scenarios. The framework exhibits excellent static pose estimation performance for articulated objects, contributing to the advancement of category-level pose estimation. Codes will be made publicly available. Li Zhang 0104, Lin Wu 0001, Linlin Ou, Liu Liu 0012 |
NeurIPS | 3 |
| 2024 | Rethinking 3D Convolution in $\ell_p$-norm SpaceabstractConvolution is a fundamental operation in the 3D backbone. However, under certain conditions, the feature extraction ability of traditional convolution methods may be weakened. In this paper, we introduce a new convolution method based on $\ell_p$-norm.
For theoretical support, we prove the universal approximation theorem for $\ell_p$-norm based convolution, and analyze the robustness and feasibility of $\ell_p$-norms in 3D point cloud tasks. Concretely, $\ell_{\infty}$-norm based convolution is prone to feature loss. $\ell_2$-norm based convolution is essentially a linear transformation of the traditional convolution. $\ell_1$-norm based convolution is an economical and effective feature extractor. We propose customized optimization strategies to accelerate the training process of $\ell_1$-norm based Nets and enhance the performance. Besides, a theoretical guarantee is given for the convergence by \textit{regret} argument. We apply our methods to classic networks and conduct related experiments. Experimental results indicate that our approach exhibits competitive performance with traditional CNNs, with lower energy consumption and instruction latency. Li Zhang 0104, Yan Zhong 0001, Zhe Min, RujingWang, Liu Liu 0012 |
NeurIPS | 1 |
| 2022 | Memory-Augmented Model-Driven Network for Pansharpening
Man Zhou 0003, Li Zhang 0104, Chengjun Xie |
ECCV (19) | 3 |
| 2022 | Model-Guided Multi-Contrast Deep Unfolding Network for MRI Super-resolution ReconstructionabstractMagnetic resonance imaging (MRI) with high resolution (HR) provides more detailed information for accurate diagnosis and quantitative image analysis. Despite the significant advances, most existing super-resolution (SR) reconstruction network for medical images has two flaws: 1) All of them are designed in a black-box principle, thus lacking sufficient interpretability and further limiting their practical applications. Interpretable neural network models are of significant interest since they enhance the trustworthiness required in clinical practice when dealing with medical images. 2) most existing SR reconstruction approaches only use a single contrast or use a simple multi-contrast fusion mechanism, neglecting the complex relationships between different contrasts that are critical for SR improvement. To deal with these issues, in this paper, a novel Model-Guided interpretable Deep Unfolding Network (MGDUN) for medical image SR reconstruction is proposed. The Model-Guided image SR reconstruction approach solves manually designed objective functions to reconstruct HR MRI. We show how to unfold an iterative MGDUN algorithm into a novel model-guided deep unfolding network by taking the MRI observation matrix and explicit multi-contrast relationship matrix into account during the end-to-end optimization. Extensive experiments on the multi-contrast IXI dataset and BraTs 2019 dataset demonstrate the superiority of our proposed model. Li Zhang 0104, Man Zhou 0003, Aiping Liu, Xun Chen 0001, Zhiwei Xiong, Feng Wu 0001 |
ACM Multimedia | 2 |
| 2014 | An interval weighed fuzzy c-means clustering by genetically guided alternating optimization
Liyong Zhang, Witold Pedrycz, Wei Lu 0005, Xiaodong Liu 0001, Li Zhang 0104 |
Expert Syst. Appl. | 5 |