Xinzhu Ma

dblp:191/3902 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
33since 2021 · last 2026
0000-0003-0504-0186ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 5 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 17 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EMS-GL: Adaptive Evict-then-Merge Strategy for KV Cache Compression Based on Global-Local Importance
Yingxin Li, Ye Li 0016, Xinzhu Ma, Zihan Geng, Shutao Xia, Zhi Wang 0001
KSEM (1)4
2026 PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference
abstract
The troublesome model size and quadratic computational complexity associated with token quantity pose significant deployment challenges for Vision Transformers (ViTs) in practical applications. Despite recent advancements in model pruning and token reduction techniques speed up the inference speed of ViTs, these approaches either adopt a fixed sparsity ratio or overlook the meaningful interplay between architectural optimization and token selection. Consequently, this static and single-dimension compression often leads to pronounced accuracy degradation under aggressive compression rates, as they fail to fully explore redundancies across these two orthogonal dimensions. Therefore, we introduce PRANCE, a framework which can jointly optimize activated channels and tokens on a per-sample basis, aiming to accelerate ViTs' inference process from a unified data and architectural perspective. However, the joint framework poses challenges to both architectural and decision-making aspects. First, while ViTs inherently support variable-token inference, they do not facilitate dynamic computations for variable channels. To overcome this limitation, we propose a meta-network using weight-sharing techniques to support arbitrary channels of the Multi-Head Self-Attention (MHSA) and Multi-Layer Perceptron (MLP) layers, serving as a foundational model for architectural decision-making. Second, simultaneously optimizing the model structure and input data constitutes a combinatorial optimization problem with an extremely large decision space, reaching up to around $10^{14}$1014, making supervised learning infeasible. To this end, we design a lightweight selector employing Proximal Policy Optimization algorithm (PPO) for efficient decision-making. Furthermore, we introduce a novel "Result-to-Go" training mechanism that models ViTs' inference process as a Markov decision process, significantly reducing action space and mitigating delayed-reward issues during training. Additionally, our framework simultaneously supports different kinds of token optimization methods such as pruning, merging, and sequential pruning-merging strategies. Extensive experiments demonstrate the effectiveness of PRANCE in reducing FLOPs by approximately 50%, retaining only about 10% of tokens while achieving lossless Top-1 accuracy.
Ye Li 0016, Jiajun Fan, Zenghao Chai, Xinzhu Ma, Zhi Wang 0001, Wenwu Zhu 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Towards unbiased source-free object detection via vision foundation models
Zhi Cai, Yingjie Gao 0001, Yanan Zhang 0005, Xinzhu Ma, Di Huang 0001
Pattern Recognit.4
2026 UniAlign: A Universal Cross-Modality Knowledge Alignment Framework for Fine-Grained Action Recognition
abstract
The key to fine-grained video action recognition is identifying subtle differences between action categories. Relying solely on visual features supervised by action labels makes it challenging to characterize robust and discriminative action dynamics from videos. With significant advancements in human pose estimation and the powerful capabilities of Vision-Language Models (VLMs), obtaining reliable and cost-free human pose data and textual semantics has become increasingly feasible, enabling their effective use in fine-grained action recognition. However, the inherent disparities in feature representations across different modalities necessitate a robust alignment strategy to achieve opti mal fusion. To address this, we propose a universal cross-modality knowledge alignment framework, namely UniAlign, to transfer the knowledge from such pre-trained multi-modal models into action recognition models. Specifically, UniAlign introduces two additional branches to extract pose features and textual semantics with the pre-trained pose encoder and VLM. To align the action relevant cues among video features, pose features, and textual semantics, we propose a Cross-Modality Similarity Aggregation module (CMSA) that utilizes the importance of different modal cues while aggregating cross-modal similarities. Additionally, we adopt a fine-tuning mechanism similar to Exponential Moving Average (EMA) to refine the textual semantics, ensuring that the semantic representations encoded by VLMs are preserved while being optimized towards the specific task preferences. Extensive experiments on widely used fine-grained action recognition benchmarks (e.g., FineGym, NTURGB-D, Diving48) and coarse-grained K400 dataset demonstrate the effectiveness of the proposed UniAlign method.
Yihan Wang 0011, Baoli Sun, Xinzhu Ma, Zhihui Wang 0001, Zhiyong Wang 0001
IEEE Trans. Multim.4
2026 Toward an Effective Action-Region Tracking Framework for Fine-Grained Video Action Recognition
abstract
Fine-grained action recognition (FGAR) aims to identify subtle and distinctive differences among fine-grained action categories. However, current recognition methods often capture coarse-grained motion patterns but struggle to identify subtle details in local regions evolving over time. In this work, we introduce the action-region tracking (ART) framework, a novel solution leveraging a query-response mechanism to discover and track the dynamics of distinctive local details, enabling distinguishing similar actions effectively. Specifically, we propose a region-specific semantic activation module that employs discriminative and text-constrained semantics serve as queries to capture the most action-related region responses in each video frame, facilitating interaction among spatial and temporal dimensions with corresponding video features. The captured region responses are then organized into action tracklets, which characterize the region-based action dynamics by linking related responses across different video frames in a coherent sequence. The text-constrained queries are designed to expressly encode nuanced semantic representations derived from the textual descriptions of action labels, as extracted by the language branches within visual language models. To optimize generated action tracklets, we design a multilevel tracklet contrastive constraint among multiple region responses at spatial and temporal levels, which can effectively distinguish individual region responses in each video frame (spatial level) and establish the correlation of similar region responses between adjacent video frames (temporal level). In addition, we implement a task-specific fine-tuning mechanism to refine textual semantics during training. This ensures that the semantic representations encoded by vision language models (VLMs) are not only preserved but also optimized for specific task preferences. Comprehensive experiments on several widely used action recognition benchmarks, i.e., FineGym, Diving48, NTURGB-D, Kinetics, and Something-Something, clearly demonstrate the superiority to previous state-of-the-art baselines.
Baoli Sun, Yihan Wang 0011, Xinzhu Ma, Zhihui Wang 0001, Kun Lu 0003, Zhiyong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
abstract
Recent advancements in diffusion models, particularly the architectural transformation from UNet-based models to Diffusion Transformers (DiTs), significantly improve the quality and scalability of image and video generation. However, despite their impressive capabilities, the substantial computational costs of these large-scale models pose significant challenges for real-world deployment. Post-Training Quantization (PTQ) emerges as a promising solution, enabling model compression and accelerated inference for pretrained models, without the costly retraining. However, research on DiT quantization remains sparse, and existing PTQ frameworks, primarily designed for traditional diffusion models, tend to suffer from biased quantization, leading to notable performance degradation. In this work, we identify that DiTs typically exhibit significant spatial variance in both weights and activations, along with temporal variance in activations. To address these issues, we propose Q-DiT, a novel approach that seamlessly integrates two key techniques: automatic quantization granularity allocation to handle the significant variance of weights and activations across input channels, and sample-wise dynamic activation quantization to adaptively capture activation changes across both timesteps and samples. Extensive experiments conducted on ImageNet and VBench demonstrate the effectiveness of the proposed Q-DiT. Specifically, when quantizing DiT-XL/2 to W6A8 on ImageNet (256 × 256), Q-DiT achieves a remarkable reduction in FID by 1.09 compared to the baseline. Under the more challenging W4A8 setting, it maintains high fidelity in image and video generation, establishing a new benchmark for efficient, high-quality quantization in DiTs.
Xinzhu Ma, Jingyan Jiang, Xin Wang 0019, Zhi Wang 0001, Wenwu Zhu 0001
CVPR4
2025 UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
abstract
Traditional spatiotemporal models generally rely on task-specific architectures, which limit their generalizability and scalability across diverse tasks due to domain-specific design requirements. In this paper, we introduce UniSTD, a unified Transformer-based framework for spatiotemporal modeling, which is inspired by advances in recent foundation models with the two-stage pretraining-then-adaption paradigm. Specifically, our work demonstrates that task-agnostic pretraining on 2D vision and vision-text datasets can build a generalizable model foundation for spatiotemporal learning, followed by specialized joint training on spatiotemporal datasets to enhance task-specific adaptability. To improve the learning capabilities across domains, our framework employs a rank-adaptive mixture-of-expert adaptation by using fractional interpolation to relax the discrete variables so that can be optimized in the continuous space. Additionally, we introduce a temporal module to incorporate temporal dynamics explicitly. We evaluate our approach on a large-scale dataset covering 10 tasks across 4 disciplines, demonstrating that a unified spatiotemporal model can achieve scalable, cross-task learning and support up to 10 tasks simultaneously within one model while reducing training costs in multi-domain applications. Code will be available at https://github.com/1hunters/UniSTD.
Xinzhu Ma, Encheng Su, Xiufeng Song, Xiaohong Liu 0001, Wei-Hong Li 0001, Lei Bai 0001, Wanli Ouyang, Xiangyu Yue 0001
CVPR2
2025 Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time Shifts
abstract
Accurate monocular 3D object detection (M3OD) is pivotal for safety-critical applications like autonomous driving, yet its reliability deteriorates significantly under real-world domain shifts caused by environmental or sensor variations. To address these shifts, Test-Time Adaptation (TTA) methods have emerged, enabling models to adapt to target distributions during inference. While prior TTA approaches recognize the positive correlation between low uncertainty and high generalization ability, they fail to address the dual uncertainty inherent to M3OD: semantic uncertainty (ambiguous class predictions) and geometric uncertainty (unstable spatial localization). To bridge this gap, we propose Dual Uncertainty Optimization (DUO), the first TTA framework designed to jointly minimize both uncertainties for robust M3OD. Through a convex optimization lens, we introduce an innovative convex structure of the focal loss and further derive a novel unsupervised version, enabling label-agnostic uncertainty weighting and balanced learning for high-uncertainty objects. In parallel, we design a semantic-aware normal field constraint that preserves geometric coherence in regions with clear semantic cues, reducing uncertainty from the unstable 3D representation. This dual-branch mechanism forms a complementary loop: enhanced spatial perception improves semantic classification, and robust semantic predictions further refine spatial understanding. Extensive experiments demonstrate the superiority of DUO over existing methods across various datasets and domain shift types.
Xinzhu Ma, Shixiang Tang, Wenhan Yang, Ling-Yu Duan
ICCV3
2025 RobAVA: A Large-Scale Dataset and Baseline Towards Video Based Robotic Arm Action Understanding
Baoli Sun, Xinzhu Ma, Anqi Zou, Chuixuan Fan, Zhihui Wang 0001, Kun Lu 0003, Zhiyong Wang 0001
ICCV3
2025 CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
abstract
While accurate and user-friendly Computer-Aided Design (CAD) is crucial for industrial design and manufacturing, existing methods still struggle to achieve this due to their over-simplified representations or architectures incapable of supporting multimodal design requirements. In this paper, we attempt to tackle this problem from both methods and datasets aspects. First, we propose a cascade MAR with topology predictor (CMT), the first multimodal framework for CAD generation based on Boundary Representation (B-Rep). Specifically, the cascade MAR can effectively capture the ``edge-counters-surface'' priors that are essential in B-Reps, while the topology predictor directly estimates topology in B-Reps from the compact tokens in MAR. Second, to facilitate large-scale training, we develop a large-scale multimodal CAD dataset, mmABC, which includes over 1.3 million B-Rep models with multimodal annotations, including point clouds, text descriptions, and multi-view images. Extensive experiments show the superior of CMT in both conditional and unconditional CAD generation tasks. For example, we improve Coverage and Valid ratio by +10.68% and +10.3%, respectively, compared to state-of-the-art methods on ABC in unconditional generation. CMT also improves +4.01 Chamfer on image conditioned CAD generation on mmABC.
Yizhou Wang 0007, Xiangyu Yue 0001, Xinzhu Ma, Jinyang Guo 0002, Dongzhan Zhou, Wanli Ouyang, Shixiang Tang
ICCV4
2025 Revisiting Convolution Architecture in the Realm of DNA Foundation Models
abstract
In recent years, A variety of methods based on Transformer and state space model (SSM) architectures have been proposed, advancing foundational DNA language models. However, there is a lack of comparison between these recent approaches and the classical architecture—convolutional networks (CNNs)—on foundation model benchmarks. This raises the question: are CNNs truly being surpassed by these recent approaches based on transformer and SSM architectures? In this paper, we develop a simple but well-designed CNN-based method, termed ConvNova. ConvNova identifies and proposes three effective designs: 1) dilated convolutions, 2) gated convolutions, and 3) a dual-branch framework for gating mechanisms. Through extensive empirical experiments, we demonstrate that ConvNova significantly outperforms recent methods on more than half of the tasks across several foundation model benchmarks. For example, in histone-related tasks, ConvNova exceeds the second-best method by an average of 5.8\%, while generally utilizing fewer parameters and enabling faster computation. In addition, the experiments observed findings that may be related to biological characteristics. This indicates that CNNs are still a strong competitor compared to Transformers and SSMs. We anticipate that this work will spark renewed interest in CNN-based methods for DNA foundation models.
Yu Bo, Weian Mao, Yanjun Shao, Weiqiang Bai, Peng Ye 0006, Xinzhu Ma, Hao Chen 0041, Chunhua Shen
ICLR6
2025 CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
abstract
Competitive programming is widely used to evaluate the coding and reasoning abilities of large language models. However, the growing presence of duplicate or highly similar problems raises concerns not only about competition fairness, but also about the validity of competitive programming as a benchmark for model evaluation. We introduce a retrieval-oriented benchmark suite for competitive programming, covering four retrieval tasks—two code-centric (Text-to-Code, Code-to-Code) and two newly proposed problem-centric tasks (Problem-to-Duplicate, Simplified-to-Full)—built from a combination of automatically crawled problem–solution data and manually curated annotations. Our contribution includes both high-quality training data and temporally separated test sets for reliable evaluation. We develop two task-specialized retrievers based on this dataset: CPRetriever-Code, trained with a novel Group-InfoNCE loss for problem–code alignment, and CPRetriever-Prob, fine-tuned for problem-level similarity. Both models achieve strong results and are open-sourced for local use. Finally, we analyze LiveCodeBench and find that high-similarity problems inflate model pass rates and reduce differentiation, underscoring the need for similarity-aware evaluation in future benchmarks.
Shixiang Tang, Wanli Ouyang, Xinzhu Ma
NeurIPS5
2025 Venus-MAXWELL: Efficient Learning of Protein-Mutation Stability Landscapes using Protein Language Models
abstract
In-silico prediction of protein mutant stability, measured by the difference in Gibbs free energy change ($\Delta \Delta G$), is fundamental for protein engineering. Current sequence-to-label methods typically employ two-stage pipelines: (i) encoding mutant sequences using neural networks (e.g., transformers), followed by (ii) the $\Delta \Delta G$ regression from the latent representations. Although these methods have demonstrated promising performance, their dependence on specialized neural network encoders significantly increases the complexity. Additionally, the requirement to compute latent representations individually for each mutant sequence negatively impacts computational efficiency and poses the risk of overfitting. This work proposes the Venus-MAXWELL framework, which reformulates mutation $\Delta \Delta G$ prediction as a sequence-to-landscape task. In Venus-MAXWELL, mutations of a protein and their corresponding $\Delta \Delta G$ values are organized into a landscape matrix, allowing our framework to learn the $\Delta \Delta G$ landscape of a protein with a single forward and backward pass during training. To this end, we curated a new $\Delta \Delta G$ benchmark dataset with strict controls on data leakage and redundancy to ensure robust evaluation. Leveraging the zero-shot scoring capability of protein language models (PLMs), Venus-MAXWELL effectively utilizes the evolutionary patterns learned by PLMs during pre-training. More importantly, Venus-MAXWELL is compatible with multiple protein language models. For example, when integrated with the ESM-IF, Venus-MAXWELL achieves higher accuracy than ThermoMPNN with 10$\times$ faster in inference speed (despite having 50$\times$ more parameters than ThermoMPNN). The training codes, model weights, and datasets are publicly available at https://github.com/ai4protein/Venus-MAXWELL.
Yuanxi Yu, Fan Jiang 0013, Xinzhu Ma, Bozitao Zhong, Wanli Ouyang, Guisheng Fan, Huiqun Yu
NeurIPS3
2025 Point2skh: End-to-end Parametric Primitive Inference from Point Clouds with Improved Denoising Transformer
abstract
Recovering the CAD command sequence from the point cloud is an essential component in CAD reverse engineering. In this paper, we strive to solve this problem from both the perspectives of artificial intelligence and the procedures of procedural CAD models. We propose a CAD reconstruction method based on an end-to-end point-to-sketch network (Point2Skh) that can produce the CAD modeling sequence from the input geometrical point cloud by recovering the inverse sketch-and-extrude process. The point cloud is first segmented into point sets corresponding to the same extrusion. The modeling sequence can then be recovered by combining the network prediction of each point set. The proposed Point2Skh can detect and infer command vectors of sketch curves (line, arc, and circle) and the extrusion operation from the input point cloud of a single extrusion. By directly representing the sketch with its curves and inferring the command parameters, accurate sketch reconstruction is produced, which further leads to precise CAD reconstruction with sharp edges. The produced CAD modeling sequence is human-interpretable and can be readily edited by importing it into CAD tools. Experiments show that the Chamfer Distance (CD) between the predicted results and the ground truth is 0.312, and the primitive type and parameter accuracy are 93.87% and 83.24%, respectively. • We propose a CAD reconstruction method based on an point-to-sketch network (Point2Skh). • The CAD modeling sequence can be directly predicted from the input point cloud and imported into CAD tools. • The Point2Skh can directly infer the CAD command type and parameters from the extrusion point cloud.
Cheng Wang 0026, Wenyu Sun, Xinzhu Ma
Comput. Aided Des.3
2025 3DAxisPrompt: Promoting the 3D grounding and reasoning in GPT-4o
Dingning Liu, Cheng Wang 0026, Peng Gao 0007, Renrui Zhang, Xinzhu Ma, Zhihui Wang 0001
Neurocomputing5
2025 GUPNet++: Geometry Uncertainty Propagation Network for Monocular 3D Object Detection
abstract
Geometry plays a significant role in monocular 3D object detection. It can be used to estimate object depth by using the perspective projection between object's physical size and 2D projection in the image plane, which can introduce mathematical priors into deep models. However, this projection process also introduces error amplification, where the error of the estimated height is amplified and reflected into the projected depth. It leads to unreliable depth inferences and also impairs training stability. To tackle this problem, we propose a novel Geometry Uncertainty Propagation Network (GUPNet++) by modeling geometry projection in a probabilistic manner. This ensures depth predictions are well-bounded and associated with a reasonable uncertainty. The significance of introducing such geometric uncertainty is two-fold: (1). It models the uncertainty propagation relationship of the geometry projection during training, improving the stability and efficiency of the end-to-end model learning. (2). It can be derived to a highly reliable confidence to indicate the quality of the 3D detection result, enabling more reliable detection inference. Experiments show that the proposed approach not only obtains (state-of-the-art) SOTA performance in image-based monocular 3D detection but also demonstrates superiority in efficacy with a simplified framework. The code and model will be released at https://github.com/SuperMHP/GUPNet_Plus.
Yan Lu 0001, Xinzhu Ma, Lei Yang 0045, Tianzhu Zhang 0001, Qi Chu 0001, Tong He 0001, Yonghui Li 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Propagating Sparse Depth via Depth Foundation Model for Out-of-Distribution Depth Completion
abstract
Depth completion is a pivotal challenge in computer vision, aiming at reconstructing the dense depth map from a sparse one, typically with a paired RGB image. Existing learning-based models rely on carefully prepared but limited data, leading to significant performance degradation in out-of-distribution (OOD) scenarios. Recent foundation models have demonstrated exceptional robustness in monocular depth estimation through large-scale training, and using such models to enhance the robustness of depth completion models is a promising solution. In this work, we propose a novel depth completion framework that leverages depth foundation models to attain remarkable robustness without large-scale training. Specifically, we leverage a depth foundation model to extract environmental cues, including structural and semantic context, from RGB images to guide the propagation of sparse depth information into missing regions. We further design a dual-space propagation approach, without any learnable parameters, to effectively propagate sparse depth in both 3D and 2D spaces to maintain geometric structure and local consistency. To refine the intricate structure, we introduce a learnable correction module to progressively adjust the depth prediction towards the real depth. We train our model on the NYUv2 and KITTI datasets as in-distribution datasets and extensively evaluate the framework on 16 other datasets. Our framework performs remarkably well in the OOD scenarios and outperforms existing state-of-the-art depth completion methods. Our models are released in https://github.com/shenglunch/PSD.
Shenglun Chen, Xinzhu Ma, Hong Zhang 0011, Zhihui Wang 0001
IEEE Trans. Image Process.2
2025 Referring Video Object Segmentation With Cross-Modality Proxy Queries
abstract
Referring video object segmentation (RVOS) is an emerging cross-modality task that aims to generate pixel-level maps of the target objects referred by given textual expressions. The main concept involves learning an accurate alignment visual elements and language expressions within a semantic space. Recent approaches address cross-modality alignment through conditional queries, tracking the target object using a queryresponse based mechanism built upon transformer structure. However, they exhibit two limitations: (1) these conditional queries, identifying the same object across different frames through the same query, lack inter-frame dependency and variation modeling, making accurate target tracking challenging amid significant frame-to-frame variations; and (2) they handle the temporal feature of a video and build visual-language interaction sequentially, integrating textual constraints belatedly, which may cause the video features potentially focus on the non-referred objects. Therefore, we propose a novel RVOS architecture called ProxyFormer, which introduces a set of proxy queries to integrate visual and text semantics and facilitate the flow of semantics between them. By progressively updating and propagating proxy queries across multiple stages of video feature encoder, ProxyFormer ensures that the video features are as focused as much possible on the object of interest. This dynamic evolution of the queries across video also enables the proxy queries to establish inter-frame dependencies, enhancing the accuracy and coherence of object tracking throughout the video sequence. To mitigate the high computational costs associated with full spatio-temporal interactions between video and proxy queries, we propose to decouple cross-modality interactions into their temporal and spatial dimensions, respectively. Additionally, we design a Joint Semantic Consistency (JSC) training strategy to align semantic consensus between the proxy queries and the combined videotext pairs. Comprehensive experiments on four widely used RVOS benchmarks, i.e., Ref-Youtube-VOS, Ref-DAVIS17, A2D-Sentences and JHMDB-Sentences, clearly demonstrate the superiority of our ProxyFormer to the state-of-the-art methods
Baoli Sun, Xinzhu Ma, Zhihui Wang 0001, Zhiyong Wang 0001
IEEE Trans. Multim.2
2025 P$^{2}$M: Progressive Perspective Mining for Referring Video Object Segmentation
Yihan Wang 0011, Baoli Sun, Xinzhu Ma, Hong-Wei Ge, Jiulin Fan
IEEE Trans. Multim.3
2025 Real-Time Depth Completion With Multimodal Feature Alignment
abstract
As a key problem in computer vision, depth completion aims to recover dense depth maps from sparse ones [generally derived from light detection and ranging (LiDAR)]. Most methods introduce synchronous RGB images and leverage multimodal fusion to integrate multimodal features from these modalities to describe the complete scene. However, their different natural characteristics lead to inconsistency in features, potentially impacting the effectiveness of multimodal feature fusion. To address this issue, we propose a feature alignment network (FANet) that introduces an alignment scheme to enhance the consistency between multimodal features. This scheme aligns the modality-invariant semantic context, which is invariant to changes in modality and represents the correlation between a pixel and its surroundings. Specifically, we first design an asymmetric context extraction (ACE) module to extract modality-invariant semantic contexts from multimodal features within limited GPU memory, and then pull them closer to improve consistency. Crucially, our alignment scheme is only applied during the training phase, and no additional computation cost is incurred in the inference phase. Moreover, we introduce a simple yet effective refinement module to refine estimated results via residual learning based on intermediate depth maps and sparse depth maps. Extensive experiments on KITTI and VOID datasets demonstrate that our method achieves competitive performance against typical real-time methods. In addition, we embed the proposed alignment scheme and refinement module into other methods to demonstrate their effectiveness.
Shenglun Chen, Xinzhu Ma, Hong Zhang 0011, Baoli Sun, Zhihui Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Retraining-free Model Quantization via One-Shot Weight-Coupling Learning
abstract
Quantization is of significance for compressing the over-parameterized deep neural models and deploying them on resource-limited devices. Fixed-precision quantization suf-fers from performance drop due to the limited numerical representation ability. Conversely, mixed-precision quan-tization (MPQ) is advocated to compress the model ef-fectively by allocating heterogeneous bit-width for layers. MPQ is typically organized into a searching-retraining two-stage process. Previous works only focus on determining the optimal bit-width configuration in the first stage effi-ciently, while ignoring the considerable time costs in the second stage and thus hindering deployment efficiency sig-nificantly. In this paper, we devise a one-shot training-searching paradigm for mixed-precision model compression. Specifically, in the first stage, all potential bit-width configurations are coupled and thus optimized simultane-ously within a set of shared weights. However, our ob-servations reveal a previously unseen and severe bit-width interference phenomenon among highly coupled weights during optimization, leading to considerable performance degradation under a high compression ratio. To tackle this problem, we first design a bit-width scheduler to dy-namically freeze the most turbulent bit-width of layers during training, to ensure the rest bit-widths converged prop-erly. Then, taking inspiration from information theory, we present an information distortion mitigation technique to align the behaviour of the bad-performing bit-widths to the well-performing ones. In the second stage, an inference-only greedy search scheme is devised to evaluate the good-ness of configurations without introducing any additional training costs. Extensive experiments on three representative models and three datasets demonstrate the effective-ness of the proposed method. Code can be available on https://github.com/1hunters/retraining-free-quantization.
Shuzhao Xie, Rongwei Lu, Xinzhu Ma, Zhi Wang 0001, Wenwu Zhu 0001
CVPR6
2024 ProSST: Protein Language Modeling with Quantized Structure and Disentangled Attention
abstract
Protein language models (PLMs) have shown remarkable capabilities in various protein function prediction tasks. However, while protein function is intricately tied to structure, most existing PLMs do not incorporate protein structure information. To address this issue, we introduce ProSST, a Transformer-based protein language model that seamlessly integrates both protein sequences and structures. ProSST incorporates a structure quantization module and a Transformer architecture with disentangled attention. The structure quantization module translates a 3D protein structure into a sequence of discrete tokens by first serializing the protein structure into residue-level local structures and then embeds them into dense vector space. These vectors are then quantized into discrete structure tokens by a pre-trained clustering model. These tokens serve as an effective protein structure representation. Furthermore, ProSST explicitly learns the relationship between protein residue token sequences and structure token sequences through the sequence-structure disentangled attention. We pre-train ProSST on millions of protein structures using a masked language model objective, enabling it to learn comprehensive contextual representations of proteins. To evaluate the proposed ProSST, we conduct extensive experiments on the zero-shot mutation effect prediction and several supervised downstream tasks, where ProSST achieves the state-of-the-art performance among all baselines. Our code and pre-trained models are publicly available.
Yang Tan 0001, Xinzhu Ma, Bozitao Zhong, Huiqun Yu, Ziyi Zhou 0002, Wanli Ouyang, Bingxin Zhou, Pan Tan
NeurIPS3
2024 Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA
abstract
Foundation models have made significant strides in understanding the genomic language of DNA sequences. However, previous models typically adopt the tokenization methods designed for natural language, which are unsuitable for DNA sequences due to their unique characteristics. In addition, the optimal approach to tokenize DNA remains largely under-explored, and may not be intuitively understood by humans even if discovered. To address these challenges, we introduce MxDNA, a novel framework where the model autonomously learns an effective DNA tokenization strategy through gradient decent. MxDNA employs a sparse Mixture of Convolution Experts coupled with a deformable convolution to model the tokenization process, with the discontinuous, overlapping, and ambiguous nature of meaningful genomic segments explicitly considered. On Nucleotide Transformer Benchmarks and Genomic Benchmarks, MxDNA demonstrates superior performance to existing methods with less pretraining data and time, highlighting its effectiveness. Finally, we show that MxDNA learns unique tokenization strategy distinct to those of previous methods and captures genomic functionalities at a token level during self-supervised pretraining. Our MxDNA aims to provide a new perspective on DNA tokenization, potentially offering broad applications in various domains and yielding profound insights. Code is available at https://github.com/qiaoqiaoLF/MxDNA.
Lifeng Qiao, Peng Ye 0006, Yuchen Ren 0001, Weiqiang Bai, Chaoqi Liang, Xinzhu Ma, Nanqing Dong, Wanli Ouyang
NeurIPS6
2024 BEACON: Benchmark for Comprehensive RNA Tasks and Language Models
abstract
RNA plays a pivotal role in translating genetic instructions into functional outcomes, underscoring its importance in biological processes and disease mechanisms. Despite the emergence of numerous deep learning approaches for RNA, particularly universal RNA language models, there remains a significant lack of standardized benchmarks to assess the effectiveness of these methods. In this study, we introduce the first comprehensive RNA benchmark BEACON BEnchmArk for COmprehensive RNA Task and Language Models).First, BEACON comprises 13 distinct tasks derived from extensive previous work covering structural analysis, functional studies, and engineering applications, enabling a comprehensive assessment of the performance of methods on various RNA understanding tasks. Second, we examine a range of models, including traditional approaches like CNNs, as well as advanced RNA foundation models based on language models, offering valuable insights into the task-specific performances of these models. Third, we investigate the vital RNA language model components from the tokenizer and positional encoding aspects. Notably, our findings emphasize the superiority of single nucleotide tokenization and the effectiveness of Attention with Linear Biases (ALiBi) over traditional positional encoding methods. Based on these insights, a simple yet strong baseline called BEACON-B is proposed, which can achieve outstanding performance with limited data and computational resources. The datasets and source code of our benchmark are available at https://github.com/terry-r123/RNABenchmark.
Yuchen Ren 0001, Lifeng Qiao, Hongtai Jing, Peng Ye 0006, Xinzhu Ma, Hongliang Yan, Wanli Ouyang, Xihui Liu
NeurIPS8
2024 3D Object Detection From Images for Autonomous Driving: A Survey
abstract
3D object detection from images, one of the fundamental and challenging problems in autonomous driving, has received increasing attention from both industry and academia in recent years. Benefiting from the rapid development of deep learning technologies, image-based 3D detection has achieved remarkable progress. Particularly, more than 200 works have studied this problem from 2015 to 2021, encompassing a broad spectrum of theories, algorithms, and applications. However, to date no recent survey exists to collect and organize this knowledge. In this paper, we fill this gap in the literature and provide the first comprehensive survey of this novel and continuously growing research field, summarizing the most commonly used pipelines for image-based 3D detection and deeply analyzing each of their components. Additionally, we also propose two new taxonomies to organize the state-of-the-art methods into different categories, with the intent of providing a more systematic review of existing methods and facilitating fair comparisons with future works. In retrospect of what has been achieved so far, we also analyze the current challenges in the field and discuss future directions for image-based 3D detection research.
Xinzhu Ma, Wanli Ouyang, Andrea Simonelli, Elisa Ricci 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Learning Pixel-Wise Continuous Depth Representation via Clustering for Depth Completion
abstract
Depth completion is a long-standing challenge in computer vision, where classification-based methods have made tremendous progress in recent years. However, most existing classification-based methods rely on pre-defined pixel-shared and discrete depth values as depth categories. This representation fails to capture the continuous depth values that conform to the real depth distribution, leading to depth smearing in boundary regions. To address this issue, we revisit depth completion from the clustering perspective and propose a novel clustering-based framework called CluDe which focuses on learning the pixel-wise and continuous depth representation. The key idea of CluDe is to iteratively update the pixel-shared and discrete depth representation to its corresponding pixel-wise and continuous counterpart, driven by the real depth distribution. Specifically, CluDe first utilizes depth value clustering to learn a set of depth centers as the depth representation. While these depth centers are pixel-shared and discrete, they are more in line with the real depth distribution compared to pre-defined depth categories. Then, CluDe estimates offsets for these depth centers, enabling their dynamic adjustment along the depth axis of the depth distribution to generate the pixel-wise and continuous depth representation. Extensive experiments demonstrate that CluDe successfully reduces depth smearing around object boundaries by utilizing pixel-wise and continuous depth representation. Furthermore, CluDe achieves state-of-the-art performance on the VOID datasets and outperforms classification-based methods on the KITTI dataset.
Shenglun Chen, Hong Zhang 0011, Xinzhu Ma, Zhihui Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Push-and-Pull: A General Training Framework With Differential Augmentor for Domain Generalized Point Cloud Classification
abstract
As a fundamental task of 3D perception, point cloud recognition has shown significant progress in recent years. However, existing methods still face challenges when dealing with geometry differences, resulting in performance degradation when a distribution gap exists between the training and testing data, also known as domain generalization. In this work, we focus on this problem and propose a general training framework, named Push-and-Pull, aimed at effectively improving the generalization ability of models on unseen target domains. Specifically, our framework first introduces a learnable 3D data augmentor to generate new training point clouds, which helps to reduce the domain bias and enrich the source training set. Also, an adversarial training strategy is proposed topushthe augmented samples away from the original ones in the latent space and meanwhile keep the geometric structure. Second, based on the original and augmented samples, a dual-level consistency regularization strategy on logits and feature spaces is designed topullthe deviated representations back to their original space as close as possible, and promote discriminative and domain-agnostic representations. These two steps are iteratively optimized to enhance the overall performance. Extensive experiments on the PointDA-10 and Sim2Real benchmarks consistently demonstrate the effectiveness of our proposed framework.
Xinzhu Ma, Lin Zhang 0055, Bo Zhang 0069, Tao Chen 0003
IEEE Trans. Circuits Syst. Video Technol.2
2023 Towards Fair and Comprehensive Comparisons for Image-Based 3D Object Detection
abstract
In this work, we build a modular-designed codebase, formulate strong training recipes, design an error diagnosis toolbox, and discuss current methods for image-based 3D object detection. In particular, different from other highly mature tasks, e.g., 2D object detection, the community of image-based 3D object detection is still evolving, where methods often adopt different training recipes and tricks resulting in unfair evaluations and comparisons. What is worse, these tricks may overwhelm their proposed designs in performance, even leading to wrong conclusions. To address this issue, we build a module-designed code-base and formulate unified training standards for the community. Furthermore, we also design an error diagnosis toolbox to measure the detailed characterization of detection models. Using these tools, we analyze current methods in-depth under varying settings and provide discussions for some open questions, e.g., discrepancies in conclusions on KITTI-3D and nuScenes datasets, which have led to different dominant methods for these datasets. We hope that this work will facilitate future research in image-based 3D object detection. Our codes will be released at https://github.com/OpenGVLab/3dodi.
Xinzhu Ma, Yongtao Wang, Yinmin Zhang, Zhiyi Xia, Zhihui Wang 0001, Wanli Ouyang
ICCV1
2023 Better Teacher Better Student: Dynamic Prior Knowledge for Knowledge Distillation
Martin Zong, Zengyu Qiu, Xinzhu Ma, Chunya Liu, Shuai Yi, Wanli Ouyang
ICLR3
2022 Scale-Prior Deformable Convolution for Exemplar-Guided Class-Agnostic Counting
Wei Lin 0018, Xinzhu Ma, Junyu Gao 0001, Lingbo Liu, Shinan Liu, Shuai Yi, Antoni B. Chan
BMVC3
2022 MonoDistill: Learning Spatial Features for Monocular 3D Object Detection
Zhiyu Chong, Xinzhu Ma, Hong Zhang 0011, Yuxin Yue, Zhihui Wang 0001, Wanli Ouyang
ICLR2
2021 Delving Into Localization Errors for Monocular 3D Object Detection
abstract
Estimating 3D bounding boxes from monocular images is an essential component in autonomous driving, while accurate 3D object detection from this kind of data is very challenging. In this work, by intensive diagnosis experiments, we quantify the impact introduced by each sub-task and found the ‘localization error’ is the vital factor in restricting monocular 3D detection. Besides, we also investigate the underlying reasons behind localization errors, analyze the issues they might bring, and propose three strategies. First, we revisit the misalignment between the center of the 2D bounding box and the projected center of the 3D object, which is a vital factor leading to low localization accuracy. Second, we observe that accurately localizing distant objects with existing technologies is almost impossible, while those samples will mislead the learned network. To this end, we propose to remove such samples from the training set for improving the overall performance of the detector. Lastly, we also propose a novel 3D IoU oriented loss for the size estimation of the object, which is not affected by ‘localization error’. We conduct extensive experiments on the KITTI dataset, where the proposed method achieves real-time detection and outperforms previous methods by a large margin. The code will be made available at: https://github.com/xinzhuma/monodle.
Xinzhu Ma, Yinmin Zhang, Dan Xu 0002, Dongzhan Zhou, Shuai Yi, Wanli Ouyang
CVPR1
2021 Geometry Uncertainty Projection Network for Monocular 3D Object Detection
abstract
Geometry Projection is a powerful depth estimation method in monocular 3D object detection. It estimates depth dependent on heights, which introduces mathematical priors into the deep model. But projection process also introduces the error amplification problem, in which the error of the estimated height will be amplified and reflected greatly at the output depth. This property leads to uncontrollable depth inferences and also damages the training efficiency. In this paper, we propose a Geometry Uncertainty Projection Network (GUP Net) to tackle the error amplification problem at both inference and training stages. Specifically, a GUP module is proposed to obtains the geometry-guided uncertainty of the inferred depth, which not only provides high reliable confidence for each depth but also benefits depth learning. Furthermore, at the training stage, we propose a Hierarchical Task Learning strategy to reduce the instability caused by error amplification. This learning algorithm monitors the learning situation of each task by a proposed indicator and adaptively assigns the proper loss weights for different tasks according to their pre-tasks situation. Based on that, each task starts learning only when its pre-tasks are learned well, which can significantly improve the stability and efficiency of the training process. Extensive experiments demonstrate the effectiveness of the proposed method. The overall model can infer more reliable object depth than existing methods and outperforms the state-of-the-art image-based monocular 3D detectors by 3.74% and 4.7% AP40of the car and pedestrian categories on the KITTI benchmark. The code and model will be released at https://github.com/SuperMHP/GUPNet.
Yan Lu 0001, Xinzhu Ma, Lei Yang 0045, Tianzhu Zhang 0001, Qi Chu 0001, Wanli Ouyang
ICCV2
2020 Rethinking Pseudo-LiDAR Representation
Xinzhu Ma, Shinan Liu, Zhiyi Xia, Hongwen Zhang 0001, Xingyu Zeng, Wanli Ouyang
ECCV (13)1
2019 Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous Driving
abstract
In this paper, we propose a monocular 3D object detection framework in the domain of autonomous driving. Unlike previous image-based methods which focus on RGB feature extracted from 2D images, our method solves this problem in the reconstructed 3D space in order to exploit 3D contexts explicitly. To this end, we first leverage a stand-alone module to transform the input data from 2D image plane to 3D point clouds space for a better input representation, then we perform the 3D detection using PointNet backbone net to obtain objects' 3D locations, dimensions and orientations. To enhance the discriminative capability of point clouds, we propose a multi-modal feature fusion module to embed the complementary RGB cue into the generated point clouds representation. We argue that it is more effective to infer the 3D bounding boxes from the generated 3D scene space (i.e., X,Y, Z space) compared to the image plane (i.e., R,G,B image plane). Evaluation on the challenging KITTI dataset shows that our approach boosts the performance of state-of-the-art monocular approach by a large margin.
Xinzhu Ma, Zhihui Wang 0001, Wanli Ouyang, Xin Fan 0001
ICCV1
2019 Self-Adaption Multi-classifier Fusion Networks for Image Recognition
abstract
Recently, many visual recognition related studies have proved that making full use of different levels of features can effectively enhance the representational ability of convolutional neural networks (CNNs). Different from other CNN architecture which are devoted to aggregate features of different scales, we proposed a multi-classifier network (MCN) to make more effective use of these feature maps. Specifically, MCN can directly make full use of features of different levels and fuse intermediate results in a self-adaption way. Note that the auxiliary classifiers not only can optimize the internal features of CNNs directly, but also bring additional gradient which further solves the problem of vanishing-gradient. In addition, MCN is a very flexible architecture and can be combined with existing state-of-the-art networks (ResNet, DenseNet, ResNeXt, etc.) easily. Extensive experiments on three highly competitive benchmark datasets, CIFAR-10, CIFAR-100 and ImageNet, clearly demonstrate superior performance of the proposed MCN over state-of-the-arts.
Zengyuan Guo, Xinzhu Ma, Zhihui Wang 0001
ICME2
2019 Learning to Segment Unseen Category Objects using Gradient Gaussian Attention
abstract
Existing semantic segmentation models are trapped in the categories of training sets. Unfortunately, public datasets provide pixel-level annotations only for a small quantity of images and few categories. In this paper, we propose a novel cross-category supervised object segmentation network to explore the similar features among different categories, which can transfer the learned segmentation knowledge from categories with mask annotations to unseen categories that only have bounding boxes. Specifically, we fuse gaussian attention map of an object with guided gradient back-propagation map as an extra input, which gives localizable and discriminative prior cues to obtain precise object segmentation from the bounding box. In addition to computing a segment for each box, we also fuse segments to generate pixel-level labels. Then, without modifying the segmentation training program, the generated labels are still sufficient and achieve about 98.2% of the fully supervised model, in the case of only 50% categories with pixel-level annotations on PASCAL VOC 2012. The proposed method is also effective for interactive segmentation and salient object detection.
Zhihui Wang 0001, Xinzhu Ma, Jianjun Li 0007
ICME3
2018 User-Guided Deep Anime Line Art Colorization with Conditional Adversarial Networks
abstract
Scribble colors based line art colorization is a challenging computer vision problem since neither greyscale values nor semantic information is presented in line arts, and the lack of authentic illustration-line art training pairs also increases difficulty of model generalization. Recently, several Generative Adversarial Nets (GANs) based methods have achieved great success. They can generate colorized illustrations conditioned on given line art and color hints. However, these methods fail to capture the authentic illustration distributions and are hence perceptually unsatisfying in the sense that they often lack accurate shading. To address these challenges, we propose a novel deep conditional adversarial architecture for scribble based anime line art colorization. Specifically, we integrate the conditional framework with WGAN-GP criteria as well as the perceptual loss to enable us to robustly train a deep network that makes the synthesized images more natural and real. We also introduce a local features network that is independent of synthetic data. With GANs conditioned on features from such network, we notably increase the generalization capability over "in the wild" line arts. Furthermore, we collect two datasets that provide high-quality colorful illustrations and authentic line arts for training and benchmarking. With the proposed model trained on our illustration dataset, we demonstrate that images synthesized by the presented approach are considerably more realistic and precise than alternative approaches.
Yuanzheng Ci, Xinzhu Ma, Zhihui Wang 0001, Zhongxuan Luo
ACM Multimedia2
2018 Disparity-Based Robust Unstructured Terrain Segmentation
Xinzhu Ma, Zhihui Wang 0001, Zhongxuan Luo
PRCV (4)2