EDBT 2026 Demo / reviewers in the wild / expert
Xu Ma 0005
dblp:77/9370-5
· DBLP profile ↗
21ranked-venue papers
10as first author
16since 2021 · last 2025
0000-0002-5794-119XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accessing Vision Foundation Models via ImageNet-1KabstractVision foundation models are renowned for the generalization ability due to massive training data. Nevertheless, they demand tremendous training resources, and the training data is often inaccessible, e.g., CLIP, DINOv2, posing great challenges to developing derivatives that could facilitate the research. In this work, we offer a very simple and general solution, named Proteus, to distill foundation models into smaller equivalents on ImageNet-1K without access to the original training data. Specifically, we remove the designs from conventional knowledge distillation settings that result in dataset bias and present three levels of training objectives, i.e., token, patch, and feature, to maximize the efficacy of knowledge transfer. In this manner, Proteus is trained at ImageNet-level costs with surprising ability, facilitating the accessibility of training foundation models for the broader research community. When leveraging DINOv2-g/14 as the teacher, Proteus-L/14 matches the performance of the Oracle method DINOv2-L/14 (142M training data) across 19 benchmarks and outperforms other vision foundation models including CLIP-L/14 (400M), OpenCLIP-L/14 (400M/2B) and SynCLR-L/14 (600M) with a significantly smaller training set of 1.2M images. Xu Ma 0005, Yun Fu 0001 |
ICLR | 2 |
| 2024 | Rewrite the StarsabstractRecent studies have drawn attention to the untapped potential of the “star operation” (element-wise multiplication) in network design. While intuitive explanations abound, the foundational rationale behind its application remains largely unexplored. Our study attempts to reveal the star operation's ability of mapping inputs into high-dimensional, non-linear feature spaces-akin to kernel tricks-without widening the network. We further introduce StarNet, a simple yet pow-erful prototype, demonstrating impressive performance and low latency under compact network structure and efficient budget. Like stars in the sky, the star operation appears unremarkable but holds a vast universe of potential. Our work encourages further exploration across tasks, with codes available at https://github.com/ma-xu/Rewrite-the-Stars. Xu Ma 0005, Xiyang Dai, Yizhou Wang 0006, Yun Fu 0001 |
CVPR | 1 |
| 2024 | Efficient Modulation for Vision NetworksabstractIn this work, we present efficient modulation, a novel design for efficient vision networks. We revisit the modulation mechanism, which operates input through convolutional context modeling and feature projection layers, and fuses features via element-wise multiplication and an MLP block. We demonstrate that the abstracted modulation mechanism is particularly well suited for efficient networks and further tailor the modulation design by proposing the efficient modulation (EfficientMod) block, which is considered the essential building block for our networks. Bene- fiting from the prominent representational ability of modulation mechanism and the efficiency of efficient modulation design, our network can accomplish better accuracy-efficiency trade-offs and set new state-of-the-art performance for efficient networks. When integrating EfficientMod block with the vanilla self-attention block, we obtain the hybrid architecture and further improve the performance without sacrificing the efficiency. We carry out comprehensive experiments to verify EfficientMod’s performance. With fewer parameters, our EfficientMod-s performs 0.6 top-1 accuracy better than the prior state-of-the-art approach EfficientFormerV2-s2 without any training tricks and is 25% faster on GPU. Additionally, our method presents a notable improvement in downstream tasks, outperforming EfficientFormerV2-s by 3.6 mIoU on the ADE20K benchmark. Code and checkpoints are available at https://github.com/ma-xu/EfficientMod. Xu Ma 0005, Xiyang Dai, Bin Xiao 0004, Yinpeng Chen, Yun Fu 0001, Lu Yuan 0001 |
ICLR | 1 |
| 2024 | Slicing Vision Transformer for Flexible InferenceabstractVision Transformers (ViT) is known for its scalability. In this work, we target to scale down a ViT to fit in an environment with dynamic-changing resource constraints. We observe that smaller ViTs are intrinsically the sub-networks of a larger ViT with different widths. Thus, we propose a general framework, named Scala, to enable a single network to represent multiple smaller ViTs with flexible inference capability, which aligns with the inherent design of ViT to vary from widths. Concretely, Scala activates several subnets during training, introduces Isolated Activation to disentangle the smallest sub-network from other subnets, and leverages Scale Coordination to ensure each sub-network receives simplified, steady, and accurate learning objectives. Comprehensive empirical validations on different tasks demonstrate that with only one-shot training, Scala learns slimmable representation without modifying the original ViT structure and matches the performance of Separate Training. Compared with the prior art, Scala achieves an average improvement of 1.6% on ImageNet-1K with fewer parameters. Huseyin Coskun, Xu Ma 0005, Xi Stephen Chen, Derek Hao Hu, Yun Fu 0001 |
NeurIPS | 3 |
| 2023 | Efficient Converted Spiking Neural Network for 3D and 2D ClassificationabstractSpiking Neural Networks (SNNs) have attracted enormous research interest due to their low-power and biologically plausible nature. Existing ANN-SNN conversion methods can achieve lossless conversion by converting a well-trained Artificial Neural Network (ANN) into an SNN. However, converted SNN requires a large amount of time steps to achieve competitive performance with the well-trained ANN, which means a large latency. In this paper, we propose an efficient unified ANN-SNN conversion method for point cloud classification and image classification to significantly reduce the time step to meet the fast and lossless ANN-SNN transformation. Specifically, we first adaptively adjust the threshold according to the activation state of spiking neurons, ensuring a certain proportion of spiking neurons are activated at each time step to reduce the time for accumulation of membrane potential. Next, we use an adaptive firing mechanism to enlarge the range of spiking output, getting more discrimination features in short time steps. Extensive experimental results on challenging point cloud and image datasets demonstrate that the suggested approach significantly outmatches state-of-the-art ANN-SNN conversion based methods. Yuxiang Lan, Yachao Zhang 0001, Xu Ma 0005, Yanyun Qu, Yun Fu 0001 |
ICCV | 3 |
| 2023 | BEV-DG: Cross-Modal Learning under Bird's-Eye View for Domain Generalization of 3D Semantic SegmentationabstractCross-modal Unsupervised Domain Adaptation aims to exploit the complementarity of 2D-3D data to overcome the lack of annotation in an unknown domain. However, the training of these methods relies on access to target samples, meaning the trained model only works in a specific target domain. In light of this, we propose cross-modal learning under bird’s-eye view for Domain Generalization (DG) of 3D semantic segmentation, called BEV-DG. DG is more challenging because the model cannot access the target domain during training, meaning it needs to rely on cross-modal learning to alleviate the domain gap. Since 3D semantic segmentation requires the classification of each point, existing cross-modal learning is directly conducted point-to-point, which is sensitive to the misalignment in projections between pixels and points. To this end, our approach aims to optimize domain-irrelevant representation modeling with the aid of cross-modal learning under bird’s-eye view. We propose BEV-based Area-to-area Fusion (BAF) to conduct cross-modal learning under bird’s-eye view, which has a higher fault tolerance for point-level misalignment. Furthermore, to model domain-irrelevant representations, we propose BEV-driven Domain Contrastive Learning (BDCL) with the help of cross-modal learning under bird’s-eye view. We design three domain generalization settings based on three 3D datasets, and BEV-DG significantly outperforms state-of-the-art competitors with tremendous margins in all settings. Miaoyu Li, Yachao Zhang 0001, Xu Ma 0005, Yanyun Qu, Yun Fu 0001 |
ICCV | 3 |
| 2023 | Image as Set of Points
Xu Ma 0005, Yuqian Zhou, Huan Wang 0014, Can Qin, Bin Sun 0002, Chang Liu 0022, Yun Fu 0001 |
ICLR | 1 |
| 2022 | Towards Layer-wise Image VectorizationabstractImage rasterization is a mature technique in computer graphics, while image vectorization, the reverse path of rasterization, remains a major challenge. Recent advanced deep learning-based models achieve vectorization and semantic interpolation of vector graphs and demonstrate a better topology of generating new figures. However, deep models cannot be easily generalized to out-of-domain testing data. The generated SVGs also contain complex and redundant shapes that are not quite convenient for further editing. Specifically, the crucial layer-wise topology and fundamental semantics in images are still not well understood and thus not fully explored. In this work, we propose Layer-wise Image Vectorization, namely LIVE, to convert raster images to SVGs and simultaneously maintain its image topology. LIVE can generate compact SVG forms with layer-wise structures that are semantically consistent with human perspective. We progressively add new bezier paths and optimize these paths with the layer-wise framework, newly designed loss functions, and component-wise path initialization technique. Our experiments demonstrate that LIVE presents more plausible vectorized forms than prior works and can be generalized to new images. With the help of this newly learned topology, LIVE initiates human editable SVGs for both designers and other downstream applications. Codes are made available at https://github.com/Picsart-AI-Research/LIVE-Layerwise-Image-Vectorization. Xu Ma 0005, Yuqian Zhou, Xingqian Xu, Bin Sun 0002, Valerii Filev, Nikita Orlov, Yun Fu 0001, Humphrey Shi |
CVPR | 1 |
| 2022 | Making Reconstruction-based Method Great Again for Video Anomaly DetectionabstractAnomaly detection in videos is a significant yet challenging problem. Previous approaches based on deep neural networks employ either reconstruction-based or prediction-based approaches. Nevertheless, existing reconstruction-based methods 1) rely on old-fashioned convolutional autoencoders and are poor at modeling temporal dependency; 2) are prone to overfit the training samples, leading to indistinguishable reconstruction errors of normal and abnormal frames during the inference phase. To address such issues, firstly, we get inspiration from transformer and propose Spatio-Temporal Auto-Trans-Encoder, dubbed as STATE, as a new autoencoder model for enhanced consecutive frame reconstruction. Our STATE is equipped with a specifically designed learnable convolutional attention module for efficient temporal learning and reasoning. Secondly, we put forward a novel reconstruction-based input perturbation technique during testing to further differentiate anomalous frames. With the same perturbation magnitude, the testing reconstruction error of the normal frames lowers more than that of the abnormal frames, which contributes to mitigating the overfitting problem of reconstruction. Owing to the high relevance of the frame abnormality and the objects in the frame, we conduct object-level reconstruction using both the raw frame and the corresponding optical flow patches. Finally, the anomaly score is designed based on the combination of the raw and motion reconstruction errors using perturbed inputs. Extensive experiments on benchmark video anomaly detection datasets demonstrate that our approach outperforms previous reconstruction-based methods by a notable margin, and achieves state-of-the-art anomaly detection performance consistently. The code is available at https://github.com/wyzjack/MRMGA4VAD. Yizhou Wang 0006, Can Qin, Yi Xu 0005, Xu Ma 0005, Yun Fu 0001 |
ICDM | 5 |
| 2022 | Rethinking Network Design and Local Geometry in Point Cloud: A Simple Residual MLP Framework
Xu Ma 0005, Can Qin, Haoxuan You, Haoxi Ran, Yun Fu 0001 |
ICLR | 1 |
| 2022 | MLCNN: Cross-Layer Cooperative Optimization and Accelerator Architecture for Speeding Up Deep Learning ApplicationsabstractThe ever-increasing number of layers, millions of parameters, and large data volume make deep learning workloads resource-intensive and power-hungry. In this paper, we develop a convolutional neural network (CNN) acceleration framework, named MLCNN, which explores algorithm-hardware co-design to achieve cross-layer cooperative optimization and acceleration. MLCNN dramatically reduces computation and on-off chip communication, improving CNN's performance. To achieve this, MLCNN reorders the position of nonlinear activation layers and pooling layers, which we prove results in a negligible accuracy loss; then the convolutional layer and pooling layer are co-optimized by means of redundant multiplication elimination, local addition reuse, and global addition reuse. To the best of our knowledge, MLCNN is the first of its kind that incorporates cooperative optimization across convolutional, activation, and pooling layers. We further customize the MLCNN accelerator to take full advantage of cross-layer CNN optimization to reduce both computation and on-off chip communication. Our analysis shows that MLCNN can significantly reduce (up to 98%) multiplications and additions. We have implemented a prototype of MLCNN and evaluated its performance on several widely used CNN models using both an accelerator-level cycle and energy model and RTL implementation. Experimental results show that MLCNN achieves 3.2x speedup and 2.9x energy efficiency compared with dense CNNs. MLCNN's optimization methods are orthogonal to other CNN acceleration techniques, such as quantization and pruning. Combined with quantization, our quantized MLCNN gains a 12.8x speedup and 11.3x energy efficiency compared with DCNN. Beilei Jiang, Xianwei Cheng, Sihai Tang, Xu Ma 0005, Zhaochen Gu, Song Fu, Qing Yang 0003, Mingxiong Liu |
IPDPS | 4 |
| 2022 | Parameter-Efficient Masking NetworksabstractA deeper network structure generally handles more complicated non-linearity and performs more competitively. Nowadays, advanced network designs often contain a large number of repetitive structures (e.g., Transformer). They empower the network capacity to a new level but also increase the model size inevitably, which is unfriendly to either model restoring or transferring. In this study, we are the first to investigate the representative potential of fixed random weights with limited unique values by learning diverse masks and introduce the Parameter-Efficient Masking Networks (PEMN). It also naturally leads to a new paradigm for model compression to diminish the model size. Concretely, motivated by the repetitive structures in modern neural networks, we utilize one random initialized layer, accompanied with different masks, to convey different feature mappings and represent repetitive network modules. Therefore, the model can be expressed as \textit{one-layer} with a bunch of masks, which significantly reduce the model storage cost. Furthermore, we enhance our strategy by learning masks for a model filled by padding a given random weights vector. In this way, our method can further lower the space complexity, especially for models without many repetitive architectures. We validate the potential of PEMN learning masks on random weights with limited unique values and test its effectiveness for a new compression paradigm based on different network architectures.Code is available at \href{https://github.com/yueb17/PEMN}{\textcolor{magenta}{https://github.com/yueb17/PEMN}}. Huan Wang 0014, Xu Ma 0005, Zhiqiang Tao, Yun Fu 0001 |
NeurIPS | 3 |
| 2021 | APCNN: Explore Multi-Layer Cooperation for CNN Optimization and Acceleration on FPGAabstractIn this paper, we introduce APCNN, which explores algorithm-hardware co-design and provides a CNN acceleration framework with multi-layer cooperative optimization and customized design on FPGA. In terms of the algorithm design, the pooling layer is moved before the non-linear activation function and normalization in APCNN, which we prove causes negligible accuracy loss; the pooling layer is then co-optimized with the convolutional layer by means of redundant multiplication elimination, local addition reuse, and global addition reuse. We further design a dedicated accelerator to take full advantage of convolutional-pooling cross-layer optimization to not only accelerate computation but also reduce on-off chip data communication on FPGA. We demonstrate that our novel APCNN can achieve 75% multiplication and 75% addition reduction in the best case. For on-off chip data communication, a max{Row,Col} /(Row x Col) percent of memory footprint can be eliminated, where Row and Col are the number of rows and columns in the activation feature map respectively. We have implemented a prototype of APCNN and evaluated its performance on LeNet-5 and VGG16 using both an accelerator-level cycle and energy model and an RTL implementation. Our experimental results show that APCNN achieves a 2.5× speedup and 4.7× energy efficiency compared with the dense CNN. (This research was supported in part by NSF grants CCF-1563750, OAC-2017564, and CNS-2037982.) Beilei Jiang, Xianwei Cheng, Sihai Tang, Xu Ma 0005, Zhaochen Gu, Hui Zhao 0013, Song Fu |
FPGA | 4 |
| 2021 | Learning Connected Attentions for Convolutional Neural NetworksabstractWhile self-attention mechanism has shown promising results for many vision tasks, it only considers the current features at a time. We show that such a manner cannot take full advantage of the attention mechanism. In this paper, we present Deep Connected Attention Network (DCANet), a novel design that boosts attention modules in a CNN model without any modification of the internal structure. To achieve this, we interconnect adjacent attention blocks, making information flow among attention blocks possible. With DCANet, all attention blocks in a CNN model are trained jointly, which improves the ability of attention learning. Our DCANet is generic. It is not limited to a specific attention module or base network architecture. Experimental results on ImageNet and MS COCO benchmarks show that DCANet consistently outperforms the state-of-the-art attention modules with a minimal additional computational overhead in all test cases. The code is available at: https://github.com/13952522076/DCANet. Xu Ma 0005, Jingda Guo, Sihai Tang, Zhinan Qiao, Qi Chen 0018, Qing Yang 0003, Song Fu, Paparao Palacharla, Nannan Wang 0003, Xi Wang 0001 |
ICME | 1 |
| 2021 | CoConv: Learning Dynamic Cooperative Convolution for Image RecognitionabstractIn this paper, we present a conceptually simple, yet powerful method for image recognition. The method, called Cooperative Dynamic Convolution (CoConv), introduces a cooperative learning of dynamic convolution from multiple convolutional experts. CoConv can be used as a substitute for the traditional static convolution, and can be seamlessly integrated in various visual models. Moreover, CoConv is easy to train with only a minimal computational overhead introduced in the inference phase. CoConv is trained by using multiple convolutional experts simultaneously, and the convolutional weights are merged by a weighted summation before convolutional operations for efficiency during inference. Results from extensive experiments show that CoConv leads to consistent improvement for image classification on various datasets, independent of the choice of the base convolutional network. Remarkably, CoConv improves the top-1 classification accuracy of ResNet18 by 3.06% on ImageNet. The code is available at: https://github.com/Nyquixt/CoConv. Kien X. Nguyen 0002, Tiffany Ryu, Jocelyn Zhang, Xu Ma 0005, Qing Yang 0003, Song Fu, Paparao Palacharla, Nannan Wang 0003, Xi Wang 0001 |
ICME | 4 |
| 2021 | Spatial Pyramid Attention for Deep Convolutional Neural NetworksabstractAttention mechanisms have shown great success in computer vision. However, the commonly used global average pooling in some implementations aggregates a three-dimensional feature map to a one-dimensional attention map, leading a significant loss of structural information in the attention learning. In this article, we present a novel Spatial Pyramid Attention Network (SPANet), which exploits the structural information and channel relationships for better feature representation. SPANet enhances a base network by adding Spatial Pyramid Attention (SPA) blocks laterally. By rethinking the self-attention mechanism design, we further present three topology structures of attention path connection for our SPANet. They can be flexibly applied to various CNN architectures. SPANet is conceptually simple but practically powerful. It uses both structural regularization and structural information to achieve better learning capability. We have comprehensively evaluated the performance of SPANet on four benchmark datasets for different visual tasks. The experimental results show that SPANet significantly improves the recognition accuracy without adding much computation overhead. Using SPANet, we achieve an improvement of 1.6% top-1 classification accuracy on the ImageNet 2012 benchmark based on ResNet50, and SPANet outperforms SENet and other attention methods. SPANet also significantly improves the object detection performance by a clear margin with negligible additional computation overhead. When applying SPANet to RetinaNet based on the ResNet50 backbone, we improve the performance of the baseline model by 2.3 mAP and the enhanced model outperforms SENet and GCNet by 1.1 mAP and 1.7 mAP respectively. The code of SPANet is made publicly available.11[Online]. Available:https://github.com/13952522076/SPANet_TMM Xu Ma 0005, Jingda Guo, Andrew Sansom, Mara McGuire, Andrew Kalaani, Qi Chen 0018, Sihai Tang, Qing Yang 0003, Song Fu |
IEEE Trans. Multim. | 1 |
| 2020 | Cascaded Context Dependency: An Extremely Lightweight Module For Deep Convolutional Neural NetworksabstractIn this paper, we present a cascaded context dependency module, which is a highly lightweight module that can improve the performance of deep convolutional neural networks for various visual tasks. Inspired by the feature pyramid work in object detection and the context dependency work in image recognition, we consider to cascade the contexts of multiscale feature maps to aggregate the locality and globality in a local region. We further extract the dependency between original input and cascaded contexts for feature recalibration. Without employing learnable layers, our method introduces almost no additional parameters and computations. Furthermore, Our module can be seamlessly plugged into many existing CNN architectures to improve the performance. Experiments on ImageNet and MS COCO benchmarks indicate that our method can achieve results on par with or better than related work. Qualitatively, we achieve an absolute 1.42% (77.3137% vs. 75.8974%) top-1 classification accuracy improvement based on ResNet50 on ImageNet 2012 validation set with negligible computational overhead. Besides, our method yields significant gains on the MS COCO benchmark for the object detection task. All codes and models are made publicly available1.1We submit all the codes, pre-trained models, and training log files to https://github.com/13952522076/ParameterFree. Xu Ma 0005, Zhinan Qiao, Jingda Guo, Sihai Tang, Qi Chen 0018, Qing Yang 0003, Song Fu |
ICIP | 1 |
| 2020 | Spanet: Spatial Pyramid Attention Network for Enhanced Image RecognitionabstractAttention mechanism has shown great success in computer vision. In this paper, we introduce Spatial Pyramid Attention Network (SPANet) to investigate the role of attention block for image recognition. Our SPANet is conceptually simple but practically powerful. It enhances the base network by adding Spatial Pyramid Attention (SPA) Blocks laterally. In contrast to other attention based networks that leverage global average pooling, our proposed SPANet considers both structural regularization and structural information. Furthermore, we investigate the topology structure of attention path connection and present three SPANet structures. SPA block is flexible to be deployed to various convolutional neural network (CNN) architectures. The experimental results show that our SPANet significantly improves the recognition accuracy without introducing much computation overhead compared with other CNN models. Codes are made publicly available11https://github.com/13952522076/SPANet. Jingda Guo, Xu Ma 0005, Andrew Sansom, Mara McGuire, Andrew Kalaani, Qi Chen 0018, Sihai Tang, Qing Yang 0003, Song Fu |
ICME | 2 |
| 2020 | Attention Meets Normalization and BeyondabstractTo make Convolutional Neural Networks (CNNs) more efficient and accurate, various lightweight self-attention modules have been proposed. In this paper, we systematically study state-of-the-art attention modules in CNNs and discover that self-attention mechanism can be closely related to normalization. Based on this observation, we propose a novel attention module, named Normalization-Attention module (NA module in short), which is almost parameter-free. The NA module calculates the mean and standard deviation of intermediate feature maps and processes the feature context with normalization, which makes a CNN model easier to be trained and more responsive to informative features. Our proposed Normalization-Attention module can be integrated into various base CNN architectures, and used for many computer vision tasks, including image recognition, object detection, and more. Experimental results on ImageNet and MS COCO benchmarks show that our method outperforms state-of-the-art works using fewer parameters. Codes are made publicly available. Xu Ma 0005, Jingda Guo, Qi Chen 0018, Sihai Tang, Qing Yang 0003, Song Fu |
ICME | 1 |
| 2020 | Position-Aware Recalibration Module: Learning From Feature Semantics and Feature PositionabstractWe present a new method to improve the representational power of the features in Convolutional Neural Networks (CNNs). By studying traditional image processing methods and recent CNN architectures, we propose to use positional information in CNNs for effective exploration of feature dependencies. Rather than considering feature semantics alone, we incorporate spatial positions as an augmentation for feature semantics in our design. From this vantage, we present a Position-Aware Recalibration Module (PRM in short) which recalibrates features leveraging both feature semantics and position. Furthermore, inspired by multi-head attention, our module is capable of performing multiple recalibrations where results are concatenated as the output. As PRM is efficient and easy to implement, it can be seamlessly integrated into various base networks and applied to many position-aware visual tasks. Compared to original CNNs, our PRM introduces a negligible number of parameters and FLOPs, while yielding better performance. Experimental results on ImageNet and MS COCO benchmarks show that our approach surpasses related methods by a clear margin with less computational overhead. For example, we improve the ResNet50 by absolute 1.75% (77.65% vs. 75.90%) on ImageNet 2012 validation dataset, and 1.5%~1.9% mAP on MS COCO validation dataset with almost no computational overhead. Codes are made publicly available. Xu Ma 0005, Song Fu |
IJCAI | 1 |
| 2020 | Robust discriminant feature selection via joint L2, 1-norm distance minimization and maximization
Zhangjing Yang, Qiaolin Ye, Qiao Chen 0004, Xu Ma 0005, Liyong Fu, Guowei Yang 0002, Fan Liu 0003 |
Knowl. Based Syst. | 4 |