Xue Geng

dblp:149/3281 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Low-Rank Guided Attention with Wavelet Augmentation for Sequential Recommendation
Mingxing Shao, Tiancheng Zhang 0001, Minghe Yu 0001, Xue Geng, Ge Yu 0001
DASFAA (1)5
2026 Calibrating distributions, not just networks: The PR-Q method for low-bit quantization
Xue He, Xue Geng, Tiancheng Zhang 0001, Minghe Yu 0001, Yuhai Zhao, Xulei Yang, Min Wu 0008, Ge Yu 0001
Neurocomputing2
2026 MeLoRA : Probability measures-based low-rank adaptation with Gaussian variational inference
Xue He, Xue Geng, Tiancheng Zhang 0001, Minghe Yu 0001, Yuhai Zhao, Xulei Yang, Min Wu 0008, Ge Yu 0001
Knowl. Based Syst.2
2026 LB-PTQ: Effective Low-Bit Post-Training Quantization for Vision Transformers
abstract
Recently, Vision Transformers (ViTs) have become the state-of-the-art architecture on various computer vision tasks including image classification, object detection and semantic segmentation. However, such success in high-accuracy performance comes at the price of high computational complexity, with typically tens of millions of or even more parameters in a Vision Transformer (ViT) model. Such a large volume of parameters makes it very difficult to deploy ViT models on mobile devices and cumbers their applications. In this paper, we present a novel post-training quantization approach that is able to quantize ViT models to very low bit widths, without the need of re-training. Prior works on post-training quantization for ViTs optimize the quantization of each layer separately thus leading to sub-optimal results. In contrast, we propose a unified learning framework that jointly optimizes the quantization of all layers to directly reduce the overall output error of the network. Moreover, we explore an important property of ViTs, i.e., the additivity property, revealing that the output error caused by the quantization of multiple layers equals the sum of the output error due to the quantization of each layer. Utilizing this property, we present a very efficient algorithm to solve the joint optimization problem with linear time complexity. We performed extensive experiments on the large-scale ImageNet dataset to evaluate the effectiveness of our approach. Empirical results show that our approach improves state-of-the-art noticeably on various ViT models and lowers the bit width from 8-bit to 6-bit without hurting the accuracy. Specifically, at 4 bits, our approach significantly outperforms existing works by 1.72%, 11.49%, 6.15%, and 3.54% on ViT-S, ViT-B, DeiT-S, and DeiT-B, respectively. In the end, we evaluate the performance when deploying our quantized models on hardware. Our approach achieves $1.5\times $ to $1.7\times $ speedups for the inference on NVIDIA A100 GPU.
Zhe Wang 0019, Kaixin Xu, Xue Geng, Jie Lin 0001, Mohamed M. Sabry, Min Wu 0008, Xiaoli Li 0001, Weisi Lin
IEEE Trans. Image Process.3
2025 Exploiting Temporal State Space Sharing for Video Semantic Segmentation
abstract
Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant computations, and heavy memory requirements. To this end, we introduce a Temporal Video State Space Sharing (TV3S) architecture to leverage Mamba state space models for temporal feature sharing. Our model features a selective gating mechanism that efficiently propagates relevant information across video frames, eliminating the need for a memory-heavy feature pool. By processing spatial patches independently and incorporating shifted operation, TV3S supports highly parallel computation in both training and inference stages, which reduces the delay in sequential state space processing and improves the scalability for long video sequences. Moreover, TV3S incorporates information from prior frames during inference, achieving long-range temporal coherence and superior adaptability to extended sequences. Evaluations on the VSPW and Cityscapes datasets reveal that our approach outperforms current state-of-the-art methods, establishing a new standard for VSS with consistent results across long video sequences. By achieving a good balance between accuracy and efficiency, TV3S shows a significant advancement in spatiotemporal modeling, paving the way for efficient video analysis. The code is publicly available at https://github.com/Ashesham/TV3S.git.
Syed Ariff Syed Hesham, Yun Liu 0011, Guolei Sun, Henghui Ding, Ender Konukoglu, Xue Geng, Xudong Jiang 0001
CVPR7
2025 A probabilistic optimizer for binary neural networks
Xue He, Xue Geng, Tiancheng Zhang 0001, Minghe Yu 0001, Min Wu 0008, Ge Yu 0001, Yuhai Zhao
Neurocomputing2
2025 Efficient Distortion-Minimized Layerwise Pruning
abstract
In this paper, we propose a post-training pruning framework that jointly optimizes layerwise pruning to minimize model output distortion. Through theoretical and empirical analysis, we discover an important additivity property of output distortion from pruning weights/channels in DNNs. Leveraging this property, we reformulate pruning optimization as a combinatorial problem and solve it with dynamic programming, achieving linear time complexity and making the algorithm very fast on CPUs. Furthermore, we optimize additivity-derived distortions using Hessian-based Taylor approximation to enhance pruning efficiency, accompanied by fine-grained complexity reduction techniques. Our method is evaluated on various DNN architectures, including CNNs, ViTs, and object detectors, and on vision tasks such as image classification on CIFAR-10 and ImageNet, and 3D object detection and various datasets. We achieve SoTA with significant FLOPs reductions without accuracy loss. Specifically, on CIFAR-10, we achieve up to $27.9\times$27.9×, $29.2\times$29.2×, and $14.9\times$14.9× FLOPs reductions on ResNet-32, VGG-16, and DenseNet-121, respectively. On ImageNet, we observe no accuracy loss with $1.69\times$1.69× and $2\times$2× FLOPs reductions on ResNet-50 and DeiT-Base, respectively. For 3D object detection, we achieve $\mathbf {3.89}\times, \mathbf {3.72}\times$3.89×,3.72× FLOPs reductions on CenterPoint and PVRCNN models. These results demonstrate the effectiveness and practicality of our approach for improving model performance through layer-adaptive weight pruning.
Kaixin Xu, Zhe Wang 0019, Runtao Huang, Xue Geng, Jie Lin 0001, Xulei Yang, Min Wu 0008, Xiaoli Li 0001, Weisi Lin
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 PSRR-MaxpoolNMS++: Fast Non-Maximum Suppression With Discretization and Pooling
abstract
Non-maximum suppression (NMS) is an essential post-processing step for object detection. The de-facto standard for NMS, namely GreedyNMS, is not parallelizable and could thus be the performance bottleneck in object detection pipelines. MaxpoolNMS is introduced as a fast and parallelizable alternative to GreedyNMS. However, MaxpoolNMS is only capable of replacing the GreedyNMS at the first stage of two-stage detectors like Faster R-CNN. To address this issue, we observe that MaxpoolNMS employs the process of box coordinate discretization followed by local score argmax calculation, to discard the nested-loop pipeline in GreedyNMS to enable parallelizable implementations. In this paper, we introduce a simple Relationship Recovery module and a Pyramid Shifted MaxpoolNMS module to improve the above two stages, respectively. With these two modules, our PSRR-MaxpoolNMS is a generic and parallelizable approach, which can completely replace GreedyNMS at all stages in all detectors. Furthermore, we extend PSRR-MaxpoolNMS to the more powerful PSRR-MaxpoolNMS++. As for box coordinate discretization, we propose Density-based Discretization for better adherence to the target density of the suppression. As for local score argmax calculation, we propose an Adjacent Scale Pooling scheme for mining out the duplicated box pairs more accurately and efficiently. Extensive experiments demonstrate that both our PSRR-MaxpoolNMS and PSRR-MaxpoolNMS++ outperform MaxpoolNMS by a large margin. Additionally, PSRR-MaxpoolNMS++ not only surpasses PSRR-MaxpoolNMS but also attains competitive accuracy and much better efficiency when compared with GreedyNMS. Therefore, PSRR-MaxpoolNMS++ is a parallelizable NMS solution that can effectively replace GreedyNMS at all stages in all detectors.
Tianyi Zhang 0004, Chunyun Chen, Yun Liu 0011, Xue Geng, Mohamed M. Sabry, Jie Lin 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 From Algorithm to Hardware: A Survey on Efficient and Safe Deployment of Deep Neural Networks
abstract
Deep neural networks (DNNs) have been widely used in many artificial intelligence (AI) tasks. However, deploying them brings significant challenges due to the huge cost of memory, energy, and computation. To address these challenges, researchers have developed various model compression techniques such as model quantization and model pruning. Recently, there has been a surge in research on compression methods to achieve model efficiency while retaining performance. Furthermore, more and more works focus on customizing the DNN hardware accelerators to better leverage the model compression techniques. In addition to efficiency, preserving security and privacy is critical for deploying DNNs. However, the vast and diverse body of related works can be overwhelming. This inspires us to conduct a comprehensive survey on recent research toward the goal of high-performance, cost-efficient, and safe deployment of DNNs. Our survey first covers the mainstream model compression techniques, such as model quantization, model pruning, knowledge distillation, and optimizations of nonlinear operations. We then introduce recent advances in designing hardware accelerators that can adapt to efficient model compression approaches. In addition, we discuss how homomorphic encryption can be integrated to secure DNN deployment. Finally, we discuss several issues, such as hardware evaluation, generalization, and integration of various compression approaches. Overall, we aim to provide a big picture of efficient DNNs from algorithm to hardware accelerators and security perspectives.
Xue Geng, Zhe Wang 0019, Chunyun Chen, Qing Xu 0015, Kaixin Xu, Jin Chao, Manas Gupta, Xulei Yang, Zhenghua Chen, Mohamed M. Sabry, Jie Lin 0001, Min Wu 0008, Xiaoli Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 CRS-Based Joint CFO and Channel Estimation Using Deep Learning in OFDM-Based Vehicular Communication Systems
abstract
Vehicular communication in high mobility environments has been widely explored in recent years. However, due to the existence of carrier frequency offset (CFO) and dynamic channel, the performance of vehicular communication systems over fast time-varying scenes drops severely. To address this problem, in this paper, we propose a simple and effective CRS-based joint CFO and channel estimation baseline method using deep learning (DL) for orthogonal frequency division multiplexing (OFDM) systems. Concretely, we construct a joint neural network (NN) architecture consisting of a CFO estimation network (CFOENet) based on fully connected layers and a channel estimation network (CENet) composed of convolutional layers. The proposed NN architecture can fully exploit the correlation of the cell reference signal (CRS), while learning the CFO characteristics and channel state information (CSI) changes simultaneously, which highly improves the pilot usage efficiency. We conduct adequate simulation experiments, and the results demonstrate that the proposed DL-based scheme can achieve better performance in terms of CFO estimation, channel estimation and overall system performance than conventional methods, while our method has stronger robustness and generalization ability under various channel conditions. The proposed joint CFO and channel estimation scheme has great potential in the field of the Internet of Vehicles.
Xue Geng, Yingxin Zhao
IEEE Trans. Wirel. Commun.4
2024 LPViT: Low-Power Semi-structured Pruning for Vision Transformers
Kaixin Xu, Zhe Wang 0019, Chunyun Chen, Xue Geng, Jie Lin 0001, Xulei Yang, Min Wu 0008, Xiaoli Li 0001, Weisi Lin
ECCV (71)4
2024 Training Binary Neural Networks via Gaussian Variational Inference and Low-Rank Semidefinite Programming
abstract
Current methods for training Binarized Neural Networks (BNNs) heavily rely on the heuristic straight-through estimator (STE), which crucially enables the application of SGD-based optimizers to the combinatorial training problem. Although the STE heuristics and their variants have led to significant improvements in BNN performance, their theoretical underpinnings remain unclear and relatively understudied. In this paper, we propose a theoretically motivated optimization framework for BNN training based on Gaussian variational inference. In its simplest form, our approach yields a non-convex linear programming formulation whose variables and associated gradients motivate the use of latent weights and STE gradients. More importantly, our framework allows us to formulate semidefinite programming (SDP) relaxations to the BNN training task. Such formulations are able to explicitly models pairwise correlations between weights during training, leading to a more accurate optimization characterization of the training problem. As the size of such formulations grows quadratically in the number of weights, quickly becoming intractable for large networks, we apply the Burer-Monteiro approach and only optimize over linear-size low-rank SDP solutions. Our empirical evaluation on CIFAR-10, CIFAR-100, Tiny-ImageNet and ImageNet datasets shows our method consistently outperforming all state-of-the-art algorithms for training BNNs.
Lorenzo Orecchia, Xue He, Wang Mark, Xulei Yang, Min Wu 0008, Xue Geng
NeurIPS7
2024 Evolving filter criteria for randomly initialized network pruning in image classification
Chenjing Liu, Peng Hu 0002, Jie Lin 0001, Yunhong Gong, Yingke Chen, Dezhong Peng, Xue Geng
Neurocomputing8
2024 Attention Guided Multi-Task Network for Joint CFO and Channel Estimation in OFDM Systems
abstract
The existence of high carrier frequency offset (CFO) and fading channels degrades the performance of communication systems in high mobility environments significantly. To address these challenges, in this paper, we propose an attention guided multi-task network for joint CFO and channel estimation in orthogonal frequency division multiplexing (OFDM) systems. Specifically, considering the correlation between CFO estimation and channel estimation, we construct a multi-task neural network framework which consists of a shared branch (SB), a CFO estimation branch (CFOEB) and a channel estimation branch (CEB). Among them, the SB extracts common features of two tasks, and the CFOEB and CEB make more detailed CFO and channel estimation, respectively. To further suppress the influence of CFO estimation error on channel estimation, we introduce an attention guided module (AGM) which completes the re-calibration of the channel dimension on the basis of different CFO characteristics. We train the whole network by an end-to-end paradigm and construct a suitable multi-task loss function to balance the losses between different tasks. Simulation results demonstrate that the proposed attention guided multi-task learning based joint CFO and channel estimation scheme outperforms the conventional methods and has greater robustness under various conditions.
Zhuo Chen 0010, Xue Geng, Yingxin Zhao
IEEE Trans. Wirel. Commun.3
2023 Efficient Joint Optimization of Layer-Adaptive Weight Pruning in Deep Neural Networks
abstract
In this paper, we propose a novel layer-adaptive weight-pruning approach for Deep Neural Networks (DNNs) that addresses the challenge of optimizing the output distortion minimization while adhering to a target pruning ratio constraint. Our approach takes into account the collective influence of all layers to design a layer-adaptive pruning scheme. We discover and utilize a very important additivity property of output distortion caused by pruning weights on multiple layers. This property enables us to formulate the pruning as a combinatorial optimization problem and efficiently solve it through dynamic programming. By decomposing the problem into sub-problems, we achieve linear time complexity, making our optimization algorithm fast and feasible to run on CPUs. Our extensive experiments demonstrate the superiority of our approach over existing methods on the ImageNet and CIFAR-10 datasets. On CIFAR-10, our method achieves remarkable improvements, outperforming others by up to 1.0% for ResNet-32, 0.5% for VGG-16, and 0.7% for DenseNet-121 in terms of top-1 accuracy. On ImageNet, we achieve up to 4.7% and 4.6% higher top-1 accuracy compared to other methods for VGG-16 and ResNet-50, respectively. These results highlight the effectiveness and practicality of our approach for enhancing DNN performance through layer-adaptive weight pruning. Code will be available on https://github.com/Akimoto-Cris/RD_VIT_PRUNE.
Kaixin Xu, Zhe Wang 0019, Xue Geng, Min Wu 0008, Xiaoli Li 0001, Weisi Lin
ICCV3
2022 CRAFT: Cross-Attentional Flow Transformer for Robust Optical Flow
abstract
Optical flow estimation aims to find the 2D motion field by identifying corresponding pixels between two images. Despite the tremendous progress of deep learning-based optical flow methods, it remains a challenge to accurately estimate large displacements with motion blur. This is mainly because the correlation volume, the basis of pixel matching, is computed as the dot product of the convolutional features of the two images. The locality of convolutional features makes the computed correlations susceptible to various noises. On large displacements with motion blur, noisy correlations could cause severe errors in the estimated flow. To overcome this challenge, we propose a new architecture “CRoss-Attentional Flow Trans-former” (CRAFT), aiming to revitalize the correlation volume computation. In CRAFT, a Semantic Smoothing Trans-former layer transforms the features of one frame, making them more global and semantically stable. In addition, the dot-product correlations are replaced with trans-former Cross-Frame Attention. This layer filters out feature noises through the Query and Key projections, and computes more accurate correlations. On Sintel (Final) and KITTI (foreground) benchmarks, CRAFT has achieved new state-of-the-art performance. Moreover, to test the robust-ness of different models on large motions, we designed an image shifting attack that shifts input images to generate large artificial motions. Under this attack, CRAFT per-forms much more robustly than two representative meth-ods, RAFT and GMA. The code of CRAFT is is available at https://github.com/askerlee/craft.
Xiuchao Sui, Shaohua Li 0003, Xue Geng, Yan Wu 0002, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh, Hongyuan Zhu 0002
CVPR3
2022 RDO-Q: Extremely Fine-Grained Channel-Wise Quantization via Rate-Distortion Optimization
Zhe Wang 0019, Jie Lin 0001, Xue Geng, Mohamed M. Sabry, Vijay Chandrasekhar 0001
ECCV (12)3
2020 Cascaded Mixed-Precision Networks
abstract
There has been a vast literature on Neural Network Compression, either by quantizing network variables to low precision numbers or pruning redundant connections from the network architecture. However, these techniques experience performance degradation when the compression ratio is increased to an extreme extent. In this paper, we propose Cascaded Mixed-precision Networks (CMNs), which are compact yet efficient neural networks without incurring performance drop. CMN is designed as a cascade framework by concatenating a group of neural networks with sequentially increased bitwidth. The execution flow of CMN is conditional on the difficulty of input samples, i.e., easy examples will be correctly classified by going through extremely low-bitwidth networks, and hard examples will be handled by high-bitwidth networks, so that the average compute is reduced. In addition, weight pruning is incorporated into the cascaded framework and jointly optimized with the mixed-precision quantization. To validate this method, we implemented a 2-stage CMN consisting of a binary neural network and a multi-bit (e.g. 8 bits) neural network. Empirical results on CIFAR-100 and ImageNet demonstrate that CMN performs better than state-of-the-art methods, in terms of accuracy and compute.
Xue Geng, Jie Lin 0001
ICIP1
2019 Dataflow-Based Joint Quantization for Deep Neural Networks
abstract
This paper addresses a challenging problem - how to reduce energy consumption without incurring performance drop when deploying deep neural networks (DNNs) at the inference stage. In order to alleviate the computation and storage burdens, we propose a novel dataflow-based joint quantization approach with the hypothesis that a fewer number of quantization operations would incur less information loss and thus improve the final performance. It first introduces a quantization scheme with efficient bit-shifting and rounding operations to represent network parameters and activations in low precision. Then it re-structures the network architectures to form unified modules for optimization on the quantized model. Extensive experiments on ImageNet and KITTI validate the effectiveness of our model, demonstrating that state-of-the-art results for various tasks can be achieved by this quantized model. Besides, we designed and synthesized an RTL model to measure the hardware costs among various quantization methods. For each quantization operation, it reduces area cost by about 15 times and energy consumption by about 9 times, compared to a strong baseline.
Xue Geng, Jie Fu 0001, Jie Lin 0001, Mohamed M. Sabry, Christopher Joseph Pal, Vijay Chandrasekhar 0001
DCC1
2018 Hardware-Aware Softmax Approximation for Deep Neural Networks
Xue Geng, Jie Lin 0001, Anmin Kong, Mohamed M. Sabry, Vijay Chandrasekhar 0001
ACCV (4)1
2015 Learning Image and User Features for Recommendation in Social Networks
abstract
Good representations of data do help in many machine learning tasks such as recommendation. It is often a great challenge for traditional recommender systems to learn representative features of both users and images in large social networks, in particular, social curation networks, which are characterized as the extremely sparse links between users and images, and the extremely diverse visual contents of images. To address the challenges, we propose a novel deep model which learns the unified feature representations for both users and images. This is done by transforming the heterogeneous user-image networks into homogeneous low-dimensional representations, which facilitate a recommender to trivially recommend images to users by feature similarity. We also develop a fast online algorithm that can be easily scaled up to large networks in an asynchronously parallel way. We conduct extensive experiments on a representative subset of Pinterest, containing 1,456,540 images and 1,000,000 users. Results of image recommendation experiments demonstrate that our feature learning approach significantly outperforms other state-of-the-art recommendation methods.
Xue Geng, Hanwang Zhang, Jingwen Bian, Tat-Seng Chua
ICCV1
2014 One of a Kind: User Profiling by Social Curation
abstract
Social Curation Service (SCS) is a new type of emerging social media platform, where users can select, organize and keep track of multimedia contents they like. In this paper, we take advantage of this great opportunity and target at the very starting point in social media: user profiling, which supports fundamental applications such as personalized search and recommendation. As compared to other profiling methods in conventional Social Network Services (SNS), our work benefits from the two distinguishable characteristics of SCS: a) organized multimedia user-generated contents, and b) content-centric social network. Based on these two characteristics, we are able to deploy the state-of-the-art multimedia analysis techniques to establish content-based user profiles by extracting user preferences and their social relations. First, we automatically construct a content-based user preference ontology and learn the ontological models to generate comprehensive user profiles. In particular, we propose a new deep learning strategy called multi-task convolutional neural network (mtCNN) to learn profile models and profile-related visual features simultaneously. Second, we propose to model the multi-level social relations offered by SCS to refine the user profiles in a low-rank recovery framework. To the best of our knowledge, our work is the first that explores how social curation can help in content-based social media technologies, taking user profiling as an example. Extensive experiments on 1,293 users and 1.5 million images collected from Pinterest in fashion domain demonstrate that recommendation methods based on the proposed user profiles are considerably more effective than other state-of-the-art recommendation strategies.
Xue Geng, Hanwang Zhang, Yang Yang 0002, Huan-Bo Luan, Tat-Seng Chua
ACM Multimedia1