Ke Xu 0011

dblp:181/2626-11 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-6266-4257ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 4 · 1 first-author
YearPublicationVenuePosition
2026 DeFT-LoRA: Decoupled and Fused Tuning with LoRA Experts for Universal Cross-Domain Retrieval
abstract
Universal Cross-Domain Retrieval (UCDR) aims to retrieve images across unseen domains and categories, a critical capability for real-world applications. While large-scale Vision-Language Models (VLMs) like CLIP offer strong zero-shot category generalization, they struggle with domain shifts. Existing methods often improve domain robustness at the cost of high computational overhead or by compromising the VLM's inherent knowledge. To address this, we propose Decoupled and Fused Tuning with LoRA (DeFT-LoRA), a novel and parameter-efficient framework that integrates Low-Rank Adaptation (LoRA) with a Mixture-of-Experts (MoE) mechanism. This approach resolves the intrinsic conflict between domain-invariant and domain-specific knowledge in a single adapter, enabling our model to construct a domain adapters for each input image. We propose a three-stage training strategy, which first learns a shared Base LoRA for domain-invariant features, then derives Domain-Specific Experts to capture specific styles, and finally fuses them dynamically with a lightweight gating network. Extensive experiments on three UCDR benchmarks demonstrate that DeFT-LoRA achieves comparable or superior performance to state-of-the-art methods while requiring only 1.46 percent of CLIP's image-encoder parameters and reducing computational overhead, thereby establishing an exceptional balance between accuracy and efficiency.
Ke Xu 0011, Xiaozheng Shen, Shanshan Wang 0008, Mengzhu Wang, Xun Yang 0001
AAAI1
2026 Towards personalized long-term learning modeling in knowledge tracing
Shanshan Wang 0008, Jianqi Qiu, Jiaxin Pang, Xun Yang 0001, Ke Xu 0011, Zhangling Duan, Yuanhong Zhong, Xingyi Zhang 0001
Expert Syst. Appl.5
2026 Feature Responsive LoRA: Toward Parameter-Efficient Transfer Learning for Self-Supervised Visual Models
abstract
Low-Rank Adaptation (LoRA) is a widely utilized technique in topic of Parameter-Efficient Transfer Learning (PETL) which could use a limited number of trainable parameters to adapt the model to various downstream tasks. However, the setting of the locations and low-rank sizes in traditional LoRA relies heavily on the fixed and empirical values, which may hinder adaptability and lead to sharply decreasing performance, especially on some self-supervised pre-trained models. To alleviate this dilemma, we introduce a feature responsive LoRA (ResLoRA) method, a resource-efficient algorithm that automatically determines the LoRA modules’ required size based on the downstream task’s response. Firstly, we propose a Feature Decomposition loss (FD-loss) which leverages the feature singular values to mine the corresponding features of different downstream tasks, making the model parameters able to adequately represent downstream tasks. Subsequently, we leverage the Taylor expansion to measure the salience of the model parameters, then some high-efficient parameters with high significance could be leveraged to design a dynamically responsive LoRA. Specifically, the location and low-rank sizes of LoRA are determined based on the response parameters of the features for downstream tasks. Extensive experiments show that our ResLoRA achieves state-of-the-art performance, especially in the transfer capability of self-supervised models based on MoCo v3 and MAE. Our code is available at: https://github.com/wildboarman/ResLoRA.
Shanshan Wang 0008, Xiaozheng Shen, Xun Yang 0001, Ke Xu 0011, Xingyi Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Style-Aware Contrastive Test-Time Adaptation: A Dual-Cache Model for Robust Vision-Language Alignment
abstract
Test-time adaptation (TTA) has emerged as a key strategy to enhance vision-language models (VLMs) under real-world distribution changes. However, existing methods always face two problems: 1) The fundamental trade-off dilemma: parameter-free TTA retains inference efficiency but fails to correct modality misalignment, while prompt tuning adapts to shifts, incurs high computational costs, and lacks knowledge retention. 2) Discriminative collapse also exists in TTA when faced with fine-grained downstream tasks. To alleviate these two bottlenecks, we introduce Style-aware Contrastive Test-Time Adaptation (SCTTA), a novel framework that jointly addresses modality misalignment and discriminative collapse. Firstly, we introduce Style-aware Embedding Adaptation (SEA), which dynamically refines text embeddings by incorporating domain-specific style attributes, improving alignment between visual and textual modalities. Secondly, we propose Fine-grained Contrastive Adaptation (FCA), which enhances feature separation by enforcing contrastive learning with adaptive prototypes, reducing inter-class feature overlap in fine-grained tasks. In addition, we introduce Dual-Cache Model (DCM), which extends prior unimodal cache model to a multimodal cache for the first time. Eventually, it accumulates adaptation knowledge through a visual-cache (capturing evolving domain styles) and a textual-cache (retaining discriminative semantics), enabling long-term adaptation without additional overhead. Extensive experiments on 15 datasets demonstrate that our approach achieves state-of-the-art performance for both fine-grained and out-of-distribution dataset benchmarks. Furthermore, SCTTA continuously improves as more test samples accumulate, validating its sustainable adaptation capacity. Our code is available at https://github.com/alusi123/SCTTA.
Shanshan Wang 0008, ALuSi, Xun Yang 0001, Pichao Wang, Ke Xu 0011, Xingyi Zhang 0001
IEEE Trans. Image Process.5
2024 Boosting Neural Cognitive Diagnosis with Student's Affective State Modeling
abstract
Cognitive Diagnosis Modeling aims to infer students' proficiency level on knowledge concepts from their response logs. Existing methods typically model students’ response processes as the interaction between students and exercises or concepts based on hand-crafted or deeply-learned interaction functions. Despite their promising achievements, they fail to consider the relationship between students' cognitive states and affective states in learning, e.g., the feelings of frustration, boredom, or confusion with the learning content, which is insufficient for comprehensive cognitive diagnosis in intelligent education. To fill the research gap, we propose a novel Affect-aware Cognitive Diagnosis (ACD) model which can effectively diagnose the knowledge proficiency levels of students by taking into consideration the affective factors. Specifically, we first design a student affect perception module under the assumption that the affective state is jointly influenced by the student's affect trait and the difficulty of the exercise. Then, our inferred affective distribution is further used to estimate the student's subjective factors, i.e., guessing and slipping, respectively. Finally, we integrate the estimated guessing and slipping parameters with the basic neural cognitive diagnosis framework based on the DINA model, which facilitates the modeling of complex exercising interactions in a more accurate and interpretable fashion. Besides, we also extend our affect perception module in an unsupervised learning setting based on contrastive learning, thus significantly improving the compatibility of our ACD. To the best of our knowledge, we are the first to unify the cognition modeling and affect modeling into the same framework for student cognitive diagnosis. Extensive experiments on real-world datasets clearly demonstrate the effectiveness of our ACD. Our code is available at https://github.com/zeng-zhen/ACD.
Shanshan Wang 0008, Xun Yang 0001, Ke Xu 0011, Xingyi Zhang 0001
AAAI4
2024 PTMQ: Post-training Multi-Bit Quantization of Neural Networks
abstract
The ability of model quantization with arbitrary bit-width to dynamically meet diverse bit-width requirements during runtime has attracted significant attention. Recent research has focused on optimizing large-scale training methods to achieve robust bit-width adaptation, which is a time-consuming process requiring hundreds of GPU hours. Furthermore, converting bit-widths requires recalculating statistical parameters of the norm layers, thereby impeding real-time switching of the bit-width. To overcome these challenges, we propose an efficient Post-Training Multi-bit Quantization (PTMQ) scheme that requires only a small amount of calibration data to perform block-wise reconstruction of multi-bit quantization errors. It eliminates the influence of statistical parameters by fusing norm layers, and supports real-time switching bit-widths in uniform quantization and mixed-precision quantization. To improve quantization accuracy and robustness, we propose a Multi-bit Feature Mixer technique (MFM) for fusing features of different bit-widths to enhance robustness across varying bit-widths. Moreover, we introduced the Group-wise Distillation Loss (GD-Loss) to enhance the correlation between different bit-width groups and further improve the overall performance of PTMQ. Extensive experiments demonstrate that PTMQ achieves comparable performance to existing state-of-the-art post-training quantization methods, while optimizing it speeds up by 100$\times$ compared to recent multi-bit quantization works. Code can be available at https://github.com/xuke225/PTMQ.
Ke Xu 0011, Zhongcheng Li, Shanshan Wang 0008, Xingyi Zhang 0001
AAAI1
2024 Dual-stream Feature Augmentation for Domain Generalization
Shanshan Wang 0008, ALuSi, Xun Yang 0001, Ke Xu 0011, Huibin Tan, Xingyi Zhang 0001
ACM Multimedia4
2024 BNN-SAM: Improving generalization of binary object detector by Seeking Flat Minima
Han Pu, Dezheng Zhang 0004, Ke Xu 0011, Ruchan Mo, Zhihong Yan, Dong Wang 0040
Appl. Intell.3
2023 EQ-Net: Elastic Quantization Neural Networks
abstract
Current model quantization methods have shown their promising capability in reducing storage space and computation complexity. However, due to the diversity of quantization forms supported by different hardware, one limitation of existing solutions is that usually require repeated optimization for different scenarios. How to construct a model with flexible quantization forms has been less studied. In this paper, we explore a one-shot network quantization regime, named Elastic Quantization Neural Networks (EQ-Net), which aims to train a robust weight-sharing quantization supernet. First of all, we propose an elastic quantization space (including elastic bit-width, granularity, and symmetry) to adapt to various mainstream quantitative forms. Secondly, we propose the Weight Distribution Regularization Loss (WDR-Loss) and Group Progressive Guidance Loss (GPG-Loss) to bridge the inconsistency of the distribution for weights and output logits in the elastic quantization space gap. Lastly, we incorporate genetic algorithms and the proposed Conditional Quantization-Aware Accuracy Predictor (CQAP) as an estimator to quickly search mixed-precision quantized neural networks in supernet. Extensive experiments demonstrate that our EQ-Net is close to or even better than its static counterparts as well as state-of-the-art robust bit-width methods. Code can be available at https://github.com/xuke225/EQ-Net.
Ke Xu 0011, Ye Tian 0009, Shangshang Yang, Xingyi Zhang 0001
ICCV1
2022 MultiQuant: Training Once for Multi-bit Quantization of Neural Networks
abstract
Quantization has become a popular technique to compress deep neural networks (DNNs) and reduce computational costs, but most prior work focuses on training DNNs at each individual fixed bit-width and accuracy trade-off point. How to produce a model with flexible precision is largely unexplored. This work proposes a multi-bit quantization framework (MultiQuant) to make the learned DNNs robust for different precision configuration during inference by adopting Lowest-Random-Highest bit-width co-training method. Meanwhile, we propose an online adaptive label generation strategy to alleviate the problem of vicious competition under different precision caused by one-hot labels in the supernet training. The trained supernet model can be flexibly set to different bit widths to support dynamic speed and accuracy trade-off. Furthermore, we adopt the Monte Carlo sampling-based genetic algorithm search strategy with quantization-aware accuracy predictor as evaluation criterion to incorporate the mixed precision technology in our framework. Experiment results on ImageNet datasets demonstrate MultiQuant method can attain the quantization results under different bit-widths comparable with quantization-aware training without retraining.
Ke Xu 0011, Qiantai Feng, Xingyi Zhang 0001, Dong Wang 0040
IJCAI1
2022 TA-BiDet: Task-aligned binary object detector
Han Pu, Ke Xu 0011, Dezheng Zhang 0004, Dong Wang 0040
Neurocomputing2
2021 Edge-Wise One-Level Global Pruning on NAS Generated Networks
Qiantai Feng, Ke Xu 0011, Yuhai Li, Dong Wang 0040
PRCV (4)2
2021 GenExp: Multi-objective pruning for deep neural network based on genetic algorithm
Ke Xu 0011, Dezheng Zhang 0004, Jianjing An, Dong Wang 0040
Neurocomputing1
2020 DSP-Efficient Hardware Acceleration of Convolutional Neural Network Inference on FPGAs
abstract
Field-programmable gate array (FPGA)-based accelerators for convolutional neural network (CNN) inference have received significant attention in recent years. The reported designs tend to adopt a similar underlying approach based on multiplier-accumulator (MAC) arrays, which yields strong demand for the available on-chip DSP blocks, while leaving FPGA logic and memory resources underutilized. The practical outcome is that the computational roof of the accelerator is bound by the number of DSP blocks offered by the target FPGA. In addition, integrating the CNN accelerator with other functional units that may also need DSP blocks would degrade the inference performance. Leveraging the robustness of inference accuracy to limited arithmetic precision, we propose a transformation to the convolution computation, which leads to transformation of the accelerator design space and relaxes the pressure on the required DSP resources. Through analytical and empirical evaluations, we demonstrate that our approach enables us to strike a favorable balance between utilization of the FPGA on-chip memory, logic, and DSP resources, due to which, our accelerator considerably outperforms state of the art. We report the effectiveness of our approach on a variety of FPGA devices, including Cyclone-V, Stratix-V, and Arria-10, which are used in large number of applications, ranging from embedded settings to high performance computing. Our proposed technique yields 1.5x throughput improvement and 4x DSP resource reduction compared to the best frequency domain convolution-based accelerator, and 2.5x boost in raw arithmetic performance and 8.4x saving in DSPs compared to a state-of-the-art sparse convolution-based accelerator.
Dong Wang 0040, Ke Xu 0011, Jingning Guo, Soheil Ghiasi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 ABM-SpConv: A Novel Approach to FPGA-Based Acceleration of Convolutional Neural Network Inference
abstract
Hardware accelerators for convolutional neural network (CNN) inference have been extensively studied in recent years. The reported designs tend to utilize a similar underlying architecture based on multiplier-accumulator (MAC) arrays, which has the practical consequence of limiting the FPGA-based accelerator performance by the number of available on-chip DSP blocks, while leaving other resource under-utilized. To address this problem, we consider a transformation to the convolution computation, which leads to transformation of the accelerator design space and relaxes the pressure on the required DSP resources. We demonstrate that our approach enables us to strike a judicious balance between utilization of the on-chip memory, logic, and DSP resources, due to which, our accelerator considerably outperforms state of the art. We report the effectiveness of our approach on a Stratix-V GXA7 FPGA, which shows 55% throughput improvement, while using 6.25% less DSP blocks, compared to the best reported CNN accelerator on the same device.
Dong Wang 0040, Ke Xu 0011, Qun Jia, Soheil Ghiasi
DAC2
2019 A Scalable OpenCL-Based FPGA Accelerator for YOLOv2
abstract
This paper implements an OpenCL-based FPGA accelerator for YOLOv2 on Arria-10 GX1150 FPGA board. The hardware architecture adopts a scalable pipeline design to support multi-resolution input image, and improves resource utilization by full 8-bit fixed-point computation and CONV+BN+Leaky-ReLU layer fusion technology. The proposed design achieves a peak throughput of 566 GOPs under 190 MHz working frequency. The accelerator could run YOLOv2 inference with 288×288 input resolution and tiny YOLOv2 with 416×416 input resolution at the speed of 35 and 71 FPS, respectively.
Ke Xu 0011, Xiaoyun Wang 0001, Dong Wang 0040
FCCM1
2017 PipeCNN: An OpenCL-based open-source FPGA accelerator for convolution neural networks
abstract
Convolutional neural networks (CNNs) have been employed in many applications, such as image classification, video analysis and speech recognition. Being compute-intensive, CNNs are widely accelerated by GPUs with high power dissipations. Recently, studies were carried out exploiting FPGA as CNN accelerator because of its reconfigurability and advantage on energy efficiency over GPU, especially when OpenCL-based high-level synthesis tools are now available providing fast verification and implementation flows. In this paper, we demonstrate PipeCNN - an efficient FPGA accelerator that can be implemented on a variety of FPGA platforms with reconfigurable performance and cost. The PipeCNN project is openly accessible, and thus can be used either by researchers as a generic framework to explore new hardware architectures or by teachers as a off-the-self design example for any academic courses related to FPGAs.
Dong Wang 0040, Ke Xu 0011, Diankun Jiang
FPT2