VLDB 2026 Research / reviewers in the wild / expert
Xingrun Xing
dblp:245/7952
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-0716-859XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EfficientLLM: Unified Pruning-Aware Pretraining for Auto-Designed Compact Language ModelsabstractXingrun Xing, Zheng Liu, Shitao Xiao, Boyan Gao, Yiming Liang, Haokun Lin, Xianlin Zeng, Guoqi Li, Jiajun Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xingrun Xing, Shitao Xiao, Boyan Gao, Yiming Liang, Haokun Lin, Xianlin Zeng, Guoqi Li 0002 |
ACL (1) | 1 |
| 2025 | OmniGen: Unified Image GenerationabstractThe emergence of Large Language Models (LLMs) has unified language generation tasks and revolutionized human-machine interaction. However, in the realm of image generation, a unified model capable of handling various tasks within a single framework remains largely unexplored. In this work, we introduce OmniGen, a new diffusion model for unified image generation. OmniGen is characterized by the following features: 1) Unification: OmniGen not only demonstrates text-to-image generation capabilities but also inherently supports various downstream tasks, such as image editing, subject-driven generation, and visual-conditional generation. 2) Simplicity: The architecture of OmniGen is highly simplified, eliminating the need for additional plugins. Moreover, compared to existing diffusion models, it is more user-friendly and can complete complex tasks end-to-end through instructions without the need for extra intermediate steps, greatly simplifying the image generation workflow. 3) Knowledge Transfer: Benefit from learning in a unified format, OmniGen effectively transfers knowledge across different tasks, manages unseen tasks and domains, and exhibits novel capabilities. We also explore the model’s reasoning capabilities and potential applications of the chain-of-thought mechanism. This work represents the first attempt at a general-purpose image generation model, and we will release our resources at https://github.com/VectorSpaceLab/OmniGen to foster future advancements. Shitao Xiao, Yueze Wang, Junjie Zhou 0001, Huaying Yuan, Xingrun Xing, Ruiran Yan, Shuting Wang 0002, Tiejun Huang 0001, Zheng Liu 0011 |
CVPR | 5 |
| 2025 | From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual ModalitiesabstractMultimodal Large Language Models have made significant strides in integrating visual and textual information, yet they often struggle with effectively aligning these modalities. We introduce a novel image tokenizer that bridges this gap by applying the principle of Byte-Pair Encoding (BPE) to visual data. Unlike conventional approaches that rely on separate visual encoders, our method directly incorporates structural prior information into image tokens, mirroring the successful tokenization strategies used in text-only Large Language Models. This innovative approach enables Transformer models to more effectively learn and reason across modalities. Through theoretical analysis and extensive experiments, we demonstrate that our BPE Image Tokenizer significantly enhances MLLMs' multimodal understanding capabilities, even with limited training data. Leveraging this method, we develop Being-VL-0, a model that demonstrates superior performance across various benchmarks and shows promising scalability, potentially paving the way for more efficient and capable multimodal foundation models. For further details, visit our website https://github.com/BeingBeyond/Being-VL-0. Wanpeng Zhang 0002, Zilong Xie, Yicheng Feng, Yijiang Li, Xingrun Xing, Sipeng Zheng, Zongqing Lu 0002 |
ICLR | 5 |
| 2025 | SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based SpikingabstractRecent advancements in large language models (LLMs) with billions of parameters have improved performance in various applications, but their inference processes demand significant energy and computational resources. In contrast, the human brain, with approximately 86 billion neurons, is much more energy-efficient than LLMs with similar parameters. Inspired by this, we redesign 7$\sim$70 billion parameter LLMs using bio-plausible spiking mechanisms, emulating the efficient behavior of the human brain. We propose the first spiking large language model, SpikeLLM. Coupled with the proposed model, two essential approaches are proposed to improve spike training efficiency: Generalized Integrate-and-Fire (GIF) neurons to compress spike length from $T$ to $\frac{T}{L} \log_2 L$ bits, and an Optimal Brain Spiking framework to divide outlier channels and allocate different $T$ for GIF neurons, which further compresses spike length to approximate $log_2T$ bits. The necessity of spike-driven LLM is proved by comparison with quantized LLMs with similar operations. In the OmniQuant pipeline, SpikeLLM reduces 11.01\% WikiText2 perplexity and improves 2.55\% accuracy of common scene reasoning on a LLAMA-7B W4A4 model. In the GPTQ pipeline, SpikeLLM achieves direct additive in linear layers, significantly exceeding PB-LLMs. Our code is publicly available at https://github.com/Xingrun-Xing2/SpikeLLM. Xingrun Xing, Boyan Gao, David A. Clifton, Shitao Xiao, Wanpeng Zhang 0002 |
ICLR | 1 |
| 2024 | BiPFT: Binary Pre-trained Foundation Transformer with Low-Rank Estimation of Binarization Residual PolynomialsabstractPretrained foundation models offer substantial benefits for a wide range of downstream tasks, which can be one of the most potential techniques to access artificial general intelligence. However, scaling up foundation transformers for maximal task-agnostic knowledge has brought about computational challenges, especially on resource-limited devices such as mobiles. This work proposes the first Binary Pretrained Foundation Transformer (BiPFT) for natural language understanding (NLU) tasks, which remarkably saves 56 times operations and 28 times memory. In contrast to previous task-specific binary transformers, BiPFT exhibits a substantial enhancement in the learning capabilities of binary neural networks (BNNs), promoting BNNs into the era of pre-training. Benefiting from extensive pretraining data, we further propose a data-driven binarization method. Specifically, we first analyze the binarization error in self-attention operations and derive the polynomials of binarization error. To simulate full-precision self-attention, we define binarization error as binarization residual polynomials, and then introduce low-rank estimators to model these polynomials. Extensive experiments validate the effectiveness of BiPFTs, surpassing task-specific baseline by 15.4% average performance on the GLUE benchmark. BiPFT also demonstrates improved robustness to hyperparameter changes, improved optimization efficiency, and reduced reliance on downstream distillation, which consequently generalize on various NLU tasks and simplify the downstream pipeline of BNNs. Our code and pretrained models are publicly available at https://github.com/Xingrun-Xing/BiPFT. Xingrun Xing, Xianlin Zeng, Yequan Wang |
AAAI | 1 |
| 2024 | Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter MergingabstractSupervised fine-tuning (SFT) is crucial for adapting Large Language Models (LLMs) to specific tasks.In this work, we demonstrate that the order of training data can lead to significant training imbalances, potentially resulting in performance degradation.Consequently, we propose to mitigate this imbalance by merging SFT models fine-tuned with different data orders, thereby enhancing the overall effectiveness of SFT.Additionally, we introduce a novel technique, "parameter-selection merging," which outperforms traditional weightedaverage methods on five datasets.Further, through analysis and ablation studies, we validate the effectiveness of our method and identify the sources of performance improvements. Yiming Ju, Ziyi Ni, Xingrun Xing, Zhixiong Zeng, Siqi Fan 0001, Zheng Zhang 0006 |
EMNLP | 3 |
| 2024 | SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsabstractTowards energy-efficient artificial intelligence similar to the human brain, the bio-inspired spiking neural networks (SNNs) have advantages of biological plausibility, event-driven sparsity, and binary activation. Recently, large-scale language models exhibit promising generalization capability, making it a valuable issue to explore more general spike-driven models. However, the binary spikes in existing SNNs fail to encode adequate semantic information, placing technological challenges for generalization. This work proposes the first fully spiking mechanism for general language tasks, including both discriminative and generative ones. Different from previous spikes with 0,1 levels, we propose a more general spike formulation with bi-directional, elastic amplitude, and elastic frequency encoding, while still maintaining the addition nature of SNNs. In a single time step, the spike is enhanced by direction and amplitude information; in spike frequency, a strategy to control spike firing rate is well designed. We plug this elastic bi-spiking mechanism in language modeling, named SpikeLM. It is the first time to handle general language tasks with fully spike-driven models, which achieve much higher accuracy than previously possible. SpikeLM also greatly bridges the performance gap between SNNs and ANNs in language modeling. Our code is available at https://github.com/Xingrun-Xing/SpikeLM. Xingrun Xing, Zheng Zhang 0006, Ziyi Ni, Shitao Xiao, Yiming Ju, Siqi Fan 0001, Yequan Wang |
ICML | 1 |
| 2024 | KLoB: a Benchmark for Assessing Knowledge Localization Methods in Language Models
Yiming Ju, Xingrun Xing, Zhixiong Zeng |
PRICAI (2) | 3 |
| 2024 | SNN-BERT: Training-efficient Spiking Neural Networks for energy-efficient BERT
Qiaoyi Su, Shijie Mei 0001, Xingrun Xing, Man Yao, Bo Xu 0002, Guoqi Li 0002 |
Neural Networks | 3 |
| 2022 | Towards Accurate Binary Neural Networks via Modeling Contextual Dependencies
Xingrun Xing, Yangguang Li 0001, Wei Li 0022, Wenrui Ding, Yalong Jiang, Yufeng Wang 0004, Chunlei Liu 0001, Xianglong Liu 0001 |
ECCV (11) | 1 |
| 2022 | Equal Loss: A Simple Loss Function for Noise Robust LearningabstractTraining accurate deep neural networks in the presence of noisy labels is an important task. Though a number of approaches have been proposed for learning with noisy labels, many open issues remain. In this paper, we show that DNN learning with Cross Entropy is not robust to label noise and exhibits imbalance between the gradient of clean and noisy samples. We propose a new loss function, Equal Loss (EL), boosting DNN with a relaxed target probability and balanced gradient density. Both theoretical analysis and experiments on a range of benchmarks and real-world datasets show that EL outperforms state-of-the-art methods. Huan Peng, Chuming Li, Xingrun Xing |
ICASSP | 5 |
| 2022 | Binary Dense Predictors for Human Pose Estimation Based on Dynamic Thresholds and FilteringabstractBinary neural networks (BNNs) contribute a lot to the efficiency of image classification models. However, in dense predication tasks such as human pose estimation, predictions in different locations are coupled and rely on the extraction of features across entire images. As a result, more robust and adaptive binarization is required to bridge the performance gap between binarized and full precision models. We propose two approaches to conduct image-aware and pixel-aware dynamic binarization in a model for human pose estimation. Firstly, a simplified dynamic thresholding is leveraged in the backbone to determine unique binarization thresholds for each image. Secondly, in the decoder, we decouple binarization for each pixel according to the activations surrounding the pixel. Dynamic filtering modules are proposed to determine a different binarization strategy for each pixel. Compared with the strong baselines, the proposed framework improves 5.2% and 3.6% mAP on the COCO test-dev benchmark for ResNet-18/34 architectures respectively. Xingrun Xing, Yalong Jiang, Baochang Zhang 0001, Wenrui Ding, Huan Peng |
ICASSP | 1 |