EDBT 2026 Demo / reviewers in the wild / expert
Sheng Xu 0007
dblp:10/1887-7
· DBLP profile ↗
26ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0002-7742-275XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 11 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 12 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Associative Recurrent Bilinear Optimization for Domain-Generalized Binary Neural Networks
Sheng Xu 0007, Yanjing Li, Chuanjian Liu, Baochang Zhang 0001 |
Int. J. Comput. Vis. | 1 |
| 2026 | Industrial Scene Gas Leakage Detection: A Cross-Attention Based Multimodal Feature Difference Network and a New BenchmarkabstractIndustrial gas leakage detection is critically important for safety and environmental protection. While infrared imaging enables detection of invisible gases, two challenges remain: existing datasets lack realistic industrial scenarios, and current methods struggle to distinguish gas plumes from background interferences or segment discontinuous gas distributions. This paper introduces a benchmark comprising an Industrial RGB-Thermal Dataset (IRTD) with gas emission and leakage data from laboratory and industrial sites. A VLM-assisted RGBThermal detection framework with a Cross-Attention based Feature Difference (CAFD) module is designed to enhance gasspecific feature differentiation by computing inter-modal feature discrepancies. Evaluations on public datasets and IRTD demonstrate state-of-the-art results. Linlin Yang 0001, Xingyu Guo, Sheng Xu 0007, Xianbin Cao 0001, Baochang Zhang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style MimickingabstractAlthough significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video content from input text while emulating the target identity's speaking and facial movement styles, including speech characteristics, head motion, and facial dynamics. Our framework adopts a dual-branch diffusion transformer (DiT) architecture, with one branch dedicated to audio generation and the other to video synthesis.
At the shallow layers, cross-modal fusion modules are introduced to integrate information between the two modalities. In deeper layers, each modality is processed independently, with the generated audio decoded by a vocoder and the video rendered using a GAN-based high-quality visual renderer. Leveraging DiT’s in-context learning capability through a masked-infilling strategy, our model can simultaneously capture both audio and visual styles without requiring explicit style extraction modules. Thanks to the efficiency of the DiT backbone and the optimized visual renderer, OmniTalker achieves real-time inference at 25 FPS.
To the best of our knowledge, OmniTalker is the first one-shot framework capable of jointly modeling speech and facial styles in real time. Extensive experiments demonstrate its superiority over existing methods in terms of generation quality, particularly in preserving style consistency and ensuring precise audio-video synchronization, all while maintaining efficient inference. Zhongjian Wang, Peng Zhang 0080, Jinwei Qi, Sheng Xu 0007, Bang Zhang |
NeurIPS | 5 |
| 2025 | Learning Accurate Low-bit Quantization towards Efficient Computational Imaging
Sheng Xu 0007, Yanjing Li, Chuanjian Liu, Baochang Zhang 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Calibrated gradient descent of convolutional neural networks for embodied visual recognition
Sheng Xu 0007, Lian Zhuo, Baochang Zhang 0001, Yanjing Li, Guodong Guo |
Image Vis. Comput. | 2 |
| 2024 | Bi-ViT: Pushing the Limit of Vision Transformer QuantizationabstractVision transformers (ViTs) quantization offers a promising prospect to facilitate deploying large pre-trained networks on resource-limited devices. Fully-binarized ViTs (Bi-ViT) that pushes the quantization of ViTs to its limit remain largely unexplored and a very challenging task yet, due to their unacceptable performance. Through extensive empirical analyses, we identify the severe drop in ViT binarization is caused by attention distortion in self-attention, which technically stems from the gradient vanishing and ranking disorder. To address these issues, we first introduce a learnable scaling factor to reactivate the vanished gradients and illustrate its effectiveness through theoretical and experimental analyses. We then propose a ranking-aware distillation method to rectify the disordered ranking in a teacher-student framework. Bi-ViT achieves significant improvements over popular DeiT and Swin backbones in terms of Top-1 accuracy and FLOPs. For example, with DeiT-Tiny and Swin-Tiny, our method significantly outperforms baselines by 22.1% and 21.4% respectively, while 61.5x and 56.1x theoretical acceleration in terms of FLOPs compared with real-valued counterparts on ImageNet. Our codes and models are attached on https://github.com/YanjingLi0202/Bi-ViT/ . Yanjing Li, Sheng Xu 0007, Mingbao Lin, Xianbin Cao 0001, Chuanjian Liu, Baochang Zhang 0001 |
AAAI | 2 |
| 2024 | Learning 1-Bit Tiny Object Detector with Discriminative Feature Refinementabstract1-bit detectors show impressive performance comparable to their real-valued counterparts when detecting commonly sized objects while exhibiting significant performance degradation on tiny objects. The challenge stems from the fact that high-level features extracted by 1-bit convolutions seem less compelling to reveal the discriminative foreground features. To address these issues, we introduce a Discriminative Feature Refinement method for 1-bit Detectors (DFR-Det), aiming to enhance the discriminative ability of foreground representation for tiny objects in aerial images. This is accomplished by refining the feature representation using an information bottleneck (IB) to achieve a distinctive representation of tiny objects. Specifically, we introduce a new decoder with a foreground mask, aiming to enhance the discriminative ability of high-level features for the target but suppress the background impact. Additionally, our decoder is simple but effective and can be easily mounted on existing detectors without extra burden added to the inference procedure. Extensive experiments on various tiny object detection (TOD) tasks demonstrate DFR-Det’s superiority over state-of-the-art 1-bit detectors. For example, 1-bit FCOS achieved by DFR-Det achieves the 12.8% AP on AI-TOD dataset, approaching the performance of the real-valued counterpart. Sheng Xu 0007, Yanjing Li, Mingbao Lin, Baochang Zhang 0001, David S. Doermann |
ICML | 1 |
| 2024 | RSBuilding: Toward General Remote Sensing Image Building Extraction and Change Detection With Foundation ModelabstractBuildings not only constitute a significant proportion of man-made structures but also serve as a crucial component of geographic information databases, closely linked to human activities. The intelligent interpretation of buildings plays a significant role in urban planning and management, macroeconomic analysis, population dynamics, etc. Remote sensing image building interpretation primarily encompasses building extraction and change detection (CD). However, current methodologies often treat these two tasks as separate entities, thereby failing to leverage shared knowledge. Moreover, the complexity and diversity of remote sensing image scenes pose additional challenges, as most algorithms are designed to model individual small datasets, thus lacking cross-scene generalization. In this article, we propose a comprehensive remote sensing image building understanding model, termed RSBuilding, developed from the perspective of the foundation model. RSBuilding is designed to enhance cross-scene generalization and task universality. Specifically, we extract image features based on the prior knowledge of the foundation model and devise a multilevel feature sampler to augment scale information. To unify task representation and integrate image spatiotemporal clues, we introduce a cross-attention decoder with task prompts. Addressing the current shortage of datasets that incorporate annotations for both tasks, we have developed a federated training strategy to facilitate smooth model convergence even when supervision for some tasks is missing, thereby bolstering the complementarity of different tasks. Our model was trained on a dataset comprising up to 245 000 images and validated on multiple building extraction and CD datasets. The experimental results substantiate that RSBuilding can concurrently handle two structurally distinct tasks and exhibits robust zero-shot generalization capabilities. The code will be made available for open-source access athttps://github.com/Meize0729/RSBuilding. Lili Su, Cilin Yan, Sheng Xu 0007, Pengcheng Yuan, Baochang Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Resilient Binary Neural NetworkabstractBinary neural networks (BNNs) have received ever-increasing popularity for their great capability of reducing storage burden as well as quickening inference time. However, there is a severe performance drop compared with {real-valued} networks, due to its intrinsic frequent weight oscillation during training. In this paper, we introduce a Resilient Binary Neural Network (ReBNN) to mitigate the frequent oscillation for better BNNs' training. We identify that the weight oscillation mainly stems from the non-parametric scaling factor. To address this issue, we propose to parameterize the scaling factor and introduce a weighted reconstruction loss to build an adaptive training objective. For the first time, we show that the weight oscillation is controlled by the balanced parameter attached to the reconstruction loss, which provides a theoretical foundation to parameterize it in back propagation. Based on this, we learn our ReBNN by calculating the balanced parameter based on its maximum magnitude, which can effectively mitigate the weight oscillation with a resilient training process. Extensive experiments are conducted upon various network models, such as ResNet and Faster-RCNN for computer vision, as well as BERT for natural language processing. The results demonstrate the overwhelming performance of our ReBNN over prior arts. For example, our ReBNN achieves 66.9% Top-1 accuracy with ResNet-18 backbone on the ImageNet dataset, surpassing existing state-of-the-arts by a significant margin. Our code is open-sourced at https://github.com/SteveTsui/ReBNN. Sheng Xu 0007, Yanjing Li, Teli Ma, Mingbao Lin, Hao Dong 0003, Baochang Zhang 0001, Peng Gao 0007, Jinhu Lü 0001 |
AAAI | 1 |
| 2023 | Implicit Diffusion Models for Continuous Super-ResolutionabstractImage super-resolution (SR) has attracted increasing attention due to its widespread applications. However, current SR methods generally suffer from over-smoothing and artifacts, and most work only with fixed magnifications. This paper introduces an Implicit Diffusion Model (IDM) for high-fidelity continuous image super-resolution. IDM integrates an implicit neural representation and a denoising diffusion model in a unified end-to-end framework, where the implicit neural representation is adopted in the decoding process to learn continuous-resolution representation. Furthermore, we design a scale-adaptive conditioning mechanism that consists of a low-resolution (LR) conditioning network and a scaling factor. The scaling factor regulates the resolution and accordingly modulates the proportion of the LR information and generated features in the final output, which enables the model to accommodate the continuous-resolution requirement. Extensive experiments validate the effectiveness of our IDM and demonstrate its superior performance over prior arts. The source code will be available at https://github.com/Ree1s/IDM. Sicheng Gao, Xuhui Liu, Bohan Zeng, Sheng Xu 0007, Yanjing Li, Xiaoyan Luo, Jianzhuang Liu, Xiantong Zhen, Baochang Zhang 0001 |
CVPR | 4 |
| 2023 | Q-DETR: An Efficient Low-Bit Quantized Detection TransformerabstractThe recent detection transformer (DETR) has advanced object detection, but its application on resource-constrained devices requires massive computation and memory resources. Quantization stands out as a solution by representing the network in low-bit parameters and operations. However, there is a significant performance drop when performing low-bit quantized DETR (Q-DETR) with existing quantization methods. We find that the bottle-necks of Q-DETR come from the query information distortion through our empirical analyses. This paper addresses this problem based on a distribution rectification distillation (DRD). We formulate our DRD as a bi-level optimization problem, which can be derived by generalizing the information bottleneck (IB) principle to the learning of Q-DETR. At the inner level, we conduct a distribution alignment for the queries to maximize the self-information entropy. At the upper level, we introduce a new foreground-aware query matching scheme to effectively transfer the teacher information to distillation-desired features to minimize the conditional information entropy. Extensive experimental results show that our method performs much better than prior arts. For example, the 4-bit Q-DETR can theoretically accelerate DETR with ResNet-50 backbone by 6.6× and achieve 39.4% AP, with only 2.6% performance gaps than its real-valued counterpart on the COCO dataset11Code: https://github.com/SteveTsui/Q-DETR. Sheng Xu 0007, Yanjing Li, Mingbao Lin, Peng Gao 0007, Guodong Guo, Jinhu Lü 0001, Baochang Zhang 0001 |
CVPR | 1 |
| 2023 | Representation Disparity-aware Distillation for 3D Object DetectionabstractIn this paper, we focus on developing knowledge distillation (KD) for compact 3D detectors. We observe that off-the-shelf KD methods manifest their efficacy only when the teacher model and student counterpart share similar intermediate feature representations. This might explain why they are less effective in building extreme-compact 3D detectors where significant representation disparity arises due primarily to the intrinsic sparsity and irregularity in 3D point clouds. This paper presents a novel representation disparity-aware distillation (RDD) method to address the representation disparity issue and reduce performance gap between compact students and over-parameterized teachers. This is accomplished by building our RDD from an innovative perspective of information bottleneck (IB), which can effectively minimize the disparity of proposal region pairs from student and teacher in features and logits. Extensive experiments are performed to demonstrate the superiority of our RDD over existing KD methods. For example, our RDD increases mAP of CP-Voxel-S to 57.1% on nuScenes dataset, which even surpasses teacher performance while taking up only 42% FLOPs. Yanjing Li, Sheng Xu 0007, Mingbao Lin, Jihao Yin, Baochang Zhang 0001, Xianbin Cao 0001 |
ICCV | 2 |
| 2023 | Q-DM: An Efficient Low-bit Quantized Diffusion ModelabstractDenoising diffusion generative models are capable of generating high-quality data, but suffers from the computation-costly generation process, due to a iterative noise estimation using full-precision networks. As an intuitive solution, quantization can significantly reduce the computational and memory consumption by low-bit parameters and operations. However, low-bit noise estimation networks in diffusion models (DMs) remain unexplored yet and perform much worse than the full-precision counterparts as observed in our experimental studies. In this paper, we first identify that the bottlenecks of low-bit quantized DMs come from a large distribution oscillation on activations and accumulated quantization error caused by the multi-step denoising process. To address these issues, we first develop a Timestep-aware Quantization (TaQ) method and a Noise-estimating Mimicking (NeM) scheme for low-bit quantized DMs (Q-DM) to effectively eliminate such oscillation and accumulated error respectively, leading to well-performed low-bit DMs. In this way, we propose an efficient Q-DM to calculate low-bit DMs by considering both training and inference process in the same framework. We evaluate our methods on popular DDPM and DDIM models. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, the 4-bit Q-DM theoretically accelerates the 1000-step DDPM by 7.8x and achieves a FID score of 5.17, on the unconditional CIFAR-10 dataset. Yanjing Li, Sheng Xu 0007, Xianbin Cao 0001, Baochang Zhang 0001 |
NeurIPS | 2 |
| 2023 | DCP-NAS: Discrepant Child-Parent Neural Architecture Search for 1-bit CNNs
Yanjing Li, Sheng Xu 0007, Xianbin Cao 0001, Lian Zhuo, Baochang Zhang 0001, Tian Wang 0002, Guodong Guo |
Int. J. Comput. Vis. | 2 |
| 2022 | Unleashing the Potential of Vision-Language Models for Long-Tailed Visual Recognition
Teli Ma, Shijie Geng, Mengmeng Wang 0005, Sheng Xu 0007, Hongsheng Li 0001, Baochang Zhang 0001, Peng Gao 0007, Yu Qiao 0001 |
BMVC | 4 |
| 2022 | Recurrent Bilinear Optimization for Binary Neural Networks
Sheng Xu 0007, Yanjing Li, Teli Ma, Baochang Zhang 0001, Peng Gao 0007, Yu Qiao 0001, Jinhu Lü 0001, Guodong Guo |
ECCV (24) | 1 |
| 2022 | IDa-Det: An Information Discrepancy-Aware Distillation for 1-Bit Detectors
Sheng Xu 0007, Yanjing Li, Bohan Zeng, Teli Ma, Baochang Zhang 0001, Xianbin Cao 0001, Peng Gao 0007, Jinhu Lü 0001 |
ECCV (11) | 1 |
| 2022 | Q-ViT: Accurate and Fully Quantized Low-bit Vision TransformerabstractThe large pre-trained vision transformers (ViTs) have demonstrated remarkable performance on various visual tasks, but suffer from expensive computational and memory cost problems when deployed on resource-constrained devices. Among the powerful compression approaches, quantization extremely reduces the computation and memory consumption by low-bit parameters and bit-wise operations. However, low-bit ViTs remain largely unexplored and usually suffer from a significant performance drop compared with the real-valued counterparts. In this work, through extensive empirical analysis, we first identify the bottleneck for severe performance drop comes from the information distortion of the low-bit quantized self-attention map. We then develop an information rectification module (IRM) and a distribution guided distillation (DGD) scheme for fully quantized vision transformers (Q-ViT) to effectively eliminate such distortion, leading to a fully quantized ViTs. We evaluate our methods on popular DeiT and Swin backbones. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, our Q-ViT can theoretically accelerates the ViT-S by 6.14x and achieves about 80.9% Top-1 accuracy, even surpassing the full-precision counterpart by 1.0% on ImageNet dataset. Our codes and models are attached on https://github.com/YanjingLi0202/Q-ViT Yanjing Li, Sheng Xu 0007, Baochang Zhang 0001, Xianbin Cao 0001, Peng Gao 0007, Guodong Guo |
NeurIPS | 2 |
| 2022 | Towards Compact 1-bit CNNs via Bayesian Learning
Junhe Zhao, Sheng Xu 0007, Baochang Zhang 0001, Jiaxin Gu, David S. Doermann, Guodong Guo |
Int. J. Comput. Vis. | 2 |
| 2022 | Filter pruning via expectation-maximization
Sheng Xu 0007, Yanjing Li, Linlin Yang 0001, Baochang Zhang 0001, Dianmin Sun |
Neural Comput. Appl. | 1 |
| 2022 | Data-adaptive binary neural networks for efficient object detection and recognition
Junhe Zhao, Sheng Xu 0007, Runqi Wang, Baochang Zhang 0001, Guodong Guo, David S. Doermann, Dianmin Sun |
Pattern Recognit. Lett. | 2 |
| 2022 | BiRe-ID: Binary Neural Network for Efficient Person Re-IDabstractPerson re-identification (Re-ID) has been promoted by the significant success of convolutional neural networks (CNNs). However, the application of such CNN-based Re-ID methods depends on the tremendous consumption of computation and memory resources, which affects its development on resource-limited devices such as next generation AI chips. As a result, CNN binarization has attracted increasing attention, which leads to binary neural networks (BNNs). In this article, we propose a new BNN-based framework for efficient person Re-ID (BiRe-ID). In this work, we discover that the significant performance drop of binarized models for Re-ID task is caused by the degraded representation capacity of kernels and features. To address the issues, we propose the kernel and feature refinement based on generative adversarial learning (KR-GAL and FR-GAL) to enhance the representation capacity of BNNs. We first introduce an adversarial attention mechanism to refine the binarized kernels based on their real-valued counterparts. Specifically, we introduce a scale factor to restore the scale of 1-bit convolution. And we employ an effective generative adversarial learning method to train the attention-aware scale factor. Furthermore, we introduce a self-supervised generative adversarial network to refine the low-level features using the corresponding high-level semantic information. Extensive experiments demonstrate that our BiRe-ID can be effectively implemented on various mainstream backbones for the Re-ID task. In terms of the performance, our BiRe-ID surpasses existing binarization methods by significant margins, at the level even comparable with the real-valued counterparts. For example, on Market-1501, BiRe-ID achieves 64.0% mAP on ResNet-18 backbone, with an impressive 12.51× speedup in theory and 11.75× storage saving. In particular, the KR-GAL and FR-GAL methods show strong generalization on multiple tasks such as Re-ID, image classification, object detection, and 3D point cloud processing. Sheng Xu 0007, Baochang Zhang 0001, Jinhu Lü 0001, Guodong Guo, David S. Doermann |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | POEM: 1-bit Point-wise Operations based on Expectation-Maximization for Efficient Point Cloud Processing
Sheng Xu 0007, Junhe Zhao, Yanjing Li, Baochang Zhang 0001, Guodong Guo |
BMVC | 1 |
| 2021 | Layer-Wise Searching for 1-Bit Detectorsabstract1-bit detectors show great promise for resource-constrained embedded devices but often suffer from a significant performance gap compared with their real-valued counterparts. The primary reason lies in the error during binarization. This paper presents a layer-wise searching (LWS) strategy to generate 1-bit detectors that maintain a performance very close to the original real-valued model. The approach introduces angular and amplitude loss functions to increase detector capacity. At 1-bit layers, it exploits a differentiable binarization search (DBS) to minimize the angular error in a student-teacher framework. We also learn the scale factor by minimizing the amplitude loss in the same student-teacher framework. Extensive experiments show that LWS-Det outperforms state-of-the-art 1-bit detectors by a considerable margin on the PASCAL VOC and COCO datasets. For example, the LWS-Det achieves 1-bit Faster-RCNN with ResNet-34 backbone within 2.0% mAP of its real-valued counterpart on the PASCAL VOC dataset. Sheng Xu 0007, Junhe Zhao, Jinhu Lü 0001, Baochang Zhang 0001, Shumin Han, David S. Doermann |
CVPR | 1 |
| 2021 | Efficient structured pruning based on deep feature stabilization
Sheng Xu 0007, Jinhu Lü 0001, Baochang Zhang 0001 |
Neural Comput. Appl. | 1 |
| 2021 | FlashP: An Analytical Pipeline for Real-time Forecasting of Time-Series Relational DataabstractInteractive response time is important in analytical pipelines for users to explore a sufficient number of possibilities and make informed business decisions. We consider a forecasting pipeline with large volumes of high-dimensional time series data. Real-time forecasting can be conducted in two steps. First, we specify the part of data to be focused on and the measure to be predicted by slicing, dicing, and aggregating the data. Second, a forecasting model is trained on the aggregated results to predict the trend of the specified measure. While there are a number of forecasting models available, the first step is the performance bottleneck. A natural idea is to utilize sampling to obtain approximate aggregations in real time as the input to train the forecasting model. Our scalable real-time forecasting system FlashP (Flash Prediction) is built based on this idea, with two major challenges to be resolved in this paper: first, we need to figure out how approximate aggregations affect the fitting of forecasting models, and forecasting results; and second, accordingly, what sampling algorithms we should use to obtain these approximate aggregations and how large the samples are. We introduce a new sampling scheme, called GSW sampling, and analyze error bounds for estimating aggregations using GSW samples. We introduce how to construct compact GSW samples with the existence of multiple measures to be analyzed. We conduct experiments to evaluate our solution its alternatives on real data. Shuyuan Yan, Bolin Ding, Jingren Zhou 0001, Zhewei Wei, Xiaowei Jiang, Sheng Xu 0007 |
Proc. VLDB Endow. | 7 |