EDBT 2026 Demo / reviewers in the wild / expert
Bingxin Xu
dblp:78/7674
· DBLP profile ↗
23ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CBHA-DETR: multi-kernel attention and deformable fusion network for behavior recognition in classroom monitoring
Cheng Xu 0005, Bingxin Xu |
Multim. Syst. | 4 |
| 2025 | LLaVA-Prumerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsabstractLarge Multimodal Models (LMMs) have shown significant visual reasoning capabilities by connecting a visual encoder and a large language model. LMMs typically take in a fixed and large amount of visual tokens, such as the penultimate layer features in the CLIP visual encoder, as the prefix content. Recent LMMs incorporate more complex visual inputs, such as high-resolution images and videos, which further increases the number of visual tokens significantly. However, due to the inherent design of the Transformer architecture, the computational costs of these models tend to increase quadratically with the number of input tokens. To tackle this problem, we explore a token reduction mechanism that identifies significant spatial redundancy among visual tokens. In response, we propose PruMerge, a novel adaptive visual token reduction strategy that significantly reduces the number of visual tokens without compromising the performance of LMMs. Specifically, to metric the importance of each token, we exploit the sparsity observed in the visual encoder, characterized by the sparse distribution of attention scores between the class token and visual tokens. This sparsity enables us to dynamically select the most crucial visual tokens to retain. Subsequently, we cluster the selected (unpruned) tokens based on their key similarity and merge them with the unpruned tokens, effectively supplementing and enhancing their informational content. Empirically, when applied to LLaVA-1.5, our approach can compress the visual tokens by 14 times on average, and achieve comparable performance across diverse visual question-answering and reasoning tasks. Code and checkpoints are at https://llava-prumerge.github.io/. Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, Yan Yan 0002 |
ICCV | 3 |
| 2025 | YOLODF: YOLO-Based Spatial-Frequency Interaction Mining for General Deepfake Detection
Xin Li 0184, Bingxin Xu, Hongzhe Liu 0001, Weiguo Pan, Cheng Xu 0005 |
PRCV (7) | 2 |
| 2025 | Self-attention enhanced dynamic semantic multi-scale graph convolutional network for skeleton-based action recognition
Cheng Xu 0005, Songyin Dai, Nuoya Li, Weiguo Pan, Bingxin Xu, Hongzhe Liu 0001 |
Image Vis. Comput. | 6 |
| 2025 | DAN: Distortion-aware Network for fisheye image rectification using graph reasoning
Yongjia Yan, Hongzhe Liu 0001, Cheng Xu 0005, Bingxin Xu, Weiguo Pan, Songyin Dai, Yiqing Song |
Image Vis. Comput. | 5 |
| 2025 | Ihenet: an illumination invariant hierarchical feature enhancement network for low-light object detection
Nuoya Li, Weiguo Pan, Bingxin Xu, Hongzhe Liu 0001, Songyin Dai, Cheng Xu 0005 |
Multim. Syst. | 3 |
| 2025 | Generalization-oriented face forgery detection via discriminative feature analysis and normalization
Xin Li 0184, Bingxin Xu, Hongzhe Liu 0001, Weiguo Pan, Cheng Xu 0005 |
Multim. Syst. | 2 |
| 2024 | TSD-YOLO: Small traffic sign detection based on improved YOLO v8abstractAbstract Traffic sign detection is critical for autonomous driving technology. However, accurately detecting traffic signs in complex traffic environments remains challenge despite the widespread use of one‐stage detection algorithms known for their real‐time processing capabilities. In this paper, the authors propose a traffic sign detection method based on YOLO v8. Specifically, this study introduces the Space‐to‐Depth (SPD) module to address missed detections caused by multi‐scale variations of traffic signs in traffic scenes. The SPD module compresses spatial information into depth channels, expanding the receptive field and enhancing the detection capabilities for objects of varying sizes. Furthermore, to address missed detections caused by complex backgrounds such as trees, this paper employs the Select Kernel attention mechanism. This mechanism enables the model to dynamically adjust its focus and more effectively concentrate on key features. Additionally, considering the uneven distribution of training data, the authors adopted the WIoUv3 loss function, which optimizes loss calculation through a weighted approach, thereby improving the model's detection performance across various sizes and frequencies of instances. The proposed methods were validated on the CCTSDB and TT100K datasets. Experimental results demonstrate that the authors’ method achieves substantial improvements of 3.2% and 5.1% on the mAP50 metric compared to YOLOv8s, while maintaining high detection speed, significantly enhancing the overall performance of the detection system. The code for this paper is located at https://github.com/dusongjie/TSD‐YOLO‐Small‐Traffic‐Sign‐Detection‐Based‐on‐Improved‐YOLO‐v8 Songjie Du, Weiguo Pan, Nuoya Li, Songyin Dai, Bingxin Xu, Hongzhe Liu 0001, Cheng Xu 0005, Xuewei Li 0006 |
IET Image Process. | 5 |
| 2024 | SFDiff: Diffusion model with sufficient spatial-Fourier frequency information interaction for low-light image enhancementabstractAbstract Diffusion models are increasingly applied in low‐light image enhancement tasks due to their exceptional capability to model data distributions, but most current methods focus only on the original pixel space and neglect the potential of Fourier frequency information. In this article, SFDiff is proposed, a novel low‐light image enhancement method that integrates Fourier frequency information into the diffusion process. Specifically, Fourier transforms are applied at both the image and feature levels to separately enhance the amplitude and phase components, which restores global illumination degradation and positional information. Then a Spatial‐Frequency Fusion (SFF) block is used to fully integrate and interact with the information across spatial and frequency domains. Since illumination degradation is primarily manifested in the amplitude component, a loss function based on maximum likelihood learning is employed to constrain the amplitude component at each step of the sampling process, ensuring that the reverse process maintains an optimal trajectory. Owing to the streamlined network design and the fact that the Fourier transform requires no extra parameters, SFDiff achieves a reduction in parameters of over compared to several state‐of‐the‐art (SOTA) diffusion models and delivers high‐quality enhancement results on multiple real‐world datasets. The code is available at https://github.com/MrWan001/SFDiff . Bingxin Xu, Jingli Yao, Weiguo Pan, Hongzhe Liu 0001 |
IET Image Process. | 2 |
| 2024 | FSKT-GE: Feature maps similarity knowledge transfer for low-resolution gaze estimationabstractAbstract The limited of texture details information in low‐resolution facial or eye images presents a challenge for gaze estimation. To address this, FSKT‐GE (feature maps similarity knowledge transfer for low‐resolution gaze estimation) is proposed, a gaze estimation framework consisting of both a high resolution (HR) network and low resolution (LR) network with the identical structure. Rather than mere feature imitation, this issue is addressed by assessing the cosine similarity of feature layers, emphasizing the distribution similarity between the HR and LR networks. This enables the LR network to acquire richer knowledge. This framework utilizes a combination loss function, incorporating cosine similarity measurement, soft loss based on probability distribution difference and gaze direction output, along with a hard loss from the LR network output layer. This approach on low‐resolution datasets derived from Gaze360 and RT‐Gene datasets is validated, demonstrating excellent performance in low‐resolution gaze estimation. Evaluations on low‐resolution images obtained through 2×, 4×, and 8× down‐sampling are conducted on two datasets. On the Gaze360 dataset, the lowest mean angular errors of 10.97°, 11.22°, and 13.61° were achieved, while on the RT‐Gene dataset, the lowest mean angular errors of 6.73°, 6.83°, and 7.75° were obtained. Weiguo Pan, Songyin Dai, Bingxin Xu, Cheng Xu 0005, Hongzhe Liu 0001, Xuewei Li 0006 |
IET Image Process. | 4 |
| 2024 | PSC diffusion: patch-based simplified conditional diffusion model for low-light image enhancement
Bingxin Xu, Weiguo Pan, Hongzhe Liu 0001 |
Multim. Syst. | 2 |
| 2023 | Causal-DFQ: Causality Guided Data-free Network QuantizationabstractModel quantization, which aims to compress deep neural networks and accelerate inference speed, has greatly facilitated the development of cumbersome models on mobile and edge devices. There is a common assumption in quantization methods from prior works that training data is available. In practice, however, this assumption cannot always be fulfilled due to reasons of privacy and security, rendering these methods inapplicable in real-life situations. Thus, data-free network quantization has recently received significant attention in neural network compression. Causal reasoning provides an intuitive way to model causal relationships to eliminate data-driven correlations, making causality an essential component of analyzing data-free problems. However, causal formulations of data-free quantization are inadequate in the literature. To bridge this gap, we construct a causal graph to model the data generation and discrepancy reduction between the pre-trained and quantized models. Inspired by the causal understanding, we propose the Causality-guided Data-free Network Quantization method, Causal-DFQ, to eliminate the reliance on data via approaching an equilibrium of causality-driven intervened distributions. Specifically, we design a content-style-decoupled generator, synthesizing images conditioned on the relevant and irrelevant factors; then we propose a discrepancy reduction loss to align the intervened distributions of the pre-trained and quantized models. It is worth noting that our work is the first attempt towards introducing causality to data-free quantization problem. Extensive experiments demonstrate the efficacy of Causal-DFQ. The code is available at Causal-DFQ. Yuzhang Shang, Bingxin Xu, Gaowen Liu, Ramana Rao Kompella, Yan Yan 0002 |
ICCV | 2 |
| 2022 | Bayesian Pseudoinverse Learners: From Uncertainty to Deterministic LearningabstractPseudo-inverse learners (PILs) are a kind of feedforward neural network trained with the pseudoinverse learning algorithm, which can be traced back to 1995 originally. PIL is an approach for nongradient descent learning, and its main advantage is the lower computational cost and fast learning procedure, which is especially relevant in the edge computing research field. However, PIL is mostly applied to a deterministic learning problem, while in the real world, the greatest case that is of concern is the uncertainty learning problem. In this work, under the framework of the synergetic learning system (SLS), we introduce an approximated synergetic learning scheme, which can transform uncertainty learning into deterministic learning. We call this new learning framework the Bayesian PIL, and the advantages are also demonstrated in this work. Qian Yin 0001, Bingxin Xu, Kaiyan Zhou, Ping Guo 0002 |
IEEE Trans. Cybern. | 2 |
| 2019 | Multi-channel expected patch log likelihood for color image denoising
Xiuling Zhou, Bingxin Xu, Ping Guo 0002 |
Neurocomputing | 2 |
| 2018 | Pseudoinverse Learning Algorithom for Fast Sparse Autoencoder TrainingabstractSparse autoencoder is one approach to automatically learn features from unlabeled data and received significant attention during the development of deep neural networks. However, the learning algorithm of sparse autoencoder suffers from slow learning speed because of gradient descent based algorithms have many drawbacks. In this paper, a fast learning algorithm for sparse autoenceder is proposed which based on pseudoinverse learning algorithm (PIL). The proposed method calculates encoder weight matrix by truncating the pseudoinverse matrix of input data. The pseudoinverse truncation matrix is used as the weights of encoder, and then the input data is mapped to the hidden layer space through the biased ReLU activation function. The decoder weights are also can computed by the PIL. Unlike the gradient descent based algorithm, the proposed method does not require a time-consuming iterative optimization process and select many user-dependent parameters such as learning rate or momentum constant too. The experimental results indicate the superiority of proposed method which is very efficient and also can learned the sparsity of samples. Bingxin Xu, Ping Guo 0002 |
CEC | 1 |
| 2018 | Broad and Pseudoinverse Learning for AutoencoderabstractAutoencoder is one approach to automatically learn features from unlabeled data and received significant attention during the development of deep neural networks. However, the learning algorithm of autoencoder suffers from slow learning speed because of gradient descent based algorithms have many drawbacks. Pseudoinverse learning algorithm is a fast and fully automated method to train autoencoders. While when the dimension of data is far less than the number of data, the pseudoinverse learning can only obtain the optimal initial value of the autoencoder network and need further learning to achieve satisfactory results. In order to overcome the shortcomings mentioned above, we present a broad learning strategy to transform the input space to the high dimensional space through receptive function in this paper. The transformed data can be more suitable to pseudoinverse learning algorithm which can be obtained the accurate results of autoencoder efficiently. The experimental results show that the proposed method can achieve a comprehensively better performance in terms of training autoencoder efficiency and accuracy. Bingxin Xu, Ping Guo 0002 |
SMC | 1 |
| 2016 | Image representation via sub-dictionary based sparse codingabstractIn this paper, a sub-dictionary based sparse coding method is proposed for image representation. The novel sparse coding method substitutes a new regularization item for L1-norm in the sparse representation model. The proposed sparse coding method involves a series of sub-dictionaries. Each sub-dictionary contains all the training samples except for those from one particular category. For the test sample to be represented, all the sub-dictionaries should linearly represent it apart from the one that does not contain samples from that label, and this sub-dictionary is called irrelevant sub-dictionary. This new regularization item restricts the sparsity of each sub-dictionary's residual, and this restriction is helpful for classification. The experimental results demonstrate that the proposed method is superior to the previous related sparse representation based classification. Bingxin Xu, Qian Yin 0001, Ping Guo 0002, Hongzhe Liu 0001 |
IJCNN | 1 |
| 2013 | Combining affinity propagation with supervised dictionary learning for image classification
Bingxin Xu, Rukun Hu, Ping Guo 0002 |
Neural Comput. Appl. | 1 |
| 2012 | Avoiding forgetfulness: Structured English specifications for high-level robot control with implicit memoryabstractThis paper addresses the challenge of incorporating event memory into the automatic synthesis of hybrid controllers for high-level reactive robot behavior. The goal is to provide a natural, concise grammar for specifying high-level tasks that require remembering past events, and to ensure that the required memory is correctly updated during controller execution. To this end, a structured English grammar for specifying high level behavior is provided that automatically performs memory operations, without requiring explicit definition from the specification designer. This grammar admits intuitive, unambiguous specifications for tasks that implicitly use memory for purposes including non-repeated goals, strictly ordered action sequences, etc. The proposed framework also guarantees the correctness of memory operations during continuous execution. The approach is implemented within the LTLMoP toolkit for reactive mission planning. Vasumathi Raman, Bingxin Xu, Hadas Kress-Gazit |
IROS | 2 |
| 2011 | A Generalized Subspace Projection Approach for Sparse Representation Classification
Bingxin Xu, Ping Guo 0002 |
ICONIP (2) | 1 |
| 2010 | Kernel ICA applied to feature extraction for image annotationabstractIn automatic image annotation, it is often extracting low-level visual features from original image for the purpose of mapping to high level image semantic information. In this paper, we propose a novel method which integrates kernel independent component analysis (KICA) and support vector machine (SVM) for analyzing the semantic information of natural images. KICA, which contains a nonlinear kernel mapping component, is adopted to extract low-level features from the original image data. Then these feature vectors are mapped to high-level semantic words using SVM to annotate images with labels in a given semantic label set. Comparative studies have done for the performance of KICA with traditional color histogram and discrete cosine transform features. The experimental results show that the proposed method is capable of extracting the components of images as key features, and with these features to map into semantic categories, higher accuracy is achieved. Bingxin Xu, Ping Guo 0002 |
ISDA | 1 |
| 2010 | Optimization of Training Samples with Affinity Propagation Algorithm for Multi-class SVM Classification
Guangjun Lv, Qian Yin 0001, Bingxin Xu, Ping Guo 0002 |
ISNN (2) | 3 |
| 2010 | Modeling multi-source remote sensing image classifier based on the MDL principle: Experimental studiesabstractIn classification of multi-source remote sensing image, it is usually difficult to obtain higher classification accuracy. In the previous work, the modeling technique for the remote sensing image classification based on the minimum description length (MDL) principle with mixture model is analyzed theoretically. In this work, experimental studies are performed for investigating the modeling technique. With intensive experiments and sophisticated analysis, it is found that the developed modeling technique can build a robust classification system, which can avoid classifier over-fitting training data and make the learning process trade-off between bias and variance. Meanwhile, designed mixture model is more efficient to represent real multi-source remote sensing images compared to single model. Huaiying Xia, Rukun Hu, Bingxin Xu, Ping Guo 0002 |
SMC | 3 |