VLDB 2026 Research / reviewers in the wild / expert
Xuewen Zhang
dblp:16/10338
· DBLP profile ↗
28ranked-venue papers
12as first author
18since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Inference Compute-Optimal Video Vision Language ModelsabstractPeiqi Wang, ShengYun Peng, Xuewen Zhang, Hanchao Yu, Yibo Yang, Lifu Huang, Fujun Liu, Qifan Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shengyun Peng, Xuewen Zhang, Hanchao Yu, Lifu Huang, Fujun Liu, Qifan Wang 0001 |
ACL (1) | 3 |
| 2025 | PAIR-VAE: Variational Pairwise Augmentation with Label Sharing for Generalizable Drug-Target Interaction PredictionabstractDrug-target interaction prediction plays a vital role in computer-aided drug discovery, yet existing models often suffer from poor generalization to unseen drugs or targets due to distributional shifts. To address this challenge, we propose PAIR-VAE, a novel variational autoencoder framework that incorporates a pairwise augmented interaction reconstruction mechanism. PAIR-VAE reconstructs perturbed variants of drugs and targets to generate three types of synthetic interaction pairs, which are jointly trained with the original pairs under a label-sharing co-regularized objective. To enhance prediction stability, the model introduces a mean-based variational embedding strategy to mitigate the noise caused by stochastic sampling. Experimental results on two benchmark datasets demonstrate that PAIRVAE consistently outperforms existing methods across various generalization settings. Notably, in the challenging unseen-pair scenario, it achieves up to 5% and 6% improvements in AUROC and AUPRC, respectively. Moreover, PAIR-VAE maintains strong performance under a more difficult cross-domain split, where training and test sets share minimal structural similarity, further validating its real-world generalization capability. To the best of our knowledge, PAIR-VAE is the first framework to introduce variational pairwise augmentation and label-sharing supervision for DTI prediction, offering a generalizable and robust solution for real-world drug discovery. Xuewen Zhang, An-Yang Lu |
BIBM | 2 |
| 2025 | CompCap: Improving Multimodal Large Language Models with Composite CaptionsabstractHow well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as charts, posters, or screenshots, rather than being captured directly by a camera. While CIs are prevalent in real-world applications, recent MLLM developments have primarily focused on interpreting natural images (NIs). Our research reveals that current MLLMs face significant challenges in accurately understanding CIs, often struggling to extract information or perform complex reasoning based on these images. We find that existing training data for CIs are mostly formatted for question-answer tasks (e.g., in datasets like ChartQA and ScienceQA), while high-quality image-caption datasets, critical for robust vision-language alignment, are only available for NIs. To bridge this gap, we introduce Composite Captions (CompCap), a flexible framework that leverages Large Language Models (LLMs) and automation tools to synthesize CIs with accurate and detailed captions. Using CompCap, we curate CompCap-118K, a dataset containing 118K image-caption pairs across six CI types. We validate the effectiveness of CompCap-118K by supervised fine-tuning MLLMs of three sizes: xGen-MM-inst.-4B and LLaVA-NeXT-Vicuna-7B/13B. Empirical results show that CompCap-118K significantly enhances MLLMs' understanding of CIs, yielding average gains of 1.7%, 2.0%, and 2.9% across eleven benchmarks, respectively. Satya Narayan Shukla, Mahmoud Azab, Aashu Singh, Qifan Wang 0001, Shengyun Peng, Hanchao Yu, Shen Yan 0007, Xuewen Zhang, Baosheng He |
ICCV | 10 |
| 2025 | Evolving Knowledge Distillation for Lightweight Neural Machine TranslationabstractRecent advancements in Neural Machine Translation (NMT) have significantly improved translation quality. However, the increasing size and complexity of state-of-the-art models present significant challenges for deployment on resource-limited devices. Knowledge distillation (KD) is a promising approach for compressing models, but its effectiveness diminishes when there is a large capacity gap between teacher and student models. To address this issue, we propose Evolving Knowledge Distillation (EKD), a progressive training framework in which the student model learns from a sequence of teachers with gradually increasing capacities. Experiments on IWSLT14, WMT-17, and WMT-23 benchmarks show that EKD leads to consistent improvements at each stage. On IWSLT-14, the final student achieves a BLEU score of 34.24, narrowing the gap to the strongest teacher (34.32 BLEU) to just 0.08 BLEU. Similar trends are observed on other datasets. These results demonstrate that EKD effectively bridges the capacity gap, enabling compact models to achieve performance close to that of much larger teacher models.Code and models are available at https://github.com/agi-content-generation/EKD. Xuewen Zhang, Haixiao Zhang, Xinlong Huang |
ICTAI | 1 |
| 2024 | Spectral Calibration of the Advanced Hyperspectral Imager Aboard the Gaofen-5 02 SatelliteabstractThe Gaofen-5 02 (GF-5 02) satellite is equipped with two Earth observation sensors, the Advanced Hyperspectral Imager (AHSI) and the Full-Spectrum Spectral Imager (VIMI), which are used in environmental monitoring, natural resource exploration and so on. In this study, a spectral calibration method by matching the effective top-of-atmosphere (TOA) reflectance and atmospheric transmittance at the atmospheric absorption channels of 760 nm, 940 nm, 1140 nm, and 2060 nm was introduced. The spectral calibration method was validated with simulated hyperspectral data and the accuracy of center wavelength and full-width at half-maximum (FWHM) were generally better than 0.18 nm and 0.57 nm, respectively. Then, the method was applied to calibrate the center wavelength and FWHM of the GF-5 02/AHSI with a hyperspectral image collected on 7 May 2022. The spectral calibration results showed that the center wavelength and FWHM of GF-5 02/AHSI shifts with -0.375 nm and 0.35 nm, respectively. These results demonstrated a significant improvement in the surface reflectance spectrum of the inversions after spectral calibration at the Railroad Valley Playa (RVP) site. Yingxian Wang, Yaokai Liu, Qiongqiong Lan, Taoming Qi, Xuewen Zhang, Qijin Han |
IGARSS | 5 |
| 2023 | Data-Driven Linear Predictive Control of Nonlinear Processes Based on Reduced-Order Koopman OperatorabstractIn this paper, we propose an efficient data-driven predictive control approach for general nonlinear processes based on a reduced-order Koopman operator. A Kalman-based sparse identification of nonlinear dynamics method is employed to select lifting functions for Koopman identification. The selected lifting functions are used to project the original nonlinear state space into a higher-dimensional linear function space, in which Koopman-based linear models may be constructed for the underlying nonlinear process. To address the potential issue of a significant increase in the dimensionality of the resulting full-order Koopman models caused by the use of lifting functions, we propose a reduced-order Koopman modeling approach based on proper orthogonal decomposition. A computationally efficient linear robust predictive control scheme is established based on the reduced-order Koopman model. A case study on a benchmark chemical process is conducted to illustrate the proposed framework. Xuewen Zhang, Minghao Han, Xunyuan Yin |
SMC | 1 |
| 2022 | Nonparametric Forest-Structured Neural Topic ModelingabstractNeural topic models have been widely used in discovering the latent semantics from a corpus. Recently, there are several researches on hierarchical neural topic models since the relationships among topics are valuable for data analysis and exploration. However, the existing hierarchical neural topic models are limited to generate a single topic tree. In this study, we present a nonparametric forest-structured neural topic model by firstly applying the self-attention mechanism to capture parent-child topic relationships, and then build a sparse directed acyclic graph to form a topic forest. Experiments indicate that our model can automatically learn a forest-structured topic hierarchy with indefinite numbers of trees and leaves, and significantly outperforms the baseline models on topic hierarchical rationality and affinity. Xuewen Zhang, Yanghui Rao |
COLING | 2 |
| 2022 | FPD: Feature Pyramid Knowledge Distillation
Qi Wang 0089, Wenxin Yu 0001, Xuewen Zhang, Jun Gong 0001 |
ICONIP (1) | 7 |
| 2022 | Music to Dance: Motion Generation Based on Multi-Feature Fusion StrategyabstractStudies on generating dance sequences from music can greatly promote dance teaching and animation production. However, the current results of such tasks are not very satisfactory. Due to the lack of consideration of human body structure, unreal phenomena such as slippery feet and twisted joints will appear in the resulting movements. Since only a single correlation feature between music and dance is learned, there is a problem of poor matching between generated actions and music. To solve these problems, we propose a learning framework based on GAN and a multi-feature fusion strategy to realize the mapping from music to dance. We use two discriminators to constrain style and authenticity respectively to make our actions generated by the network more coherent and natural, which is essential in producing authentic dance sequences. Moreover, we extracted the three characteristics of music style, beat, and structure. Then we used the feature fusion method to obtain a comprehensive feature representation, which guarantees the consistency of the generated dance sequence with the input music to the greatest extent. The experimental quantitative and qualitative results show that our method can generate accurate, consistent, beat-matching dance movements from music. In addition, we also collected and produced a data set containing music features and corresponding pose sequences, which is convenient for pose generation based on music features. Yufei Gao 0002, Wenxin Yu 0001, Xuewen Zhang |
ISCAS | 3 |
| 2022 | Lifelong topic modeling with knowledge-enhanced adversarial network
Xuewen Zhang, Yanghui Rao, Qing Li 0001 |
World Wide Web | 1 |
| 2021 | A Progressive Image Inpainting Algorithm with a Mask Auto-update Branch
Liang Nie, Wenxin Yu 0001, Xuewen Zhang, Siyuan Li 0004, Ning Jiang 0002 |
ICANN (2) | 3 |
| 2021 | SCAN: Spatial and Channel Attention Normalization for Image Inpainting
Wenxin Yu 0001, Liang Nie, Xuewen Zhang, Siyuan Li 0004, Jun Gong 0001 |
ICONIP (6) | 4 |
| 2021 | Progressive Inpainting Strategy with Partial Convolutions Generative Networks (PPCGN)
Liang Nie, Wenxin Yu 0001, Siyuan Li 0004, Ning Jiang 0002, Xuewen Zhang, Jun Gong 0001 |
ICONIP (6) | 6 |
| 2021 | Free-Form Image Inpainting with Separable Gate Encoder-Decoder Network
Liang Nie, Wenxin Yu 0001, Xuewen Zhang, Siyuan Li 0004, Jun Gong 0001 |
ICONIP (3) | 3 |
| 2021 | QS-Hyper: A Quality-Sensitive Hyper Network for the No-Reference Image Quality Assessment
Xuewen Zhang, Yunye Zhang, Wenxin Yu 0001, Liang Nie, Ning Jiang 0002, Jun Gong 0001 |
ICONIP (4) | 1 |
| 2021 | Semi-supervised Learning with Conditional GANs for Blind Generated Image Quality Assessment
Xuewen Zhang, Yunye Zhang, Wenxin Yu 0001, Liang Nie, Jun Gong 0001 |
ICONIP (3) | 1 |
| 2021 | SPS: A Subjective Perception Score for Text-to-Image SynthesisabstractA fundamental problem of text-to-image synthesis is the lack of quality assessment for a single generated image. Quantitative indicators of this work (such as Inception Score and Fréchet Inception Distance) only affect plenty of images' feature distribution. It causes monotonous evaluation and plenty of poor-quality image results. This paper proposes a new evaluation criterion for text-to-image synthesis by the blind image quality assessment(BIQA) method. To train the model, a Multi-Metrics Quality Assessment Dataset for generated birds' images(MMQA) is proposed. Besides, the Multi-hyper model is proposed to fit our dataset better. Experiments show that our method evaluates text- to-image tasks more comprehensively and optimize their results. Xuewen Zhang, Wenxin Yu 0001, Ning Jiang 0002, Yunye Zhang |
ISCAS | 1 |
| 2021 | Time-Series Regeneration With Convolutional Recurrent Generative Adversarial Network for Remaining Useful Life EstimationabstractFor health prognostic task, ever-increasing efforts have been focused on machine learning based methods, which are capable of yielding accurate remaining useful life (RUL) estimation for industrial equipment or components without exploring the degradation mechanism. A prerequisite ensuring the success of these methods depends on a wealth of run-to-failure data; however, run-to-failure data may be insufficient in practice. That is, conducting a substantial amount of destructive experiments not only is of high cost but also may cause catastrophic consequences. Out of this consideration, an enhanced RUL framework focusing on data self-generation is put forward for both noncyclic and cyclic degradation patterns for the first time. It is designed to enrich data from a data-driven way, generating realistic-like time-series to enhance current RUL methods. First, high-quality data generation is ensured through the proposed convolutional recurrent generative adversarial network, which adopts a two-channel fusion convolutional recurrent neural network. Next, a hierarchical framework is proposed to combine generated data into current RUL estimation methods. Finally, in this article the efficacy of the proposed method is verified through both noncyclic and cyclic degradation systems. With the enhanced RUL framework, an aero-engine system following noncyclic degradation has been tested using three typical RUL models. State-of-the-art RUL estimation results are achieved by enhancing capsule network with generated time-series. Specifically, estimation errors evaluated by the index score function have been reduced by 21.77$\%$ and 32.67$\%$ for the two employed operating conditions, respectively. Besides, the estimation error is reduced to zero for the lithium-ion battery system, which presents cyclic degradation. Xuewen Zhang, Chau Yuen, Lahiru Jayasinghe, Xiang Liu 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Customizable GAN: Customizable Image Synthesis Based on Adversarial Learning
Wenxin Yu 0001, Jinjia Zhou, Xuewen Zhang, Jialiang Tang, Siyuan Li 0004, Ning Jiang 0002, Gang He 0001, Gang He 0002 |
ICONIP (4) | 4 |
| 2020 | No-Reference Quality Assessment Based on Spatial Statistic for Generated Images
Yunye Zhang, Xuewen Zhang, Wenxin Yu 0001, Ning Jiang 0002, Gang He 0002 |
ICONIP (4) | 2 |
| 2020 | Deep Feature Compatibility for Generated Images Quality Assessment
Xuewen Zhang, Yunye Zhang, Wenxin Yu 0001, Ning Jiang 0002, Gang He 0001 |
ICONIP (4) | 1 |
| 2020 | Deep Learning Based Defect Detection for Solder Joints on Industrial X-Ray Circuit Board ImagesabstractQuality control is of vital importance during electronics production. As the methods of producing electronic circuits improve, there is an increasing chance of solder defects during assembling the printed circuit board (PCB). Many technologies have been incorporated for inspecting failed soldering, such as X-ray imaging, optical imaging, and thermal imaging. With some advanced algorithms, the new technologies are expected to control the production quality based on the digital images. However, current algorithms sometimes are not accurate enough to meet the quality control. Specialists are needed to do a follow-up checking. For automated X-ray inspection, joint of interest on the X-ray image is located by region of interest (ROI) and inspected by some algorithms. Some incorrect ROIs deteriorate the inspection algorithm. The high dimension of X-ray images and the varying sizes of image dimensions also challenge the inspection algorithms. On the other hand, recent advances on deep learning shed light on image-based tasks and are competitive to human levels. In this paper, deep learning is incorporated in X-ray imaging based quality control during PCB quality inspection. Two artificial intelligence (AI) based models are proposed and compared for joint defect detection. The noised ROI problem and the varying sizes of imaging dimension problem are addressed. The efficacy of the proposed methods are verified through experimenting on a real-world 3D X-ray dataset. By incorporating the proposed methods, specialist inspection workload is largely saved. Qianru Zhang, Meng Zhang 0010, Chinthaka Gamanayake, Chau Yuen, Zehao Geng, Hirunima Jayasekaraand, Xuewen Zhang, Chia-wei Woo, Jenny Chen Ni Low, Xiang Liu 0001 |
INDIN | 7 |
| 2019 | Hie-Transformer: A Hierarchical Hybrid Transformer for Abstractive Article Summarization
Xuewen Zhang, Gongshen Liu |
ICONIP (3) | 1 |
| 2017 | Embroidery: Patching Vulnerable Binary Code of Fragmentized Android DevicesabstractThe rapid-iteration, web-style update cycle of Android helps fix revealed security vulnerabilities for its latest version. However, such security enhancements are usually only available for few Android devices released by certain manufacturers (e.g., Google's official Nexus devices). More manufactures choose to stop providing system update service for their obsolete models, remaining millions of vulnerable Android devices in use. In this situation, a feasible solution is to leverage existing source code patches to fix outdated vulnerable devices. To implement this, we introduce Embroidery, a binary rewriting based vulnerability patching system for obsolete Android devices without requiring the manufacturer's source code against Android fragmentation. Embroidery patches the known critical framework and kernel vulnerabilities in Android using both static and dynamic binary rewriting techniques. It transplants official patches (CVE source code patches) of known vulnerabilities to different devices by adopting heuristic matching strategies to deal with the code diversity introduced by Android fragmentation, and fulfills a complex dynamic memory modification to implement kernel vulnerabilities patching. We employ Embroidery to patch sophisticated Android kernel and framework vulnerabilities for various manufactures' obsolete devices ranging from Android 4.2 to 5.1. The result shows the patched devices are able to defend against known exploits and the normal functions are not affected. Xuewen Zhang, Yuanyuan Zhang 0002, Juanru Li, Yikun Hu 0003, Huayi Li, Dawu Gu |
ICSME | 1 |
| 2016 | Using superpixels to improve the efficiency of Laplacian Eigenmap based methods for target detection in hyperspectral imageryabstractLE-based methods have been shown to be effective for target detection in HSI. However, they can be slow due to the costly graph construction and eigenvector computation steps. In this paper, we proposed including a step of pre-segmenting an HSI into superpixels prior to dimensionality reduction. Carrying out experiments on an HIS from the SHARE 2012 data campaign, we show that incorporating superpixels in the BNC target detection method can yield much faster computation times without sacrificing accuracy. When incorporated in SE-based target detection, superpixels do cause a slight decrease in accuracy. Future work involves a more thorough validation on multiple datasets, and testing whether or not the inclusion of superpixels is useful for other target detection algorithms. Xuewen Zhang, Yilong Liang, Nathan D. Cahill |
IGARSS | 1 |
| 2012 | A new hierarchical classifier for hyperspctral data with similar spectrumabstractTo improve the classification accuracy of image in which many classes have the similar spectrum, this paper presents a new hierarchical classification scheme for hyperspectral images (HSI). The Spectral Angle Mapping (SAM) is firstly used to combine the similar classes into large classes. Next the hierarchical classifier classifies the image with large classes and then divides every large class into normal classes further. For every large class, the most suitable feature extraction method and classifier are chosen empirically. Meanwhile, a new band selection is proposed to help every large class find the bands which can better reflect the differences of classes according to the characteristic of spectrum. Experiments are conducted on a 103-band ROSIS image of University of Pavia. The experimental results show that the hierarchical classifier is better than the single classifier used only once. Especially when the spectra of the given classes are so similar that the traditional classifiers couldn't divide them thoroughly, the proposed classifier can make it. Moreover, the hierarchical classifier can do more efficiently because it excludes some redundant bands and concentrates on the bands with slight differences. Junping Zhang, Xuewen Zhang, Ye Zhang 0008 |
IGARSS | 2 |
| 2012 | Spectral-spatial classification of hyperspectral image based on semi-supervised and level set methodsabstractA new scheme integrating segmentation into classification to analyze hyperspectral images is presented in this paper, particularly for images with a very few number of labels and largely adjacent spatial structures. Using pixel-wise semi-supervised support vector machine, the image is classified, and segmented by modified C-V level set in this method. Afterwards, classification and segmentation images are combined with neighborhood voting. Experiments are conducted on a 200-band AVIRIS image of the Northwestern Indiana's Indian Pine site. The integration of the spatial information from the level set segmentation provides classification images with more homogeneous regions and improves the classification accuracy, comparing to the general pixel-wise supervised and semi-supervised classification. Shuang Zhou 0002, Xuewen Zhang, Junping Zhang, Hao Chen 0014 |
IGARSS | 2 |
| 2011 | Estimation of glycyrrhizic acid content using canopy spectroscopy in visible-shortwave infrared regionabstractChinese licorice, the root of Glycyrrhiza uralensis Fisch, is one of the more widely used herbal drugs in China. Glycyrrhizic acid (GA), the one of the main active ingredients in G. uralensis, is generally used as an indicator to assess the quality of G. uralensis. At present, the methods of determining GA content are laborious and time consuming and cannot be applied to remotely sensed technology. In this study, we developed methods using canopy spectral data in visible-shortwave infrared region to qualitatively classify G. uralensis samples and to quantitatively predict the GA content by a partial least square (PLS) model. The results showed that our methods provided acceptable results and implied the ability of determining GA content from remote sensing data. Xuewen Zhang, Caixiang Xie |
IGARSS | 1 |