Wenjia Xu

dblp:34/9467 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
14since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 From fuzzy intent to executable visual workflows: A multi-model orchestration approach
Yuhuan Huang, Dongdong Lu, Fei Li 0029, Zhe Zhang 0026, Wenjia Xu
Expert Syst. Appl.6
2026 A Monolithic GaN Active Gate Driver Achieving 71.6% Ringing Reduction Across Various Load Currents Using A Ringing Sensor and Simplified Digital Algorithm
Wenjia Xu, Huajun Zhang 0001, Hesheng Lin, G. Q. Zhang, Qinwen Fan
ISCAS1
2026 Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency
abstract
Abstract While Large Vision Language Models (LVLMs) demonstrate impressive capabilities, their substantial computational and memory requirements pose deployment challenges on resource-constrained edge devices. Current parameter reduction techniques primarily involve training LVLMs from small language models, but these methods offer limited flexibility and remain computationally intensive. We study a complementary route: compressing existing LVLMs by applying structured pruning to the language model backbone, followed by lightweight recovery training. Specifically, we investigate two structural pruning paradigms: layerwise and widthwise pruning, and pair them with supervised finetuning and knowledge distillation on logits and hidden states. Additionally, we assess the feasibility of conducting recovery training with only a small fraction of the available data. Our results show that widthwise pruning generally maintains better performance in low-resource scenarios, where computational resources are limited or there is insufficient finetuning data. As for the recovery training, finetuning only the multimodal projector is sufficient at small compression levels. Furthermore, a combination of supervised finetuning and hidden-state distillation yields optimal recovery across various pruning levels. Notably, effective recovery can be achieved using just 5% of the original data, while retaining over 95% of the original performance. Through empirical study on three representative LVLM families ranging from 3B to 7B parameters, this study offers actionable insights for practitioners to compress LVLMs without extensive computation resources or sufficient data.
Lukas Thede, Massimiliano Mancini, Wenjia Xu, Zeynep Akata
Int. J. Comput. Vis.4
2025 Group-Based Distinctive Image Captioning with Memory Difference Encoding and Attention
abstract
Abstract Recent advances in image captioning have focused on enhancing accuracy by substantially increasing the dataset and model size. While conventional captioning models exhibit high performance on established metrics such as BLEU, CIDEr, and SPICE, the capability of captions to distinguish the target image from other similar images is under-explored. To generate distinctive captions, a few pioneers employed contrastive learning or re-weighted the ground-truth captions. However, these approaches often overlook the relationships among objects in a similar image group (e.g., items or properties within the same album or fine-grained events). In this paper, we introduce a novel approach to enhance the distinctiveness of image captions, namely Group-based Differential Distinctive Captioning Method, which visually compares each image with other images in one similar group and highlights the uniqueness of each image. In particular, we introduce a Group-based Differential Memory Attention (GDMA) module, designed to identify and emphasize object features in an image that are uniquely distinguishable within its image group, i.e., those exhibiting low similarity with objects in other images. This mechanism ensures that such unique object features are prioritized during caption generation for the image, thereby enhancing the distinctiveness of the resulting captions. To further refine this process, we select distinctive words from the ground-truth captions to guide both the language decoder and the GDMA module. Additionally, we propose a new evaluation metric, the Distinctive Word Rate (DisWordRate), to quantitatively assess caption distinctiveness. Quantitative results indicate that the proposed method significantly improves the distinctiveness of several baseline models, and achieves state-of-the-art performance on distinctiveness while not excessively sacrificing accuracy. Moreover, the results of our user study are consistent with the quantitative evaluation and demonstrate the rationality of the new metric DisWordRate.
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. Chan
Int. J. Comput. Vis.2
2025 TS-SatMVSNet: Slope Aware Height Estimation for Large-Scale Earth Terrain Multiview Stereo
abstract
3D terrain reconstruction with satellite imagery achieves cost-effective and large-scale earth observation and is crucial for safeguarding natural disasters, monitoring ecological changes, and preserving the environment. Recently, learning-based multi-view stereo (MVS) methods have shown promise in this task. However, these methods simply modify the general learning-based MVS framework for height estimation, which overlooks the terrain characteristics and results in insufficient accuracy. Considering that the Earth’s surface generally undulates without drastic changes and can be measured by slope, integrating slope considerations into MVS frameworks could enhance the accuracy of terrain reconstruction. To this end, we propose an end-to-end slope-aware height estimation network named TS-SatMVSNet for large-scale remote sensing terrain reconstruction. To effectively obtain the slope representation, drawing from mathematical gradient concepts, we innovatively proposed a height-based slope calculation strategy to first calculate a slope map from a height map to measure the terrain undulation. To fully integrate slope information into the MVS pipeline, we separately design two slope-guided modules to enhance reconstruction outcomes. Specifically, we designed a slope-guided interval partition module for refined height estimation using slope values. And, a height correction module is proposed, using a learnable Gaussian smoothing operator to amend the inaccurate height values. Additionally, to enhance the efficacy of height estimation, we proposed a slope direction loss for implicitly optimizing height estimation results. Extensive experiments on the WHU-TLC dataset and MVS3D dataset show that our proposed method achieves state-of-the-art performance and demonstrates competitive generalization ability compared to all listed methods. Our code will be available at https://github.com/StriveZs/TS-SatMVSNet.
Zhiwei Wei, Wenjia Xu
IEEE Trans. Geosci. Remote. Sens.3
2024 Generalized Category Discovery for Remote Sensing Image Scene Classification
abstract
Deep neural networks have achieved promising progress in remote sensing (RS) image classification. However, the training process requires abundant samples for each class, and it is unrealistic to annotate labels for each RS category, especially considering that the RS target database is increasing dynamically. Therefore, we introduce an innovative prototype network tailored for Generalized Category Discovery (GCD) in remote sensing scene classification. This network consists of two essential modules: one dedicated to representation learning and the other to prototype learning. Through extensive experiments conducted on three benchmark datasets, i.e., RSS-DIVCS, NWPU-RESISC45, and AID, we demonstrate that the proposed model achieves remarkable performance gain up to 20%, effectively addressing the challenges inherent in classifying dynamically varying remote sensing images.
Wenjia Xu, Zijian Yu, Zhiwei Wei, Jiuniu Wang, Mugen Peng
IGARSS1
2023 On Distinctive Image Captioning via Comparing and Reweighting
abstract
Recent image captioning models are achieving impressive results based on popular metrics, i.e., BLEU, CIDEr, and SPICE. However, focusing on the most popular metrics that only consider the overlap between the generated captions and human annotation could result in using common words and phrases, which lacks distinctiveness, i.e., many similar images have the same caption. In this paper, we aim to improve the distinctiveness of image captions via comparing and reweighting with a set of similar images. First, we propose a distinctiveness metric-between-set CIDEr (CIDErBtw) to evaluate the distinctiveness of a caption with respect to those of similar images. Our metric reveals that the human annotations of each image in the MSCOCO dataset are not equivalent based on distinctiveness; however, previous works normally treat the human annotations equally during training, which could be a reason for generating less distinctive captions. In contrast, we reweight each ground-truth caption according to its distinctiveness during training. We further integrate a long-tailed weight strategy to highlight the rare words that contain more information, and captions from the similar image set are sampled as negative examples to encourage the generated sentence to be unique. Finally, extensive experiments are conducted, showing that our proposed approach significantly improves both distinctiveness (as measured by CIDErBtw and retrieval metrics) and accuracy (e.g., as measured by CIDEr) for a wide variety of image captioning baselines. These results are further confirmed through a user study.
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. Chan
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 ARAI-MVSNet: A multi-view stereo depth estimation network with adaptive depth range and depth interval
Wenjia Xu, Zhiwei Wei
Pattern Recognit.2
2022 VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning
abstract
Human-annotated attributes serve as powerful semantic embeddings in zero-shot learning. However, their annotation process is labor-intensive and needs expert supervision. Current unsupervised semantic embeddings, i.e., word embeddings, enable knowledge transfer between classes. However, word embeddings do not always reflect visual similarities and result in inferior zero-shot performance. We propose to discover semantic embeddings containing discriminative visual properties for zero-shot learning, without requiring any human annotation. Our model visually divides a set of images from seen classes into clusters of local image regions according to their visual similarity, and further imposes their class discrimination and semantic relatedness. To associate these clusters with previously unseen classes, we use external knowledge, e.g., word embeddings and propose a novel class relation discovery module. Through quantitative and qualitative evaluation, we demonstrate that our model discovers semantic embeddings that model the visual properties of both seen and unseen classes. Furthermore, we demonstrate on three benchmarks that our visually-grounded semantic embeddings further improve performance over word embeddings across various ZSL models by a large margin. Code is available at https://github.com/wenjiaXu/VGSE
Wenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele, Zeynep Akata
CVPR1
2022 Multi-Dimension Geospatial Feature Learning for Urban Region Function Recognition
abstract
Urban region function recognition plays a vital character in monitoring and managing the limited urban areas. Since urban functions are complex and full of social-economic properties, simply using remote sensing (RS) images equipped with physical and optical information cannot completely solve the classification task. On the other hand, with the development of mobile communication and the internet, the acquisition of geospatial big data (GBD) becomes possible. In this paper, we propose a Multi-dimension Feature Learning Model (MDFL) using high-dimensional GBD data in conjunction with RS images for urban region function recognition. When extracting multi-dimension features, our model considers the user-related information modeled by their activity, as well as the region-based information abstracted from the region graph. Furthermore, we propose a decision fusion network that integrates the decisions from several neural networks and machine learning classifiers, and the final decision is made considering both the visual cue from the RS images and the social information from the GBD data. Through quantitative evaluation, we demonstrate that our model achieves overall accuracy at 92.75%, outperforming the state-of-the-art by 10% percent.
Wenjia Xu, Jiuniu Wang, Yirong Wu
IGARSS1
2022 Learning Prototype via Placeholder for Zero-shot Recognition
abstract
Zero-shot learning (ZSL) aims to recognize unseen classes by exploiting semantic descriptions shared between seen classes and unseen classes. Current methods show that it is effective to learn visual-semantic alignment by projecting semantic embeddings into the visual space as class prototypes. However, such a projection function is only concerned with seen classes. When applied to unseen classes, the prototypes often perform suboptimally due to domain shift. In this paper, we propose to learn prototypes via placeholders, termed LPL, to eliminate the domain shift between seen and unseen classes. Specifically, we combine seen classes to hallucinate new classes which play as placeholders of the unseen classes in the visual and semantic space. Placed between seen classes, the placeholders encourage prototypes of seen classes to be highly dispersed. And more space is spared for the insertion of well-separated unseen ones. Empirically, well-separated prototypes help counteract visual-semantic misalignment caused by domain shift. Furthermore, we exploit a novel semantic-oriented fine-tuning method to guarantee the semantic reliability of placeholders. Extensive experiments on five benchmark datasets demonstrate the significant performance gain of LPL over the state-of-the-art methods.
Zaiquan Yang, Yang Liu 0357, Wenjia Xu, Lei Zhou 0008
IJCAI3
2022 Attribute Prototype Network for Any-Shot Learning
Wenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele, Zeynep Akata
Int. J. Comput. Vis.1
2021 Human Attention in Fine-grained Classification
Yao Rong 0001, Wenjia Xu, Zeynep Akata, Enkelejda Kasneci
BMVC2
2021 Group-based Distinctive Image Captioning with Memory Attention
abstract
Describing images using natural language is widely known as image captioning, which has made consistent progress due to the development of computer vision and natural language generation techniques. Though conventional captioning models achieve high accuracy based on popular metrics, i.e., BLEU, CIDEr, and SPICE, the ability of captions to distinguish the target image from other similar images is under-explored. To generate distinctive captions, a few pioneers employ contrastive learning or re-weighted the ground-truth captions, which focuses on one single input image. However, the relationships between objects in a similar image group (e.g., items or properties within the same album or fine-grained events) are neglected. In this paper, we improve the distinctiveness of image captions using a Group-based Distinctive Captioning Model (GdisCap), which compares each image with other images in one similar group and highlights the uniqueness of each image. In particular, we propose a group-based memory attention (GMA) module, which stores object features that are unique among the image group (i.e., with low similarity to objects in other images). These unique object features are highlighted when generating captions, resulting in more distinctive captions. Furthermore, the distinctive words in the ground-truth captions are selected to supervise the language decoder and GMA. Finally, we propose a new evaluation metric, distinctive word rate (DisWordRate) to measure the distinctiveness of captions. Quantitative results indicate that the proposed method significantly improves the distinctiveness of several baseline models, and achieves the state-of-the-art performance on both accuracy and distinctiveness. Results of a user study agree with the quantitative evaluation and demonstrate the rationality of the new metric DisWordRate.
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. Chan
ACM Multimedia2
2020 Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. Chan
ECCV (1)2
2020 Attribute Prototype Network for Zero-Shot Learning
abstract
From the beginning of zero-shot learning research, visual attributes have been shown to play an important role. In order to better transfer attribute-based knowledge from known to unknown classes, we argue that an image representation with integrated attribute localization ability would be beneficial for zero-shot learning. To this end, we propose a novel zero-shot representation learning framework that jointly learns discriminative global and local features using only class-level attributes. While a visual-semantic embedding layer learns global features, local features are learned through an attribute prototype network that simultaneously regresses and decorrelates attributes from intermediate features. We show that our locality augmented image representations achieve a new state-of-the-art on three zero-shot learning benchmarks. As an additional benefit, our model points to the visual evidence of the attributes in an image, e.g. for the CUB dataset, confirming the improved attribute localization ability of our image representation.
Wenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele, Zeynep Akata
NeurIPS1
2020 SCRSR: An efficient recursive convolutional neural network for fast and accurate image super-resolution
Daoyu Lin, Guangluan Xu, Wenjia Xu, Yang Wang 0056, Xian Sun 0001, Kun Fu 0001
Neurocomputing3
2020 SRQA: Synthetic Reader for Factoid Question Answering
Jiuniu Wang, Wenjia Xu, Li Jin 0001, Guangluan Xu, Yirong Wu
Knowl. Based Syst.2
2020 ASTRAL: Adversarial Trained LSTM-CNN for Named Entity Recognition
Jiuniu Wang, Wenjia Xu, Guangluan Xu, Yirong Wu
Knowl. Based Syst.2
2018 High Quality Remote Sensing Image Super-Resolution Using Deep Memory Connected Network
abstract
Single image super-resolution is an effective way to enhance the spatial resolution of remote sensing image, which is crucial for many applications such as target detection and image classification. However, existing methods based on the neural network usually have small receptive fields and ignore the image detail. We propose a novel method named deep memory connected network (DMCN) based on a convolutional neural network to reconstruct high-quality super-resolution images. We build local and global memory connections to combine image detail with environmental information. To further reduce parameters and ease time-consuming, we propose downsampling units, shrinking the spatial size of feature maps. We test DMCN on three remote sensing datasets with different spatial resolution. Experimental results indicate that our method yields promising improvements in both accuracy and visual performance over the current state-of-the-art.
Wenjia Xu, Guangluan Xu, Yang Wang 0056, Xian Sun 0001, Daoyu Lin, Yirong Wu
IGARSS1