Jiuniu Wang

dblp:224/5429 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-6113-0066ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Vision and language · 53% 3D vision · 13% Transfer learning and domain adaptation · 12%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 14 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › image captioning › fine-grained image captioning
distinctive image captioning
2.542025
Group-Based Distinctive Image Captioning with Memory Difference Encoding and Attention · Int. J. Comput. Vis. 2025
On Distinctive Image Captioning via Comparing and Reweighting · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Group-based Distinctive Image Captioning with Memory Attention · ACM Multimedia 2021
Computer vision › Vision and language
image captioning
2.542025
Group-Based Distinctive Image Captioning with Memory Difference Encoding and Attention · Int. J. Comput. Vis. 2025
On Distinctive Image Captioning via Comparing and Reweighting · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Group-based Distinctive Image Captioning with Memory Attention · ACM Multimedia 2021
Machine learning › Transfer learning and domain adaptation
zero-shot learning
1.022022
VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning · CVPR 2022
Attribute Prototype Network for Zero-Shot Learning · NeurIPS 2020
Computer vision › 3D vision
3d scene reconstruction
0.912025
Towards Scalable Spatial Intelligence Via 2D-To-3D Data Lifting · ICCV 2025
Computer vision › Vision and language › image captioning › multi-image captioning
group captioning
0.912025
Group-Based Distinctive Image Captioning with Memory Difference Encoding and Attention · Int. J. Comput. Vis. 2025
Computer vision › Vision and language › cross-modal alignment
visual-semantic embedding
0.722022
VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning · CVPR 2022
Attribute Prototype Network for Zero-Shot Learning · NeurIPS 2020
Machine learning › Generative modeling › video generation
controllable video generation
0.712023
VideoComposer: Compositional Video Synthesis with Motion Controllability · NeurIPS 2023
Machine learning › Generative modeling
video generation
0.712023
VideoComposer: Compositional Video Synthesis with Motion Controllability · NeurIPS 2023
Computer vision › Image recognition and object detection
attribute-based recognition
0.612022
Attribute Prototype Network for Any-Shot Learning · Int. J. Comput. Vis. 2022
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.612022
Attribute Prototype Network for Any-Shot Learning · Int. J. Comput. Vis. 2022
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
semantic embedding
0.612022
VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning · CVPR 2022
Machine learning › Representation and self-supervised learning › representation learning › semantic representation learning
attribute-based learning
0.412020
Attribute Prototype Network for Zero-Shot Learning · NeurIPS 2020
Machine learning › Generative modeling
diffusion model
0.212023
VideoComposer: Compositional Video Synthesis with Motion Controllability · NeurIPS 2023
Information retrieval › image retrieval
similar image search
0.112020
Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets · ECCV (1) 2020

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.9memory attention · 1.4motion vector · 1.3diffusion model · 1.3scale calibration · 0.9depth estimation · 0.9camera calibration · 0.9negative sampling · 0.7long-tailed weight strategy · 0.7CIDErBtw metric · 0.7similarity comparison · 0.4contrastive reweighting · 0.4
YearPublicationVenuePosition
2025 Towards Scalable Spatial Intelligence Via 2D-To-3D Data Lifting
abstract
Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of largescale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and laborious annotation. In this work, we present a scalable pipeline that converts single-view images into comprehensive, scale- and appearance-realistic 3D representations - including point clouds, camera poses, depth maps, and pseudo-RGBD - via integrated depth estimation, camera calibration, and scale calibration. Our method bridges the gap between the vast repository of imagery and the increasing demand for spatial scene understanding. By automatically generating authentic, scale-aware 3D data from images, we significantly reduce data collection costs and open new avenues for advancing spatial intelligence. We release two generated spatial datasets, i.e., COCO-3D and Objects365-v2-3D, and demonstrate through extensive experiments that our generated data can benefit various 3D tasks, ranging from fundamental perception to MLLMbased reasoning. These results validate our pipeline as an effective solution for developing AI systems capable of perceiving, understanding, and interacting with physical environments.
Xingyu Miao, Haoran Duan 0001, Quanhao Qian, Jiuniu Wang, Yang Long 0001, Ling Shao 0001, Deli Zhao, Gongjie Zhang
ICCV4
2025 Group-Based Distinctive Image Captioning with Memory Difference Encoding and Attention
abstract
Abstract Recent advances in image captioning have focused on enhancing accuracy by substantially increasing the dataset and model size. While conventional captioning models exhibit high performance on established metrics such as BLEU, CIDEr, and SPICE, the capability of captions to distinguish the target image from other similar images is under-explored. To generate distinctive captions, a few pioneers employed contrastive learning or re-weighted the ground-truth captions. However, these approaches often overlook the relationships among objects in a similar image group (e.g., items or properties within the same album or fine-grained events). In this paper, we introduce a novel approach to enhance the distinctiveness of image captions, namely Group-based Differential Distinctive Captioning Method, which visually compares each image with other images in one similar group and highlights the uniqueness of each image. In particular, we introduce a Group-based Differential Memory Attention (GDMA) module, designed to identify and emphasize object features in an image that are uniquely distinguishable within its image group, i.e., those exhibiting low similarity with objects in other images. This mechanism ensures that such unique object features are prioritized during caption generation for the image, thereby enhancing the distinctiveness of the resulting captions. To further refine this process, we select distinctive words from the ground-truth captions to guide both the language decoder and the GDMA module. Additionally, we propose a new evaluation metric, the Distinctive Word Rate (DisWordRate), to quantitatively assess caption distinctiveness. Quantitative results indicate that the proposed method significantly improves the distinctiveness of several baseline models, and achieves state-of-the-art performance on distinctiveness while not excessively sacrificing accuracy. Moreover, the results of our user study are consistent with the quantitative evaluation and demonstrate the rationality of the new metric DisWordRate.
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. Chan
Int. J. Comput. Vis.1
2024 Generalized Category Discovery for Remote Sensing Image Scene Classification
abstract
Deep neural networks have achieved promising progress in remote sensing (RS) image classification. However, the training process requires abundant samples for each class, and it is unrealistic to annotate labels for each RS category, especially considering that the RS target database is increasing dynamically. Therefore, we introduce an innovative prototype network tailored for Generalized Category Discovery (GCD) in remote sensing scene classification. This network consists of two essential modules: one dedicated to representation learning and the other to prototype learning. Through extensive experiments conducted on three benchmark datasets, i.e., RSS-DIVCS, NWPU-RESISC45, and AID, we demonstrate that the proposed model achieves remarkable performance gain up to 20%, effectively addressing the challenges inherent in classifying dynamically varying remote sensing images.
Wenjia Xu, Zijian Yu, Zhiwei Wei, Jiuniu Wang, Mugen Peng
IGARSS4
2023 VideoComposer: Compositional Video Synthesis with Motion Controllability
abstract
The pursuit of controllability as a higher standard of visual content creation has yielded remarkable progress in customizable image synthesis. However, achieving controllable video synthesis remains challenging due to the large variation of temporal dynamics and the requirement of cross-frame temporal consistency. Based on the paradigm of compositional generation, this work presents VideoComposer that allows users to flexibly compose a video with textual conditions, spatial conditions, and more importantly temporal conditions. Specifically, considering the characteristic of video data, we introduce the motion vector from compressed videos as an explicit control signal to provide guidance regarding temporal dynamics. In addition, we develop a Spatio-Temporal Condition encoder (STC-encoder) that serves as a unified interface to effectively incorporate the spatial and temporal relations of sequential inputs, with which the model could make better use of temporal conditions and hence achieve higher inter-frame consistency. Extensive experimental results suggest that VideoComposer is able to control the spatial and temporal patterns simultaneously within a synthesized video in various forms, such as text description, sketch sequence, reference video, or even simply hand-crafted motions. The code and models are publicly available at https://videocomposer.github.io.
Xiang Wang 0012, Hangjie Yuan, Shiwei Zhang 0001, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, Jingren Zhou 0001
NeurIPS5
2023 On Distinctive Image Captioning via Comparing and Reweighting
abstract
Recent image captioning models are achieving impressive results based on popular metrics, i.e., BLEU, CIDEr, and SPICE. However, focusing on the most popular metrics that only consider the overlap between the generated captions and human annotation could result in using common words and phrases, which lacks distinctiveness, i.e., many similar images have the same caption. In this paper, we aim to improve the distinctiveness of image captions via comparing and reweighting with a set of similar images. First, we propose a distinctiveness metric-between-set CIDEr (CIDErBtw) to evaluate the distinctiveness of a caption with respect to those of similar images. Our metric reveals that the human annotations of each image in the MSCOCO dataset are not equivalent based on distinctiveness; however, previous works normally treat the human annotations equally during training, which could be a reason for generating less distinctive captions. In contrast, we reweight each ground-truth caption according to its distinctiveness during training. We further integrate a long-tailed weight strategy to highlight the rare words that contain more information, and captions from the similar image set are sampled as negative examples to encourage the generated sentence to be unique. Finally, extensive experiments are conducted, showing that our proposed approach significantly improves both distinctiveness (as measured by CIDErBtw and retrieval metrics) and accuracy (e.g., as measured by CIDEr) for a wide variety of image captioning baselines. These results are further confirmed through a user study.
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. Chan
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot Learning
abstract
Human-annotated attributes serve as powerful semantic embeddings in zero-shot learning. However, their annotation process is labor-intensive and needs expert supervision. Current unsupervised semantic embeddings, i.e., word embeddings, enable knowledge transfer between classes. However, word embeddings do not always reflect visual similarities and result in inferior zero-shot performance. We propose to discover semantic embeddings containing discriminative visual properties for zero-shot learning, without requiring any human annotation. Our model visually divides a set of images from seen classes into clusters of local image regions according to their visual similarity, and further imposes their class discrimination and semantic relatedness. To associate these clusters with previously unseen classes, we use external knowledge, e.g., word embeddings and propose a novel class relation discovery module. Through quantitative and qualitative evaluation, we demonstrate that our model discovers semantic embeddings that model the visual properties of both seen and unseen classes. Furthermore, we demonstrate on three benchmarks that our visually-grounded semantic embeddings further improve performance over word embeddings across various ZSL models by a large margin. Code is available at https://github.com/wenjiaXu/VGSE
Wenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele, Zeynep Akata
CVPR3
2022 Multi-Dimension Geospatial Feature Learning for Urban Region Function Recognition
abstract
Urban region function recognition plays a vital character in monitoring and managing the limited urban areas. Since urban functions are complex and full of social-economic properties, simply using remote sensing (RS) images equipped with physical and optical information cannot completely solve the classification task. On the other hand, with the development of mobile communication and the internet, the acquisition of geospatial big data (GBD) becomes possible. In this paper, we propose a Multi-dimension Feature Learning Model (MDFL) using high-dimensional GBD data in conjunction with RS images for urban region function recognition. When extracting multi-dimension features, our model considers the user-related information modeled by their activity, as well as the region-based information abstracted from the region graph. Furthermore, we propose a decision fusion network that integrates the decisions from several neural networks and machine learning classifiers, and the final decision is made considering both the visual cue from the RS images and the social information from the GBD data. Through quantitative evaluation, we demonstrate that our model achieves overall accuracy at 92.75%, outperforming the state-of-the-art by 10% percent.
Wenjia Xu, Jiuniu Wang, Yirong Wu
IGARSS2
2022 Attribute Prototype Network for Any-Shot Learning
Wenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele, Zeynep Akata
Int. J. Comput. Vis.3
2021 Group-based Distinctive Image Captioning with Memory Attention
abstract
Describing images using natural language is widely known as image captioning, which has made consistent progress due to the development of computer vision and natural language generation techniques. Though conventional captioning models achieve high accuracy based on popular metrics, i.e., BLEU, CIDEr, and SPICE, the ability of captions to distinguish the target image from other similar images is under-explored. To generate distinctive captions, a few pioneers employ contrastive learning or re-weighted the ground-truth captions, which focuses on one single input image. However, the relationships between objects in a similar image group (e.g., items or properties within the same album or fine-grained events) are neglected. In this paper, we improve the distinctiveness of image captions using a Group-based Distinctive Captioning Model (GdisCap), which compares each image with other images in one similar group and highlights the uniqueness of each image. In particular, we propose a group-based memory attention (GMA) module, which stores object features that are unique among the image group (i.e., with low similarity to objects in other images). These unique object features are highlighted when generating captions, resulting in more distinctive captions. Furthermore, the distinctive words in the ground-truth captions are selected to supervise the language decoder and GMA. Finally, we propose a new evaluation metric, distinctive word rate (DisWordRate) to measure the distinctiveness of captions. Quantitative results indicate that the proposed method significantly improves the distinctiveness of several baseline models, and achieves the state-of-the-art performance on both accuracy and distinctiveness. Results of a user study agree with the quantitative evaluation and demonstrate the rationality of the new metric DisWordRate.
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. Chan
ACM Multimedia1
2020 Neighbours Matter: Image Captioning with Similar Images
Qingzhong Wang, Jiuniu Wang, Antoni B. Chan, Siyu Huang, Haoyi Xiong, Xingjian Li 0002, Dejing Dou
BMVC2
2020 Compare and Reweight: Distinctive Image Captioning Using Similar Images Sets
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. Chan
ECCV (1)1
2020 Attribute Prototype Network for Zero-Shot Learning
abstract
From the beginning of zero-shot learning research, visual attributes have been shown to play an important role. In order to better transfer attribute-based knowledge from known to unknown classes, we argue that an image representation with integrated attribute localization ability would be beneficial for zero-shot learning. To this end, we propose a novel zero-shot representation learning framework that jointly learns discriminative global and local features using only class-level attributes. While a visual-semantic embedding layer learns global features, local features are learned through an attribute prototype network that simultaneously regresses and decorrelates attributes from intermediate features. We show that our locality augmented image representations achieve a new state-of-the-art on three zero-shot learning benchmarks. As an additional benefit, our model points to the visual evidence of the attributes in an image, e.g. for the CUB dataset, confirming the improved attribute localization ability of our image representation.
Wenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele, Zeynep Akata
NeurIPS3
2020 SRQA: Synthetic Reader for Factoid Question Answering
Jiuniu Wang, Wenjia Xu, Li Jin 0001, Guangluan Xu, Yirong Wu
Knowl. Based Syst.1
2020 ASTRAL: Adversarial Trained LSTM-CNN for Named Entity Recognition
Jiuniu Wang, Wenjia Xu, Guangluan Xu, Yirong Wu
Knowl. Based Syst.1
2018 A3Net: Adversarial-and-Attention Network for Machine Reading Comprehension
Jiuniu Wang, Guangluan Xu, Yirong Wu, Li Jin 0001
NLPCC (1)1