Xinze Wang

dblp:204/0097 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0002-8233-4260ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 52% Vision and language · 24% Representation and self-supervised learning · 12%
Databases, data mining, and information retrieval
3 papers
Recommender systems · 74% Information retrieval · 18% Web and social media mining · 8%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.912025
CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
conditional generation
0.912025
CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching · NeurIPS 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
Contrastive Localized Language-Image Pre-Training · ICML 2025
Machine learning › Generative modeling
diffusion model
0.912025
CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching · NeurIPS 2025
Machine learning › Generative modeling
flow matching
0.912025
CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Contrastive Localized Language-Image Pre-Training · ICML 2025
Computer vision › Image recognition and object detection › object recognition
region-based recognition
0.912025
Contrastive Localized Language-Image Pre-Training · ICML 2025
Computer vision › Vision and language
vision-language pretraining
0.912025
Contrastive Localized Language-Image Pre-Training · ICML 2025
Recommender systems
point-of-interest recommendation
0.722019
Hierarchical Multi-Clue Modelling for POI Popularity Prediction with Heterogeneous Tourist Information · IEEE Trans. Knowl. Data Eng. 2019
POI Popularity Prediction via Hierarchical Fusion of Multiple Social Clues · SIGIR 2017
Recommender systems › multimodal recommendation
multimodal fusion
0.412019
Hierarchical Multi-Clue Modelling for POI Popularity Prediction with Heterogeneous Tourist Information · IEEE Trans. Knowl. Data Eng. 2019
Information retrieval
cross-modal retrieval
0.312025
Contrastive Localized Language-Image Pre-Training · ICML 2025
Web and social media mining
user-generated content
0.112019
Hierarchical Multi-Clue Modelling for POI Popularity Prediction with Heterogeneous Tourist Information · IEEE Trans. Knowl. Data Eng. 2019

Methods — techniques the papers use, named apart from their topics

region-text contrastive loss · 1.7promptable embeddings · 1.7captioning · 1.7reparameterization · 0.9flow matching · 0.9diffusion model · 0.9semantic knowledge injection · 0.7hierarchical multi-clue fusion · 0.4multimodal fusion · 0.3
YearPublicationVenuePosition
2025 Contrastive Localized Language-Image Pre-Training
abstract
CLIP has been a celebrated method for training vision encoders to generate image/text representations facilitating various applications. Recently, it has been widely adopted as the vision backbone of multimodal large language models (MLLMs). The success of CLIP relies on aligning web-crawled noisy text annotations at image levels. However, such criteria may be insufficient for downstream tasks in need of fine-grained vision representations, especially when understanding region-level is demanding for MLLMs. We improve the localization capability of CLIP with several advances. Our proposed pre-training method, Contrastive Localized Language-Image Pre-training (CLOC), complements CLIP with region-text contrastive loss and modules. We formulate a new concept, promptable embeddings, of which the encoder produces image embeddings easy to transform into region representations given spatial hints. To support large-scale pre-training, we design a visually-enriched and spatially-localized captioning framework to effectively generate region-text labels. By scaling up to billions of annotated images, CLOC enables high-quality regional embeddings for recognition and retrieval tasks, and can be a drop-in replacement of CLIP to enhance MLLMs, especially on referring and grounding tasks.
Hong-You Chen, Zhengfeng Lai, Haotian Zhang 0005, Xinze Wang, Marcin Eichner, Keen You, Bowen Zhang 0002, Yinfei Yang, Zhe Gan
ICML4
2025 CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching
abstract
Conditional generative modeling aims to learn a conditional data distribution from samples containing data-condition pairs. For this, diffusion and flow-based methods have attained compelling results. These methods use a learned (flow) model to transport an initial standard Gaussian noise that ignores the condition to the conditional data distribution. The model is hence required to learn both mass transport \emph{and} conditional injection. To ease the demand on the model, we propose \emph{Condition-Aware Reparameterization for Flow Matching} (CAR-Flow) -- a lightweight, learned \emph{shift} that conditions the source, the target, or both distributions. By relocating these distributions, CAR-Flow shortens the probability path the model must learn, leading to faster training in practice. On low-dimensional synthetic data, we visualize and quantify the effects of CAR-Flow. On higher-dimensional natural image data (ImageNet-256), equipping SiT-XL/2 with CAR-Flow reduces FID from 2.07 to 1.68, while introducing less than \(0.6\%\) additional parameters.
Chen Chen 0005, Pengsheng Guo, Liangchen Song, Jiasen Lu, Rui Qian 0003, Tsu-Jui Fu, Xinze Wang, Yinfei Yang, Alex Schwing 0002
NeurIPS7
2019 Hierarchical Multi-Clue Modelling for POI Popularity Prediction with Heterogeneous Tourist Information
abstract
Predicting the popularity of Point of Interest (POI) has become increasingly crucial for location-based services, such as POI recommendation. Most of the existing methods can seldom achieve satisfactory performance due to the scarcity of POI's information, which tendentiously confines the recommendation to popular scene spots, and ignores the unpopular attractions with potentially precious values. In this paper, we propose a novel approach, termed Hierarchical Multi-Clue Fusion (HMCF), for predicting the popularity of POIs. Specifically, in order to cope with the problem of data sparsity, we propose to comprehensively describe POI using various types of user generated content (UGC) (e.g., text and image) from multiple sources. Then, we devise an effective POI modelling method in a hierarchical manner, which simultaneously injects semantic knowledge as well as multi-clue representative power into POIs. For evaluation, we construct a multi-source POI dataset by collecting all the textual and visual content of several specific provinces in China from four main-stream tourism platforms during 2006 to 2017. Extensive experimental results show that the proposed method can significantly improve the performance of predicting the attractions' popularity as compared to several baseline methods.
Yang Yang 0002, Yaqian Duan, Xinze Wang, Zi Huang, Ning Xie 0003, Heng Tao Shen
IEEE Trans. Knowl. Data Eng.3
2017 POI Popularity Prediction via Hierarchical Fusion of Multiple Social Clues
abstract
Predicting the popularity of Point of Interest (POI) has become increasingly crucial for location-based services, such as POI recommendation. Most of the existing methods can seldom achieve satisfactory performance due to the scarcity of POI's information, which tendentiously confines the recommendation to popular scenic spots, and ignores the unpopular attractions with potentially precious values. In this paper, we propose a novel approach, termed Hierarchical Multi-Clue Fusion (HMCF), for predicting the popularity of POIs. Specifically, we devise an effective hierarchy to comprehensively describe POI by integrating various types of media information (e.g., image and text) from multiple social sources. For each individual POI, we simultaneously inject semantic knowledge as well as multi-clue representative power. We collect a multi-source POI dataset from four widely-used tourism platforms. Extensive experimental results show that the proposed method can significantly improve the performance of predicting the attractions' popularity as compared to several baselines.
Yaqian Duan, Xinze Wang, Yang Yang 0002, Zi Huang, Ning Xie 0003, Heng Tao Shen
SIGIR2