VLDB 2026 Research / reviewers in the wild / expert
Xinsong Zhang
dblp:04/2640
· DBLP profile ↗
12ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Vision and language · 52% Information extraction and text analysis · 20% Representation and self-supervised learning · 7% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-language pretraining |
2.6 | 4 | 2024 | X$^{2}$2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2024 Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training · ACL (1) 2023 VLUE: A Multi-Task Multi-Dimension Benchmark for Evaluating Vision-Language Pre-training · ICML 2022 |
Natural language and speech › Information extraction and text analysis
relation extraction |
1.6 | 4 | 2021 | Robust Neural Relation Extraction via Multi-Granularity Noises Reduction · IEEE Trans. Knowl. Data Eng. 2021 Fine-grained relation extraction with focal multi-task learning · Sci. China Inf. Sci. 2020 Multi-Labeled Relation Extraction with Attentive Capsule Network · AAAI 2019 |
Computer vision › Vision and language › cross-modal alignment
multi-grained alignment |
1.3 | 2 | 2024 | X$^{2}$2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2024 Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts · ICML 2022 |
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision |
0.8 | 2 | 2021 | Robust Neural Relation Extraction via Multi-Granularity Noises Reduction · IEEE Trans. Knowl. Data Eng. 2021 Neural Relation Extraction via Inner-Sentence Noise Reduction and Transfer Learning · EMNLP 2018 |
Computer vision › Vision and language
image captioning |
0.8 | 1 | 2024 | X$^{2}$2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › Vision and language › video-language understanding
video-text tasks |
0.8 | 1 | 2024 | X$^{2}$2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › 3D vision › visual localization
cross-view matching |
0.7 | 1 | 2023 | Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training · ACL (1) 2023 |
Machine learning › Generative modeling
multimodal generation |
0.7 | 1 | 2023 | Write and Paint: Generative Vision-Language Models are Unified Modal Learners · ICLR 2023 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.7 | 1 | 2023 | Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training · ACL (1) 2023 |
Computer vision › Vision and language › vision-language model › vision-language model architecture
unified vision-language model |
0.7 | 1 | 2023 | Write and Paint: Generative Vision-Language Models are Unified Modal Learners · ICLR 2023 |
Computer vision › Vision and language › cross-modal alignment › visual-semantic alignment
visual concept alignment |
0.6 | 1 | 2022 | Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts · ICML 2022 |
Natural language and speech › Speech recognition and synthesis › speech enhancement
noise reduction |
0.5 | 1 | 2021 | Robust Neural Relation Extraction via Multi-Granularity Noises Reduction · IEEE Trans. Knowl. Data Eng. 2021 |
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining |
0.2 | 1 | 2024 | X$^{2}$2-VLM: All-in-One Pre-Trained Model for Vision-Language Tasks · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Natural language and speech › Information extraction and text analysis › relation extraction
neural relation extraction |
0.1 | 1 | 2021 | Robust Neural Relation Extraction via Multi-Granularity Noises Reduction · IEEE Trans. Knowl. Data Eng. 2021 |
Machine learning › Learning paradigms
multi-task learning |
0.1 | 1 | 2020 | Fine-grained relation extraction with focal multi-task learning · Sci. China Inf. Sci. 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction |
0.1 | 1 | 2018 | Neural Relation Extraction via Inner-Sentence Noise Reduction and Transfer Learning · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 1.4transfer learning · 0.8object detection · 0.8masked language modeling · 0.7generative vision-language modeling · 0.7visual concept localization · 0.6out-of-distribution evaluation · 0.6multi-task benchmarking · 0.6multi-grained alignment · 0.6multi-instance learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prescribed Performance-Based Distributed Adaptive Predefined-Time Consensus Control for BESSsabstractThis work investigates distributed state-of-charge (SOC) balancing for battery energy storage systems (BESSs). Existing predefined-time consensus designs, despite their explicit convergence-time guarantees, may exhibit singularity, induce transient overcurrents under large initial SOC mismatches, and rely on continuous communication. To overcome these limitations, a distributed adaptive predefined-time robust control framework with prescribed performance constraints is developed.Within this framework, a dynamic error envelope is constructed to shape the convergence profile under prescribed performance bounds. In this manner, the convergence time is prescribed a priori, whereas initial transients and the associated current peaks are limited to improve the safety margins of power electronic devices. To enforce the prescribed error evolution under disturbances, an event-triggered (ET)-based distributed predefined-time sliding mode controller is designed, in which a predefined-time adaptive law is incorporated to track mismatched disturbances, thereby mitigating high-frequency chattering while ensuring nonsingular dynamics. In addition, an ET-based model-free state observer with input-to-state stability is developed to support distributed implementation without continuous communication. Predefined-time stability is rigorously established via Lyapunov analysis. Multi-scenario simulations validate the proposed method and confirm reduced communication overhead, prescribed convergence time, and improved transient safety. Chunkai Yan, Xingjian Sun, Xinsong Zhang, Juping Gu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | X$^{2}$2-VLM: All-in-One Pre-Trained Model for Vision-Language TasksabstractVision language pre-training aims to learn alignments between vision and language from a large amount of data. Most existing methods only learn image-text alignments. Some others utilize pre-trained object detectors to leverage vision language alignments at the object level. In this paper, we propose to learn multi-grained vision language alignments by a unified pre-training framework that learns multi-grained aligning and multi-grained localization simultaneously. Based on it, we present X2-VLM, an all-in-one model with a flexible modular architecture, in which we further unify image-text pre-training and video-text pre-training in one model. X2-VLM is able to learn unlimited visual concepts associated with diverse text descriptions. Experiment results show that X2-VLM performs the best on base and large scale for both image-text and video-text tasks, making a good trade-off between performance and model scale. Moreover, we show that the modular design of X2-VLM results in high transferability for it to be utilized in any language or domain. For example, by simply replacing the text encoder with XLM-R, X2-VLM outperforms state-of-the-art multilingual multi-modal pre-trained models without any multilingual pre-training. The code and pre-trained models are available athttps://github.com/zengyan-97/X2-VLM. Yan Zeng 0003, Xinsong Zhang, Hang Li 0001, Wangchunshu Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-trainingabstractIn this paper, we introduce Cross-View Language Modeling, a simple and effective pretraining framework that unifies cross-lingual and cross-modal pre-training with shared architectures and objectives.Our approach is motivated by a key observation that cross-lingual and cross-modal pre-training share the same goal of aligning two different views of the same object into a common semantic space.To this end, the cross-view language modeling framework considers both multi-modal data (i.e., image-caption pairs) and multi-lingual data (i.e., parallel sentence pairs) as two different views of the same object, and trains the model to align the two views by maximizing the mutual information between them with conditional masked language modeling and contrastive learning.We pre-train CCLM, a Crosslingual Cross-modal Language Model, with the cross-view language modeling framework.Empirical results on IGLUE, a multi-lingual multi-modal benchmark, and two multi-lingual image-text retrieval datasets show that while conceptually simpler, CCLM significantly outperforms the prior state-of-the-art with an average absolute improvement of over 10%.Moreover, CCLM is the first multi-lingual multimodal pre-trained model that surpasses the translate-test performance of representative English vision-language models by zero-shot cross-lingual transfer.1 Yan Zeng 0003, Wangchunshu Zhou, Ao Luo, Ziming Cheng, Xinsong Zhang |
ACL (1) | 5 |
| 2023 | Write and Paint: Generative Vision-Language Models are Unified Modal Learners
Shizhe Diao, Wangchunshu Zhou, Xinsong Zhang |
ICLR | 3 |
| 2022 | Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsabstractMost existing methods in vision language pre-training rely on object-centric features extracted through object detection and make fine-grained alignments between the extracted features and texts. It is challenging for these methods to learn relations among multiple objects. To this end, we propose a new method called X-VLM to perform ‘multi-grained vision language pre-training.’ The key to learning multi-grained alignments is to locate visual concepts in the image given the associated texts, and in the meantime align the texts with the visual concepts, where the alignments are in multi-granularity. Experimental results show that X-VLM effectively leverages the learned multi-grained alignments to many downstream vision language tasks and consistently outperforms state-of-the-art methods. Yan Zeng 0003, Xinsong Zhang, Hang Li 0001 |
ICML | 2 |
| 2022 | VLUE: A Multi-Task Multi-Dimension Benchmark for Evaluating Vision-Language Pre-trainingabstractRecent advances in vision-language pre-training (VLP) have demonstrated impressive performance in a range of vision-language (VL) tasks. However, there exist several challenges for measuring the community’s progress in building general multi-modal intelligence. First, most of the downstream VL datasets are annotated using raw images that are already seen during pre-training, which may result in an overestimation of current VLP models’ generalization ability. Second, recent VLP work mainly focuses on absolute performance but overlooks the efficiency-performance trade-off, which is also an important indicator for measuring progress. To this end, we introduce the Vision-Language Understanding Evaluation (VLUE) benchmark, a multi-task multi-dimension benchmark for evaluating the generalization capabilities and the efficiency-performance trade-off (“Pareto SOTA”) of VLP models. We demonstrate that there is a sizable generalization gap for all VLP models when testing on out-of-distribution test sets annotated on images from a more diverse distribution that spreads across cultures. Moreover, we find that measuring the efficiency-performance trade-off of VLP models leads to complementary insights for several design choices of VLP. We release the VLUE benchmark to promote research on building vision-language models that generalize well to images unseen during pre-training and are practical in terms of efficiency-performance trade-off. Wangchunshu Zhou, Yan Zeng 0003, Shizhe Diao, Xinsong Zhang |
ICML | 4 |
| 2021 | Robust Neural Relation Extraction via Multi-Granularity Noises ReductionabstractDistant supervision is widely used to extract relational facts with automatically labeled datasets to reduce high cost of human annotation. However, current distantly supervised methods suffer from the common problems of word-level and sentence-level noises, which come from a large proportion of irrelevant words in a sentence and inaccurate relation labels for numerous sentences. The problems lead to unacceptable precision in relation extraction and are critical for the success of using distant supervision. In this paper, we propose a novel and robust neural approach to deal with both problems by reducing influences of the multi-granularity noises. Three levels of noises from word, sentence until knowledge type are carefully considered in this work. We first initiate a question-answering based relation extractor (QARE) to remove noisy words in a sentence. Then we use multi-focus multi-instance learning (MMIL) to alleviate the effects of sentence-level noise by utilizing wrongly labeled sentences properly. Finally, to enhance our method against all the noises, we initialize parameters in our method with a priori knowledge learned from the relevant task of entity type classification by transfer learning. Extensive experiments on both existing benchmark and an improved larger dataset demonstrate that our proposed approach remarkably achieves new state-of-the-art performance. Xinsong Zhang, Pengshuai Li, Weijia Jia 0001, Hai Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Fine-grained relation extraction with focal multi-task learning
Xinsong Zhang, Weijia Jia 0001, Pengshuai Li |
Sci. China Inf. Sci. | 1 |
| 2019 | Multi-Labeled Relation Extraction with Attentive Capsule NetworkabstractTo disclose overlapped multiple relations from a sentence still keeps challenging. Most current works in terms of neural models inconveniently assuming that each sentence is explicitly mapped to a relation label, cannot handle multiple relations properly as the overlapped features of the relations are either ignored or very difficult to identify. To tackle with the new issue, we propose a novel approach for multi-labeled relation extraction with capsule network which acts considerably better than current convolutional or recurrent net in identifying the highly overlapped relations within an individual sentence. To better cluster the features and precisely extract the relations, we further devise attention-based routing algorithm and sliding-margin loss function, and embed them into our capsule network. The experimental results show that the proposed approach can indeed extract the highly overlapped features and achieve significant performance improvement for relation extraction comparing to the state-of-the-art works. Xinsong Zhang, Pengshuai Li, Weijia Jia 0001, Hai Zhao 0001 |
AAAI | 1 |
| 2018 | Neural Relation Extraction via Inner-Sentence Noise Reduction and Transfer LearningabstractExtracting relations is critical for knowledge base completion and construction in which distant supervised methods are widely used to extract relational facts automatically with the existing knowledge bases.However, the automatically constructed datasets comprise amounts of low-quality sentences containing noisy words, which is neglected by current distant supervised methods resulting in unacceptable precisions.To mitigate this problem, we propose a novel word-level distant supervised approach for relation extraction.We first build Sub-Tree Parse (STP) to remove noisy words that are irrelevant to relations.Then we construct a neural network inputting the subtree while applying the entity-wise attention to identify the important semantic features of relational words in each instance.To make our model more robust against noisy words, we initialize our network with a priori knowledge learned from the relevant task of entity classification by transfer learning.We conduct extensive experiments using the corpora of New York Times (NYT) and Freebase.Experiments show that our approach is effective and improves the area of Precision/Recall (PR) from 0.35 to 0.39 over the state-of-the-art work. Xinsong Zhang, Wanhao Zhou, Weijia Jia 0001 |
EMNLP | 2 |
| 2008 | Comparison of NIST and Wavelet Transform Test Point Selection Methods For a Programmable Gain Amplifier
Xinsong Zhang, Simon S. Ang, Chandra Carter |
J. Electron. Test. | 1 |
| 2007 | Test Point Selections for a Programmable Gain Amplifier Using NIST and Wavelet Transform MethodsabstractTest point selections for a programmable gain amplifier (PGA) using the National Institute of Standard (NIST) and wavelet transform methods are investigated. Although the wavelet transform method is an efficient method in test point selection for many mixed-signal devices, for a PGA with 31 input steps, the NIST method is shown to be more accurate in predicting the output gain responses than the wavelet transform method. Xinsong Zhang, Simon S. Ang, Chandra Carter |
ATS | 1 |