VLDB 2026 Research / reviewers in the wild / expert
Ziwang Zhao
dblp:333/0911
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-5538-3190ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Valley: Video Assistant with Large Language Model Enhanced AbilityabstractLarge Language Models (LLMs), with remarkable conversational capabilities, have emerged as AI assistants that can handle both visual and textual modalities. However, their effectiveness in joint video-language understanding has not been extensively explored. In the article, we introduce Valley , a multi-modal foundation model designed to enable enhanced video comprehension and instruction-following capabilities. To this end, we construct two datasets, namely “ Valley-702k ” and “ Valley-instruct-73k ,” to cover a diverse range of video–text alignment and video-based instruction tasks, such as multi-shot captions, long video descriptions, action recognition, causal inference, and so on. Then, we adopt ViT-L/14 as the vision encoder and explore three different temporal modeling modules to learn multifaceted features for enhanced video understanding. In addition, we implement a two-phase training approach for Valley: the first phase focuses solely on training the projection module to facilitate the LLM’s capacity to understand visual input, and the second phase jointly trains the projection module and the LLM to improve their instruction following ability. Extensive experiments demonstrate that Valley has the potential to serve as an effective video assistant, simplifying complex video understanding tasks. Our code and data are publicly available at https://github.com/RupertLuo/Valley . Ruipu Luo, Ziwang Zhao, Zheming Yang, Minghui Qiu, Zhongyu Wei, Yanhao Wang 0001, Cen Chen 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | LLM vs Small Model? Large Language Model Based Text Augmentation Enhanced Personality Detection ModelabstractPersonality detection aims to detect one's personality traits underlying in social media posts. One challenge of this task is the scarcity of ground-truth personality traits which are collected from self-report questionnaires. Most existing methods learn post features directly by fine-tuning the pre-trained language models under the supervision of limited personality labels. This leads to inferior quality of post features and consequently affects the performance. In addition, they treat personality traits as one-hot classification labels, overlooking the semantic information within them. In this paper, we propose a large language model (LLM) based text augmentation enhanced personality detection model, which distills the LLM's knowledge to enhance the small model for personality detection, even when the LLM fails in this task. Specifically, we enable LLM to generate post analyses (augmentations) from the aspects of semantic, sentiment, and linguistic, which are critical for personality detection. By using contrastive learning to pull them together in the embedding space, the post encoder can better capture the psycho-linguistic information within the post representations, thus improving personality detection. Furthermore, we utilize the LLM to enrich the information of personality labels for enhancing the detection performance. Experimental results on the benchmark datasets demonstrate that our model outperforms the state-of-the-art methods on personality detection. Linmei Hu, Duokang Wang, Ziwang Zhao, Yingxia Shao, Liqiang Nie |
AAAI | 4 |
| 2024 | Multimodal matching-aware co-attention networks with mutual knowledge distillation for fake news detection
Linmei Hu, Ziwang Zhao, Weijian Qi, Xuemeng Song, Liqiang Nie |
Inf. Sci. | 2 |
| 2024 | A Survey of Knowledge Enhanced Pre-Trained Language ModelsabstractPre-trained Language Models (PLMs) which are trained on large text corpus via self-supervised learning method, have yielded promising performance on various tasks in Natural Language Processing (NLP). However, though PLMs with huge parameters can effectively possess rich knowledge learned from massive training text and benefit downstream tasks at the fine-tuning stage, they still have some limitations such as poor reasoning ability due to the lack of external knowledge. Research has been dedicated to incorporating knowledge into PLMs to tackle these issues. In this paper, we present a comprehensive review of Knowledge Enhanced Pre-trained Language Models (KE-PLMs) to provide a clear insight into this thriving field. We introduce appropriate taxonomies respectively for Natural Language Understanding (NLU) and Natural Language Generation (NLG) to highlight these two main tasks of NLP. For NLU, we divide the types of knowledge into four categories: linguistic knowledge, text knowledge, knowledge graph (KG), and rule knowledge. The KE-PLMs for NLG are categorized into KG-based and retrieval-based methods. Finally, we point out some promising future directions of KE-PLMs. Linmei Hu, Ziwang Zhao, Lei Hou 0001, Liqiang Nie, Juan-Zi Li |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Knowledgeable Parameter Efficient Tuning Network for Commonsense Question AnsweringabstractCommonsense question answering is important for making decisions about everyday matters.Although existing commonsense question answering works based on fully fine-tuned PLMs have achieved promising results, they suffer from prohibitive computation costs as well as poor interpretability.Some works improve the PLMs by incorporating knowledge to provide certain evidence, via elaborately designed GNN modules which require expertise.In this paper, we propose a simple knowledgeable parameter efficient tuning network to couple PLMs with external knowledge for commonsense question answering.Specifically, we design a trainable parameter-sharing adapter attached to a parameter-freezing PLM to incorporate knowledge at a small cost.The adapter is equipped with both entity-and query-related knowledge via two auxiliary knowledge-related tasks (i.e., span masking and relation discrimination).To make the adapter focus on the relevant knowledge, we design gating and attention mechanisms to respectively filter and fuse the query information from the PLM.Extensive experiments on two benchmark datasets show that KPE is parameter-efficient and can effectively incorporate knowledge for improving commonsense question answering. Ziwang Zhao, Linmei Hu, Yingxia Shao, Yequan Wang |
ACL (1) | 1 |
| 2023 | Causal Inference for Leveraging Image-Text Matching Bias in Multi-Modal Fake News DetectionabstractMulti-modal fake news detection has drawn considerable attention with the development of online social media. Existing methods primarily conduct direct cross-modal fusion, while ignoring the image-text matching degree which may introduce unexpected bias. This work studies an unexplored problem in multi-modal fake news detection – how to deconfound and leverage the image-text matching bias to improve the performance of fake news detection. The key lies in two aspects: how to remove the confounding effect of the image-text matching bias during training, and how to utilize the bias in the inference stage since the news with mismatched image and text is more likely to be fake. To achieve our goal, we formulate the fake news detection task as a causal graph that reflects the cause-effect factors, and propose a novel framework –Causal Inference forLeveragingImage-textMatchingBias (CLIMB) in multi-modal fake news detection. To our best knowledge, this is the first work that considers the image-text matching degree into the fake news detection task with the approach of causal inference. CLIMB can be applied to any fake news detection models with visual and textual features as inputs. Extensive experiments on two real-world datasets validate the effectiveness of CLIMB. Linmei Hu, Ziwang Zhao, Jianhua Yin 0001, Liqiang Nie |
IEEE Trans. Knowl. Data Eng. | 3 |