VLDB 2026 Research / reviewers in the wild / expert
Daiqing Qi
dblp:229/9064
· DBLP profile ↗
9ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0001-9543-5792ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Vision and language · 42% Efficient and distributed learning · 13% Image recognition and object detection · 8% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.1 | 2 | 2025 | The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers · CVPR 2025 Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive Decoding · NeurIPS 2025 |
Computer vision › Image recognition and object detection
image aesthetics assessment |
0.9 | 1 | 2025 | The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers · CVPR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
temporal reasoning |
0.9 | 1 | 2025 | Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive Decoding · NeurIPS 2025 |
Computer vision › Vision and language › video-language model
video large language model |
0.9 | 1 | 2025 | Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive Decoding · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024 |
Computer vision › Vision and language
cross-modal retrieval |
0.8 | 1 | 2024 | Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.8 | 1 | 2024 | Generalizing to Unseen Domains via Text-Guided Augmentation: A Training-Free Approach · ECCV (69) 2024 |
Computer vision › Vision and language
vision-language pretraining |
0.8 | 1 | 2024 | Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024 |
Computer vision › Vision and language › vision-language model › multimodal large language model
visual instruction tuning |
0.8 | 1 | 2024 | Tag-grounded Visual Instruction Tuning with Retrieval Augmentation · EMNLP 2024 |
Machine learning › Efficient and distributed learning › federated learning
federated continual learning |
0.7 | 1 | 2023 | Better Generative Replay for Continual Federated Learning · ICLR 2023 |
Machine learning › Efficient and distributed learning
federated learning |
0.7 | 1 | 2023 | Better Generative Replay for Continual Federated Learning · ICLR 2023 |
Machine learning › Learning paradigms › continual learning › rehearsal-based continual learning
generative replay |
0.7 | 1 | 2023 | Better Generative Replay for Continual Federated Learning · ICLR 2023 |
Computational photography and imaging
image aesthetics |
0.3 | 1 | 2025 | The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers · CVPR 2025 |
Information retrieval › image retrieval
composed image retrieval |
0.2 | 1 | 2024 | Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024 |
Information retrieval
multimodal retrieval |
0.2 | 1 | 2024 | Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024 |
Information retrieval
retrieval augmentation |
0.2 | 1 | 2024 | Tag-grounded Visual Instruction Tuning with Retrieval Augmentation · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 2.3language-guided multi-view vision fusion · 1.7retrieval augmentation · 1.5fine-tuning · 1.5contrastive learning · 1.5temporal consistency corruption · 0.9contrastive decoding · 0.9text-guided augmentation · 0.8generative replay · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographersabstract"While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky."John Szarkowski, William Eggleston’s Guide1Photographer and curator, Szarkowski insightfully revealed one of the notable gaps between general and aesthetic visual understanding: while the former focuses on identifying the factual element in an image (sky), the latter transcends such object identification, viewing it instead as an aesthetic component—a pure color block (blue). Such fundamental distinctions between general (detection, localization, etc.) and aesthetic (color, lighting, composition, etc.) visual understanding present a significant challenge for Multimodal Large Language Models (MLLMs). Although some recent works have made initial explorations, they are often limited to general and basic aesthetic commonsense. As a result, they frequently fall short in real-world scenarios (Fig. 1), which require extensive expertise—including photographic techniques, photo pre/post-processing knowledge, and more, to provide a detailed analysis and description. To fundamentally enhance the aesthetics understanding of MLLMs, we first introduce a novel dataset, PhotoCritique, derived from extensive discussions among professional photographers and enthusiasts, and characterized by the large scale, expertise, and diversity. Then, to better learn visual aesthetics from PhotoCritique, we furthur propose a novel model, PhotoEye, featuring a language-guided multi-view vision fusion mechanism to understand image aesthetics from multiple perspectives. Finally, we present a novel benchmark, PhotoBench, a comprehensive and professional benchmark for aesthetic visual understanding. On existing benchmarks and PhotoBench, our model demonstrates clear advantages over existing models. Daiqing Qi, Handong Zhao, Jing Shi 0005, Simon Jenni, Franck Dernoncourt, Scott Cohen, Sheng Li 0001 |
CVPR | 1 |
| 2025 | Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive DecodingabstractA major distinction between video and image understanding is that the former requires reasoning over time.
Existing Video Large Language Models (VLLMs) demonstrate promising performance in general video understanding, such as brief captioning or object recognition within individual frames. However, they often struggle with temporal reasoning such as understanding continuous actions or tracking object transformations over time—which typically demands the integration of multiple frames in a temporally coherent manner.
We first explore and explain such failures in Video LLMs from the perspective of \textit{language and ``image'' priors.}
While existing research has attempted to enhance the temporal understanding of VLLMs through various training strategies, the demand for expensive computational resources and training data often presents significant barriers.
To this end, we further propose a simple yet novel idea for improving temporal reasoning in videos at no additional training cost.
Specifically, to better capture the temporal structure across multiple frames—the key to effective temporal reasoning—we distort the temporal consistency in key frames \textit{during the decoding phase}. Such corruption induces time-insensitive wrong responses from the model, which are then contrastively avoided when generating the final correct output. In this way, the model is encouraged to perform more temporally coherent reasoning.
Our method yields consistent improvements across both temporal-specific and general video understanding benchmarks, demonstrating its effectiveness and generalizability. Daiqing Qi, Dongliang Guo 0002, Hanzhang Yuan, Handong Zhao, Mengxuan Hu, Lehan Yang, Sheng Li 0001 |
NeurIPS | 1 |
| 2024 | Generalizing to Unseen Domains via Text-Guided Augmentation: A Training-Free Approach
Daiqing Qi, Handong Zhao, Aidong Zhang 0001, Sheng Li 0001 |
ECCV (69) | 1 |
| 2024 | Tag-grounded Visual Instruction Tuning with Retrieval AugmentationabstractPlease describe this photo in Daiqing Qi, Handong Zhao, Zijun Wei, Sheng Li 0001 |
EMNLP | 1 |
| 2024 | Easy Regional Contrastive Learning of Expressive Fashion RepresentationsabstractWhen learning vision-language models (VLM) for the fashion domain, most existing works design new architectures from vanilla BERT with additional objectives, or perform dense multi-task learning with fashion-specific tasks. Though progress has been made, their architecture or objectives are often intricate and the extendibility is limited.
By contrast, with simple architecture (comprising only two unimodal encoders) and just the contrastive objective, popular pre-trained VL models (e.g., CLIP) achieve superior performance in general domains, which are further easily extended to downstream tasks.
However, inheriting such benefits of CLIP in the fashion domain is non-trivial in the presence of the notable domain gap. Empirically, we find that directly finetuning on fashion data leads CLIP to frequently ignore minor yet important details such as logos and composition, which are critical in fashion tasks such as retrieval and captioning.
In this work, to maintain CLIP's simple architecture and objective while explicitly attending to fashion details, we propose $E^2$: Easy Regional Contrastive Learning of Expressive Fashion Representations.
$E^2$ introduces only a few selection tokens and fusion blocks (just 1.9\% additional parameters in total) with only contrastive losses. Despite lightweight, in our primary focus, cross-modal retrieval, $E^2$ notably outperforms existing fashion VLMs with various fashion-specific objectives.
Moreover, thanks to CLIP's widespread use in downstream tasks in general domains (e.g., zero-shot composed image retrieval and image captioning), our model can easily extend these models from general domain to the fashion domain with notable improvement.
To conduct a comprehensive evaluation, we further collect data from Amazon Reviews to build a new dataset for cross-modal retrieval in the fashion domain. Daiqing Qi, Handong Zhao, Sheng Li 0001 |
NeurIPS | 1 |
| 2024 | A Survey of Trustworthy Representation Learning Across DomainsabstractAs AI systems have obtained significant performance to be deployed widely in our daily lives and human society, people both enjoy the benefits brought by these technologies and suffer many social issues induced by these systems. To make AI systems good enough and trustworthy, plenty of researches have been done to build guidelines for trustworthy AI systems. Machine learning is one of the most important parts of AI systems, and representation learning is the fundamental technology in machine learning. How to make representation learning trustworthy in real-world application, e.g., cross domain scenarios, is very valuable and necessary for both machine learning and AI system fields. Inspired by the concepts in trustworthy AI, we proposed the first trustworthy representation learning across domains framework, which includes four concepts, i.e., robustness, privacy, fairness, and explainability, to give a comprehensive literature review on this research direction. Specifically, we first introduce the details of the proposed trustworthy framework for representation learning across domains. Second, we provide basic notions and comprehensively summarize existing methods for the trustworthy framework from four concepts. Finally, we conclude this survey with insights and discussions on future research directions. Ronghang Zhu, Dongliang Guo 0002, Daiqing Qi, Zhixuan Chu, Xiang Yu 0002, Sheng Li 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | Better Generative Replay for Continual Federated Learning
Daiqing Qi, Handong Zhao, Sheng Li 0001 |
ICLR | 1 |
| 2022 | Knowledge-Guided Article Embedding Refinement for Session-Based News RecommendationabstractPersonalized news recommendation aims to recommend news articles to customers, by exploiting the personal preferences and short-term reading interest of users. A practical challenge in personalized news recommendations is the lack of logged user interactions. Recently, the session-based news recommendation has attracted increasing attention, which tries to recommend the next news article given previous articles in an active session. Current session-based news recommendation methods mainly extract latent embeddings from news articles and user-item interactions. However, many existing methods could not exploit the semantic-level structural information among news articles. And the feature learning process simply relies on the news articles in training data, which may not be sufficient to learn semantically rich embeddings. This brief presents a context-aware graph embedding (CAGE) approach for session-based news recommendation. It employs external knowledge graphs to improve the semantic-level representations of news articles. Moreover, graph neural networks are incorporated to further enhance the article embeddings. In addition, we consider the similarity among sessions and design attention neural networks to model the short-term user preferences. Extensive results on multiple news recommendation benchmark datasets show that CAGE performs better than some competitive baselines in most cases. Heng-Shiou Sheu, Zhixuan Chu, Daiqing Qi, Sheng Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Less is more: Data-efficient complex question answering over knowledge bases
Yuncheng Hua, Yuan-Fang Li, Guilin Qi, Daiqing Qi |
J. Web Semant. | 6 |