Daiqing Qi

dblp:229/9064 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0001-9543-5792ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 42% Efficient and distributed learning · 13% Image recognition and object detection · 8%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
1.122025
The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers · CVPR 2025
Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive Decoding · NeurIPS 2025
Computer vision › Image recognition and object detection
image aesthetics assessment
0.912025
The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers · CVPR 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
temporal reasoning
0.912025
Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive Decoding · NeurIPS 2025
Computer vision › Vision and language › video-language model
video large language model
0.912025
Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive Decoding · NeurIPS 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024
Computer vision › Vision and language
cross-modal retrieval
0.812024
Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation
domain generalization
0.812024
Generalizing to Unseen Domains via Text-Guided Augmentation: A Training-Free Approach · ECCV (69) 2024
Computer vision › Vision and language
vision-language pretraining
0.812024
Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024
Computer vision › Vision and language › vision-language model › multimodal large language model
visual instruction tuning
0.812024
Tag-grounded Visual Instruction Tuning with Retrieval Augmentation · EMNLP 2024
Machine learning › Efficient and distributed learning › federated learning
federated continual learning
0.712023
Better Generative Replay for Continual Federated Learning · ICLR 2023
Machine learning › Efficient and distributed learning
federated learning
0.712023
Better Generative Replay for Continual Federated Learning · ICLR 2023
Machine learning › Learning paradigms › continual learning › rehearsal-based continual learning
generative replay
0.712023
Better Generative Replay for Continual Federated Learning · ICLR 2023
Computational photography and imaging
image aesthetics
0.312025
The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers · CVPR 2025
Information retrieval › image retrieval
composed image retrieval
0.212024
Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024
Information retrieval
multimodal retrieval
0.212024
Easy Regional Contrastive Learning of Expressive Fashion Representations · NeurIPS 2024
Information retrieval
retrieval augmentation
0.212024
Tag-grounded Visual Instruction Tuning with Retrieval Augmentation · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

vision-language model · 2.3language-guided multi-view vision fusion · 1.7retrieval augmentation · 1.5fine-tuning · 1.5contrastive learning · 1.5temporal consistency corruption · 0.9contrastive decoding · 0.9text-guided augmentation · 0.8generative replay · 0.7
YearPublicationVenuePosition
2025 The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers
abstract
"While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky."John Szarkowski, William Eggleston’s Guide1Photographer and curator, Szarkowski insightfully revealed one of the notable gaps between general and aesthetic visual understanding: while the former focuses on identifying the factual element in an image (sky), the latter transcends such object identification, viewing it instead as an aesthetic component—a pure color block (blue). Such fundamental distinctions between general (detection, localization, etc.) and aesthetic (color, lighting, composition, etc.) visual understanding present a significant challenge for Multimodal Large Language Models (MLLMs). Although some recent works have made initial explorations, they are often limited to general and basic aesthetic commonsense. As a result, they frequently fall short in real-world scenarios (Fig. 1), which require extensive expertise—including photographic techniques, photo pre/post-processing knowledge, and more, to provide a detailed analysis and description. To fundamentally enhance the aesthetics understanding of MLLMs, we first introduce a novel dataset, PhotoCritique, derived from extensive discussions among professional photographers and enthusiasts, and characterized by the large scale, expertise, and diversity. Then, to better learn visual aesthetics from PhotoCritique, we furthur propose a novel model, PhotoEye, featuring a language-guided multi-view vision fusion mechanism to understand image aesthetics from multiple perspectives. Finally, we present a novel benchmark, PhotoBench, a comprehensive and professional benchmark for aesthetic visual understanding. On existing benchmarks and PhotoBench, our model demonstrates clear advantages over existing models.
Daiqing Qi, Handong Zhao, Jing Shi 0005, Simon Jenni, Franck Dernoncourt, Scott Cohen, Sheng Li 0001
CVPR1
2025 Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive Decoding
abstract
A major distinction between video and image understanding is that the former requires reasoning over time. Existing Video Large Language Models (VLLMs) demonstrate promising performance in general video understanding, such as brief captioning or object recognition within individual frames. However, they often struggle with temporal reasoning such as understanding continuous actions or tracking object transformations over time—which typically demands the integration of multiple frames in a temporally coherent manner. We first explore and explain such failures in Video LLMs from the perspective of \textit{language and ``image'' priors.} While existing research has attempted to enhance the temporal understanding of VLLMs through various training strategies, the demand for expensive computational resources and training data often presents significant barriers. To this end, we further propose a simple yet novel idea for improving temporal reasoning in videos at no additional training cost. Specifically, to better capture the temporal structure across multiple frames—the key to effective temporal reasoning—we distort the temporal consistency in key frames \textit{during the decoding phase}. Such corruption induces time-insensitive wrong responses from the model, which are then contrastively avoided when generating the final correct output. In this way, the model is encouraged to perform more temporally coherent reasoning. Our method yields consistent improvements across both temporal-specific and general video understanding benchmarks, demonstrating its effectiveness and generalizability.
Daiqing Qi, Dongliang Guo 0002, Hanzhang Yuan, Handong Zhao, Mengxuan Hu, Lehan Yang, Sheng Li 0001
NeurIPS1
2024 Generalizing to Unseen Domains via Text-Guided Augmentation: A Training-Free Approach
Daiqing Qi, Handong Zhao, Aidong Zhang 0001, Sheng Li 0001
ECCV (69)1
2024 Tag-grounded Visual Instruction Tuning with Retrieval Augmentation
abstract
Please describe this photo in
Daiqing Qi, Handong Zhao, Zijun Wei, Sheng Li 0001
EMNLP1
2024 Easy Regional Contrastive Learning of Expressive Fashion Representations
abstract
When learning vision-language models (VLM) for the fashion domain, most existing works design new architectures from vanilla BERT with additional objectives, or perform dense multi-task learning with fashion-specific tasks. Though progress has been made, their architecture or objectives are often intricate and the extendibility is limited. By contrast, with simple architecture (comprising only two unimodal encoders) and just the contrastive objective, popular pre-trained VL models (e.g., CLIP) achieve superior performance in general domains, which are further easily extended to downstream tasks. However, inheriting such benefits of CLIP in the fashion domain is non-trivial in the presence of the notable domain gap. Empirically, we find that directly finetuning on fashion data leads CLIP to frequently ignore minor yet important details such as logos and composition, which are critical in fashion tasks such as retrieval and captioning. In this work, to maintain CLIP's simple architecture and objective while explicitly attending to fashion details, we propose $E^2$: Easy Regional Contrastive Learning of Expressive Fashion Representations. $E^2$ introduces only a few selection tokens and fusion blocks (just 1.9\% additional parameters in total) with only contrastive losses. Despite lightweight, in our primary focus, cross-modal retrieval, $E^2$ notably outperforms existing fashion VLMs with various fashion-specific objectives. Moreover, thanks to CLIP's widespread use in downstream tasks in general domains (e.g., zero-shot composed image retrieval and image captioning), our model can easily extend these models from general domain to the fashion domain with notable improvement. To conduct a comprehensive evaluation, we further collect data from Amazon Reviews to build a new dataset for cross-modal retrieval in the fashion domain.
Daiqing Qi, Handong Zhao, Sheng Li 0001
NeurIPS1
2024 A Survey of Trustworthy Representation Learning Across Domains
abstract
As AI systems have obtained significant performance to be deployed widely in our daily lives and human society, people both enjoy the benefits brought by these technologies and suffer many social issues induced by these systems. To make AI systems good enough and trustworthy, plenty of researches have been done to build guidelines for trustworthy AI systems. Machine learning is one of the most important parts of AI systems, and representation learning is the fundamental technology in machine learning. How to make representation learning trustworthy in real-world application, e.g., cross domain scenarios, is very valuable and necessary for both machine learning and AI system fields. Inspired by the concepts in trustworthy AI, we proposed the first trustworthy representation learning across domains framework, which includes four concepts, i.e., robustness, privacy, fairness, and explainability, to give a comprehensive literature review on this research direction. Specifically, we first introduce the details of the proposed trustworthy framework for representation learning across domains. Second, we provide basic notions and comprehensively summarize existing methods for the trustworthy framework from four concepts. Finally, we conclude this survey with insights and discussions on future research directions.
Ronghang Zhu, Dongliang Guo 0002, Daiqing Qi, Zhixuan Chu, Xiang Yu 0002, Sheng Li 0001
ACM Trans. Knowl. Discov. Data3
2023 Better Generative Replay for Continual Federated Learning
Daiqing Qi, Handong Zhao, Sheng Li 0001
ICLR1
2022 Knowledge-Guided Article Embedding Refinement for Session-Based News Recommendation
abstract
Personalized news recommendation aims to recommend news articles to customers, by exploiting the personal preferences and short-term reading interest of users. A practical challenge in personalized news recommendations is the lack of logged user interactions. Recently, the session-based news recommendation has attracted increasing attention, which tries to recommend the next news article given previous articles in an active session. Current session-based news recommendation methods mainly extract latent embeddings from news articles and user-item interactions. However, many existing methods could not exploit the semantic-level structural information among news articles. And the feature learning process simply relies on the news articles in training data, which may not be sufficient to learn semantically rich embeddings. This brief presents a context-aware graph embedding (CAGE) approach for session-based news recommendation. It employs external knowledge graphs to improve the semantic-level representations of news articles. Moreover, graph neural networks are incorporated to further enhance the article embeddings. In addition, we consider the similarity among sessions and design attention neural networks to model the short-term user preferences. Extensive results on multiple news recommendation benchmark datasets show that CAGE performs better than some competitive baselines in most cases.
Heng-Shiou Sheu, Zhixuan Chu, Daiqing Qi, Sheng Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2020 Less is more: Data-efficient complex question answering over knowledge bases
Yuncheng Hua, Yuan-Fang Li, Guilin Qi, Daiqing Qi
J. Web Semant.6