Danyang Hou

dblp:228/0778 · also Dan Yang Hou · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0006-6949-2703ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Event-Aware Video Corpus Moment Retrieval
Danyang Hou, Liang Pang 0001, Yanyan Lan, Huawei Shen, Xueqi Cheng 0001
ECIR (1)1
2024 Improving Video Corpus Moment Retrieval with Partial Relevance Enhancement
abstract
Video Corpus Moment Retrieval (VCMR) is a new video retrieval task aimed at retrieving a relevant moment from a large corpus of untrimmed videos using a text query. The relevance between the video and query is partial, mainly evident in two aspects: (1) Scope: The untrimmed video contains many frames, but not all are relevant to the query. Strong relevance is typically observed only within the relevant moment. (2) Modality: The relevance of the query varies with different modalities. Action descriptions align more with visual elements, while character conversations are more related to textual information. Existing methods often treat all video contents equally, leading to sub-optimal moment retrieval. We argue that effectively capturing the partial relevance between the query and video is essential for the VCMR task. To this end, we propose a Partial Relevance Enhanced Model (PREM) to improve VCMR. VCMR involves two sub-tasks: video retrieval and moment localization. To align with their distinct objectives, we implement specialized partial relevance enhancement strategies. For video retrieval, we introduce a multi-modal collaborative video retriever, generating different query representations for the two modalities by modality-specific pooling, ensuring a more effective match. For moment localization, we propose the focus-then-fuse moment localizer, utilizing modality-specific gates to capture essential content. We also introduce relevant content-enhanced training methods for both retriever and localizer to enhance the ability of model to capture relevant content. Experimental results on TVR and DiDeMo datasets show that the proposed model outperforms the baselines, achieving a new state-of-the-art of VCMR. The code is available at https://github.com/hdy007007/PREM.
Danyang Hou, Liang Pang 0001, Huawei Shen, Xueqi Cheng 0001
ICMR1
2024 Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated Images
abstract
With the application of generation models, internet is increasingly inundated with AI-generated content (AIGC), causing both real and AI-generated content indexed in corpus for search. This paper explores the impact of AI-generated images on text-image search in this scenario. Firstly, we construct a benchmark consisting of both real and AI-generated images for this study. In this benchmark, AI-generated images possess visual semantics sufficiently similar to real images. Experiments on this benchmark reveal that text-image retrieval models tend to rank the AI-generated images higher than the real images, even though the AI-generated images do not exhibit more visually relevant semantics to the queries than real images. We call this bias as invisible relevance bias. This bias is detected across retrieval models with different training data and architectures. Further exploration reveals that mixing AI-generated images into the training data of retrieval models exacerbates the invisible relevance bias. These problems cause a vicious cycle in which AI-generated images have a higher chance of exposing from massive data, which makes them more likely to be mixed into the training of retrieval models and such training makes the invisible relevance bias more and more serious. To mitigate this bias and elucidate the potential causes of the bias, firstly, we propose an effective method to alleviate this bias. Subsequently, we apply our proposed debiasing method to retroactively identify the causes of this bias, revealing that the AI-generated images induce the image encoder to embed additional information into their representation. This information makes the retriever estimate a higher relevance score. We conduct experiments to support this assertion.
Danyang Hou, Liang Pang 0001, Jingcheng Deng, Jun Xu 0001, Huawei Shen, Xueqi Cheng 0001
SIGIR2
2022 Region-based Cross-modal Retrieval
abstract
Cross-modal retrieval aims to identify relevant information from different modalities, such as image and text. Existing works build a coarse relationship between whole image and whole text. However, the retrieval in fine-grained elements is more necessary in research and has many practical applications, such as explainable and interactive retrieval. In this paper, we propose a region-based cross-modal retrieval task that focuses on finding the semantic match between fine-grained parts of text and image, e.g., sentence in paragraph and region in image. The challenge of this task is to enhance fine-grained representation with contextual knowledge. To this end, we propose a Context-Aware Region Retrieval (CARR) model. It utilizes two pre-trained models as backbones, Faster RCNN for image and BERT for text, and transformer-based encoders to obtain the contextual aware representations of two modalities. We conduct experiments on Visual Genome and Localized Narratives datasets, and the experimental results demonstrate that the proposed model outperforms the baseline methods.
Danyang Hou, Liang Pang 0001, Yanyan Lan, Huawei Shen, Xueqi Cheng 0001
IJCNN1