VLDB 2026 Research / reviewers in the wild / expert
Xinglin Hou
dblp:319/3945
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 88% Representation and self-supervised learning · 9% Language models and text generation · 3% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 100% |
Topics — the 11 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
image captioning |
1.2 | 2 | 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023 CapOnImage: Context-driven Dense-Captioning on Image · EMNLP 2022 |
Computer vision › Vision and language › video captioning
controllable video captioning |
0.8 | 1 | 2024 | Edit As You Wish: Video Caption Editing with Multi-grained User Control · ACM Multimedia 2024 |
Computer vision › Vision and language
cross-modal retrieval |
0.8 | 1 | 2024 | Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning · AAAI 2024 |
Information retrieval
multimedia analysis and retrieval |
0.8 | 1 | 2024 | Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning · AAAI 2024 |
Information retrieval › cross-modal retrieval
text-to-video retrieval |
0.8 | 1 | 2024 | Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning · AAAI 2024 |
Computer vision › Vision and language › image captioning › low-shot image captioning
few-shot image captioning |
0.7 | 1 | 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023 |
Computer vision › Vision and language › image captioning › controllable image captioning
stylized image captioning |
0.7 | 1 | 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023 |
Computer vision › Vision and language › image captioning
dense captioning |
0.6 | 1 | 2022 | CapOnImage: Context-driven Dense-Captioning on Image · EMNLP 2022 |
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining |
0.6 | 1 | 2022 | CapOnImage: Context-driven Dense-Captioning on Image · EMNLP 2022 |
Multimedia analysis and retrieval
video captioning |
0.2 | 1 | 2024 | Edit As You Wish: Video Caption Editing with Multi-grained User Control · ACM Multimedia 2024 |
Natural language and speech › Language models and text generation › language modeling
conditional language model |
0.2 | 1 | 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences · ACM Multimedia 2023 |
Methods — techniques the papers use, named apart from their topics
text-gated interaction · 1.5pearson constraint · 1.5operation-position-attribute triplet · 1.5multi-grained user control · 1.5coarse-to-fine representation learning · 1.5neighbor location embedding · 1.1multimodal pretraining · 1.1visual projection · 0.7style extractor · 0.7conditional encoder-decoder · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation LearningabstractIn recent years, text-to-video retrieval methods based on CLIP have experienced rapid development. The primary direction of evolution is to exploit the much wider gamut of visual and textual cues to achieve alignment. Concretely, those methods with impressive performance often design a heavy fusion block for sentence (words)-video (frames) interaction, regardless of the prohibitive computation complexity. Nevertheless, these approaches are not optimal in terms of feature utilization and retrieval efficiency. To address this issue, we adopt multi-granularity visual feature learning, ensuring the model's comprehensiveness in capturing visual content features spanning from abstract to detailed levels during the training phase. To better leverage the multi-granularity features, we devise a two-stage retrieval architecture in the retrieval phase. This solution ingeniously balances the coarse and fine granularity of retrieval content. Moreover, it also strikes a harmonious equilibrium between retrieval effectiveness and efficiency. Specifically, in training phase, we design a parameter-free text-gated interaction block (TIB) for fine-grained video representation learning and embed an extra Pearson Constraint to optimize cross-modal representation learning. In retrieval phase, we use coarse-grained video representations for fast recall of top-k candidates, which are then reranked by fine-grained video representations. Extensive experiments on four benchmarks demonstrate the efficiency and effectiveness. Notably, our method achieves comparable performance with the current state-of-the-art methods while being nearly 50 times faster. Kaibin Tian, Yanhua Cheng, Xinglin Hou, Quan Chen 0006, Han Li 0005 |
AAAI | 4 |
| 2024 | Edit As You Wish: Video Caption Editing with Multi-grained User ControlabstractAutomatically narrating videos in natural language complying with user requests, i.e. Controllable Video Captioning task, can help people manage massive videos with desired intentions. However, existing works suffer from two shortcomings: 1) the control signal is single-grained which can not satisfy diverse user intentions; 2) the video description is generated in a single round which can not be further edited to meet dynamic needs. In this paper, we propose a novel Video Caption Editing (VCE) task to automatically revise an existing video description guided by multi-grained user requests. Inspired by human writing-revision habits, we design the user command as a pivotal triplet {operation, position, attribute} to cover diverse user needs from coarse-grained to fine-grained. To facilitate the VCE task, we automatically construct an open-domain benchmark dataset named VATEX-EDIT and manually collect an e-commerce dataset called EMMAD-EDIT. We further propose a specialized small-scale model (i.e., OPA) compared with two generalist Large Multi-modal Models to perform an exhaustive analysis of the novel task. For evaluation, we adopt comprehensive metrics considering caption fluency, command-caption consistency, and video-caption alignment. Experiments reveal the task challenges of fine-grained multi-modal semantics understanding and processing. Our datasets, codes, and evaluation tools are available at https://github.com/yaolinli/VCE. Linli Yao, Yuanmeng Zhang, Xinglin Hou, Tiezheng Ge, Yuning Jiang 0001, Xu Sun 0001, Qin Jin |
ACM Multimedia | 4 |
| 2023 | Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized SentencesabstractStylized visual captioning aims to generate image or video descriptions with specific styles, making them more attractive and emotionally appropriate. One major challenge with this task is the lack of paired stylized captions for visual content, so most existing works focus on unsupervised methods that do not rely on parallel datasets. However, these approaches still require training with sufficient examples that have style labels, and the generated captions are limited to predefined styles. To address these limitations, we explore the problem of Few-Shot Stylized Visual Captioning, which aims to generate captions in any desired style, using only a few examples as guidance during inference, without requiring further training. We propose a framework called FS-StyleCap for this task, which utilizes a conditional encoder-decoder language model and a visual projection module. Our two-step training scheme proceeds as follows: first, we train a style extractor to generate style representations on an unlabeled text-only corpus. Then, we freeze the extractor and enable our decoder to generate stylized descriptions based on the extracted style vector and projected visual content vectors. During inference, our model can generate desired stylized captions by deriving the style representation from user-supplied examples. Our automatic evaluation results for few-shot sentimental visual captioning outperform state-of-the-art approaches and are comparable to models that are fully trained on labeled style corpora. Human evaluations further confirm our model's ability to handle multiple styles. Dingyi Yang, Hongyu Chen 0005, Xinglin Hou, Tiezheng Ge, Yuning Jiang 0001, Qin Jin |
ACM Multimedia | 3 |
| 2022 | CapOnImage: Context-driven Dense-Captioning on ImageabstractExisting image captioning systems are dedicated to generating narrative captions for images, which are spatially detached from the image in presentation.However, texts can also be used as decorations on the image to highlight the key points and increase the attractiveness of images.In this work, we introduce a new task called captioning on image (CapOn-Image) 1 , which aims to generate dense captions at different locations of the image based on contextual information.For this new task, we introduce a large-scale benchmark called CapOn-Image2M, which contains 2.1 million product images, each with an average of 4.8 spatially localized captions.To fully exploit the surrounding visual context to generate the most suitable caption for each location, we propose a multimodal pre-training model with multi-level pretraining tasks that progressively learn the correspondence between texts and image locations from easy to hard.To avoid generating redundant captions for nearby locations, we further enhance the location embedding with neighbor locations .Compared with other image captioning model variants, our model achieves the best results in both captioning accuracy and diversity aspects. Yiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge, Yuning Jiang 0001, Peng Wang 0015 |
EMNLP | 2 |
| 2022 | Dual-Level Decoupled Transformer for Video CaptioningabstractVideo captioning aims to understand the spatio-temporal semantic concept of the video and generate descriptive sentences. The de-facto approach to this task dictates a text generator to learn from offline-extracted motion or appearance features from pre-trained vision models. However, these methods may suffer from the so-called "couple" drawbacks on both video spatio-temporal representation and sentence generation. For the former, "couple" means learning spatio-temporal representation in a single model(3DCNN), resulting the problems named disconnection in task/pre-train domain and hard for end-to-end training. As for the latter, "couple" means treating the generation of visual semantic and syntax-related words equally. To this end, we present D2 - a dual-level decoupled transformer pipeline to solve the above drawbacks: (i) for video spatio-temporal representation, we decouple the process of it into "first-spatial-then-temporal" paradigm, releasing the potential of using dedicated model(e.g. image-text pre-training) to connect the pre-training and downstream tasks, and makes the entire model end-to-end trainable. (ii) for sentence generation, we propose Syntax-Aware Decoder to dynamically measure the contribution of visual semantic and syntax-related words. Extensive experiments on three widely-used benchmarks (MSVD, MSR-VTT and VATEX) have shown great potential of the proposed D2 and surpassed the previous methods by a large margin in the task of video captioning. Yiqi Gao, Xinglin Hou, Wei Suo, Mengyang Sun, Tiezheng Ge, Yuning Jiang 0001, Peng Wang 0015 |
ICMR | 2 |