VLDB 2026 Research / reviewers in the wild / expert
Weiran Chen 0001
dblp:178/6451-1
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2025
0009-0008-1559-2498ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MedKI: Knowledge Dual Injections for Medical Visual Question AnsweringabstractMedical Visual Question Answering (Med VQA) is a challenging task for the sake of diverse medical image and multidisciplinary knowledge. Nowadays, the visual and language pretraining-finetuning framework is widely used in Med VQA task. However, most methods neglect the potential semantics and clinical information of image-text pairs, resulting in an inability to accurately match question semantics with image information. To address this, we propose a method called MedKI with dual injections of clinical and semantic knowledge, which is based on the pretraining and finetuning framework. Specifically, during pretraining, we inject clinical knowledge into the alignment module. Here, clinical knowledge is composed of the structural and the conceptual features that are extracted from the graph structure and entity definitions of the expert domain knowledge graph, respectively. In the finetuning stage, we retrieve similar texts from the pretraining corpus and encode them as semantic knowledge. Then, the knowledge is injected into the semantic knowledge fusion module. Extensive experimental results on both VQA-RAD dataset and SLAKE dataset demonstrate the validity of our proposed method. Hongyi Ren, Weiran Chen 0001, Chunping Liu, Yi Ji 0001, Ying Li 0065 |
ICIP | 2 |
| 2025 | DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid IntegrationabstractFew-shot font generation aims to create new fonts with a limited number of glyph references. It can be used to significantly reduce the labor cost of manual font design. However, due to the variety and complexity of font styles, the results generated by existing methods often suffer from visible defects, such as stroke errors, artifacts and blurriness. To address these issues, we propose DA-Font, a novel framework which integrates a Dual-Attention Hybrid Module (DAHM). Specifically, we introduce two synergistic attention blocks: the component attention block that leverages component information from content images to guide the style transfer process, and the relation attention block that further refines spatial relationships through interacting the content feature with both original and stylized component-wise representations. These two blocks collaborate to preserve accurate character shapes and stylistic textures. Moreover, we also design a corner consistency loss and an elastic mesh feature loss to better improve geometric alignment. Extensive experiments show that our DA-Font outperforms the state-of-the-art methods across diverse font styles and characters, demonstrating its effectiveness in enhancing structural integrity and local fidelity. The source code can be found at https://github.com/wrchen2001/DA-Font. Weiran Chen 0001, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
ACM Multimedia | 1 |
| 2025 | SiamHCC: a novel siamese network for quality evaluation of handwritten Chinese characters
Weiran Chen 0001, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
Multim. Syst. | 1 |
| 2024 | TARN-VIST: Topic Aware Reinforcement Network for Visual StorytellingabstractAs a cross-modal task, visual storytelling aims to generate a story for an ordered image sequence automatically. Different from the image captioning task, visual storytelling requires not only modeling the relationships between objects in the image but also mining the connections between adjacent images. Recent approaches primarily utilize either end-to-end frameworks or multi-stage frameworks to generate relevant stories, but they usually overlook latent topic information. In this paper, in order to generate a more coherent and relevant story, we propose a novel method, Topic Aware Reinforcement Network for VIsual StoryTelling (TARN-VIST). In particular, we pre-extracted the topic information of stories from both visual and linguistic perspectives. Then we apply two topic-consistent reinforcement learning rewards to identify the discrepancy between the generated story and the human-labeled story so as to refine the whole generation process. Extensive experimental results on the VIST dataset and human evaluation demonstrate that our proposed model outperforms most of the competitive models across multiple evaluation metrics. Weiran Chen 0001, Jiaqi Su, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
LREC/COLING | 1 |
| 2024 | MAGIC: Multi-prompt Any Length Video Generation Model with Controllable Inter-frame Correlation and Low Barrier
Weiran Chen 0001, Lingbing Xu, Yi Ji 0001, Ying Li 0065, Chunping Liu |
ICANN (3) | 2 |
| 2024 | Glocal Cascading Network for Topic Enhanced Visual StorytellingabstractAs a cross-modal task, visual storytelling aims to generate a semantically coherent story for an ordered image sequence. Despite significant achievements in existing methods for this task, few works focus on improving the conception ability which humans usually use when writing stories. In this work, we propose a framework called GLocal Cascading Network for Topic Enhanced Visual Storytelling which explores the conception ability by pre-modeling a latent topic for each image during story telling. Inspired by the global-local (glocal) ideology, we firstly propose a hierarchical latent-topic decoder consisting of two levels of topic generator which respectively focus on different levels of topic information. Then we propose a topic-aware loss which encourages the model to focus on the topic information of the story. With these two novel modules, our framework can effectively utilize the topic information and improve the informativeness and consistency of stories. Our model has been proven highly competitive across multiple metrics through extensive experiments conducted on the VIST dataset. Jiaqi Su, Weiran Chen 0001, Yi Ji 0001, Chunping Liu |
ICASSP | 2 |
| 2024 | Uncertainty-Aware with Negative Samples for Video-Text Retrieval
Weiran Chen 0001, Yi Ji 0001, Ying Li 0065, Chunping Liu |
PRCV (5) | 2 |
| 2024 | Quality evaluation methods of handwritten Chinese characters: a comprehensive survey
Weiran Chen 0001, Jiaqi Su, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
Multim. Syst. | 1 |
| 2022 | Chinese Character Style Transfer Model Based on Convolutional Neural Network
Weiran Chen 0001, Chunping Liu, Yi Ji 0001 |
ICANN (4) | 1 |