EDBT 2026 Demo / reviewers in the wild / expert
Ying Li 0065
dblp:22/1805-65
· DBLP profile ↗
23ranked-venue papers
3as first author
21since 2021 · last 2025
0009-0004-1669-1878ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MedKI: Knowledge Dual Injections for Medical Visual Question AnsweringabstractMedical Visual Question Answering (Med VQA) is a challenging task for the sake of diverse medical image and multidisciplinary knowledge. Nowadays, the visual and language pretraining-finetuning framework is widely used in Med VQA task. However, most methods neglect the potential semantics and clinical information of image-text pairs, resulting in an inability to accurately match question semantics with image information. To address this, we propose a method called MedKI with dual injections of clinical and semantic knowledge, which is based on the pretraining and finetuning framework. Specifically, during pretraining, we inject clinical knowledge into the alignment module. Here, clinical knowledge is composed of the structural and the conceptual features that are extracted from the graph structure and entity definitions of the expert domain knowledge graph, respectively. In the finetuning stage, we retrieve similar texts from the pretraining corpus and encode them as semantic knowledge. Then, the knowledge is injected into the semantic knowledge fusion module. Extensive experimental results on both VQA-RAD dataset and SLAKE dataset demonstrate the validity of our proposed method. Hongyi Ren, Weiran Chen 0001, Chunping Liu, Yi Ji 0001, Ying Li 0065 |
ICIP | 5 |
| 2025 | DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid IntegrationabstractFew-shot font generation aims to create new fonts with a limited number of glyph references. It can be used to significantly reduce the labor cost of manual font design. However, due to the variety and complexity of font styles, the results generated by existing methods often suffer from visible defects, such as stroke errors, artifacts and blurriness. To address these issues, we propose DA-Font, a novel framework which integrates a Dual-Attention Hybrid Module (DAHM). Specifically, we introduce two synergistic attention blocks: the component attention block that leverages component information from content images to guide the style transfer process, and the relation attention block that further refines spatial relationships through interacting the content feature with both original and stylized component-wise representations. These two blocks collaborate to preserve accurate character shapes and stylistic textures. Moreover, we also design a corner consistency loss and an elastic mesh feature loss to better improve geometric alignment. Extensive experiments show that our DA-Font outperforms the state-of-the-art methods across diverse font styles and characters, demonstrating its effectiveness in enhancing structural integrity and local fidelity. The source code can be found at https://github.com/wrchen2001/DA-Font. Weiran Chen 0001, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
ACM Multimedia | 3 |
| 2025 | SiamHCC: a novel siamese network for quality evaluation of handwritten Chinese characters
Weiran Chen 0001, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
Multim. Syst. | 3 |
| 2024 | TARN-VIST: Topic Aware Reinforcement Network for Visual StorytellingabstractAs a cross-modal task, visual storytelling aims to generate a story for an ordered image sequence automatically. Different from the image captioning task, visual storytelling requires not only modeling the relationships between objects in the image but also mining the connections between adjacent images. Recent approaches primarily utilize either end-to-end frameworks or multi-stage frameworks to generate relevant stories, but they usually overlook latent topic information. In this paper, in order to generate a more coherent and relevant story, we propose a novel method, Topic Aware Reinforcement Network for VIsual StoryTelling (TARN-VIST). In particular, we pre-extracted the topic information of stories from both visual and linguistic perspectives. Then we apply two topic-consistent reinforcement learning rewards to identify the discrepancy between the generated story and the human-labeled story so as to refine the whole generation process. Extensive experimental results on the VIST dataset and human evaluation demonstrate that our proposed model outperforms most of the competitive models across multiple evaluation metrics. Weiran Chen 0001, Jiaqi Su, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
LREC/COLING | 5 |
| 2024 | MAGIC: Multi-prompt Any Length Video Generation Model with Controllable Inter-frame Correlation and Low Barrier
Weiran Chen 0001, Lingbing Xu, Yi Ji 0001, Ying Li 0065, Chunping Liu |
ICANN (3) | 6 |
| 2024 | Uncertainty-Aware with Negative Samples for Video-Text Retrieval
Weiran Chen 0001, Yi Ji 0001, Ying Li 0065, Chunping Liu |
PRCV (5) | 5 |
| 2024 | Quality evaluation methods of handwritten Chinese characters: a comprehensive survey
Weiran Chen 0001, Jiaqi Su, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
Multim. Syst. | 6 |
| 2024 | Gazing After Glancing: Edge Information Guided Perception Network for Video Moment RetrievalabstractVideo Moment Retrieval (VMR) is a challenging task aimed at locating video segments in untrimmed videos through semantic matching of the given queries. Due to the fact that most existing methods neglect the valuable clues of edge information, it is difficult to precisely pinpoint the target segment as the target moment is complex. To this end, this paper proposes a novel perception network,GazingAfterGlancing(GAG), to utilize edge information. Inspired by human reading habits, we propose a localization strategy of glancing and gazing, and using this strategy, we divide the proposed VMR task with the perceptual network into two stages, “glancing” and “gazing”. The glancing stage utilizes a commonly used coarse-grained feature encoder and an edge-guided span predictor to locate the approximate area. The gazing stage leverages the edge information extracted from the result of “glancing” to recalibrate the query feature. Specifically, we propose an edge-guided highlighting block to recalibrate the encoded query feature according to the visual edge semantic information. Then the refined query feature and visual feature are utilized by the edge-guided span predictor. Moreover, we employ the distillation to enhance the ability of the coarse-grained feature encoder. Experimental results on two widely used ActivityNet Captions and TACoS datasets show that the proposed edge information guided two-stage VMR method effectively improves the localization accuracy. Zhanghao Huang, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IEEE Signal Process. Lett. | 3 |
| 2023 | Regional Consistency for Semi-Supervised Segmentation of 3D Medical ImagesabstractSemi-supervised medical image segmentation (SSMIS)is a research hotspot.However,existing consistency regular-ization methods do not adequately consider the robustness gains obtained by the model overcoming perturbations in the network structure and the spatial context. To address this problem, we propose a regional consistency strategy for SSMIS. Specifically, we construct network structure perturbations by making the two networks use different downsampling strategies. As for spatial contextual perturbations, we perform two random crops for each 3D medical image and feed different sub-image to different networks. We introduce entropy minimization to encourage both networks to produce consistent, high-confidence predictions for intersecting regions. A weighted combination of supervised and unsupervised losses optimizes the networks. We conducted extensive experiments on two datasets, and the results show that introducing network structure perturbations and spatial environmental perturbations can improve various metrics and demonstrate the effectiveness of our method Shidi Liu, Chunping Liu, Yi Ji 0001, Ying Li 0065 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Semi-supervised Domain Adaptation for Dependency Parsing with Dynamic Matching NetworkabstractSupervised parsing models have achieved impressive results on in-domain texts.However, their performances drop drastically on out-ofdomain texts due to the data distribution shift.The shared-private model has shown its promising advantages for alleviating this problem via feature separation, whereas prior works pay more attention to enhancing shared features but neglect the in-depth relevance of specific ones.To address this issue, we for the first time apply a dynamic matching network on the shared-private model for semi-supervised crossdomain dependency parsing.Meanwhile, considering the scarcity of target-domain labeled data, we leverage unlabeled data from two aspects, i.e., designing a new training strategy to improve the capability of the dynamic matching network and fine-tuning BERT to obtain domain-related contextualized representations.Experiments on benchmark datasets show that our proposed model consistently outperforms various baselines, leading to new state-of-theart results on all domains.Detailed analysis on different matching strategies demonstrates that it is essential to learn suitable matching weights to emphasize useful features and ignore useless or even harmful ones.Besides, our proposed model can be directly extended to multi-source domain adaptation and achieves best performances among various baselines, further verifying the effectiveness and robustness. Ying Li 0065, Shuaike Li, Min Zhang 0005 |
ACL (1) | 1 |
| 2022 | Parallel Data Augmentation for Text-based Person Re-identificationabstractGiven textual descriptions, text-based person reidentification aims at retrieving the matched target person in a large-scale image pool. In contrast to the traditional person re-identification (Re-ID) task, text-based person Re-ID requires extra extracted discriminative textual representations and then aligns two modal features to narrow down the semantic gap between linguistic domain and visual domain. A majority of previous works design complex network structures and concatenate multi-branch features while failing to pay much attention to problems with the dataset, which requires more parameters learning and might lead to over-fitting. Hence, in this paper, we propose a Parallel Data Augmentation method (PDA) to reduce over-fitting and make the model occlusion resistant without increasing the number of training parameters. Specifically, prior to the training, for an image, we randomly choose a rectangular region of variable size and erase the region with a constant value. Similar to image processing, we randomly add a mask of random length words to a sentence, then the processed data is sent to the TIPCB framework for training. Extensive experimentations on the large-scale CUHK-PEDES dataset show the effectiveness of our method and verify that our method exceeds the state-of-the-art methods. Hanqing Cai, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 4 |
| 2022 | Attention-based Neighbor Selective Aggregation Network for Camouflaged Object DetectionabstractCamouflaged Object Detection (COD) aims to discover objects that are finely disguised in the environment. Its challenge is that the targets generally have similar textures and colors to their surroundings. In this paper, we propose a novel network, named attention-based neighbor selective aggregation network (ANSA-Net), which can effectively and efficiently detect camouflaged objects. Specifically, our ANSA-Net contains two novel modules, namely, neighbor selective aggregation (NSA) and high-level feature guided attention (HLGA). The NSA is designed to locate concealed targets by fusing multi-scale features adaptively. Furthermore, the HLGA is designed to improve the semantic information of features by employing attention maps derived from high-level features. Experiments show that ANSA-Net exhibits relatively accurate detection performance on four COD datasets, outperforming existing state-of-the-art methods. Hao-Zhou Hao, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 4 |
| 2022 | Multi-enhanced Adaptive Attention Network for RGB- T Salient Object DetectionabstractNowadays, Salient object detection (SOD) on RGB images has achieved remarkable success. However, the performance of this single-modal SOD will be considerably reduced when faced with complicated situations. To deal with these challenges, the fusion of RGB and thermal infrared images, termed as RGB- T SOD, becomes a new SOD research direction recently. Thermal images can supply the essential additional information to RGB because they are immune to illumination and weather conditions. Though in this area, existing methods don't take full advantage of multi-level encoded features to generate global context. In addition, these approaches feed unprocessed encoded features that contain interference such as background directly to the decoder and don't explicitly establish the correlation of two heterogeneous modalities. In this paper, we proposed a multi-enhanced adaptive attention network (MEAANet) to solve aforementioned problems. Specifically, we use a multi-modal multi-level feature fusion (MMFF) module to fuse low-level and high-level encoded features to enhance the global context. Then, we design the thermal adaptive attention module (TAAM) to enhance encoded features while reducing noise interference. Moreover, to explore the correlation between the two modalities, we utilize the cross-enhanced integration module (CIM) to learn the shared features of two modalities. The comprehensive experimental results demonstrate the effectiveness of our proposed approach against the state-of-the-art. Hao-Zhou Hao, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 4 |
| 2022 | Spatio-Temporal Graph-based Semantic Compositional Network for Video CaptioningabstractVideo Captioning aims to generate natural language descriptions for given videos and is one of the challenging problems in computer vision's high-level understanding tasks. Existing methods are relatively lacking in the mining of object-level spatio-temporal relationships, which is important for generating captions with accurate object information. In this paper, we improve the existing SCN-LSTM method from the perspective of modeling spatio-temporal relationships and propose the Spatio-Temporal Graph-based Semantic Compositional Network for Video Captioning (STG-SCN). In terms of spatial-temporal relationships modeling, we propose the Spatial Relation Graph (SRG) and the Temporal Relation Graph (TRG) based on the Graph Attention Network, respectively. SRG is employed to establish the spatial relationships between spatially Neighboring objects within each keyframe conditioned on their correlation with the current keyframe. TRG is used to model the temporal relationship between all the objects at different time steps and incorporates the object-level information into frame-level features. Based on the proposed Semantics Guided Decoder, visual representations enhanced by object-level information are dynamically fused with high-level semantic concepts to generate captions that not only consider the global visual content but also have stronger language expressiveness. Extended experiments show that our proposed method achieves significant performance gains on Microsoft Video Description (MSVD) and Microsoft Research Video-to-Text (MSR-VTT) datasets, outperforming existing methods. Zefan Zhang, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 4 |
| 2022 | Semantic Image Synthesis via Hierarchical Structure FeaturesabstractSemantic image synthesis, which converts semantic masks into photo-realistic images, is essentially a special form of a label-to-image task. In this area, previous work has made great progress, but we found that their models usually lose certain semantic information during the generation process, and the metrics of each generated result have a certain degree of fluctuation. So how to generate stable and high-quality images is still a challenge for this task. In this paper, we propose a Hierarchical Feature Block (HF-Block) from the perspective of improving the stability of generation. It generates different hierarchical features through a Hierarchical Feature Encoder (HF-Encoder) and merges them into the generator. We conducted extensive experiments on several very challenging datasets: ADE20K, Deepfashion, and Deepfashion2 datasets. Compared with the state-of-the-art methods, ours can provide more stable and high-quality images. Jun-Jie Tao, Guo-Ying Zhu, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 4 |
| 2022 | Selective and Representative Sequence Feature Alignment for Domain Adaptive Detection TransformerabstractRecently, several studies have applied the Unsupervised Domain Adaptation (UDA) method on detection transformers to improve their cross-domain detection performance. However, the majority of them directly apply adversarial alignment on expatiatory token sequences, which will introduce too much background information and disturb the alignment process. To tackle the problem, we propose a domain adaption method focused on the detection transformer named selective and representative sequence feature alignment (SR-SFA). Specifically, our SR-SFA contains two modules: self-guided weight map generation module (SWG) and classification-guided domain query generation module (CQG). The SWG module takes full advantage of transformer detection capability to locate the foreground parts of the token sequences for local alignment. The other CQG module introduces an image-level multi-label classification task as an auxiliary task to capture the representative information of the whole image for global level alignment. Therefore, more effective feature alignment is performed in a local and global fashion. Experiments on two adaptation scenarios demonstrate our method gets better performance compared with other approaches. Zhi-Yuan Yang, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 3 |
| 2022 | Scene graph generation with award-punishment strategy
Haiyan Gao, Dibo Shi, Tianling Jiang, Zefan Zhang, Yi Ji 0001, Ying Li 0065, Chunping Liu |
Knowl. Based Syst. | 7 |
| 2021 | Video Captioning with External Knowledge Assistance and Multi-feature Fusion
Jiao-Wei Miao, Huan Shao 0001, Yi Ji 0001, Ying Li 0065, Chunping Liu |
ICONIP (6) | 4 |
| 2021 | Do We Really Reduce Bias for Scene Graph Generation?abstractFor a given image, the corresponding scene graph is a kind of structural expression which benefits to high-level tasks. To generate a meaningful and useful one, the existing models pay more attention on reducing the bias from long-tail distribution of dataset. However, they overlook the unimodal bias and evaluation bias from models themselves. In this paper, we construct an unbiased solution called Balanced Label and Vision for Multilabel Classification (BLVMC). BLVMC consists of two modules, label-vision grounding module (LVGM) and no graph constraint (NGC). Specially, the LVGM aims to be in equilibrium for label and vision by introducing visual information into label branch. This module reduces unimodal bias from previous models and makes them more stable. The NGC views the Scene Graph Generation (SGG) as a multilabel classification task instead of multiclass classification. Besides, the NGC uses the corresponding NGC mR@K to evaluate models. This module allows each subject-object pair to retain multi-predicates, which relieves evaluation bias. The quantitative and qualitative experiments on Visual Genome (VG) dataset demonstrate the proposed BLVMC effectively eliminates the above two biases and outperforms previous state-of-the-art models. Haiyan Gao, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 4 |
| 2021 | BDFPN: Bi-Direction Feature Pyramid Network for Scene Text DetectionabstractScene text detection in the natural environment is widely used in real-world applications, ranging from autonomous driving, image search and assistance for the blind. However, a vast of the existing methods have limited ability to detect text instances in challenging scenes such as texts with low contrast or blur. To address the problem, we propose a novel Bi-Direction Feature Pyramid Network (BDFPN), which draws inspiration from the two-way visual information processing mechanism of human beings. Specifically, the bottom-up path is data-driven for fine details and the top-down path is task-driven for obtaining semantic information. In the top-down path, the Feature Alignment Module (FAM) is proposed to narrow the semantic differences that exist in features of adjacent levels. To combine features from two paths, we propose a novel fusion strategy named Attention Fusion Module (AFM). We conduct extensive experiments on ICDAR2015, Totaltext and MSRA-TD500 to demonstrate the effectiveness and robustness of BDFPN. Hailin Shao, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 3 |
| 2021 | Deep Content Guidance Network for Arbitrary Style TransferabstractArbitrary style transfer refers to generate a new image based on any set of existing images. Meanwhile, the generated image retains the content structure of one and the style pattern of another. In terms of content retention and style transfer, the recent arbitrary style transfer algorithms normally perform well in one, but it is difficult to find a trade-off between the two. In this paper, we propose the Deep Content Guidance Network (DCGN) which is stacked by content guidance (CG) layers. And each CG layer involves one position self-attention (pSA) module, one channel self-attention (cSA) module and one content guidance attention (cGA) module. Specially, the pSA module extracts more effective content information on the spatial layout of content images and the cSA module makes the style representation of style images in the channel dimension richer. And in the non-local view, the cGA module utilizes content information to guide the distribution of style features, which obtains a more detailed style expression. Moreover, we introduce a new permutation loss to generalize feature expression, so as to obtain abundant feature expressions while maintaining content structure. Qualitative and quantitative experiments verify that our approach can transform into better stylized images than the state-of-the-art methods. Dibo Shi, Huan Xie 0007, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 4 |
| 2020 | Semi-supervised Domain Adaptation for Dependency Parsing via Improved Contextualized Word RepresentationsabstractIn recent years, parsing performance is dramatically improved on in-domain texts thanks to the rapid progress of deep neural network models.The major challenge for current parsing research is to improve parsing performance on out-of-domain texts that are very different from the indomain training data when there is only a small-scale out-domain labeled data.To deal with this problem, we propose to improve the contextualized word representations via adversarial learning and fine-tuning BERT processes.Concretely, we apply adversarial learning to three representative semi-supervised domain adaption methods, i.e., direct concatenation (CON), feature augmentation (FA), and domain embedding (DE) with two useful strategies, i.e., fused targetdomain word representations and orthogonality constraints, thus enabling to model more pure yet effective domain-specific and domain-invariant representations.Simultaneously, we utilize a large-scale target-domain unlabeled data to fine-tune BERT with only the language model loss, thus obtaining reliable contextualized word representations that benefit for the cross-domain dependency parsing.Experiments on a benchmark dataset show that our proposed adversarial approaches achieve consistent improvements, and fine-tuning BERT further boosts the parsing accuracy by a large margin.Our single model achieves the same state-of-the-art performance as the top submitted system in the NLPCC-2019 shared task, which uses ensemble models and BERT. Ying Li 0065, Zhenghua Li, Min Zhang 0005 |
COLING | 1 |
| 2019 | Self-attentive Biaffine Dependency ParsingabstractThe current state-of-the-art dependency parsing approaches employ BiLSTMs to encode input sentences.Motivated by the success of the transformer-based machine translation, this work for the first time applies the self-attention mechanism to dependency parsing as the replacement of the BiLSTM-based encoders, leading to competitive performance on both English and Chinese benchmark data. Based on the detailed error analysis, we then combine the power of both BiLSTM and self-attention via model ensembles, demonstrating their complementary capability of capturing contextual information. Finally, we explore the recently proposed contextualized word representations as extra input features, and further improve the parsing performance. Ying Li 0065, Zhenghua Li, Min Zhang 0005, Rui Wang 0005, Sheng Li 0017, Luo Si |
IJCAI | 1 |