EDBT 2026 Demo / reviewers in the wild / expert
Yi Ji 0001
dblp:16/4272-1
· DBLP profile ↗
60ranked-venue papers
3as first author
35since 2021 · last 2025
0000-0001-6965-4158ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Navigating the Unseen: Zero-shot Scene Graph Generation via Capsule-Based Equivariant FeaturesabstractIn scene graph generation (SGG), the accurate prediction of unseen triplets is essential for its effectiveness in downstream vision-language tasks. We hypothesize that the predicates of unseen triplets can be viewed as transformations of seen predicates in feature space, and the essence of the zero-shot task is to bridge the gap caused by this transformation. Traditional models, however, have difficulty addressing this challenge, which we attribute to their inability to model the predicates equivariant. To overcome this limitation, we introduce a novel framework based on capsule networks (CAPSGG). We propose a Three-Stream Pipeline that generates modality-specific representations for predicates, while building low-level predicate capsules of these modalities. Then, these capsules are aggregated into high-level predicate capsules using a Routing Capsule Layer. In addition, we introduce GroupLoss to aggregate capsules with the same predicate label into groups. This replaces the global loss with the intra-group loss, effectively balancing the learning of predicate invariant and equivariant features while mitigating the impact of the severe long-tail distribution of the predicate categories. Our extensive experiments demonstrate the notable superiority of our approach over state-of-the-art methods, with zero-shot indicators outperforming up to 132.26% on the SGCls task than the T-CAR [21]. Our code will be available upon publication. Wenhuan Huang, Yi Ji 0001, Guiqian Zhu, Li Ying, Chunping Liu |
CVPR | 2 |
| 2025 | MedKI: Knowledge Dual Injections for Medical Visual Question AnsweringabstractMedical Visual Question Answering (Med VQA) is a challenging task for the sake of diverse medical image and multidisciplinary knowledge. Nowadays, the visual and language pretraining-finetuning framework is widely used in Med VQA task. However, most methods neglect the potential semantics and clinical information of image-text pairs, resulting in an inability to accurately match question semantics with image information. To address this, we propose a method called MedKI with dual injections of clinical and semantic knowledge, which is based on the pretraining and finetuning framework. Specifically, during pretraining, we inject clinical knowledge into the alignment module. Here, clinical knowledge is composed of the structural and the conceptual features that are extracted from the graph structure and entity definitions of the expert domain knowledge graph, respectively. In the finetuning stage, we retrieve similar texts from the pretraining corpus and encode them as semantic knowledge. Then, the knowledge is injected into the semantic knowledge fusion module. Extensive experimental results on both VQA-RAD dataset and SLAKE dataset demonstrate the validity of our proposed method. Hongyi Ren, Weiran Chen 0001, Chunping Liu, Yi Ji 0001, Ying Li 0065 |
ICIP | 4 |
| 2025 | DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid IntegrationabstractFew-shot font generation aims to create new fonts with a limited number of glyph references. It can be used to significantly reduce the labor cost of manual font design. However, due to the variety and complexity of font styles, the results generated by existing methods often suffer from visible defects, such as stroke errors, artifacts and blurriness. To address these issues, we propose DA-Font, a novel framework which integrates a Dual-Attention Hybrid Module (DAHM). Specifically, we introduce two synergistic attention blocks: the component attention block that leverages component information from content images to guide the style transfer process, and the relation attention block that further refines spatial relationships through interacting the content feature with both original and stylized component-wise representations. These two blocks collaborate to preserve accurate character shapes and stylistic textures. Moreover, we also design a corner consistency loss and an elastic mesh feature loss to better improve geometric alignment. Extensive experiments show that our DA-Font outperforms the state-of-the-art methods across diverse font styles and characters, demonstrating its effectiveness in enhancing structural integrity and local fidelity. The source code can be found at https://github.com/wrchen2001/DA-Font. Weiran Chen 0001, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
ACM Multimedia | 4 |
| 2025 | SiamHCC: a novel siamese network for quality evaluation of handwritten Chinese characters
Weiran Chen 0001, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
Multim. Syst. | 4 |
| 2024 | TARN-VIST: Topic Aware Reinforcement Network for Visual StorytellingabstractAs a cross-modal task, visual storytelling aims to generate a story for an ordered image sequence automatically. Different from the image captioning task, visual storytelling requires not only modeling the relationships between objects in the image but also mining the connections between adjacent images. Recent approaches primarily utilize either end-to-end frameworks or multi-stage frameworks to generate relevant stories, but they usually overlook latent topic information. In this paper, in order to generate a more coherent and relevant story, we propose a novel method, Topic Aware Reinforcement Network for VIsual StoryTelling (TARN-VIST). In particular, we pre-extracted the topic information of stories from both visual and linguistic perspectives. Then we apply two topic-consistent reinforcement learning rewards to identify the discrepancy between the generated story and the human-labeled story so as to refine the whole generation process. Extensive experimental results on the VIST dataset and human evaluation demonstrate that our proposed model outperforms most of the competitive models across multiple evaluation metrics. Weiran Chen 0001, Jiaqi Su, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
LREC/COLING | 6 |
| 2024 | MAGIC: Multi-prompt Any Length Video Generation Model with Controllable Inter-frame Correlation and Low Barrier
Weiran Chen 0001, Lingbing Xu, Yi Ji 0001, Ying Li 0065, Chunping Liu |
ICANN (3) | 5 |
| 2024 | Glocal Cascading Network for Topic Enhanced Visual StorytellingabstractAs a cross-modal task, visual storytelling aims to generate a semantically coherent story for an ordered image sequence. Despite significant achievements in existing methods for this task, few works focus on improving the conception ability which humans usually use when writing stories. In this work, we propose a framework called GLocal Cascading Network for Topic Enhanced Visual Storytelling which explores the conception ability by pre-modeling a latent topic for each image during story telling. Inspired by the global-local (glocal) ideology, we firstly propose a hierarchical latent-topic decoder consisting of two levels of topic generator which respectively focus on different levels of topic information. Then we propose a topic-aware loss which encourages the model to focus on the topic information of the story. With these two novel modules, our framework can effectively utilize the topic information and improve the informativeness and consistency of stories. Our model has been proven highly competitive across multiple metrics through extensive experiments conducted on the VIST dataset. Jiaqi Su, Weiran Chen 0001, Yi Ji 0001, Chunping Liu |
ICASSP | 3 |
| 2024 | Transition in Focus of Prediction Tasks for Skeleton Graph Component Detection with TransformerabstractRecent advancements in skeleton extraction have significantly improved the process by simplifying the skeleton regression task into graph component detection. Despite the advancements in skeleton topology, accuracy in detailing skeletal parts remains challenging, with specific issues such as jagged edges in high-resolution images. This paper identifies the limitations of current detection models that can adapt during the decomposition and reconstruction phases, which impacts the overall precision of the extraction. In response, we propose an approach that revises the primary focus of the detection tasks. Inspired by the success of pixel-wise binary classification methods, we propose a gradual transition in focus from a coordinate localization regression task to a classification task of predicting points during the training process. This transition can be achieved by adjusting the number of object queries in the Transformer model. Theoretical and experimental evaluations validate the effectiveness of our approach. Our method yields significant improvements in performance over the baseline across various shape and image datasets (e.g., 0.836 vs. 0.826 for BlumNet on the SK1491 dataset). Zeyd Boukhers, Wei Sui, Yi Ji 0001, Chunping Liu |
MMAsia | 6 |
| 2024 | Uncertainty-Aware with Negative Samples for Video-Text Retrieval
Weiran Chen 0001, Yi Ji 0001, Ying Li 0065, Chunping Liu |
PRCV (5) | 4 |
| 2024 | SPARK: Cross-Guided Knowledge Distillation with Spatial Position Augmentation for Medical Image Segmentation
Lingbing Xu, Yi Ji 0001, Chunping Liu |
PRCV (14) | 4 |
| 2024 | Quality evaluation methods of handwritten Chinese characters: a comprehensive survey
Weiran Chen 0001, Jiaqi Su, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu |
Multim. Syst. | 7 |
| 2024 | Gazing After Glancing: Edge Information Guided Perception Network for Video Moment RetrievalabstractVideo Moment Retrieval (VMR) is a challenging task aimed at locating video segments in untrimmed videos through semantic matching of the given queries. Due to the fact that most existing methods neglect the valuable clues of edge information, it is difficult to precisely pinpoint the target segment as the target moment is complex. To this end, this paper proposes a novel perception network,GazingAfterGlancing(GAG), to utilize edge information. Inspired by human reading habits, we propose a localization strategy of glancing and gazing, and using this strategy, we divide the proposed VMR task with the perceptual network into two stages, “glancing” and “gazing”. The glancing stage utilizes a commonly used coarse-grained feature encoder and an edge-guided span predictor to locate the approximate area. The gazing stage leverages the edge information extracted from the result of “glancing” to recalibrate the query feature. Specifically, we propose an edge-guided highlighting block to recalibrate the encoded query feature according to the visual edge semantic information. Then the refined query feature and visual feature are utilized by the edge-guided span predictor. Moreover, we employ the distillation to enhance the ability of the coarse-grained feature encoder. Experimental results on two widely used ActivityNet Captions and TACoS datasets show that the proposed edge information guided two-stage VMR method effectively improves the localization accuracy. Zhanghao Huang, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IEEE Signal Process. Lett. | 2 |
| 2023 | Associative Learning Network for Coherent Visual StorytellingabstractVisual storytelling task aims to generate relevant and coherent story for an ordered stream of images. Although visual storytelling methods have made promising improvement in recent years, existing methods pay little attention to the association ability and divergent thinking of the model, which are essential for humanistic stories. This paper introduces a novel Associative Learning Network for Coherent Visual Storytelling to explore the model’s association ability while telling a new story. Specifically, we first build a graph based on the pointwise mutual information and learn association degree of word pairs with Graph Convolutional Network. Besides, an auxiliary hierarchical decoder is designed to combine the words together to generate coherent story. In this way, our model can recall information using associative mem-ory, enhancing the coherence and informativeness of the generated story. Extensive experiments on VIST dataset demonstrate that the proposed framework substantially outperforms the state-of-the-art methods across multiple evaluation metrics. Chunping Liu, Yi Ji 0001 |
ICASSP | 3 |
| 2023 | Knowledge-Aware Causal Inference Network for Visual DialogabstractThe effective knowledge and interaction within multi-modalities are key to Visual Dialog. Classic graph-based framework with the direct connection between history dialog and answer fails to give the right answer for the spurious guidance and strong bias induced from history dialog. Recent causal inference framework without this direct connection improves the generalization while worse accuracy. In this work, we propose a novel Knowledge-Aware Causal Inference framework(KACI-Net) in which the commonsense knowledge is introduced into the causal inference framework to achieve both high accuracy and generalization. Specifically, the commonsense knowledge is first generated according to the entities extracted from the question and fused with language and visual features with the co-attention to get the final answer. Comparisons with knowledge-unaware framework and graph-based knowledge-aware framework on VisDial v1.0 dataset show the superiority of our proposed framework and verify the effectiveness the usage of the commonsense knowledge for a good reasoning in Visual Dialog. Both high NDCG and MRR metrics indicate a good trade-off between accuracy and generalization. Zefan Zhang, Yi Ji 0001, Chunping Liu |
ICMR | 2 |
| 2023 | Infer unseen from seen: Relation regularized zero-shot visual dialog
Zefan Zhang, Yi Ji 0001, Chunping Liu |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | Multi-view semantic understanding for visual dialog
Tianling Jiang, Zefan Zhang, Yi Ji 0001, Chunping Liu |
Knowl. Based Syst. | 4 |
| 2023 | Regional Consistency for Semi-Supervised Segmentation of 3D Medical ImagesabstractSemi-supervised medical image segmentation (SSMIS)is a research hotspot.However,existing consistency regular-ization methods do not adequately consider the robustness gains obtained by the model overcoming perturbations in the network structure and the spatial context. To address this problem, we propose a regional consistency strategy for SSMIS. Specifically, we construct network structure perturbations by making the two networks use different downsampling strategies. As for spatial contextual perturbations, we perform two random crops for each 3D medical image and feed different sub-image to different networks. We introduce entropy minimization to encourage both networks to produce consistent, high-confidence predictions for intersecting regions. A weighted combination of supervised and unsupervised losses optimizes the networks. We conducted extensive experiments on two datasets, and the results show that introducing network structure perturbations and spatial environmental perturbations can improve various metrics and demonstrate the effectiveness of our method Shidi Liu, Chunping Liu, Yi Ji 0001, Ying Li 0065 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Chinese Character Style Transfer Model Based on Convolutional Neural Network
Weiran Chen 0001, Chunping Liu, Yi Ji 0001 |
ICANN (4) | 3 |
| 2022 | Emotion Aware Reinforcement Network for Visual Storytelling
Hanqing Cai, Tianling Jiang, Chunping Liu, Yi Ji 0001 |
ICANN (2) | 5 |
| 2022 | Coupling Attention and Convolution for Heuristic Network in Visual DialogabstractVisual Dialog is a typical AI-agent task on images, in which the agent interprets information from heterogeneous modalities and provides the correct answer. In this area, most approaches are based on the attention mechanism. When the agent enjoys the large-capacity advantage of attention, the lack of the right inductive bias compared with convolution hinders its success. Therefore, in order to utilize their advantages and compensate for their respective shortcomings, inspired by the paraventricular thalamus (PVT) in the brain, we couple convolution and attention, termed as Attention Convolution Enhanced (ACE) method to enhance the agent’s activation of key features and strengthening the semantic understanding of visual and textual data. Meanwhile, we propose Heuristic Adjustment (HA) module to globally strengthen the agent’s semantic understanding and reduce language bias that is easy to occur after using the enhanced features. Finally, we concatenate the ACE and the HA in our Coupling Attention and Convolution for Heuristic Network (CACH-Net) to train the agent for better semantic comprehension and generalization ability. Extensive experiments on the VisDial v1.0 benchmark show that our CACH-Net has a better performance. Zefan Zhang, Tianling Jiang, Chunping Liu, Yi Ji 0001 |
ICIP | 4 |
| 2022 | Parallel Data Augmentation for Text-based Person Re-identificationabstractGiven textual descriptions, text-based person reidentification aims at retrieving the matched target person in a large-scale image pool. In contrast to the traditional person re-identification (Re-ID) task, text-based person Re-ID requires extra extracted discriminative textual representations and then aligns two modal features to narrow down the semantic gap between linguistic domain and visual domain. A majority of previous works design complex network structures and concatenate multi-branch features while failing to pay much attention to problems with the dataset, which requires more parameters learning and might lead to over-fitting. Hence, in this paper, we propose a Parallel Data Augmentation method (PDA) to reduce over-fitting and make the model occlusion resistant without increasing the number of training parameters. Specifically, prior to the training, for an image, we randomly choose a rectangular region of variable size and erase the region with a constant value. Similar to image processing, we randomly add a mask of random length words to a sentence, then the processed data is sent to the TIPCB framework for training. Extensive experimentations on the large-scale CUHK-PEDES dataset show the effectiveness of our method and verify that our method exceeds the state-of-the-art methods. Hanqing Cai, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 3 |
| 2022 | Attention-based Neighbor Selective Aggregation Network for Camouflaged Object DetectionabstractCamouflaged Object Detection (COD) aims to discover objects that are finely disguised in the environment. Its challenge is that the targets generally have similar textures and colors to their surroundings. In this paper, we propose a novel network, named attention-based neighbor selective aggregation network (ANSA-Net), which can effectively and efficiently detect camouflaged objects. Specifically, our ANSA-Net contains two novel modules, namely, neighbor selective aggregation (NSA) and high-level feature guided attention (HLGA). The NSA is designed to locate concealed targets by fusing multi-scale features adaptively. Furthermore, the HLGA is designed to improve the semantic information of features by employing attention maps derived from high-level features. Experiments show that ANSA-Net exhibits relatively accurate detection performance on four COD datasets, outperforming existing state-of-the-art methods. Hao-Zhou Hao, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 3 |
| 2022 | Multi-enhanced Adaptive Attention Network for RGB- T Salient Object DetectionabstractNowadays, Salient object detection (SOD) on RGB images has achieved remarkable success. However, the performance of this single-modal SOD will be considerably reduced when faced with complicated situations. To deal with these challenges, the fusion of RGB and thermal infrared images, termed as RGB- T SOD, becomes a new SOD research direction recently. Thermal images can supply the essential additional information to RGB because they are immune to illumination and weather conditions. Though in this area, existing methods don't take full advantage of multi-level encoded features to generate global context. In addition, these approaches feed unprocessed encoded features that contain interference such as background directly to the decoder and don't explicitly establish the correlation of two heterogeneous modalities. In this paper, we proposed a multi-enhanced adaptive attention network (MEAANet) to solve aforementioned problems. Specifically, we use a multi-modal multi-level feature fusion (MMFF) module to fuse low-level and high-level encoded features to enhance the global context. Then, we design the thermal adaptive attention module (TAAM) to enhance encoded features while reducing noise interference. Moreover, to explore the correlation between the two modalities, we utilize the cross-enhanced integration module (CIM) to learn the shared features of two modalities. The comprehensive experimental results demonstrate the effectiveness of our proposed approach against the state-of-the-art. Hao-Zhou Hao, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 3 |
| 2022 | Spatio-Temporal Graph-based Semantic Compositional Network for Video CaptioningabstractVideo Captioning aims to generate natural language descriptions for given videos and is one of the challenging problems in computer vision's high-level understanding tasks. Existing methods are relatively lacking in the mining of object-level spatio-temporal relationships, which is important for generating captions with accurate object information. In this paper, we improve the existing SCN-LSTM method from the perspective of modeling spatio-temporal relationships and propose the Spatio-Temporal Graph-based Semantic Compositional Network for Video Captioning (STG-SCN). In terms of spatial-temporal relationships modeling, we propose the Spatial Relation Graph (SRG) and the Temporal Relation Graph (TRG) based on the Graph Attention Network, respectively. SRG is employed to establish the spatial relationships between spatially Neighboring objects within each keyframe conditioned on their correlation with the current keyframe. TRG is used to model the temporal relationship between all the objects at different time steps and incorporates the object-level information into frame-level features. Based on the proposed Semantics Guided Decoder, visual representations enhanced by object-level information are dynamically fused with high-level semantic concepts to generate captions that not only consider the global visual content but also have stronger language expressiveness. Extended experiments show that our proposed method achieves significant performance gains on Microsoft Video Description (MSVD) and Microsoft Research Video-to-Text (MSR-VTT) datasets, outperforming existing methods. Zefan Zhang, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 3 |
| 2022 | Semantic Image Synthesis via Hierarchical Structure FeaturesabstractSemantic image synthesis, which converts semantic masks into photo-realistic images, is essentially a special form of a label-to-image task. In this area, previous work has made great progress, but we found that their models usually lose certain semantic information during the generation process, and the metrics of each generated result have a certain degree of fluctuation. So how to generate stable and high-quality images is still a challenge for this task. In this paper, we propose a Hierarchical Feature Block (HF-Block) from the perspective of improving the stability of generation. It generates different hierarchical features through a Hierarchical Feature Encoder (HF-Encoder) and merges them into the generator. We conducted extensive experiments on several very challenging datasets: ADE20K, Deepfashion, and Deepfashion2 datasets. Compared with the state-of-the-art methods, ours can provide more stable and high-quality images. Jun-Jie Tao, Guo-Ying Zhu, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 3 |
| 2022 | Selective and Representative Sequence Feature Alignment for Domain Adaptive Detection TransformerabstractRecently, several studies have applied the Unsupervised Domain Adaptation (UDA) method on detection transformers to improve their cross-domain detection performance. However, the majority of them directly apply adversarial alignment on expatiatory token sequences, which will introduce too much background information and disturb the alignment process. To tackle the problem, we propose a domain adaption method focused on the detection transformer named selective and representative sequence feature alignment (SR-SFA). Specifically, our SR-SFA contains two modules: self-guided weight map generation module (SWG) and classification-guided domain query generation module (CQG). The SWG module takes full advantage of transformer detection capability to locate the foreground parts of the token sequences for local alignment. The other CQG module introduces an image-level multi-label classification task as an auxiliary task to capture the representative information of the whole image for global level alignment. Therefore, more effective feature alignment is performed in a local and global fashion. Experiments on two adaptation scenarios demonstrate our method gets better performance compared with other approaches. Zhi-Yuan Yang, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 2 |
| 2022 | Scene graph generation with award-punishment strategy
Haiyan Gao, Dibo Shi, Tianling Jiang, Zefan Zhang, Yi Ji 0001, Ying Li 0065, Chunping Liu |
Knowl. Based Syst. | 6 |
| 2021 | User-Guided Line Art Flat Filling With Split Filling MechanismabstractFlat filling is a critical step in digital artistic content creation with the objective of filling line arts with flat colors. We present a deep learning framework for user-guided line art flat filling that can compute the "influence areas" of the user color scribbles, i.e., the areas where the user scribbles should propagate and influence. This framework explicitly controls such scribble influence areas for artists to manipulate the colors of image details and avoid color leakage/contamination between scribbles, and simultaneously, leverages data-driven color generation to facilitate content creation. This framework is based on a Split Filling Mechanism (SFM), which first splits the user scribbles into individual groups and then independently processes the colors and influence areas of each group with a Convolutional Neural Network (CNN). Learned from more than a million illustrations, the framework can estimate the scribble influence areas in a content-aware manner, and can smartly generate visually pleasing colors to assist the daily works of artists. We show that our proposed framework is easy to use, allowing even amateurs to obtain professional-quality results on a wide variety of line arts. Lvmin Zhang, Chengze Li, Edgar Simo-Serra, Yi Ji 0001, Tien-Tsin Wong, Chunping Liu |
CVPR | 4 |
| 2021 | Generating Manga From Illustrations via Mimicking Manga Creation WorkflowabstractWe present a framework to generate manga from digital illustrations. In professional mange studios, the manga create workflow consists of three key steps: (1) Artists use line drawings to delineate the structural outlines in manga storyboards. (2) Artists apply several types of regular screentones to render the shading, occlusion, and object materials. (3) Artists selectively paste irregular screen textures onto the canvas to achieve various background layouts or special effects. Motivated by this workflow, we propose a data-driven framework to convert a digital illustration into three corresponding components: manga line drawing, regular screen-tone, and irregular screen texture. These components can be directly composed into manga images and can be further retouched for more plentiful manga creations. To this end, we create a large-scale dataset with these three components annotated by artists in a human-in-the-loop manner. We conduct both perceptual user study and qualitative evaluation of the generated manga, and observe that our generated image layers for these three components are practically usable in the daily works of manga artists. We provide 60 qualitative results and 15 additional comparisons in the supplementary material. We will make our presented manga dataset publicly available to assist related applications. Lvmin Zhang, Qingnan Fan, Yi Ji 0001, Chunping Liu |
CVPR | 4 |
| 2021 | SmartShadow: Artistic Shadow Drawing Tool for Line DrawingsabstractSmartShadow is a deep learning application for digital painting artists to draw shadows on line drawings, with three proposed tools. (1) Shadow brush: artists can draw scribbles to coarsely indicate the areas inside or outside their wanted shadows, and the application will generate the shadows in real-time. (2) Shadow boundary brush: this brush can precisely control the boundary of any specific shadow. (3) Global shadow generator: this tool can estimate the global shadow direction from input brush scribbles, and then consistently propagate local shadows to the entire image. These three tools can not only speed up the shadow drawing process (by 3.1× as experiments validate), but also allow for the flexibility to achieve various shadow effects and facilitate richer artistic creations. To this end, we train Convolutional Neural Networks (CNNs) with a collected large-scale dataset of both real and synthesized data, and especially, we collect 1670 shadow samples drawn by real artists. Both qualitative analysis and user study show that our approach can generate high-quality shadows that are practically usable in the daily works of digital painting artists. We present 30 additional results and 15 visual comparisons in the supplementary materiel. Lvmin Zhang, Jinyue Jiang, Yi Ji 0001, Chunping Liu |
ICCV | 3 |
| 2021 | Video Captioning with External Knowledge Assistance and Multi-feature Fusion
Jiao-Wei Miao, Huan Shao 0001, Yi Ji 0001, Ying Li 0065, Chunping Liu |
ICONIP (6) | 3 |
| 2021 | Do We Really Reduce Bias for Scene Graph Generation?abstractFor a given image, the corresponding scene graph is a kind of structural expression which benefits to high-level tasks. To generate a meaningful and useful one, the existing models pay more attention on reducing the bias from long-tail distribution of dataset. However, they overlook the unimodal bias and evaluation bias from models themselves. In this paper, we construct an unbiased solution called Balanced Label and Vision for Multilabel Classification (BLVMC). BLVMC consists of two modules, label-vision grounding module (LVGM) and no graph constraint (NGC). Specially, the LVGM aims to be in equilibrium for label and vision by introducing visual information into label branch. This module reduces unimodal bias from previous models and makes them more stable. The NGC views the Scene Graph Generation (SGG) as a multilabel classification task instead of multiclass classification. Besides, the NGC uses the corresponding NGC mR@K to evaluate models. This module allows each subject-object pair to retain multi-predicates, which relieves evaluation bias. The quantitative and qualitative experiments on Visual Genome (VG) dataset demonstrate the proposed BLVMC effectively eliminates the above two biases and outperforms previous state-of-the-art models. Haiyan Gao, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 3 |
| 2021 | BDFPN: Bi-Direction Feature Pyramid Network for Scene Text DetectionabstractScene text detection in the natural environment is widely used in real-world applications, ranging from autonomous driving, image search and assistance for the blind. However, a vast of the existing methods have limited ability to detect text instances in challenging scenes such as texts with low contrast or blur. To address the problem, we propose a novel Bi-Direction Feature Pyramid Network (BDFPN), which draws inspiration from the two-way visual information processing mechanism of human beings. Specifically, the bottom-up path is data-driven for fine details and the top-down path is task-driven for obtaining semantic information. In the top-down path, the Feature Alignment Module (FAM) is proposed to narrow the semantic differences that exist in features of adjacent levels. To combine features from two paths, we propose a novel fusion strategy named Attention Fusion Module (AFM). We conduct extensive experiments on ICDAR2015, Totaltext and MSRA-TD500 to demonstrate the effectiveness and robustness of BDFPN. Hailin Shao, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 2 |
| 2021 | Deep Content Guidance Network for Arbitrary Style TransferabstractArbitrary style transfer refers to generate a new image based on any set of existing images. Meanwhile, the generated image retains the content structure of one and the style pattern of another. In terms of content retention and style transfer, the recent arbitrary style transfer algorithms normally perform well in one, but it is difficult to find a trade-off between the two. In this paper, we propose the Deep Content Guidance Network (DCGN) which is stacked by content guidance (CG) layers. And each CG layer involves one position self-attention (pSA) module, one channel self-attention (cSA) module and one content guidance attention (cGA) module. Specially, the pSA module extracts more effective content information on the spatial layout of content images and the cSA module makes the style representation of style images in the channel dimension richer. And in the non-local view, the cGA module utilizes content information to guide the distribution of style features, which obtains a more detailed style expression. Moreover, we introduce a new permutation loss to generalize feature expression, so as to obtain abundant feature expressions while maintaining content structure. Qualitative and quantitative experiments verify that our approach can transform into better stylized images than the state-of-the-art methods. Dibo Shi, Huan Xie 0007, Yi Ji 0001, Ying Li 0065, Chunping Liu |
IJCNN | 3 |
| 2021 | Aligning vision-language for graph inference in visual dialog
Tianling Jiang, Hailin Shao, Yi Ji 0001, Chunping Liu |
Image Vis. Comput. | 4 |
| 2020 | Visual-Textual Alignment for Graph Inference in Visual DialogabstractAs a conversational intelligence task, visual dialog entails answering a series of questions grounded in an image, using the dialog history as context.To generate correct answers, the comprehension of the semantic dependencies among implicit visual and textual contents is critical.Prior works usually ignored the underlying relation and failed to infer it reasonably.In this paper, we propose a Visual-Textual Alignment for Graph Inference (VTAGI) network.Compared with other approaches, it makes up the lack of structural inference in visual dialog.The whole system consists of two modules, Visual and Textual Alignment (VTA) and Visual Graph Attended by Text (VGAT).Specially, the VTA module aims at representing an image with a set of integrated visual regions and corresponding textual concepts, reflecting certain semantics.The VGAT module views the visual features with semantic information as observed nodes and each node learns the relationship with others in visual graph.We also qualitatively and quantitatively evaluate the model on VisDial v1.0 dataset, showing our VTAGI outperforms previous state-of-the-art models. Tianling Jiang, Yi Ji 0001, Chunping Liu, Hailin Shao |
COLING | 2 |
| 2020 | DanbooRegion: An Illustration Region Dataset
Lvmin Zhang, Yi Ji 0001, Chunping Liu |
ECCV (13) | 2 |
| 2020 | Erasing Appearance Preservation in Optimization-Based Smoothing
Lvmin Zhang, Chengze Li, Yi Ji 0001, Chunping Liu, Tien-Tsin Wong |
ECCV (6) | 3 |
| 2020 | Selective Complementary Features For Multi-Person Pose EstimationabstractMulti-person pose estimation is a fundamental yet challenging research topic for many computer vision applications. It is difficult to achieve accurate localization results due to occlusion and complex background. In this paper, we propose a novel multi-person pose estimation approach with information complement and attention refinement residual module. To recover occlusion, the complementary features with multi-scale semantics information are extracted by our proposed Information Complement Module (ICM). To effectively discover the channel relationship and selectively highlight task-related regions in the feature maps, we design an Attention Refinement Residual Bottleneck (ARRB) module, which is an extension of residual unit with attention mechanism. We conduct ablation studies to investigate the efficacy of our method and compare it with the state-of-the-art methods on the COCO keypoint benchmark. Experimental results demonstrate that the selective complementary features are effective for multi-person pose estimation. Buwei Li, Yi Ji 0001, Jianyu Yang 0002, Chunping Liu |
ICIP | 3 |
| 2020 | Image Generation with the Enhanced Latent Code and Sub-pixel Sampling
Dibo Shi, Yi Ji 0001, Chunping Liu |
ICONIP (1) | 3 |
| 2020 | Integrating Historical States and Co-attention Mechanism for Visual DialogabstractVisual dialog is a typical multi-modal task which involves both vision and language. Nowadays, it faces two major difficulties. In this paper, we propose the Integrating Historical States and Co-attention (HSCA) for visual dialog to solve them. It includes two main modules, Co-ATT and MATCH. Specifically, the main purpose of the Co-ATT module is to guide the image with questions and answers in the early stage to get more specific objects. It tackles the first difficulty of the temporal sequence issue in historical information which may influence the precise answer for multi-round questions. The MATCH module is, based on a question with pronouns, to retrieve the best matching historical information block. It overcomes the second difficulty of the visual reference problem which requires to solve pronouns referring to unknowns in the text message and then to locate the objects in the given image. We quantitatively and qualitatively evaluate our model on VisDial v1.0, at the same time, ablation studies are carried out. The experimental results demonstrate that HSCA outperforms the state-of-the-art methods in many aspects. Tianling Jiang, Yi Ji 0001, Chunping Liu |
ICPR | 2 |
| 2020 | Generating Digital Painting Lighting Effects via RGB-space GeometryabstractWe present an algorithm to generate digital painting lighting effects from a single image. Our algorithm is based on a key observation: Artists use many overlapping strokes to paint lighting effects, i.e., pixels with dense stroke history tend to gather more illumination strokes. Based on this observation, we design an algorithm to both estimate the density of strokes in a digital painting using color geometry and then generate novel lighting effects by mimicking artists’ coarse-to-fine workflow. Coarse lighting effects are first generated using a wave transform and then retouched according to the stroke density of the original illustrations into usable lighting effects. Our algorithm is content-aware, with generated lighting effects naturally adapting to image structures, and can be used as an interactive tool to simplify current labor-intensive workflows for generating lighting effects for digital and matte paintings. In addition, our algorithm can also produce usable lighting effects for photographs or three-dimensional rendered images. We evaluate our approach with both an in-depth qualitative and a quantitative analysis that includes a perceptual user study. Results show that our proposed approach is not only able to produce favorable lighting effects with respect to existing approaches, but also that it is able to significantly reduce the needed interaction time. Lvmin Zhang, Edgar Simo-Serra, Yi Ji 0001, Chunping Liu |
ACM Trans. Graph. | 3 |
| 2020 | Spatio-Temporal Deep Residual Network with Hierarchical Attentions for Video Event RecognitionabstractEvent recognition in surveillance video has gained extensive attention from the computer vision community. This process still faces enormous challenges due to the tiny inter-class variations that are caused by various facets, such as severe occlusion, cluttered backgrounds, and so forth. To address these issues, we propose a spatio-temporal deep residual network with hierarchical attentions (STDRN-HA) for video event recognition. In the first attention layer, the ResNet fully connected feature guides the Faster R-CNN feature to generate object-based attention (O-attention) for target objects. In the second attention layer, the O-attention further guides the ResNet convolutional feature to yield the holistic attention (H-attention) in order to perceive more details of the occluded objects and the global background. In the third attention layer, the attention maps use the deep features to obtain the attention-enhanced features. Then, the attention-enhanced features are input into a deep residual recurrent network, which is used to mine more event clues from videos. Furthermore, an optimized loss function named softmax-RC is designed, which embeds the residual block regularization and center loss to solve the vanishing gradient in a deep network and enlarge the distance between inter-classes. We also build a temporal branch to exploit the long- and short-term motion information. The final results are obtained by fusing the outputs of the spatial and temporal streams. Experiments on the four realistic video datasets, CCV, VIRAT 1.0, VIRAT 2.0, and HMDB51, demonstrate that the proposed method has good performance and achieves state-of-the-art results. Chunping Liu, Yi Ji 0001, Shengrong Gong, Haibao Xu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2019 | Referring Expression Comprehension via Co-attention and Visual Context
Youming Gao, Yi Ji 0001, Chunping Liu |
ICANN (3) | 2 |
| 2019 | Intra-Modality Feature Interaction Using Self-attention for Visual Question Answering
Huan Shao 0001, Yi Ji 0001, Jianyu Yang 0002, Chunping Liu |
ICONIP (5) | 3 |
| 2019 | Context Gating with Short Temporal Information for Video Captioning
Jinlei Xu, Chunping Liu, Yi Ji 0001 |
IJCNN | 5 |
| 2019 | Trajectory-Pooled Spatial-Temporal Architecture of Deep Convolutional Neural Networks for Video Event DetectionabstractNowadays content-based video event detection faces great challenges due to complex scenes and blurred actions in surveillance videos. To alleviate these challenges, we propose a novel spatial-temporal architecture of deep convolutional neural networks for this task. By taking advantage of spatial-temporal information, we fine-tune two-stream networks, and then, fuse spatial and temporal features at convolution layers using a 2D pooling fusion method to enforce the consistence of spatial-temporal information. Based on the two-stream networks and spatial-temporal layer, a triple-channel model is obtained. Furthermore, we implement trajectory-constrained pooling to deep features and hand-crafted features to combine their merits. A fusion method on triple-channel yields the final detection result. The experiments on two benchmark surveillance video data sets including VIRAT 1.0 and VIRAT 2.0, which involve a suit of challenging events, such as person loading an object to a vehicle or person opening a vehicle trunk, manifest that the proposed method can achieve superior performance compared with the state-of-the-art methods on these event benchmarks. Rui Ge 0005, Yi Ji 0001, Shengrong Gong, Chunping Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Scene Graph Generation Based on Node-Relation Context Module
Chunping Liu, Yi Ji 0001, Jianyu Yang 0002 |
ICONIP (2) | 4 |
| 2018 | Person re-identification by enhanced local maximal occurrence representation and generalized similarity metric learning
Husheng Dong, Chunping Liu, Yi Ji 0001, Shengrong Gong |
Neurocomputing | 5 |
| 2018 | Person re-identification by kernel null space marginal Fisher analysis
Husheng Dong, Chunping Liu, Yi Ji 0001, Shengrong Gong |
Pattern Recognit. Lett. | 4 |
| 2018 | Two-stage sketch colorizationabstractSketch or line art colorization is a research field with significant market demand. Different from photo colorization which strongly relies on texture information, sketch colorization is more challenging as sketches may not have texture. Even worse, color, texture, and gradient have to be generated from the abstract sketch lines. In this paper, we propose a semi-automatic learning-based framework to colorize sketches with proper color, texture as well as gradient. Our framework consists of two stages. In the first drafting stage, our model guesses color regions and splashes a rich variety of colors over the sketch to obtain a color draft. In the second refinement stage, it detects the unnatural colors and artifacts, and try to fix and refine the result. Comparing to existing approaches, this two-stage design effectively divides the complex colorization task into two simpler and goal-clearer subtasks. This eases the learning and raises the quality of colorization. Our model resolves the artifacts such as water-color blurring, color distortion, and dull textures. We build an interactive software based on our model for evaluation. Users can iteratively edit and refine the colorization. We evaluate our learning model and the interactive system through an extensive user study. Statistics shows that our method outperforms the state-of-art techniques and industrial applications in several aspects including, the visual quality, the ability of user control, user experience, and other metrics. Lvmin Zhang, Chengze Li, Tien-Tsin Wong, Yi Ji 0001, Chunping Liu |
ACM Trans. Graph. | 4 |
| 2018 | Learning Multiple Kernel Metrics for Iterative Person Re-IdentificationabstractIn person re-identification most metric learning methods learn from training data only once, and then they are deployed for testing. Although impressive performance has been achieved, the discriminative information from successfully identified test samples are ignored. In this work, we present a novel re-identification framework termed Iterative Multiple Kernel Metric Learning (IMKML). Specifically, there are two main modules in IMKML. In the first module, multiple metrics are learned via a new derived Kernel Marginal Nullspace Learning (KMNL) algorithm. Taking advantage of learning a discriminative nullspace from neighborhood manifold, KMNL can well tackle the Small Sample Size (SSS) problem in re-identification distance metric learning. The second module is to construct a pseudo training set by performing re-identification on the testing set. The pseudo training set, which consists of the test image pairs that are highly probable correct matches, is then inserted into the labeled training set to retrain the metrics. By iteratively alternating between the two modules, many more samples will be involved for training and significant performance gains can be achieved. Experiments on four challenging datasets, including VIPeR, PRID450S, CUHK01, and Market-1501, show that the proposed method performs favorably against the state-of-the-art approaches, especially on the lower ranks. Husheng Dong, Chunping Liu, Yi Ji 0001, Shengrong Gong |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Co-saliency Detection Based on Superpixel Clustering
Guiqian Zhu, Yi Ji 0001, Xianjin Jiang, Zenan Xu, Chunping Liu |
KSEM | 2 |
| 2017 | Large margin relative distance learning for person re-identificationabstractDistance metric learning has achieved great success in person re‐identification. Most existing methods that learn metrics from pairwise constraints suffer the problem of imbalanced data. In this study, the authors present a large margin relative distance learning (LMRDL) method which learns the metric from triplet constraints, so that the problem of imbalanced sample pairs can be bypassed. Different from existing triplet‐based methods, LMRDL employs an improved triplet loss that enforces penalisation on the triplets with minimal inter‐class distance, and this leads to a more stringent constraint to guide the learning. To suppress the large variations of pedestrian's appearance in different camera views, the authors propose to learn the metric over the intra‐class subspace. The proposed method is formulated as a logistic metric learning problem with positive semi‐definite constraint, and the authors derive an efficient optimisation scheme to solve it based on the accelerated proximal gradient approach. Experimental results show that the proposed method achieves state‐of‐the‐art performance on three challenging datasets (VIPeR, PRID450S, and GRID). Husheng Dong, Shengrong Gong, Chunping Liu, Yi Ji 0001 |
IET Comput. Vis. | 4 |
| 2015 | Fusion of spatially constrained attributes with kernelized ranking for person re-identificationabstractThe task of matching persons across non-overlapping camera views, known as person re-identification, is rather challenging due to strong visual similarity and large appearance changes caused by illumination, pose and occlusion. Most approaches rely on low-level features that are both discriminative and invariant. In this work, we propose a novel method to address this problem by fusing mid-level semantic attributes with kernelized ranking. First, a kernelized ranking model is learned, and it gives the initial ranking scores. Next, an adaptive similarity model based on spatially constrained attributes is used to refine the ranking list. Fusion of the two models leads to much better performance than each individual alone. Experiments demonstrate complements of the two models and the results achieve new state-of-the-art performance on two benchmark datasets. Husheng Dong, Chunping Liu, Yi Ji 0001, Shengrong Gong |
AVSS | 3 |
| 2015 | Robust Dynamic Background Model with Adaptive Region Based on T2FS and GMMabstractFor many tracking and surveillance applications, Gaussian mixture model (GMM) provides an effective mean to segment the foreground from background. Though, because of insufficient and noisy data in complex dynamic scenes, the estimated parameters of the GMM, which are based on the assumption that the pixel process meets multi-modal Gaussian distribution, may not accurately reflect the underlying distribution of the observations. And the existing block-based GMM (BGMM) method may be able to segment only rough foreground objects with time-consuming calculations. To solve these difficulties, this paper proposes to use type-2 fuzzy sets (T2FSs) to handle GMM’s uncertain parameters (T2GMM). Furthermore, this paper also introduces a novel representation of contextual spatial information including the color, edge and texture features for each block which is faster and almost lossless (T2BGMM). Experimental results demonstrate the efficiency of the proposed methods. Yun Guo, Yi Ji 0001, Jutao Zhang, Shengrong Gong, Chunping Liu |
KSEM | 2 |
| 2015 | Learning topic of dynamic scene using belief propagation and weighted visual words approach
Chunping Liu, Shengrong Gong, Yi Ji 0001, Quan Liu 0004 |
Soft Comput. | 4 |
| 2012 | Automatic facial expression recognition based on spatiotemporal descriptors
Yi Ji 0001, Khalid Idrissi |
Pattern Recognit. Lett. | 1 |
| 2010 | Using Moments on Spatiotemporal Plane for Facial Expression RecognitionabstractIn this paper, we propose a novel approach to capture the dynamic deformation caused by facial expressions. The proposed method is concentrated on the spatiotemporal plane which is not well explored. It uses the moments as features to describe the movements of essential components such as eyes and mouth on vertical time plane. The system we developed can automatically recognize the expression on images as well as on image sequences. The experiments are performed on 348 sequences from 95 subjects in Cohn-Kanade database and obtained good results as high as 96.1% in 7-class recognition for frames and 98.5% in 6-class for sequences. Yi Ji 0001, Khalid Idrissi |
ICPR | 1 |
| 2009 | Object categorization using boosting within Hierarchical Bayesian modelabstractIn this paper we address the problem of generative object categorization in computer vision. We propose a Bayesian model using hierarchical Dirichlet processes mixing AdaBoost learning. Although previous methods trained HDP model for one or two latent themes, our proposed approach uses small-patch-independent-words of appearance-based descriptor and shape information to train a set of intermediate components which are the mixture of visualwords. We then employ AdaBoost weaker learner to find the most related components for classification to handle the variance in intraclass and inter-class information. We show that it performs well for Caltech datasets and with the potential to connect the visual concepts with semantic concepts. Yi Ji 0001, Khalid Idrissi, Atilla Baskurt |
ICIP | 1 |