Tao Zhang 0022

dblp:15/4777-22 · DBLP profile ↗
← Back
47ranked-venue papers
0as first author
29since 2021 · last 2026
0000-0001-7561-0143ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 17 since 2021Artificial intelligence and machine learning · 16 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Time-frequency contrastive fusion for denoising sequential recommendation
Zhenzhen Sheng, Tao Zhang 0022
Expert Syst. Appl.4
2025 Sketch-based Point Cloud Generation with Diffusion Model and Pre-training Enhancement
abstract
Diffusion models, known for their success in various generative tasks like image generation and super-resolution, are applied in this study for point cloud generation, a field that has not been extensively explored due to the complexity of point clouds. We propose a novel method using a diffusion model to generate high-quality 3D point clouds from 2D sketches. This method employs a self-supervised contrastive learning scheme to align sketch and point cloud modalities. Additionally, it incorporates a specific partition mixing strategy to integrate edge information during pre-training. Evaluated on two benchmark datasets, our method outperforms existing state-of-the-art approaches, showcasing the potential of diffusion models in point cloud generation and setting a new direction for future research.
Yangdong Chen, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP5
2025 Minimizing Disparities between Real and Pseudo Queries for Unsupervised Visual Grounding
abstract
Visual grounding involves the identification and localization of image regions given textual descriptions. To reduce the manual labeling effort on region-text pairs, unsupervised visual grounding aims to generate pseudo bounding box and query pairs for training grounding models. However, there exists significant disparities between real and pseudo queries in terms of object, attribute distributions, and textual formats, limiting the generalization performance of unsupervised grounding methods. To address this challenge, we propose a novel unsupervised visual grounding framework. During training, we prompt Multimodal Large Language Models to generate pseudo queries, in which the entities are beyond the object detector’s pre-defined limited categories, and are associated with richer attributes. We further devise a Modifier Tree structure to bridge the gap of textual format between real and pseudo queries. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art unsupervised approaches on public benchmark datasets, particularly when dealing with complex queries.
Changkai Ji, Jilan Xu, Yanhao Zhu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP7
2025 TGSAM-2: Text-Guided Medical Image Segmentation Using Segment Anything Model 2
Runtian Yuan, Ling Zhou 0002, Jilan Xu, Qingqiu Li, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
MICCAI (10)8
2025 Text-Promptable Propagation for Referring Medical Image Sequence Segmentation
abstract
Referring Medical Image Sequence Segmentation (Ref-MISS) is a novel and challenging task that aims to segment anatomical structures in medical image sequences (e.g., endoscopy, ultrasound, CT, and MRI) based on natural language descriptions. Existing 2D and 3D segmentation models struggle to explicitly track objects of interest across medical image sequences, and lack support for interactive, text-driven guidance. To address these limitations, we propose Text-Promptable Propagation (TPP), which enables the recognition of referred objects through cross-modal referring interaction, and maintains continuous tracking across the sequence via Transformer-based triple propagation, using text embeddings as queries. To support this task, we curate a large-scale benchmark, Ref-MISS-Bench, which covers 4 imaging modalities and 20 different organs and lesions. Experimental results on this benchmark demonstrate that TPP consistently outperforms state-of-the-art methods in both medical segmentation and referring video object segmentation. Code and data are available at https://github.com/yuanruntian/TPP.
Runtian Yuan, Mohan Chen 0001, Jilan Xu, Ling Zhou 0002, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ACM Multimedia8
2025 EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues
abstract
Qiming Feng, Qiujie Xie, Xiaolong Wang, Qingqiu Li, Yuejie Zhang, Rui Feng, Tao Zhang, Shang Gao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qiming Feng, Qiujie Xie, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
NAACL (Long Papers)7
2024 Fine-Granularity Face Sketch Synthesis
abstract
Generative Adversarial Networks (GANs) are often used in face sketch synthesis due to their powerful ability in image generation. However, most GAN based synthesis methods took the entire face as the minimum unit. Differently, we propose a novel fine-granularity face sketch synthesis framework in this paper. The core idea is to first capture local information at a fine granularity (i.e., facial component), and then generate a complete face sketch based on the fine-grained information. Specifically, we partition the face sketch into multiple components, and then train a parallel network for each component. A condition enhanced detail repair network is further designed to correct the mismatches and deformations produced during parallel generation. Extensive experiments show that our approach outperforms state-of-the-art methods from both the qualitative and quantitative perspectives.
Yangdong Chen, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICASSP5
2024 ControlCap: Controllable Captioning via No-Fuss Lexicon
abstract
Controllable captioning has received much attention in recent years. Although substantial progress has been made, existing methods still face challenges such as high training costs, intricate control signals and limited control capabilities. To address these issues, we propose a straightforward and unified framework called ControlCap. It uses a no-fuss lexicon as control signal and controls the style and content of visual descriptions through Soft Guidance (a global guide to the caption distribution) and Hard Force (integrating signals without additional training). Extensive experiments, both quantitative and qualitative, have been conducted on three benchmark captioning tasks. Results demonstrate the control ability of ControlCap: it can produce controlled captions that are coherent and diverse while keeping the core content intact.
Qiujie Xie, Qiming Feng, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP5
2024 Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning
abstract
Video paragraph captioning task aims to generate a detailed, fluent and relevant paragraph for a given video. Prior studies often focus on isolating visual objects (potential main components in a sentence) from the overall video content. They rarely explore the latent semantic relations between objects and high-level video concepts, resulting in dull or even incorrect descriptions. To create fine-grained and contextually relevant paragraph captions, we propose a novel framework that constructs a concept graph from a commonsense knowledge base and infers richer semantic meaning from the visual objects. Moreover, we employ a Vision-Guided Concept Selection Network that incorporates an under-sentence supervision mechanism to align the external knowledge with the visual information. Through extensive experiments on ActivityNet captions and YouCook2, the effectiveness of our method is demonstrated compared to state-of-the-art methods.
Guorui Yu, Yimin Hu, Yiqian Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP6
2024 Temporal Feature Aggregation for Efficient 2D Video Grounding
abstract
Video grounding aims to locate the target video moment in an untrimmed video based on a text query. Most existing methods employ 3D CNNs as the video feature extractor, incurring substantial computational costs. Only a few methods use 2D backbones for video feature extraction, and they suffer from diminished accuracy due to the inherent lack of temporal information within 2D features. To address this problem, we propose a novel 2D video grounding method called TFA that improves accuracy while minimizing computational costs. Our approach involves a query-guided temporal feature aggregation module designed to explicitly capture temporal information. We disentangle time intervals of input video frames and prediction spans to reduce computational overhead. Additionally, we introduce deformable attention into the multi-modal encoder for further enhancement. Extensive experiments on two public datasets demonstrate that our method outperforms previous 2D video grounding methods and achieves competitive results with most 3D methods at significantly reduced costs.
Mohan Chen 0001, Yiren Zhang, Jueqi Wei, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICME6
2024 Memory-Augmented Transformer for Efficient End-to-End Video Grounding
abstract
Video grounding aims to localize a specific segment corresponding to a text query in an untrimmed video. Due to the tremendous computational cost required to process the video frames, the de facto paradigm of video grounding is to extract video features using pretrained video encoders. The parameters of the video encoders are fixed during training, which limits the performance of the localization model. To solve this problem, we propose a Memory-Augmented Transformer (MAT) model. Specifically, each video is split into non-overlapping clips, and our MAT processes videos in a clip-by-clip manner while caching video features into FIFO cached memory queues. By enabling early return, our MAT outperforms previous methods with only less than 60% frames seen. Extensive experimental results on three public benchmark datasets demonstrate that our MAT can achieve competitive performance while being much more efficient than currently prevailing two-stage methods. Code is available at https://github.com/xuyw1997/MAT.
Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICME5
2023 Semi-MedSeq: Semi-supervised Semantic Segmentation for Medical Image Sequences
abstract
In clinical practice, medical imaging techniques include 2D video-based examinations that capture sequential scans, and 3D volumetric imaging that forms a comprehensive 3D representation from a stack of 2D slices. The medical image sequences produced by the above techniques provide valuable spatio-temporal characteristics for analysis and segmentation, but the annotation of image sequences is extremely time-consuming and labor-intensive. To exploit the coherence and address the scarcity of labeled data, we propose a novel semi-supervised semantic segmentation framework for medical image sequences, which consists of a conditional network and a denoising network. Specifically, we embed a Sequential Feature Reconstruction module into both networks. This module reconstructs the target frame from contiguous frames and captures their shared visual features. Guided by the context-enhancing information from the conditioning network, the denoising network suppresses background noise via a Diffusion-based Noise Elimination module. Extensive experiments are conducted on 2D and 3D tasks, including cardiac segmentation, polyp segmentation, placenta vessel segmentation and abdomen multi-organ segmentation. The results show our method is superior to existing semi-supervised methods and exhibits advantages over fully-supervised medical image segmentation methods with only 1/2 labeled data, validating its effectiveness and generalization ability.
Runtian Yuan, Jilan Xu, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM7
2023 Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal Knowledge
abstract
Video paragraph captioning task aims at generating a fine-grained, coherent and relevant paragraph for a video. Different from the images where objects are static, the temporal states of objects are changing in videos. The dynamic information could be contributed to understanding the whole video content. Existing works rarely put focus on modeling the dynamic changing state of the objects in the videos, causing the activities occurred in videos are poorly or wrongly depicted in paragraphs. To address this problem, we propose a novel Object State Tracking Network, which can capture the temporal state change of objects. However, due to the similarity of the consecutive frames in the videos, the information of the video is redundant and noisy. We further propose a semantic alignment mechanism, and enable the sentence information to refine the visual information. Extensive experiments on ActivityNet Captions demonstrate the effectiveness of our method.
Yimin Hu, Guorui Yu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICASSP5
2023 Boosting Fine-Grained Sketch-Based Image Retrieval with Self-Supervised Learning
abstract
Fine-grained sketch-based image retrieval (FG-SBIR) aims at aligning images and sketches at the instance level. It is a challenging task as there are significant differences between sketch and image. Existing methods usually produce less desired performance due to the lack of large-scale fine-grained image-sketch datasets and the strong dependence on the classification models pretrained on ImageNet. In this paper, we propose a better self-supervised pre-trained FG-SBIR model which does not depend on large-scale annotated datasets. Only images and their corresponding edge maps are used at the pre-training stage. Mixed modal transformation is designed to generate different mixed-up views. The FG-SBIR model is pre-trained by minimizing the distance between the views of the same instance and then fine-tuned by a simple triplet loss. With a plain downstream network, it achieves generally better performance than state-of-the-art models on three widely used FG-SBIR datasets.
Zhaolong Zhang, Yangdong Chen, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022
ICASSP5
2023 Video Captioning via Relation-Aware Graph Learning
abstract
Recent neural models for video captioning usually employed an encoder-decoder framework. However, most approaches either neglected the spatial and temporal interactions between objects in a video or implicitly modelled the interactions, resulting in less desired performance. In this paper, we propose a novel relation-aware graph learning framework. It explicitly models both spatial and temporal relations for objects. In particular, a relation-aware graph is designed to depict the spatial relations between different objects in a scene. Parallelly, a temporal graph network is designed to perform relational reasoning for the same objects in adjacent frames. Features of both types of relations are learned and fused for the follow-up language decoder. Experiments on two bench-mark datasets show the effectiveness of our framework. It achieves state-of-the-art performance with CIDEr scores on MSVD and MSR-VTT.
Heming Jing, Qiujie Xie, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP6
2023 CAMG: Context-Aware Moment Graph Network for Multimodal Temporal Activity Localization via Language
Yuelin Hu, Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
NLPCC (1)5
2023 Deep cross-modal hashing with fine-grained similarity
Yangdong Chen, Jiaqi Quan, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022
Appl. Intell.5
2023 Enhanced graph neural network for session-based recommendation
Zhenzhen Sheng, Tao Zhang 0022, Yuejie Zhang, Shang Gao 0003
Expert Syst. Appl.2
2022 Single-Modality Endoscopic Polyp Segmentation via Random Color Reversal Synthesis and Two-Branched Learning
abstract
Endoscopic polyp segmentation plays a fundamental role in the diagnosis and treatment of colorectal cancer. However, polyp segmentation often suffers from limited accuracy due to its large variations in appearance, blurry boundary and severe imbalanced illumination. In this paper, we propose a novel Translation Assisted Segmentation Network (TASNet) for polyp segmentation of single-modality endoscopic images. It consists of two branches, i.e. an image-to-image translation branch and an image segmentation branch. These two branches communicate via a shared encoder. For the image-to-image translation branch, a Color Reversal Strategy is established to treat the original image as source image and synthesize target images. Moreover, we introduce a Random Color Reversal Synthesis module for progressive segmentation. Extensive experiments show that our framework achieves superior performance than state-of-the-art methods on five widely-used endoscopic image datasets.
Mingzhu Chen, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM7
2022 MedSeq: Semantic Segmentation for Medical Image Sequences
abstract
Medical image segmentation plays a critical role in computer-aided diagnosis, while the diversity and complexity of medical images make it difficult to segment precisely. In practice, medical images of specific modalities (e.g. Magnetic Resonance Imaging, Colonoscopy and Ultrasonography) are collected as sequences independently for every patient. However, 1) there exists few works exploiting sequence information among successive frames, neglecting inter-frame relationships that are useful to locate target objects; 2) the performance of medical image segmentation is limited to the low contrast or blurry boundary of medical images, and intra-frame dependencies are not fully explored. Thus in this paper, we propose MedSeq for segmenting objects of interest in medical image sequences. Following the “locate-then-refine” paradigm, we locate target regions by modeling cross-frame relationships and then perform refinement on coarse masks. More specifically, we design a Cross-frame Attention module to learn correlations among frames, taking advantages of their similar appearances. For refinement, we propose a novel Boundary-aware Transformer to improve the segmentation of boundary patches. Extensive experiments are conducted on benchmark datasets of Cardiac Segmentation and Video Polyp Segmentation. Our method achieves superior performance over the state-of-the-art methods.
Runtian Yuan, Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM7
2022 CREAM: Weakly Supervised Object Localization via Class RE-Activation Mapping
abstract
Weakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) de-rived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomplete localization problem). In this paper, we empirically prove that this problem is associated with the mixup of the activation values between less discrimi-native foreground regions and the background. To address it, we propose Class RE-Activation Mapping (CREAM), a novel clustering-based approach to boost the activation values of the integral object regions. To this end, we in-troduce class-specific foreground and background context embeddings as cluster centroids. A CAM-guided momen-tum preservation strategy is developed to learn the context embeddings during training. At the inference stage, the re-activation mapping is formulated as a parameter es-timation problem under Gaussian Mixture Model, which can be solved by deriving an unsupervised Expectation- Maximization based soft-clustering algorithm. By simply integrating CREAM into various WSOL approaches, our method significantly improves their performance. CREAM achieves the state-of-the-art performance on CUB, ILSVRC and OpenImages benchmark datasets. Code will be avail-able at https://github.com/lazzcharles/CREAM.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
CVPR6
2022 Semantic-Driven Saliency-Context Separation for Video Captioning
abstract
Video captioning aims at generating a natural language de-scription for a given video clip including not only salient sce-narios but also contextual scenarios. The former reveal the highlight of a video and are usually the focus of most existing captioning methods. The latter, however, are not well ex-plored and even ignored easily, though they may provide cer-tain detailed and latent information that can help with a better understanding of the video. To effectively exploit the infor-mation contained in both, a novel video captioning network is proposed. It has two key modules: Cross-Modality Selection (CMS) and Saliency-Context Adaptive Decoder (SCAD). Specifically, CMS mainly focuses on utilizing the semantic information to distinguish saliency and context. Meanwhile, SCAD adaptively identifies both the saliency and context to generate more detailed and precise captions. Experiments on two benchmark datasets, i.e., MSVD and MSR-VTT, demon-strate the effectiveness of our model through the comparison with state-of-the-art methods.
Heming Jing, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICME5
2022 STDNet: Spatio-Temporal Decomposed Network for Video Grounding
abstract
Previous methods for video grounding treated either the query or the video as a whole, while neglecting their respective semantics in the orthogonal space and time dimensions. Since spatial semantics appears frequently in a video, temporal semantics is more discriminative and deserves more attention. Based on such considerations, we propose a novel Spatio-Temporal Decomposed Network (STDNet) which decomposes the query and the video into their spatial and temporal semantics, respectively. Specifically, spatial and temporal words are selected from the query, and the video is split into two pathways. Spatial cross-modal attention is computed first and serves as prior knowledge for temporal attention. A new localization strategy is also devised which regresses the segment's start conditioned on the end and essentially breaks the independence assumption made in previous methods. Experimental results on three public benchmark datasets show that our STDNet outperforms the state-of-the-art methods.
Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICME5
2022 TCCNet: Temporally Consistent Context-Free Network for Semi-supervised Video Polyp Segmentation
abstract
Automatic video polyp segmentation (VPS) is highly valued for the early diagnosis of colorectal cancer. However, existing methods are limited in three respects: 1) most of them work on static images, while ignoring the temporal information in consecutive video frames; 2) all of them are fully supervised and easily overfit in presence of limited annotations; 3) the context of polyp (i.e., lumen, specularity and mucosa tissue) varies in an endoscopic clip, which may affect the predictions of adjacent frames. To resolve these challenges, we propose a novel Temporally Consistent Context-Free Network (TCCNet) for semi-supervised VPS. It contains a segmentation branch and a propagation branch with a co-training scheme to supervise the predictions of unlabeled image. To maintain the temporal consistency of predictions, we design a Sequence-Corrected Reverse Attention module and a Propagation-Corrected Reverse Attention module. A Context-Free Loss is also proposed to mitigate the impact of varying contexts. Extensive experiments show that even trained under 1/15 label ratio, TCCNet is comparable to the state-of-the-art fully supervised methods for VPS. Also, TCCNet surpasses existing semi-supervised methods for natural image and other medical image segmentation tasks.
Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
IJCAI6
2022 AE-Net: Fine-grained sketch-based image retrieval via attention-enhanced network
Yangdong Chen, Zhaolong Zhang, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.6
2022 Stacked Multimodal Attention Network for Context-Aware Video Captioning
abstract
Recent neural models for video captioning usually employ an attention-based encoder-decoder framework. However, current approaches mainly attend to the motion features and object features of the video when generating the caption, but ignore the potential but useful historical information. Besides, exposure bias and vanishing gradients problems always exist in current caption generation models. In this paper, we propose a novel video captioning framework, named Stacked Multimodal Attention Network (SMAN). It adopts additional visual and textual historical information during caption generation as context features, employs a stacked architecture to process different features gradually, and utilizes the Reinforcement Learning method and coarse-to-fine training strategy to further improve the generated results. Both quantitative and qualitative experiments on the benchmark datasets ofMSVDandMSR-VTTshow the effectiveness and feasibility of our framework. The codes are available onhttps://github.com/zhengyi123456/SMAN.
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
IEEE Trans. Circuits Syst. Video Technol.4
2021 HTDA: Hierarchical time-based directional attention network for sequential user behavior modeling
Zhenzhen Sheng, Tao Zhang 0022, Yuejie Zhang
Neurocomputing2
2021 Cross-modal retrieval with dual multi-angle self-attention
abstract
Abstract In recent years, cross‐modal retrieval has been a popular research topic in both fields of computer vision and natural language processing. There is a huge semantic gap between different modalities on account of heterogeneous properties. How to establish the correlation among different modality data faces enormous challenges. In this work, we propose a novel end‐to‐end framework named Dual Multi‐Angle Self‐Attention (DMASA) for cross‐modal retrieval. Multiple self‐attention mechanisms are applied to extract fine‐grained features for both images and texts from different angles. We then integrate coarse‐grained and fine‐grained features into a multimodal embedding space, in which the similarity degrees between images and texts can be directly compared. Moreover, we propose a special multistage training strategy, in which the preceding stage can provide a good initial value for the succeeding stage and make our framework work better. Very promising experimental results over the state‐of‐the‐art methods can be achieved on three benchmark datasets of Flickr8k, Flickr30k, and MSCOCO.
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
J. Assoc. Inf. Sci. Technol.5
2021 Deep Cross-Modal Face Naming for People News Retrieval
abstract
How to integrate multimodal information sources for face naming in multimodal news is a hot and yet challenging problem. A novel deep cross-modal face naming scheme is developed in this paper to facilitate more effective people news retrieval for large-scale multimodal news. This scheme integrates deep multimodal analysis, cross-modal correlation learning, and multimodal information mining, in which the efficient naming mechanism aims to cluster the deep features of different modalities into a common space to explore their inter-related correlations, and a special Web mining pattern is designed to optimize the name-face matching for rare non-celebrity. Such a cross-modal face naming model can be treated as a problem of bi-media semantic mapping and modeled as an inter-related correlation distribution over deep representations of multimodal news, in which the most important is to create more effective cross-modal name-face correlation and measure to what degree they are correlated. The experiments on a large number of public data from Yahoo! News have obtained very positive results and demonstrated the effectiveness of the proposed model.
Lian Zhou, Yuejie Zhang, Tao Zhang 0022, Weiguo Fan
IEEE Trans. Knowl. Data Eng.4
2020 Zero-Shot Sketch-Based Image Retrieval via Graph Convolution Network
abstract
Zero-Shot Sketch-based Image Retrieval (ZS-SBIR) has been proposed recently, putting the traditional Sketch-based Image Retrieval (SBIR) under the setting of zero-shot learning. Dealing with both the challenges in SBIR and zero-shot learning makes it become a more difficult task. Previous works mainly focus on utilizing one kind of information, i.e., the visual information or the semantic information. In this paper, we propose a SketchGCN model utilizing the graph convolution network, which simultaneously considers both the visual information and the semantic information. Thus, our model can effectively narrow the domain gap and transfer the knowledge. Furthermore, we generate the semantic information from the visual information using a Conditional Variational Autoencoder rather than only map them back from the visual space to the semantic space, which enhances the generalization ability of our model. Besides, feature loss, classification loss, and semantic loss are introduced to optimize our proposed SketchGCN model. Our model gets a good performance on the challenging Sketchy and TU-Berlin datasets.
Zhaolong Zhang, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
AAAI4
2020 Label Generation Network based on Self-selected Historical Information for Multiple Disease Classification on Chest Radiography
abstract
Deep learning has made significant break through's in image classification, but accurate diagnosis on chest radiography remains challenging due to a variety of potential diseases contained in one scan. Complex relations among diseases have significant clinical meanings, but are always ignored in most of previous work. Thus in this paper, we propose a novel Label Generation Network (LGN) which treats the label sequence as the caption of a radiology image and utilizes RNN to generate the disease labels according to the semantic relations and co-occurrence dependency among them. However, the sequential generation process of RNN makes it hard to capture the complex topological relations among diseases. To mitigate this problem, a Historical Information Module (HIM) is especially introduced to LGN, in which all the generated labels are fully considered when generating a new label. Moreover, a specific self-attention mechanism is applied in HIM to learn the topological disease relations and utilize them to select useful historical information which can provide positive guidance to the prediction of new label. Very positive results have been obtained in our experiments on the benchmark dataset of Chest X-ray14, which significantly outperform the state-of-the-art methods.
Yuelin Hu, Yuejie Zhang, Tao Zhang 0022, Shang Gao 0003, Weiguo Fan
BIBM3
2020 Data-Efficient Histopathology Image Analysis with Deformation Representation Learning
abstract
Histopathological examination of tissue biopsies plays a fundamental role in disease assessment. Automatic histopathology image analysis requires substantial task-specific annotations, which are often expensive and laborious in realworld scenarios. This insufficient annotation of data limits the generalization ability of supervised learning models. To address this challenge, we propose a self-supervised Deformation Representation Learning (DRL) framework to learn semantic features from unlabeled data. As a novel paradigm, our approach utilizes deformation as supervisory signals based on two critical features, i.e., local structure heterogeneity and global context homogeneity. Given an original histopathology image and its deformed counterpart, there exists a moderate difference in local structures. In contrast, due to the transformation-invariance, both images share a similar global context compared with other images. Specifically, an encoder network is trained to distinguish the local inconsistency by measuring the mutual information and maintain the global consistency with noise contrastive estimation. Extensive experiments on public histopathology image datasets show that the learned representations are generalizable for various downstream tasks, such as transfer learning on segmentation and semi-supervised classification. Our approach achieves superior results over other self-supervised methods and the ImageNet pre-trained model, and it reveals the ability as a novel pre-training scheme in histopathology image analysis.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Chunyang Ruan, Tao Zhang 0022, Weiguo Fan
BIBM6
2020 Video Captioning With Temporal And Region Graph Convolution Network
abstract
Video captioning aims to generate a natural language description for a given video clip that includes not only spatial information but also temporal information. To better exploit such spatial-temporal information attached to videos, we propose a novel video captioning framework with Temporal Graph Network (TGN) and Region Graph Network (RGN). TGN mainly focuses on utilizing the sequential information of frames that most of existing methods ignore. RGN is designed to explore the relationships among salient objects. Different from previous work, we introduce Graph Convolution Network (GCN) to encode frames with their sequential information and build a region graph for utilizing object information. We also particularly adopt a stack GRU decoder with a coarse-to-fine structure for caption generation. Very promising experimental results on two benchmark datasets (MSVD and MSR-VTT) show the effectiveness of our model.
Xinlong Xiao, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003, Weiguo Fan
ICME4
2020 Deep cascaded cross-modal correlation learning for fine-grained sketch-based image retrieval
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.5
2020 Deep reinforcement hashing with redundancy elimination for effective image retrieval
Juexu Yang, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.4
2020 Re-Caption: Saliency-Enhanced Image Captioning Through Two-Phase Learning
abstract
Visual and semantic saliency are important in image captioning. However, single-phase image captioning benefits little from limited saliency without a saliency predictor. In this paper, a novel saliency-enhanced re-captioning framework via two-phase learning is proposed to enhance the single-phase image captioning. In the framework, visual saliency and semantic saliency are distilled from the first-phase model and fused with the second-phase model for model self-boosting. The visual saliency mechanism can generate a saliency map and a saliency mask for an image without learning a saliency map predictor. The semantic saliency mechanism sheds some lights on the properties of words with part-of-speech Noun in a caption. Besides, another type of saliency, sample saliency is proposed to explicitly compute the saliency degree of each sample, which helps for more robust image captioning. In addition, how to combine the above three types of saliency for further performance boost is also examined. Our framework can treat an image captioning model as a saliency extractor, which may benefit other captioning models and related tasks. The experimental results on both the Flickr30k and MSCOCO datasets show that the saliency-enhanced models can obtain promising performance gains.
Lian Zhou, Yuejie Zhang, Yu-Gang Jiang 0001, Tao Zhang 0022, Weiguo Fan
IEEE Trans. Image Process.4
2019 Fully Convolutional Video Captioning with Coarse-to-Fine and Inherited Attention
abstract
Automatically generating natural language description for video is an extremely complicated and challenging task. To tackle the obstacles of traditional LSTM-based model for video captioning, we propose a novel architecture to generate the optimal descriptions for videos, which focuses on constructing a new network structure that can generate sentences superior to the basic model with LSTM, and establishing special attention mechanisms that can provide more useful visual information for caption generation. This scheme discards the traditional LSTM, and exploits the fully convolutional network with coarse-to-fine and inherited attention designed according to the characteristics of fully convolutional structure. Our model cannot only outperform the basic LSTM-based model, but also achieve the comparable performance with those of state-of-the-art methods
Kuncheng Fang, Lian Zhou, Cheng Jin 0001, Yuejie Zhang, Kangnian Weng, Tao Zhang 0022, Weiguo Fan
AAAI6
2019 A cutting plane method for risk-constrained traveling salesman problem with random arc costs
Zhouchun Huang, Qipeng Phil Zheng, Eduardo L. Pasiliao, Vladimir Boginski, Tao Zhang 0022
J. Glob. Optim.5
2018 Sketch-based image retrieval with deep visual semantic descriptor
Cheng Jin 0001, Yuejie Zhang, Kangnian Weng, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.5
2017 Towards sketch-based image retrieval with deep cross-modal correlation learning
abstract
A novel scheme with deep cross-modal correlation learning is developed in this paper to facilitate more effective Sketch-based Image Retrieval (SBIR) for large-scale annotated images. It integrates the deep multimodal feature generation, deep cross-modal correlation learning and similarity search optimization through mining all the beneficial multimodal information sources in sketches and images, which can be treated as an inter-related correlation distribution over deep representations of sketches and images. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ICME4
2017 A Hierarchical Multimodal Attention-based Neural Network for Image Captioning
abstract
A novel hierarchical multimodal attention-based model is developed in this paper to generate more accurate and descriptive captions for images. Our model is an "end-to-end" neural network which contains three related sub-networks: a deep convolutional neural network to encode image contents, a recurrent neural network to identify the objects in images sequentially, and a multimodal attention-based recurrent neural network to generate image captions. The main contribution of our work is that the hierarchical structure and multimodal attention mechanism is both applied, thus each caption word can be generated with the multimodal attention on the intermediate semantic objects and the global visual content. Our experiments on two benchmark datasets have obtained very positive results.
Lian Zhou, Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
SIGIR6
2017 Deep Multimodal Embedding Model for Fine-grained Sketch-based Image Retrieval
abstract
Fine-grained Sketch-based Image Retrieval (Fine-grained SBIR), which uses hand-drawn sketches to search the target object images, has been an emerging topic over the last few years. The difficulties of this task not only come from the ambiguous and abstract characteristics of sketches with less useful information, but also the cross-modal gap at both visual and semantic level. However, images on the web are always exhibited with multimodal contents. In this paper, we consider Fine-grained SBIR as a cross-modal retrieval problem and propose a deep multimodal embedding model that exploits all the beneficial multimodal information sources in sketches and images. In our experiment with large quantity of public data, we show that the proposed method outperforms the state-of-the-art methods for Fine-grained SBIR.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
SIGIR5
2016 A Novel Cross-Modal Topic Correlation Model for Cross-Media Retrieval
abstract
A novel cross-modal topic correlation model CMTCM is developed in this paper to facilitate more effective cross-modal analysis and cross-media retrieval for large-scale multimodal document collections. It can be modeled as a cross-modal topic correlation model which explores the inter-related correlation distribution over the deep representations of multimodal documents. It integrates the deep multimodal document representation, relational topic correlation modeling, and cross-modal topic correlation learning, which aims to characterize the correlations between the heterogeneous topic distributions of inter-related visual images and semantic texts, and measure their association degree more precisely. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ECAI5
2016 Enhancing Sketch-Based Image Retrieval via Deep Discriminative Representation
abstract
In this paper we aim to employ deep learning to enhance SBIR via deep discriminative representation. Our main contributions focus on: 1) The deep discriminative representation is established to bridge both the visual appearance gap and the semantic gap between sketches and images; 2) The deep learning pattern is applied to our SBIR model through training on our transformed sketch-like images to overcome the rarity of training sketches. Our experiments on a large number of public sketch and image data have obtained very positive results.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ECAI5
2016 Sketch-Based Image Retrieval with a Novel BoVW Representation
Cheng Jin 0001, Chenjie Li, Zheming Wang, Yuejie Zhang, Tao Zhang 0022
MMM (1)5
2015 People News Search via Name-Face Association Analysis
abstract
By integrating multimodal information in multimodal news, a novel scheme is developed in this paper for facilitating more effective people news search via name-face association analysis. It is treated as a problem of bi-media multimodal semantic mapping on multimodal news, and modeled as an inter-related correlation distribution over multimodal semantic representations of name-face associations. Very positive results have been obtained in our experiments using a large quantity of public multimodal news data.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ICMR6
2015 Cross-Modal Image-Tag Relevance Learning for Social Images
abstract
A new algorithm is developed in this paper to support more effective cross-modal image-tag relevance learning for large-scale social images, which integrates the multimodal feature representation, multimodal relevance measurement, and cross- modal relevance fusion. The main contribution of our work is that we provide a more reasonable base to learn cross-modal relevance among social images, which can be acquired from integrating multimodal image and tag relevance with multiple features in different modalities. Very positive results were obtained in our experiments using a large quantity of public social image data.
Zhengxiang Cai, Rui Feng 0001, Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ACM Multimedia6