Yuejie Zhang

dblp:09/5786 · DBLP profile ↗
← Back
92ranked-venue papers
5as first author
64since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 54 · 4 first-author · 34 since 2021Artificial intelligence and machine learning · 35 · 3 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 12 since 2021Databases, data management, data science and information retrieval · 7 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 EviMMQA: Multimodal question answering for medical evidence extraction in systematic reviews
Changkai Ji, Yingwen Wang, Ying Cheng 0005, Yuejie Zhang, Rui Feng 0001
Pattern Recognit.6
2025 RoBGuard: Enhancing LLMs to Assess Risk of Bias in Clinical Trial Documents
abstract
Randomized Controlled Trials (RCTs) are rigorous clinical studies crucial for reliable decision-making, but their credibility can be compromised by bias. The Cochrane Risk of Bias tool (RoB 2) assesses this risk, yet manual assessments are time-consuming and labor-intensive. Previous approaches have employed Large Language Models (LLMs) to automate this process. However, they typically focus on manually crafted prompts and a restricted set of simple questions, limiting their accuracy and generalizability. Inspired by the human bias assessment process, we propose RoBGuard, a novel framework for enhancing LLMs to assess the risk of bias in RCTs. Specifically, RoBGuard integrates medical knowledge-enhanced question reformulation, multimodal document parsing, and multi-expert collaboration to ensure both completeness and accuracy. Additionally, to address the lack of suitable datasets, we introduce two new datasets: RoB-Item and RoB-Domain. Experimental results demonstrate RoBGuard’s effectiveness on the RoB-Item dataset, outperforming existing methods.
Changkai Ji, Yingwen Wang, Yuejie Zhang, Ying Cheng 0005, Rui Feng 0001
COLING5
2025 Sketch-based Point Cloud Generation with Diffusion Model and Pre-training Enhancement
abstract
Diffusion models, known for their success in various generative tasks like image generation and super-resolution, are applied in this study for point cloud generation, a field that has not been extensively explored due to the complexity of point clouds. We propose a novel method using a diffusion model to generate high-quality 3D point clouds from 2D sketches. This method employs a self-supervised contrastive learning scheme to align sketch and point cloud modalities. Additionally, it incorporates a specific partition mixing strategy to integrate edge information during pre-training. Evaluated on two benchmark datasets, our method outperforms existing state-of-the-art approaches, showcasing the potential of diffusion models in point cloud generation and setting a new direction for future research.
Yangdong Chen, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP3
2025 Uncertainty-Aware Dynamic Fusion for Multimodal Clinical Prediction Tasks
abstract
Multimodal fusion offers significant potential for enhancing medical diagnosis, particularly in the Intensive Care Unit (ICU), where integrating diverse data sources is crucial. Traditional static fusion models often fail to account for sample-wise variations in modality importance, which can impact prediction accuracy. To address this issue, we propose a dynamic Uncertainty-Aware Weighting (UAW) strategy that adaptively adjusts the importance of different modalities based on their reliability. This strategy is coupled with an Expert Ensemble Fusion (EEF) module, which leverages self-attention mechanisms and modality-specific FeedForward Networks (FFNs) to preserve and integrate critical information from various modalities. The proposed method demonstrates its efficacy through extensive experiments on phenotype classification and mortality prediction tasks, showing improved accuracy and robustness in handling diverse clinical data.
Ying Cheng 0005, Yuejie Zhang, Rui Feng 0001
ICASSP4
2025 Minimizing Disparities between Real and Pseudo Queries for Unsupervised Visual Grounding
abstract
Visual grounding involves the identification and localization of image regions given textual descriptions. To reduce the manual labeling effort on region-text pairs, unsupervised visual grounding aims to generate pseudo bounding box and query pairs for training grounding models. However, there exists significant disparities between real and pseudo queries in terms of object, attribute distributions, and textual formats, limiting the generalization performance of unsupervised grounding methods. To address this challenge, we propose a novel unsupervised visual grounding framework. During training, we prompt Multimodal Large Language Models to generate pseudo queries, in which the entities are beyond the object detector’s pre-defined limited categories, and are associated with richer attributes. We further devise a Modifier Tree structure to bridge the gap of textual format between real and pseudo queries. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art unsupervised approaches on public benchmark datasets, particularly when dealing with complex queries.
Changkai Ji, Jilan Xu, Yanhao Zhu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP5
2025 Human Simulacra: Benchmarking the Personification of Large Language Models
abstract
Large Language Models (LLMs) are recognized as systems that closely mimic aspects of human intelligence. This capability has attracted the attention of the social science community, who see the potential in leveraging LLMs to replace human participants in experiments, thereby reducing research costs and complexity. In this paper, we introduce a benchmark for LLMs personification, including a strategy for constructing virtual characters' life stories from the ground up, a Multi-Agent Cognitive Mechanism capable of simulating human cognitive processes, and a psychology-guided evaluation method to assess human simulations from both self and observational perspectives. Experimental results demonstrate that our constructed simulacra can produce personified responses that align with their target characters. We hope this work will serve as a benchmark in the field of human simulation, paving the way for future research.
Qiujie Xie, Qiming Feng, Qingqiu Li, Linyi Yang, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003, Yue Zhang 0004
ICLR6
2025 An Empirical Analysis of Uncertainty in Large Language Model Evaluations
abstract
As LLM-as-a-Judge emerges as a new paradigm for assessing large language models (LLMs), concerns have been raised regarding the alignment, bias, and stability of LLM evaluators. While substantial work has focused on alignment and bias, little research has concentrated on the stability of LLM evaluators. In this paper, we conduct extensive experiments involving 9 widely used LLM evaluators across 2 different evaluation settings to investigate the uncertainty in model-based LLM evaluations. We pinpoint that LLM evaluators exhibit varying uncertainty based on model families and sizes. With careful comparative analyses, we find that employing special prompting strategies, whether during inference or post-training, can alleviate evaluation uncertainty to some extent. By utilizing uncertainty to enhance LLM's reliability and detection capability in Out-Of-Distribution (OOD) data, we further fine-tune an uncertainty-aware LLM evaluator named ConfiLM using a human-annotated fine-tuning set and assess ConfiLM's OOD evaluation ability on a manually designed test set sourced from the 2024 Olympics. Experimental results demonstrate that incorporating uncertainty as additional information during the fine-tuning phase can largely improve the model's evaluation performance in OOD scenarios. The code and data are released at: https://github.com/hasakiXie123/LLM-Evaluator-Uncertainty.
Qiujie Xie, Qingqiu Li, Zhuohao Yu 0001, Yuejie Zhang, Yue Zhang 0004, Linyi Yang
ICLR4
2025 EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
abstract
Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and textual instructions, the goal is to generate future frames of the ego-centric video. Inspired by the notion that hand-object interactions (HOI) in ego-centric videos represent the primary intentions and actions of the current actor, we present EgoExo-Gen that explicitly models the hand-object dynamics for cross-view video prediction. EgoExo-Gen consists of two stages. First, we design a cross-view HOI mask prediction model that anticipates the HOI masks in future ego-frames by modeling the spatio-temporal ego-exo correspondence. Next, we employ a video diffusion model to predict future ego-frames using the first ego-frame and textual instructions, while incorporating the HOI masks as structural guidance to enhance prediction quality. To facilitate training, we develop a fully automated pipeline to generate pseudo HOI masks for both ego- and exo-videos by exploiting vision foundation models. Extensive experiments demonstrate that our proposed EgoExo-Gen achieves better prediction performance compared to previous video prediction models on the public Ego-Exo4D and H2O benchmark datasets, with the HOI masks significantly improving the generation of hands and interactive objects in the ego-centric videos.
Jilan Xu, Yifei Huang 0002, Baoqi Pei, Junlin Hou, Qingqiu Li, Guo Chen 0006, Yuejie Zhang, Rui Feng 0001, Weidi Xie
ICLR7
2025 TGSAM-2: Text-Guided Medical Image Segmentation Using Segment Anything Model 2
Runtian Yuan, Ling Zhou 0002, Jilan Xu, Qingqiu Li, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
MICCAI (10)6
2025 Open-Set Image Tagging with Multi-Grained Text Supervision
abstract
This paper introduces the Recognize Anything Plus Model (RAM++), an open-set image tagging model effectively leveraging multi-grained text supervision. Previous approaches (e.g., CLIP) primarily utilize global text supervision paired with images, leading to sub-optimal performance in recognizing multiple individual semantic tags. In contrast, RAM++ seamlessly integrates individual tag supervision with global text supervision, all within a unified alignment framework. This integration not only ensures efficient recognition of predefined tag categories, but also enhances generalization capabilities for diverse open-set categories. Furthermore, RAM++ employs large language models (LLMs) to convert semantically constrained tag supervision into more expansive tag description supervision, thereby enriching the scope of open-set visual description concepts. Comprehensive evaluations on various image recognition benchmarks demonstrate RAM++ exceeds existing state-of-the-art (SOTA) open-set image tagging models on most aspects. Specifically, for predefined commonly used tag categories, RAM++ showcases 10.2 mAP and 15.4 mAP enhancements over CLIP on OpenImages and ImageNet. For open-set categories beyond predefined, RAM++ records improvements of 5.0 mAP and 6.4 mAP over CLIP and RAM respectively on OpenImages. For diverse human-object interaction phrases, RAM++ achieves 7.8 mAP and 4.7 mAP improvements on the HICO benchmark.
Yi-Jie Huang, Youcai Zhang, Rui Feng 0001, Yuejie Zhang, Yanchun Xie, Lei Zhang 0001
ACM Multimedia6
2025 Text-Promptable Propagation for Referring Medical Image Sequence Segmentation
abstract
Referring Medical Image Sequence Segmentation (Ref-MISS) is a novel and challenging task that aims to segment anatomical structures in medical image sequences (e.g., endoscopy, ultrasound, CT, and MRI) based on natural language descriptions. Existing 2D and 3D segmentation models struggle to explicitly track objects of interest across medical image sequences, and lack support for interactive, text-driven guidance. To address these limitations, we propose Text-Promptable Propagation (TPP), which enables the recognition of referred objects through cross-modal referring interaction, and maintains continuous tracking across the sequence via Transformer-based triple propagation, using text embeddings as queries. To support this task, we curate a large-scale benchmark, Ref-MISS-Bench, which covers 4 imaging modalities and 20 different organs and lesions. Experimental results on this benchmark demonstrate that TPP consistently outperforms state-of-the-art methods in both medical segmentation and referring video object segmentation. Code and data are available at https://github.com/yuanruntian/TPP.
Runtian Yuan, Mohan Chen 0001, Jilan Xu, Ling Zhou 0002, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ACM Multimedia6
2025 EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues
abstract
Qiming Feng, Qiujie Xie, Xiaolong Wang, Qingqiu Li, Yuejie Zhang, Rui Feng, Tao Zhang, Shang Gao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qiming Feng, Qiujie Xie, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
NAACL (Long Papers)5
2025 AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation
abstract
Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding, current MLMMs still face two major challenges: (1) insufficient region-level understanding and interaction, and (2) limited accuracy and interpretability due to single-step prediction. In this paper, we address these challenges by empowering MLMMs with anatomy-centric reasoning capabilities to enhance their interactivity and explainability. Specifically, we propose an Anatomical Ontology-Guided Reasoning (AOR) framework that accommodates both textual and optional visual prompts, centered on region-level information to enable multimodal multi-step reasoning. We also develop AOR-Instruction, a large instruction dataset for MLMMs training, under the guidance of expert physicians. Our experiments demonstrate AOR's superior performance in both Visual Question Answering (VQA) and report generation tasks. Code and data are available at: https://github.com/Liqq1/AOR.
Qingqiu Li, Zihang Cui, Seongsu Bae, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen, Shang Gao 0003, Junjun He
NeurIPS6
2025 Interannual Multicrop Identification in Large Area Based on Optimized Monthly Tile Classification Model With Spatio-Temporal Distance Features Fusion
abstract
Accurate crop identification is crucial for agricultural trade, market risk management, and food security. Current research on automatic interannual sample extraction for crop mapping often emphasizes multisource feature fusion but overlooks the importance of feature distance differences in crop seeding processes. This study focuses on crop mapping in the Hetao Plain from 2020 to 2023, using high-resolution (HR) Sentinel-1 and Sentinel-2 remote sensing data. We introduce a method called BGSI-DFF-WMRF, which combines the bidirectional global selection index (BGSI) and distance feature fusion (DFF) under the interannual weaving net month probability random forest (WMRF). BGSI captures the coupling between phenological and spatial factors necessary for crop growth in different regions under feature fusion. Additionally, sample migration under feature fusion enhances the accuracy and representativeness of sample points across different years. WMRF integrates a monthly classifier with multisource feature distance interpolation. The BGSI-DFF-WMRF method achieved over 82% classification accuracy for crops like wheat, corn, and sunflower in the Hetao Irrigation District (HID) region, with an accuracy of 91.81% in the western area. Field samples from 2023 were successfully applied to previous years (2020–2022) through feature fusion expression (FFE) transfer. The method outperformed existing products and local statistical data, particularly for corn, demonstrating high accuracy and robustness in crop mapping. Coupling spatio-temporal factors enhances large-scale crop identification and holds great significance for the advancement and widespread adoption of large-scale crop identification techniques.
Sijing Tian, Guo Zhang 0001, Hao Cui 0002, Yuejie Zhang, Qinghong Sheng
IEEE Trans. Geosci. Remote. Sens.5
2024 DeepPointMap: Advancing LiDAR SLAM with Unified Neural Descriptors
abstract
Point clouds have shown significant potential in various domains, including Simultaneous Localization and Mapping (SLAM). However, existing approaches either rely on dense point clouds to achieve high localization accuracy or use generalized descriptors to reduce map size. Unfortunately, these two aspects seem to conflict with each other. To address this limitation, we propose an unified architecture, DeepPointMap, achieving excellent preference on both aspects. We utilize neural network to extract highly representative and sparse neural descriptors from point clouds, enabling memory-efficient map representation and accurate multi-scale localization tasks (e.g., odometry and loop-closure). Moreover, we showcase the versatility of our framework by extending it to more challenging multi-agent collaborative SLAM. The promising results obtained in these scenarios further emphasize the effectiveness and potential of our approach.
Xiaze Zhang, Ziheng Ding, Yuejie Zhang, Wenchao Ding 0001, Rui Feng 0001
AAAI4
2024 Retrieval-Augmented Egocentric Video Captioning
abstract
Understanding human actions from videos offirst-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper, (1) we develop EgoInstructor, a retrieval-augmented multimodal captioning model that automatically retrieves semantically relevant third-person instructional videos to enhance the video captioning of egocentric videos, (2) for training the cross-view retrieval module, we devise an au-tomatic pipeline to discover ego-exo video pairs from distinct large-scale egocentric and exocentric datasets, (3) we train the cross-view retrieval module with a novel EgoEx-oNCE loss that pulls egocentric and exocentric video features closer, by aligning them to shared text features that describe similar actions, (4) through extensive experiments, our cross-view retrieval module demonstrates superior performance across seven benchmarks. Regarding egocen-tric video captioning, EgoInstructor exhibits significant improvements by leveraging third-person videos as references.
Jilan Xu, Yifei Huang 0002, Junlin Hou, Guo Chen 0006, Yuejie Zhang, Rui Feng 0001, Weidi Xie
CVPR5
2024 Fine-Granularity Face Sketch Synthesis
abstract
Generative Adversarial Networks (GANs) are often used in face sketch synthesis due to their powerful ability in image generation. However, most GAN based synthesis methods took the entire face as the minimum unit. Differently, we propose a novel fine-granularity face sketch synthesis framework in this paper. The core idea is to first capture local information at a fine granularity (i.e., facial component), and then generate a complete face sketch based on the fine-grained information. Specifically, we partition the face sketch into multiple components, and then train a parallel network for each component. A condition enhanced detail repair network is further designed to correct the mismatches and deformations produced during parallel generation. Extensive experiments show that our approach outperforms state-of-the-art methods from both the qualitative and quantitative perspectives.
Yangdong Chen, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICASSP3
2024 ControlCap: Controllable Captioning via No-Fuss Lexicon
abstract
Controllable captioning has received much attention in recent years. Although substantial progress has been made, existing methods still face challenges such as high training costs, intricate control signals and limited control capabilities. To address these issues, we propose a straightforward and unified framework called ControlCap. It uses a no-fuss lexicon as control signal and controls the style and content of visual descriptions through Soft Guidance (a global guide to the caption distribution) and Hard Force (integrating signals without additional training). Extensive experiments, both quantitative and qualitative, have been conducted on three benchmark captioning tasks. Results demonstrate the control ability of ControlCap: it can produce controlled captions that are coherent and diverse while keeping the core content intact.
Qiujie Xie, Qiming Feng, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP3
2024 Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning
abstract
Video paragraph captioning task aims to generate a detailed, fluent and relevant paragraph for a given video. Prior studies often focus on isolating visual objects (potential main components in a sentence) from the overall video content. They rarely explore the latent semantic relations between objects and high-level video concepts, resulting in dull or even incorrect descriptions. To create fine-grained and contextually relevant paragraph captions, we propose a novel framework that constructs a concept graph from a commonsense knowledge base and infers richer semantic meaning from the visual objects. Moreover, we employ a Vision-Guided Concept Selection Network that incorporates an under-sentence supervision mechanism to align the external knowledge with the visual information. Through extensive experiments on ActivityNet captions and YouCook2, the effectiveness of our method is demonstrated compared to state-of-the-art methods.
Guorui Yu, Yimin Hu, Yiqian Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP4
2024 Tag2Text: Guiding Vision-Language Model via Image Tagging
abstract
This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a limited detector, our approach utilizes tags parsed from its paired text to learn an image tagger and meanwhile provides guidance to vision-language models. Given that, Tag2Text can utilize large-scale annotation-free image tags in accordance with image-text pairs, and provides more diverse tag categories beyond objects. Strikingly, Tag2Text showcases the ability of a foundational image tagging model, with superior zero-shot performance even comparable to full supervision manner. Moreover, by leveraging tagging guidance, Tag2Text effectively enhances the performance of vision-language models on both generation-based and alignment-based tasks. Across a wide range of downstream benchmarks, Tag2Text achieves state-of-the-art results with similar model sizes and data scales, demonstrating the efficacy of the proposed tagging guidance.
Youcai Zhang, Jinyu Ma, Rui Feng 0001, Yuejie Zhang, Yandong Guo, Lei Zhang 0001
ICLR6
2024 Temporal Feature Aggregation for Efficient 2D Video Grounding
abstract
Video grounding aims to locate the target video moment in an untrimmed video based on a text query. Most existing methods employ 3D CNNs as the video feature extractor, incurring substantial computational costs. Only a few methods use 2D backbones for video feature extraction, and they suffer from diminished accuracy due to the inherent lack of temporal information within 2D features. To address this problem, we propose a novel 2D video grounding method called TFA that improves accuracy while minimizing computational costs. Our approach involves a query-guided temporal feature aggregation module designed to explicitly capture temporal information. We disentangle time intervals of input video frames and prediction spans to reduce computational overhead. Additionally, we introduce deformable attention into the multi-modal encoder for further enhancement. Extensive experiments on two public datasets demonstrate that our method outperforms previous 2D video grounding methods and achieves competitive results with most 3D methods at significantly reduced costs.
Mohan Chen 0001, Yiren Zhang, Jueqi Wei, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICME4
2024 Memory-Augmented Transformer for Efficient End-to-End Video Grounding
abstract
Video grounding aims to localize a specific segment corresponding to a text query in an untrimmed video. Due to the tremendous computational cost required to process the video frames, the de facto paradigm of video grounding is to extract video features using pretrained video encoders. The parameters of the video encoders are fixed during training, which limits the performance of the localization model. To solve this problem, we propose a Memory-Augmented Transformer (MAT) model. Specifically, each video is split into non-overlapping clips, and our MAT processes videos in a clip-by-clip manner while caching video features into FIFO cached memory queues. By enabling early return, our MAT outperforms previous methods with only less than 60% frames seen. Extensive experimental results on three public benchmark datasets demonstrate that our MAT can achieve competitive performance while being much more efficient than currently prevailing two-stage methods. Code is available at https://github.com/xuyw1997/MAT.
Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICME3
2024 Anatomical Structure-Guided Medical Vision-Language Pre-training
Qingqiu Li, Xiaohan Yan, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen
MICCAI (11)5
2024 CT2C-QA: Multimodal Question Answering over Chinese Text, Table and Chart
abstract
Multimodal Question Answering (MMQA) is crucial as it enables comprehensive understanding and accurate responses by integrating insights from diverse data representations such as tables, charts, and text. Most existing researches in MMQA only focus on two modalities such as image-text QA, table-text QA and chart-text QA, and there remains a notable scarcity in studies that investigate the joint analysis of text, tables, and charts. In this paper, we present CT2C-QA, a pioneering Chinese reasoning-based QA dataset that includes an extensive collection of text, tables, and charts, meticulously compiled from 200 selectively sourced webpages. Our dataset simulates real webpages and serves as a great test for the capability of the model to analyze and reason with multimodal data, because the answer to a question could appear in various modalities, or even potentially not exist at all. Additionally, we present AED (Allocating, Expert and Decision), a multi-agent system implemented through collaborative deployment, information interaction, and collective decision-making among different agents. Specifically, the Assignment Agent is in charge of selecting and activating expert agents, including those proficient in text, tables, and charts. The Decision Agent bears the responsibility of delivering the final verdict, drawing upon the analytical insights provided by these expert agents. We execute a comprehensive analysis, comparing AED with various state-of-the-art models in MMQA, including GPT-4. The experimental outcomes demonstrate that current methodologies, including GPT-4, are yet to meet the benchmarks set by our dataset.
Tianhao Cheng, Yuejie Zhang, Ying Cheng 0005, Rui Feng 0001
ACM Multimedia3
2023 Enhanced Knowledge Injection for Radiology Report Generation
abstract
Automatic generation of radiology reports holds crucial clinical value, as it can alleviate substantial workload on radiologists and remind less experienced ones of potential anomalies. Despite the remarkable performance of various image captioning methods in the natural image field, generating accurate reports for medical images still faces challenges, i.e., disparities in visual and textual data, and lack of accurate domain knowledge. To address these issues, we propose an enhanced knowledge injection framework, which utilizes two branches to extract different types of knowledge. The Weighted Concept Knowledge (WCK) branch is responsible for introducing clinical medical concepts weighted by TF-IDF scores. The Multimodal Retrieval Knowledge (MRK) branch extracts triplets from similar reports, emphasizing crucial clinical information related to entity positions and existence. By integrating this finer-grained and well-structured knowledge with the current image, we are able to leverage the multi-source knowledge gain to ultimately facilitate more accurate report generation. Extensive experiments have been conducted on two public benchmarks, demonstrating that our method achieves superior performance over other state-of-the-art methods. Ablation studies further validate the effectiveness of two extracted knowledge sources.
Qingqiu Li, Jilan Xu, Runtian Yuan, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
BIBM5
2023 Semi-MedSeq: Semi-supervised Semantic Segmentation for Medical Image Sequences
abstract
In clinical practice, medical imaging techniques include 2D video-based examinations that capture sequential scans, and 3D volumetric imaging that forms a comprehensive 3D representation from a stack of 2D slices. The medical image sequences produced by the above techniques provide valuable spatio-temporal characteristics for analysis and segmentation, but the annotation of image sequences is extremely time-consuming and labor-intensive. To exploit the coherence and address the scarcity of labeled data, we propose a novel semi-supervised semantic segmentation framework for medical image sequences, which consists of a conditional network and a denoising network. Specifically, we embed a Sequential Feature Reconstruction module into both networks. This module reconstructs the target frame from contiguous frames and captures their shared visual features. Guided by the context-enhancing information from the conditioning network, the denoising network suppresses background noise via a Diffusion-based Noise Elimination module. Extensive experiments are conducted on 2D and 3D tasks, including cardiac segmentation, polyp segmentation, placenta vessel segmentation and abdomen multi-organ segmentation. The results show our method is superior to existing semi-supervised methods and exhibits advantages over fully-supervised medical image segmentation methods with only 1/2 labeled data, validating its effectiveness and generalization ability.
Runtian Yuan, Jilan Xu, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM4
2023 Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision
abstract
This paper considers the problem of open-vocabulary semantic segmentation (OVS), that aims to segment objects of arbitrary classes beyond a pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor, which only exploits web-crawled imagetext pairs for pre-training without using any mask annotations. OVSegmentor assembles the image pixels into a set of learnable group tokens via a slotattention based binding module, then aligns the group tokens to corresponding caption embeddings. Second, we propose two proxy tasks for training, namely masked entity completion and cross-image mask consistency. The former aims to infer all masked entities in the caption given group tokens, that enables the model to learn fine-grained alignment between visual groups and text entities. The latter enforces consistent mask predictions between images that contain shared entities, encouraging the model to learn visual invariance. Third, we construct CC4M dataset for pre-training by filtering CC12M with frequently appeared entities, which significantly improves training efficiency. Fourth, we perform zero-shot transfer on four benchmark datasets, PASCAL VOC, PASCAL Context, COCO Object, and ADE20K. OVSegmentor achieves superior results over state-of-the-art approaches on PASCAL VOC using only 3% data (4M vs 134M) for pre-training.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Yi Wang 0074, Yu Qiao 0001, Weidi Xie
CVPR3
2023 Large Language Models are Complex Table Parsers
abstract
With the Generative Pre-trained Transformer 3.5 (GPT-3.5)exhibiting remarkable reasoning and comprehension abilities in Natural Language Processing (NLP), most Question Answering (QA) research has primarily centered around general QA tasks based on GPT, neglecting the specific challenges posed by Complex Table QA.In this paper, we propose to incorporate GPT-3.5 to address such challenges, in which complex tables are reconstructed into tuples and specific prompt designs are employed for dialogues.Specifically, we encode each cell's hierarchical structure, position information, and content as a tuple.By enhancing the prompt template with an explanatory description of the meaning of each tuple and the logical reasoning process of the task, we effectively improve the hierarchical structure awareness capability of GPT-3.5 to better parse the complex tables.Extensive experiments and results on Complex Table QA datasets, i.e., the open-domain dataset HiTAB and the aviation domain dataset AIT-QA show that our approach significantly outperforms previous work on both datasets, leading to state-of-theart (SOTA) performance.
Changkai Ji, Yuejie Zhang, Yingwen Wang, Rui Feng 0001
EMNLP3
2023 Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal Knowledge
abstract
Video paragraph captioning task aims at generating a fine-grained, coherent and relevant paragraph for a video. Different from the images where objects are static, the temporal states of objects are changing in videos. The dynamic information could be contributed to understanding the whole video content. Existing works rarely put focus on modeling the dynamic changing state of the objects in the videos, causing the activities occurred in videos are poorly or wrongly depicted in paragraphs. To address this problem, we propose a novel Object State Tracking Network, which can capture the temporal state change of objects. However, due to the similarity of the consecutive frames in the videos, the information of the video is redundant and noisy. We further propose a semantic alignment mechanism, and enable the sentence information to refine the visual information. Extensive experiments on ActivityNet Captions demonstrate the effectiveness of our method.
Yimin Hu, Guorui Yu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICASSP3
2023 SCSGNet: Spatial-Correlated and Shape-Guided Network for Breast Mass Segmentation
abstract
Automatic and accurate breast mass segmentation plays a crucial role in the early diagnosis of breast cancer. However, it has been a challenging task for two main reasons: (1) Breast masses are diverse; and (2) The boundaries of masses are ambiguous. To address these problems, we propose a Spatial-Correlated and Shape-Guided Network (SCSGNet), which combines global context extraction with local boundary refinement. Specifically, the high-level features are aggregated to produce a global map as the initial guidance area, and a Series-Parallel Feature Fusion (SPFF) module is added to capture masses of different shapes and sizes. Besides, we design a Dynamic Long-range Correlation Capture (DLCC) module to capture the spatial correlation of masses at different positions. Finally, we devise a Triplet Attention Guide (TAG) module to iteratively update the feature map and refine the boundary. Experiments on two public datasets demonstrate that our method achieves superior performance over other state-of-the-art methods.
Qingqiu Li, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001
ICASSP4
2023 Boosting Fine-Grained Sketch-Based Image Retrieval with Self-Supervised Learning
abstract
Fine-grained sketch-based image retrieval (FG-SBIR) aims at aligning images and sketches at the instance level. It is a challenging task as there are significant differences between sketch and image. Existing methods usually produce less desired performance due to the lack of large-scale fine-grained image-sketch datasets and the strong dependence on the classification models pretrained on ImageNet. In this paper, we propose a better self-supervised pre-trained FG-SBIR model which does not depend on large-scale annotated datasets. Only images and their corresponding edge maps are used at the pre-training stage. Mixed modal transformation is designed to generate different mixed-up views. The FG-SBIR model is pre-trained by minimizing the distance between the views of the same instance and then fine-tuned by a simple triplet loss. With a plain downstream network, it achieves generally better performance than state-of-the-art models on three widely used FG-SBIR datasets.
Zhaolong Zhang, Yangdong Chen, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022
ICASSP3
2023 Video Captioning via Relation-Aware Graph Learning
abstract
Recent neural models for video captioning usually employed an encoder-decoder framework. However, most approaches either neglected the spatial and temporal interactions between objects in a video or implicitly modelled the interactions, resulting in less desired performance. In this paper, we propose a novel relation-aware graph learning framework. It explicitly models both spatial and temporal relations for objects. In particular, a relation-aware graph is designed to depict the spatial relations between different objects in a scene. Parallelly, a temporal graph network is designed to perform relational reasoning for the same objects in adjacent frames. Features of both types of relations are learned and fused for the follow-up language decoder. Experiments on two bench-mark datasets show the effectiveness of our framework. It achieves state-of-the-art performance with CIDEr scores on MSVD and MSR-VTT.
Heming Jing, Qiujie Xie, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP4
2023 Class-aware Variational Auto-encoder for Open Set Recognition
abstract
Compared with traditional classification models trained under the closed world assumption, Open Set Recognition (OSR) requires accurate classification for known classes as well as rejection for unknown ones. By modeling the distribution of each known class, Conditional Variational Auto-encoder (CVAE) has achieved great success in OSR, even though it was originally proposed for image generation. In this paper, we propose a novel two-stage learning framework, Class-aware Variational Auto-encoder (CA-VAE) to better adapt CVAE to the OSR task. Pre-derived attention images are taken as the objective target for reconstruction, thus model is implicitly directed to focus on the class-discriminative regions of the image. In this way, the learned latent representation is de-biased towards class-aware. Experiments on standard image datasets demonstrate the outperformance of the proposed method over existing ones, which achieves new state-of-the-art results. Codes are available at https://github.com/roywang021/CA-VAE.
Ling Su, Yingzi Ye, Yuejie Zhang, Rui Feng 0001
ICME7
2023 Conditional Video-Text Reconstruction Network with Cauchy Mask for Weakly Supervised Temporal Sentence Grounding
abstract
Temporal sentence grounding aims to detect the target segment most related to a given query in an untrimmed video. To alleviate the expensive annotation cost for temporal labels, researchers paid more attention to weakly supervised setting. Prior studies neglected the utilization of video representation reconstruction, which led to an unbalanced alignment learning. Moreover, they used different strategies to generate proposals which ignored the temporal structure in a query. In this paper, we propose a novel Conditional Video-Text Reconstruction Network (CVTRN). It supports conditional reconstruction of video and text representation. Specifically, video and text features are fused to compute semantic alignment, which is the condition of reconstruction. A new mask strategy for mask conditioned sentence reconstruction is also devised. This strategy focuses more on boundary regions than the widely used Gaussian mask in previous methods. Experimental results on two public benchmark datasets show that our CVTRN outperforms the state-of-the-art methods.
Jueqi Wei, Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
ICME4
2023 SPTNET: Span-based Prompt Tuning for Video Grounding
abstract
When a Pre-trained Language Model (PLM) is adopted in video grounding task, it usually acts as a text encoder without having its knowledge fully utilized. Also, there exists an inconsistency problem between the pre-training and downstream objectives. To solve the issues, we propose a new paradigm, named Span-based Prompt Tuning (SPTNet). It can convert the video grounding task into a cloze form. Specifically, a query is first changed into a form with mask token by a template, then the video and the query embeddings are integrated through a cross-modal transformer. The start and end points of the query matching time span are predicted with the embedding of the mask token. Experimental results on two public benchmarks ActivityNet Captions and Charades-STA show that our SPTNet achieves surpassing performance compared with state-of-the-art methods.
Yiren Zhang, Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
ICME4
2023 CAMG: Context-Aware Moment Graph Network for Multimodal Temporal Activity Localization via Language
Yuelin Hu, Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
NLPCC (1)3
2023 Enhancing Open-Set Object Detection via Uncertainty-Boxes Identification
Wei Ji 0009, Dongqin Wu, Weijia Fu, Yingwen Wang, Yuejie Zhang, Rui Feng 0001
PRCV (8)6
2023 Deep cross-modal hashing with fine-grained similarity
Yangdong Chen, Jiaqi Quan, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022
Appl. Intell.3
2023 Enhanced graph neural network for session-based recommendation
Zhenzhen Sheng, Tao Zhang 0022, Yuejie Zhang, Shang Gao 0003
Expert Syst. Appl.3
2023 Graph classification via discriminative edge feature learning
abstract
Spectral graph convolutional neural networks (GCNNs) have been producing encouraging results in graph classification tasks. However, most spectral GCNNs utilize fixed graphs when aggregating node features, while omitting edge feature learning and failing to get an optimal graph structure. Moreover, many existing graph datasets do not provide initialized edge features, further restraining the ability of learning edge features via spectral GCNNs. In this paper, we try to address this issue by designing an edge feature scheme and an add-on layer between every two stacked graph convolution layers in spectral GCNN. Both are lightweight while effective in filling the gap between edge feature learning and performance enhancement of graph classification. The edge feature scheme makes edge features adapt to node representations at different spectral graph convolution layers. The add-on layers help adjust the edge features to an optimal graph structure. To test the effectiveness of our method, we take Euclidean positions as initial node features and extract graphs with semantic information from point cloud objects. The node features of our extracted graphs are more scalable for edge feature learning than most existing graph datasets (in one-hot encoded label format). Three new graph datasets are constructed based on ModelNet40, ModelNet10 and ShapeNet Part datasets. Experimental results show that our method outperforms state-of-the-art graph classification methods on the new datasets. Our code and the constructed graph datasets will be released to the community.
Xuequan Lu, Shang Gao 0003, Antonio Robles-Kelly, Yuejie Zhang
Pattern Recognit.5
2022 Single-Modality Endoscopic Polyp Segmentation via Random Color Reversal Synthesis and Two-Branched Learning
abstract
Endoscopic polyp segmentation plays a fundamental role in the diagnosis and treatment of colorectal cancer. However, polyp segmentation often suffers from limited accuracy due to its large variations in appearance, blurry boundary and severe imbalanced illumination. In this paper, we propose a novel Translation Assisted Segmentation Network (TASNet) for polyp segmentation of single-modality endoscopic images. It consists of two branches, i.e. an image-to-image translation branch and an image segmentation branch. These two branches communicate via a shared encoder. For the image-to-image translation branch, a Color Reversal Strategy is established to treat the original image as source image and synthesize target images. Moreover, we introduce a Random Color Reversal Synthesis module for progressive segmentation. Extensive experiments show that our framework achieves superior performance than state-of-the-art methods on five widely-used endoscopic image datasets.
Mingzhu Chen, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM5
2022 Cross-Field Transformer for Diabetic Retinopathy Grading on Two-field Fundus Images
abstract
Automatic diabetic retinopathy (DR) grading based on fundus photography has been widely explored to benefit the routine screening and early treatment. Existing researches generally focus on single-field fundus images, which have limited field of view for precise eye examinations. In clinical applications, ophthalmologists adopt two-field fundus photography as the dominating tool, where the information from each field (i.e., macula-centric and optic disc-centric) is highly correlated and complementary, and benefits comprehensive decisions. However, automatic DR grading based on two-field fundus photography remains a challenging task due to the lack of publicly available datasets and effective fusion strategies. In this work, we first construct a new benchmark dataset (DRTiD) for DR grading, consisting of 3,100 two-field fundus images. To the best of our knowledge, it is the largest public DR dataset with diverse and high-quality two-field images. Then, we propose a novel DR grading approach, namely Cross-Field Transformer (CrossFiT), to capture the correspondence between two fields as well as the long-range spatial correlations within each field. Considering the inherent two-field geometric constraints, we particularly define aligned position embeddings to preserve relative consistent position in fundus. Besides, we perform masked cross-field attention during interaction to filter the noisy relations between fields. Extensive experiments on our DRTiD dataset and a public DeepDRiD dataset demonstrate the effectiveness of our CrossFiT network. The new dataset and the source code of CrossFiT will be publicly available at https://github.com/DU-VTS/DRTiD.
Junlin Hou, Jilan Xu, Yuejie Zhang, Haidong Zou, Lina Lu, Wenwen Xue, Rui Feng 0001
BIBM5
2022 An Interpretable Causal Approach for Bronchopulmonary Dysplasia Prediction
abstract
In this paper, we focus on the Bronchopulmonary Dysplasia (BPD) prediction task, which aims to identify the BPD in premature infants based on the given medical images. The existing methods for this task sometimes learn the spurious relations (confounders) between the image and label while ignore the causal features due to the limited size of the dataset, which largely influences the interpretability and robustness of the model. To address this challenge problem, we propose an interpretable causal method for BPD prediction, which can eliminate the irrelevant features and capture the causal features. We term our method as Causal Intervention by semantic instrumental Variable (CisiV). First, we design a causal structure graph modeling module to learn the representations of instrumental variables and confounders automatically. The constraints of confounders further guarantee the instrumental variables validity. Meanwhile, we add congenital attributes of premature infants to enable the learned instrumental variables to have medical semantics via a semantic matching module. We apply our model to BPD prediction in premature infants and achieve promising results. Extensive experiments indicate that CisiV could facilitate early intervention and provide support for clinical decision-making. Codes will be released to the community.
Wei Ji 0009, Liangfeng Tang, Yuejie Zhang, Rui Feng 0001
BIBM5
2022 Multi-contrast High Quality MR Image Super-Resolution with Dual Domain Knowledge Fusion
abstract
Multi-contrast high quality high-resolution (HR) Magnetic Resonance (MR) images enrich available information for diagnosis and analysis. Deep convolutional neural network methods have shown promising ability for MR image super-resolution (SR) given low-resolution (LR) MR images. Methods taking HR images as references (Ref) have made progress to enhance the effect of MR images SR. However, existing multi-contrast MR image SR approaches are based on contrasting-expanding backbones, which lose high frequency information of Ref image during downsampling. They also failed to transfer textures of Ref image into target domain. In this paper, we propose Edge Mask Transformer UNet (EMFU) for accelerating MR images SR. We propose Edge Mask Transformer (EMF) to generate global details and texture representation of target domain. Dual domain fusion module in UNet aggregates semantic information of the representation and LR image of target domain. Specifically, we extract and encode edge masks to guide the attention in EMF by re-distributing the embedding tensors, so that the network allocates more attention to image edge area. We also design a dual domain fusion module with self-attention and cross-attention to deeply fuse semantic information of multiple protocols for MRI. Extensive experiments show the effectiveness of our proposed EMFU, which surpasses state-of-the-art methods on benchmarks quantitatively and visually. Codes will be released to the community.
Runhan Wang, Weijia Fu, Yuejie Zhang, Rui Feng 0001
BIBM5
2022 MedSeq: Semantic Segmentation for Medical Image Sequences
abstract
Medical image segmentation plays a critical role in computer-aided diagnosis, while the diversity and complexity of medical images make it difficult to segment precisely. In practice, medical images of specific modalities (e.g. Magnetic Resonance Imaging, Colonoscopy and Ultrasonography) are collected as sequences independently for every patient. However, 1) there exists few works exploiting sequence information among successive frames, neglecting inter-frame relationships that are useful to locate target objects; 2) the performance of medical image segmentation is limited to the low contrast or blurry boundary of medical images, and intra-frame dependencies are not fully explored. Thus in this paper, we propose MedSeq for segmenting objects of interest in medical image sequences. Following the “locate-then-refine” paradigm, we locate target regions by modeling cross-frame relationships and then perform refinement on coarse masks. More specifically, we design a Cross-frame Attention module to learn correlations among frames, taking advantages of their similar appearances. For refinement, we propose a novel Boundary-aware Transformer to improve the segmentation of boundary patches. Extensive experiments are conducted on benchmark datasets of Cardiac Segmentation and Video Polyp Segmentation. Our method achieves superior performance over the state-of-the-art methods.
Runtian Yuan, Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM4
2022 CREAM: Weakly Supervised Object Localization via Class RE-Activation Mapping
abstract
Weakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) de-rived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomplete localization problem). In this paper, we empirically prove that this problem is associated with the mixup of the activation values between less discrimi-native foreground regions and the background. To address it, we propose Class RE-Activation Mapping (CREAM), a novel clustering-based approach to boost the activation values of the integral object regions. To this end, we in-troduce class-specific foreground and background context embeddings as cluster centroids. A CAM-guided momen-tum preservation strategy is developed to learn the context embeddings during training. At the inference stage, the re-activation mapping is formulated as a parameter es-timation problem under Gaussian Mixture Model, which can be solved by deriving an unsupervised Expectation- Maximization based soft-clustering algorithm. By simply integrating CREAM into various WSOL approaches, our method significantly improves their performance. CREAM achieves the state-of-the-art performance on CUB, ILSVRC and OpenImages benchmark datasets. Code will be avail-able at https://github.com/lazzcharles/CREAM.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
CVPR3
2022 Semantic-Driven Saliency-Context Separation for Video Captioning
abstract
Video captioning aims at generating a natural language de-scription for a given video clip including not only salient sce-narios but also contextual scenarios. The former reveal the highlight of a video and are usually the focus of most existing captioning methods. The latter, however, are not well ex-plored and even ignored easily, though they may provide cer-tain detailed and latent information that can help with a better understanding of the video. To effectively exploit the infor-mation contained in both, a novel video captioning network is proposed. It has two key modules: Cross-Modality Selection (CMS) and Saliency-Context Adaptive Decoder (SCAD). Specifically, CMS mainly focuses on utilizing the semantic information to distinguish saliency and context. Meanwhile, SCAD adaptively identifies both the saliency and context to generate more detailed and precise captions. Experiments on two benchmark datasets, i.e., MSVD and MSR-VTT, demon-strate the effectiveness of our model through the comparison with state-of-the-art methods.
Heming Jing, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICME2
2022 Self-Supervised Video Representation Learning with Motion-Contrastive Perception
abstract
Visual-only self-supervised learning has achieved significant improvement in video representation learning. Existing related methods encourage models to learn video representations by utilizing contrastive learning or designing specific pretext tasks. However, some models are likely to focus on the background, which is unimportant for learning video representations. To alleviate this problem, we propose a new view called long-range residual frame to obtain more motion-specific information. Based on this, we propose the Motion-Contrastive Perception Network (MCPNet), which consists of two branches, namely, Motion Information Perception (MIP) and Contrastive Instance Perception (CIP), to learn generic video representations by focusing on the changing areas in videos. Specifically, the MIP branch aims to learn fine-grained motion features, and the CIP branch performs contrastive learning to learn overall semantics information for each instance. Experiments on two benchmark datasets UCF-101 and HMDB-51 show that our method outperforms current state-of-the-art visual-only self-supervised approaches.
Ying Cheng 0005, Yuejie Zhang, Rui Feng 0001
ICME3
2022 STDNet: Spatio-Temporal Decomposed Network for Video Grounding
abstract
Previous methods for video grounding treated either the query or the video as a whole, while neglecting their respective semantics in the orthogonal space and time dimensions. Since spatial semantics appears frequently in a video, temporal semantics is more discriminative and deserves more attention. Based on such considerations, we propose a novel Spatio-Temporal Decomposed Network (STDNet) which decomposes the query and the video into their spatial and temporal semantics, respectively. Specifically, spatial and temporal words are selected from the query, and the video is split into two pathways. Spatial cross-modal attention is computed first and serves as prior knowledge for temporal attention. A new localization strategy is also devised which regresses the segment's start conditioned on the end and essentially breaks the independence assumption made in previous methods. Experimental results on three public benchmark datasets show that our STDNet outperforms the state-of-the-art methods.
Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICME2
2022 TCCNet: Temporally Consistent Context-Free Network for Semi-supervised Video Polyp Segmentation
abstract
Automatic video polyp segmentation (VPS) is highly valued for the early diagnosis of colorectal cancer. However, existing methods are limited in three respects: 1) most of them work on static images, while ignoring the temporal information in consecutive video frames; 2) all of them are fully supervised and easily overfit in presence of limited annotations; 3) the context of polyp (i.e., lumen, specularity and mucosa tissue) varies in an endoscopic clip, which may affect the predictions of adjacent frames. To resolve these challenges, we propose a novel Temporally Consistent Context-Free Network (TCCNet) for semi-supervised VPS. It contains a segmentation branch and a propagation branch with a co-training scheme to supervise the predictions of unlabeled image. To maintain the temporal consistency of predictions, we design a Sequence-Corrected Reverse Attention module and a Propagation-Corrected Reverse Attention module. A Context-Free Loss is also proposed to mitigate the impact of varying contexts. Extensive experiments show that even trained under 1/15 label ratio, TCCNet is comparable to the state-of-the-art fully supervised methods for VPS. Also, TCCNet surpasses existing semi-supervised methods for natural image and other medical image segmentation tasks.
Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
IJCAI3
2022 IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training
abstract
Vision-Language Pre-training (VLP) with large-scale image-text pairs has demonstrated superior performance in various fields. However, the image-text pairs co-occurrent on the Internet typically lack explicit alignment information, which is suboptimal for VLP. Existing methods proposed to adopt an off-the-shelf object detector to utilize additional image tag information. However, the object detector is time-consuming and can only identify the pre-defined object categories, limiting the model capacity. Inspired by the observation that the texts incorporate incomplete fine-grained image information, we introduce IDEA, which stands for increasing text diversity via online multi-label recognition for VLP. IDEA shows that multi-label learning with image tags extracted from the texts can be jointly optimized during VLP. Moreover, IDEA can identify valuable image tags online to provide more explicit textual supervision. Comprehensive experiments demonstrate that IDEA can significantly boost the performance on multiple downstream datasets with a small extra computational cost.
Youcai Zhang, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang, Yandong Guo
ACM Multimedia7
2022 MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video Parsing
abstract
Recognizing and localizing events in videos is a fundamental task for video understanding. Since events may occur in auditory and visual modalities, multimodal detailed perception is essential for complete scene comprehension. Most previous works attempted to analyze videos from a holistic perspective. However, they do not consider semantic information at multiple scales, which makes the model difficult to localize events in different lengths. In this paper, we present a Multimodal Pyramid Attentional Network (MM-Pyramid ) for event localization. Specifically, we first propose the attentive feature pyramid module. This module captures temporal pyramid features via several stacking pyramid units, each of them is composed of a fixed-size attention block and dilated convolution block. We also design an adaptive semantic fusion module, which leverages a unit-level attention block and a selective fusion block to integrate pyramid features interactively. Extensive experiments on audio-visual event localization and weakly-supervised audio-visual video parsing tasks verify the effectiveness of our approach.
Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang
ACM Multimedia5
2022 Modality-aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence Detection
abstract
Weakly-supervised audio-visual violence detection aims to distinguish snippets containing multimodal violence events with video-level labels. Many prior works perform audio-visual integration and interaction in an early or intermediate manner, yet overlooking the modality heterogeneousness over the weakly-supervised setting. In this paper, we analyze the modality asynchrony and undifferentiated instances phenomena of the multiple instance learning (MIL) procedure, and further investigate its negative impact on weakly-supervised audio-visual learning. To address these issues, we propose a modality-aware contrastive instance learning with self-distillation (MACIL-SD) strategy . Specifically, we leverage a lightweight two-stream network to generate audio and visual bags, in which unimodal background, violent, and normal instances are clustered into semi-bags in an unsupervised way. Then audio and visual violent semi-bag representations are assembled as positive pairs, and violent semi-bags are combined with background and normal instances in the opposite modality as contrastive negative pairs. Furthermore, a self-distillation module is applied to transfer unimodal visual knowledge to the audio-visual model, which alleviates noises and closes the semantic gap between unimodal and multimodal features. Experiments show that our framework outperforms previous methods with lower complexity on the large-scale XD-Violence dataset. Results also demonstrate that our proposed approach can be used as plug-in modules to enhance other networks. Codes are available at https://github.com/JustinYuu/MACIL_SD.
Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang
ACM Multimedia5
2022 AE-Net: Fine-grained sketch-based image retrieval via attention-enhanced network
Yangdong Chen, Zhaolong Zhang, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.4
2022 Balanced single-shot object detection using cross-context attention-guided network
Shuyu Miao, Shanshan Du, Rui Feng 0001, Yuejie Zhang, Tianbi Liu, Weiguo Fan
Pattern Recognit.4
2022 Stacked Multimodal Attention Network for Context-Aware Video Captioning
abstract
Recent neural models for video captioning usually employ an attention-based encoder-decoder framework. However, current approaches mainly attend to the motion features and object features of the video when generating the caption, but ignore the potential but useful historical information. Besides, exposure bias and vanishing gradients problems always exist in current caption generation models. In this paper, we propose a novel video captioning framework, named Stacked Multimodal Attention Network (SMAN). It adopts additional visual and textual historical information during caption generation as context features, employs a stacked architecture to process different features gradually, and utilizes the Reinforcement Learning method and coarse-to-fine training strategy to further improve the generated results. Both quantitative and qualitative experiments on the benchmark datasets ofMSVDandMSR-VTTshow the effectiveness and feasibility of our framework. The codes are available onhttps://github.com/zhengyi123456/SMAN.
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
IEEE Trans. Circuits Syst. Video Technol.2
2021 CellDet: Dual-Task Cell Detection Network for IHC-Stained Image Analysis
abstract
Cell detection on immunohistochemistry stained (IHC-stained) images plays an essential role in computer assisted prediction of tumor progression and treatment response. Currently available cell detection datasets provide either point level or bounding box level annotations for deep object detection network training. And these widely used networks usually employ standard pyramid structured multi-scale feature fusion. However, we find that these methods have obvious limitation when facing large amounts of cells in similar scale with severe overlapping. To address this problem, we propose a novel CellDet network with (1) Scale Consistency Feature Fusion Module (SCFFM) and (2) Dual Task Detection Module to simultaneously exploit the complementary information from both point and bounding box annotations. In order to verify the effectiveness our proposed method, We make efforts to relabel the public SHIDC-B-Ki-67 dataset with bounding box annotations. Extensive experimental results show that the proposed CellDet outperforms other state-of-the-art cell detection methods with a remarkable margin. We will release our source code and dataset in https://github.com/JiweiMaster/celldet.
Wei Ji 0009, Wenbin Pan, Rui Feng 0001, Yuejie Zhang, Gang Jin
BIBM7
2021 Semi-supervised Medical Image Segmentation with Distribution Calibration and Non-local Semantic Constraint
abstract
The available medical images with accurate segmentation masks are usually limited due to the expensive and time-consuming annotation cost. Many semi-supervised approaches tried to exploit a large amount of unlabeled data together with the small number of labeled data. However, their learned segmentation models can easily become overfitted on the biased labeled data, mainly because of the misalignment between the labeled and unlabeled data distribution. To address this challenging problem, we propose a novel semi-supervised model with distribution calibration and non-local semantic constraint for medical image segmentation. In specific, we explicitly calibrate the learned feature distributions of the labeled and unlabeled data to make them aligned. Meanwhile, we add a special nonlocal semantic loss to encourage the learned features to be more discriminative for the segmentation task at the same time. Consequently, our final segmentation networks have the advantage to better generalize on the unlabeled data in both training and test set. Experimental results on three popular medical image segmentation benchmarks demonstrate that our proposed model achieves superior performance over other state-of-the-art methods. We will release our source code in this URL.
Junlin Hou, Rui Feng 0001, Yuejie Zhang
BIBM5
2021 Load Balancing and User Association Based on Historical Data
abstract
With the rapid increase of demand on mobile data traffic of user equipment (UE), network operators have begun to deploy abundant heterogeneous base stations (BSs) to ensure the quality of service (QoS) of UEs, which will cause new problems such as network congestion and load imbalance. If the pattern of user association (UA) can be adjusted in accordance with the results of traffic prediction, the performance of system will be greatly improved. Therefore, a new neural network approach based on spatial and temporal characteristics of traffic data is proposed for traffic prediction. The fluctuations of traffic in the future week are predicted by the proposed method. Then, UA is represented as a problem of maximizing the utility function of load balancing index, and a dynamic user association based on load prediction algorithm (DUALP) which aims to achieve a proactive load balancing is proposed. The QoS of UEs is ensured and the long-term stability of the system is achieved by DUALP. Experimental results show that compared to the classic UA strategies, the most optimal load distribution is realized by DUALP.
Yuejie Zhang, Kai Sun 0003, Xueliang Gao, Wei Huang 0038, Haijun Zhang 0001
GLOBECOM1
2021 Exploring Logical Reasoning for Referring Expression Comprehension
abstract
Referring expression comprehension aims to localize the target object in an image referred by a natural language expression. Most existing approaches neglect the implicit logical correlations among fine-grained cues, e.g., categories, attributes, which are beneficial for distinguishing objects. In this paper, we propose a logic-guided approach to explore logical knowledge for referring expression comprehension in a hierarchical modular-based framework. Specifically, we propose to extract fine-grained cues in visual and textual domains and perform logical reasoning over them with explicit logical expressions to regularize the matching process without extra parameters. Besides, we propose to improve existing modular-based methods by introducing context information of objects in the relationship module. Extensive experiments are conducted on three referring expression datasets, and the results demonstrate that our model can produce more consistent predictions and further achieve superior performance compared with previous methods.
Ying Cheng 0005, Jiashuo Yu, Yuejie Zhang, Rui Feng 0001
ACM Multimedia5
2021 HTDA: Hierarchical time-based directional attention network for sequential user behavior modeling
Zhenzhen Sheng, Tao Zhang 0022, Yuejie Zhang
Neurocomputing3
2021 Cross-modal retrieval with dual multi-angle self-attention
abstract
Abstract In recent years, cross‐modal retrieval has been a popular research topic in both fields of computer vision and natural language processing. There is a huge semantic gap between different modalities on account of heterogeneous properties. How to establish the correlation among different modality data faces enormous challenges. In this work, we propose a novel end‐to‐end framework named Dual Multi‐Angle Self‐Attention (DMASA) for cross‐modal retrieval. Multiple self‐attention mechanisms are applied to extract fine‐grained features for both images and texts from different angles. We then integrate coarse‐grained and fine‐grained features into a multimodal embedding space, in which the similarity degrees between images and texts can be directly compared. Moreover, we propose a special multistage training strategy, in which the preceding stage can provide a good initial value for the succeeding stage and make our framework work better. Very promising experimental results over the state‐of‐the‐art methods can be achieved on three benchmark datasets of Flickr8k, Flickr30k, and MSCOCO.
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
J. Assoc. Inf. Sci. Technol.3
2021 Periphery-aware COVID-19 diagnosis with contrastive representation enhancement
Junlin Hou, Jilan Xu, Longquan Jiang 0003, Shanshan Du, Rui Feng 0001, Yuejie Zhang, Xiangyang Xue 0001
Pattern Recognit.6
2021 Deep Cross-Modal Face Naming for People News Retrieval
abstract
How to integrate multimodal information sources for face naming in multimodal news is a hot and yet challenging problem. A novel deep cross-modal face naming scheme is developed in this paper to facilitate more effective people news retrieval for large-scale multimodal news. This scheme integrates deep multimodal analysis, cross-modal correlation learning, and multimodal information mining, in which the efficient naming mechanism aims to cluster the deep features of different modalities into a common space to explore their inter-related correlations, and a special Web mining pattern is designed to optimize the name-face matching for rare non-celebrity. Such a cross-modal face naming model can be treated as a problem of bi-media semantic mapping and modeled as an inter-related correlation distribution over deep representations of multimodal news, in which the most important is to create more effective cross-modal name-face correlation and measure to what degree they are correlated. The experiments on a large number of public data from Yahoo! News have obtained very positive results and demonstrated the effectiveness of the proposed model.
Lian Zhou, Yuejie Zhang, Tao Zhang 0022, Weiguo Fan
IEEE Trans. Knowl. Data Eng.3
2020 Zero-Shot Sketch-Based Image Retrieval via Graph Convolution Network
abstract
Zero-Shot Sketch-based Image Retrieval (ZS-SBIR) has been proposed recently, putting the traditional Sketch-based Image Retrieval (SBIR) under the setting of zero-shot learning. Dealing with both the challenges in SBIR and zero-shot learning makes it become a more difficult task. Previous works mainly focus on utilizing one kind of information, i.e., the visual information or the semantic information. In this paper, we propose a SketchGCN model utilizing the graph convolution network, which simultaneously considers both the visual information and the semantic information. Thus, our model can effectively narrow the domain gap and transfer the knowledge. Furthermore, we generate the semantic information from the visual information using a Conditional Variational Autoencoder rather than only map them back from the visual space to the semantic space, which enhances the generalization ability of our model. Besides, feature loss, classification loss, and semantic loss are introduced to optimize our proposed SketchGCN model. Our model gets a good performance on the challenging Sketchy and TU-Berlin datasets.
Zhaolong Zhang, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
AAAI2
2020 Label Generation Network based on Self-selected Historical Information for Multiple Disease Classification on Chest Radiography
abstract
Deep learning has made significant break through's in image classification, but accurate diagnosis on chest radiography remains challenging due to a variety of potential diseases contained in one scan. Complex relations among diseases have significant clinical meanings, but are always ignored in most of previous work. Thus in this paper, we propose a novel Label Generation Network (LGN) which treats the label sequence as the caption of a radiology image and utilizes RNN to generate the disease labels according to the semantic relations and co-occurrence dependency among them. However, the sequential generation process of RNN makes it hard to capture the complex topological relations among diseases. To mitigate this problem, a Historical Information Module (HIM) is especially introduced to LGN, in which all the generated labels are fully considered when generating a new label. Moreover, a specific self-attention mechanism is applied in HIM to learn the topological disease relations and utilize them to select useful historical information which can provide positive guidance to the prediction of new label. Very positive results have been obtained in our experiments on the benchmark dataset of Chest X-ray14, which significantly outperform the state-of-the-art methods.
Yuelin Hu, Yuejie Zhang, Tao Zhang 0022, Shang Gao 0003, Weiguo Fan
BIBM2
2020 Data-Efficient Histopathology Image Analysis with Deformation Representation Learning
abstract
Histopathological examination of tissue biopsies plays a fundamental role in disease assessment. Automatic histopathology image analysis requires substantial task-specific annotations, which are often expensive and laborious in realworld scenarios. This insufficient annotation of data limits the generalization ability of supervised learning models. To address this challenge, we propose a self-supervised Deformation Representation Learning (DRL) framework to learn semantic features from unlabeled data. As a novel paradigm, our approach utilizes deformation as supervisory signals based on two critical features, i.e., local structure heterogeneity and global context homogeneity. Given an original histopathology image and its deformed counterpart, there exists a moderate difference in local structures. In contrast, due to the transformation-invariance, both images share a similar global context compared with other images. Specifically, an encoder network is trained to distinguish the local inconsistency by measuring the mutual information and maintain the global consistency with noise contrastive estimation. Extensive experiments on public histopathology image datasets show that the learned representations are generalizable for various downstream tasks, such as transfer learning on segmentation and semi-supervised classification. Our approach achieves superior results over other self-supervised methods and the ImageNet pre-trained model, and it reveals the ability as a novel pre-training scheme in histopathology image analysis.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Chunyang Ruan, Tao Zhang 0022, Weiguo Fan
BIBM3
2020 Learning Class-Based Graph Representation for Object Detection
Shuyu Miao, Rui Feng 0001, Yuejie Zhang, Weiguo Fan
ECAI3
2020 Representation Reconstruction Head for Object Detection
abstract
There are two kinds of detection heads in object detection frameworks. Between them, the heads based on full connection contribute to mapping the learned feature representation to the sample label space, while the heads based on full convolution facilitate preserving location sensitivity information. However, to enjoy the benefits from both detection heads is still underexplored. In this paper, we propose a generalized Representation Reconstruction Head (RRHead) to break through the limitation that most detection heads focus on unilateral self-advantage while ignoring another one. RRHead enhances multi scale feature representation for better feature mapping, and employs location sensitivity representation for better location preservation. These optimize fully-convolutional-based heads and fully-connected-based heads separately. RRHead can be embedded in existing detection frameworks to heighten the rationality and reliability of the detection head representation without any additional modification. Extensive experiments show that our proposed RRHead improves the detection performance of the existing frameworks by a large margin on several challenging benchmarks, and achieves new state-of-the-art performance.
Shuyu Miao, Rui Feng 0001, Yuejie Zhang
ICIP3
2020 Video Captioning With Temporal And Region Graph Convolution Network
abstract
Video captioning aims to generate a natural language description for a given video clip that includes not only spatial information but also temporal information. To better exploit such spatial-temporal information attached to videos, we propose a novel video captioning framework with Temporal Graph Network (TGN) and Region Graph Network (RGN). TGN mainly focuses on utilizing the sequential information of frames that most of existing methods ignore. RGN is designed to explore the relationships among salient objects. Different from previous work, we introduce Graph Convolution Network (GCN) to encode frames with their sequential information and build a region graph for utilizing object information. We also particularly adopt a stack GRU decoder with a coarse-to-fine structure for caption generation. Very promising experimental results on two benchmark datasets (MSVD and MSR-VTT) show the effectiveness of our model.
Xinlong Xiao, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003, Weiguo Fan
ICME2
2020 Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation Learning
abstract
When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be utilized as free supervised information to train a neural network by solving the pretext task of audio-visual synchronization. In this paper, we propose a novel self-supervised framework with co-attention mechanism to learn generic cross-modal representations from unlabelled videos in the wild, and further benefit downstream tasks. Specifically, we explore three different co-attention modules to focus on discriminative visual regions correlated to the sounds and introduce the interactions between them. Experiments show that our model achieves state-of-the-art performance on the pretext task while having fewer parameters compared with existing methods. To further evaluate the generalizability and transferability of our approach, we apply the pre-trained model on two downstream tasks, i.e., sound source localization and action recognition. Extensive experiments demonstrate that our model provides competitive results with other self-supervised methods, and also indicate that our approach can tackle the challenging scenes which contain multiple sound sources.
Ying Cheng 0005, Zhihao Pan, Rui Feng 0001, Yuejie Zhang
ACM Multimedia5
2020 Deep cascaded cross-modal correlation learning for fine-grained sketch-based image retrieval
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.3
2020 Deep reinforcement hashing with redundancy elimination for effective image retrieval
Juexu Yang, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.2
2020 Re-Caption: Saliency-Enhanced Image Captioning Through Two-Phase Learning
abstract
Visual and semantic saliency are important in image captioning. However, single-phase image captioning benefits little from limited saliency without a saliency predictor. In this paper, a novel saliency-enhanced re-captioning framework via two-phase learning is proposed to enhance the single-phase image captioning. In the framework, visual saliency and semantic saliency are distilled from the first-phase model and fused with the second-phase model for model self-boosting. The visual saliency mechanism can generate a saliency map and a saliency mask for an image without learning a saliency map predictor. The semantic saliency mechanism sheds some lights on the properties of words with part-of-speech Noun in a caption. Besides, another type of saliency, sample saliency is proposed to explicitly compute the saliency degree of each sample, which helps for more robust image captioning. In addition, how to combine the above three types of saliency for further performance boost is also examined. Our framework can treat an image captioning model as a saliency extractor, which may benefit other captioning models and related tasks. The experimental results on both the Flickr30k and MSCOCO datasets show that the saliency-enhanced models can obtain promising performance gains.
Lian Zhou, Yuejie Zhang, Yu-Gang Jiang 0001, Tao Zhang 0022, Weiguo Fan
IEEE Trans. Image Process.2
2019 Fully Convolutional Video Captioning with Coarse-to-Fine and Inherited Attention
abstract
Automatically generating natural language description for video is an extremely complicated and challenging task. To tackle the obstacles of traditional LSTM-based model for video captioning, we propose a novel architecture to generate the optimal descriptions for videos, which focuses on constructing a new network structure that can generate sentences superior to the basic model with LSTM, and establishing special attention mechanisms that can provide more useful visual information for caption generation. This scheme discards the traditional LSTM, and exploits the fully convolutional network with coarse-to-fine and inherited attention designed according to the characteristics of fully convolutional structure. Our model cannot only outperform the basic LSTM-based model, but also achieve the comparable performance with those of state-of-the-art methods
Kuncheng Fang, Lian Zhou, Cheng Jin 0001, Yuejie Zhang, Kangnian Weng, Tao Zhang 0022, Weiguo Fan
AAAI4
2018 Sketch-based image retrieval with deep visual semantic descriptor
Cheng Jin 0001, Yuejie Zhang, Kangnian Weng, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.3
2017 Towards sketch-based image retrieval with deep cross-modal correlation learning
abstract
A novel scheme with deep cross-modal correlation learning is developed in this paper to facilitate more effective Sketch-based Image Retrieval (SBIR) for large-scale annotated images. It integrates the deep multimodal feature generation, deep cross-modal correlation learning and similarity search optimization through mining all the beneficial multimodal information sources in sketches and images, which can be treated as an inter-related correlation distribution over deep representations of sketches and images. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ICME3
2017 A Hierarchical Multimodal Attention-based Neural Network for Image Captioning
abstract
A novel hierarchical multimodal attention-based model is developed in this paper to generate more accurate and descriptive captions for images. Our model is an "end-to-end" neural network which contains three related sub-networks: a deep convolutional neural network to encode image contents, a recurrent neural network to identify the objects in images sequentially, and a multimodal attention-based recurrent neural network to generate image captions. The main contribution of our work is that the hierarchical structure and multimodal attention mechanism is both applied, thus each caption word can be generated with the multimodal attention on the intermediate semantic objects and the global visual content. Our experiments on two benchmark datasets have obtained very positive results.
Lian Zhou, Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
SIGIR5
2017 Deep Multimodal Embedding Model for Fine-grained Sketch-based Image Retrieval
abstract
Fine-grained Sketch-based Image Retrieval (Fine-grained SBIR), which uses hand-drawn sketches to search the target object images, has been an emerging topic over the last few years. The difficulties of this task not only come from the ambiguous and abstract characteristics of sketches with less useful information, but also the cross-modal gap at both visual and semantic level. However, images on the web are always exhibited with multimodal contents. In this paper, we consider Fine-grained SBIR as a cross-modal retrieval problem and propose a deep multimodal embedding model that exploits all the beneficial multimodal information sources in sketches and images. In our experiment with large quantity of public data, we show that the proposed method outperforms the state-of-the-art methods for Fine-grained SBIR.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
SIGIR4
2016 A Novel Cross-Modal Topic Correlation Model for Cross-Media Retrieval
abstract
A novel cross-modal topic correlation model CMTCM is developed in this paper to facilitate more effective cross-modal analysis and cross-media retrieval for large-scale multimodal document collections. It can be modeled as a cross-modal topic correlation model which explores the inter-related correlation distribution over the deep representations of multimodal documents. It integrates the deep multimodal document representation, relational topic correlation modeling, and cross-modal topic correlation learning, which aims to characterize the correlations between the heterogeneous topic distributions of inter-related visual images and semantic texts, and measure their association degree more precisely. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ECAI4
2016 Enhancing Sketch-Based Image Retrieval via Deep Discriminative Representation
abstract
In this paper we aim to employ deep learning to enhance SBIR via deep discriminative representation. Our main contributions focus on: 1) The deep discriminative representation is established to bridge both the visual appearance gap and the semantic gap between sketches and images; 2) The deep learning pattern is applied to our SBIR model through training on our transformed sketch-like images to overcome the rarity of training sketches. Our experiments on a large number of public sketch and image data have obtained very positive results.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ECAI4
2016 Sketch-Based Image Retrieval with a Novel BoVW Representation
Cheng Jin 0001, Chenjie Li, Zheming Wang, Yuejie Zhang, Tao Zhang 0022
MMM (1)4
2015 Cross-Modal Image Clustering via Canonical Correlation Analysis
abstract
A new algorithm via Canonical Correlation Analysis (CCA) is developed in this paper to support more effective cross-modal image clustering for large-scale annotated image collections. It can be treated as a bi-media multimodal mapping problem and modeled as a correlation distribution over multimodal feature representations. It integrates the multimodal feature generation with the Locality Linear Coding (LLC) and co-occurrence association network, multimodal feature fusion with CCA, and accelerated hierarchical k-means clustering, which aims to characterize the correlations between the inter-related visual features in images and semantic features in captions, and measure their association degree more precisely. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Wenhui Mao, Yuejie Zhang, Xiangyang Xue 0001
AAAI4
2015 People News Search via Name-Face Association Analysis
abstract
By integrating multimodal information in multimodal news, a novel scheme is developed in this paper for facilitating more effective people news search via name-face association analysis. It is treated as a problem of bi-media multimodal semantic mapping on multimodal news, and modeled as an inter-related correlation distribution over multimodal semantic representations of name-face associations. Very positive results have been obtained in our experiments using a large quantity of public multimodal news data.
Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ICMR5
2015 A Novel Visual-Region-Descriptor-based Approach to Sketch-based Image Retrieval
abstract
A novel Visual-Region-Descriptor-based approach is developed in this paper to facilitate more effective Sketch-based Image Retrieval (SBIR), which can be treated as a problem of bilateral visual mapping and modeled as an inter-related correlation distribution over visual semantic representations of sketches and images. For crossing the matching barrier between binary query sketches and full color natural images, we focus on constructing a visual pre-analysis via the sketch-like representation transformation to improve the general sketch-image resemblance, creating a special visual region descriptor to obtain better visual feature generation for sketches and images, and a dynamic sketch-image matching scheme to achieve more precise characterization of the correlations between sketches and images. Such a visual-region-descriptor-based SBIR pattern can not only enable users to present whatever they imagine in their mind on the sketch query panel but also return the most similar images to the picture in users' mind. Very positive results were obtained in our experiments using a large quantity of public data.
Cheng Jin 0001, Zheming Wang, Qinen Zhu, Yuejie Zhang
ICMR5
2015 Cross-Modal Image-Tag Relevance Learning for Social Images
abstract
A new algorithm is developed in this paper to support more effective cross-modal image-tag relevance learning for large-scale social images, which integrates the multimodal feature representation, multimodal relevance measurement, and cross- modal relevance fusion. The main contribution of our work is that we provide a more reasonable base to learn cross-modal relevance among social images, which can be acquired from integrating multimodal image and tag relevance with multiple features in different modalities. Very positive results were obtained in our experiments using a large quantity of public social image data.
Zhengxiang Cai, Rui Feng 0001, Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ACM Multimedia5
2014 Sketch-Based Image Retrieval via Adaptive Weighting
abstract
As touch devices become more and more popular these days, it would be convenient if the user could draw a sketch and then use the sketch as the input for an image retrieval system. Although Sketch-Based Image Retrieval (SBIR) had been studied since 1990s, how to measure the similarity between a sketch and an image with high precision is still a challenging problem. In this paper, a novel adaptive weighting method is proposed for the matching process of SBIR. We integrate a cost aggregation step into the matching process, both the neighborhood and multi-scale information are taken into account. The experiments on the public image dataset show that our method can yield promising results.
Cheng Jin 0001, Yuejie Zhang
ICMR4
2013 Automatic Name-Face Alignment to Enable Cross-Media News Retrieval
Yuejie Zhang, Cheng Jin 0001, Xiangyang Xue 0001, Jianping Fan 0001
IJCAI1
2011 Learning Inter-Related Statistical Query Translation Models for English-Chinese Bi-Directional CLIR
abstract
To support more precise query translation for English-Chinese Bi-Directional Cross-Language Information Retrieval (CLIR), we have developed a novel framework by integrating a semantic network to characterize the correlations between multiple inter-related text terms of interest and learn their inter-related statistical query translation models. First, a semantic network is automatically generated from large-scale English-Chinese bilingual parallel corpora to characterize the correlations between a large number of text terms of interest. Second, the semantic network is exploited to learn the statistical query translation models for such text terms of interest. Finally, these inter-related query translation models are used to translate the queries more precisely and achieve more effective CLIR. Our experiments on a large number of official public data have obtained very positive results.
Yuejie Zhang, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001, Jianping Fan 0001
IJCAI1
2011 Fusion of Multiple Features and Supervised Learning for Chinese OOV Term Detection and POS Guessing
abstract
In this paper, to support more precise Chinese Out-of-Vocabulary (OOV) term detection and Part-of-Speech (POS) guessing, a unified mechanism is proposed and formulated based on the fusion of multiple features and supervised learning. Besides all the traditional features, the new features for statistical information and global contexts are introduced, as well as some constraints and heuristic rules, which reveal the relationships among OOV term candidates. Our experiments on the Chinese corpora from both People’s Daily and SIGHAN 2005 have achieved the consistent results, which are better than those acquired by pure rule-based or statistics-based models. From the experimental results for combining our model with Chinese monolingual retrieval on the data sets of TREC-9, it is found that the obvious improvement for the retrieval performance can also be obtained.
Yuejie Zhang, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001
IJCAI1
2010 Bilingual query translation and expansion for supporting more effective cross-language image retrieval
abstract
To support more effective Cross-Language Image Retrieval (ImageCLIR), a novel algorithm is developed by integrating a bilingual semantic network to achieve more precise bilingual query translation and expansion. An English-Chinese bilingual parallel corpus is used to construct the bilingual semantic network for determining more meaningful text terms and characterizing the inter-term correlations and similarity contexts between multiple inter-related text terms more precisely. Our experiments on CWMT2009 and CLEF have provided very promising results.
Yuejie Zhang, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001
ACM Multimedia1
2008 CRF-based Hybrid Model for Word Segmentation, NER and even POS Tagging
Zhiting Xu 0002, Xian Qian, Yuejie Zhang, Yaqian Zhou 0001
IJCNLP3