Rui Feng 0001

dblp:28/4423-1 · DBLP profile ↗
← Back
106ranked-venue papers
0as first author
85since 2021 · last 2026
0000-0002-4747-0574ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 62 · 50 since 2021Artificial intelligence and machine learning · 40 · 29 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 16 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EviMMQA: Multimodal question answering for medical evidence extraction in systematic reviews
Changkai Ji, Yingwen Wang, Ying Cheng 0005, Yuejie Zhang, Rui Feng 0001
Pattern Recognit.7
2025 AS-Det: Active Sampling for Adaptive 3D Object Detection in Point Clouds
abstract
3D object detection in point clouds is critical in 3D computer vision, autonomous driving, and robotics. Existing point-based detectors, tailored to handle unstructured raw point clouds, often rely on simplistic sampling strategies to select a subset of points for local representation learning and detection. However, the diverse patterns exhibited by multiple types of point cloud data present a significant challenge to the universality of current detectors, particularly those captured by varied sensors (e.g., LiDAR and 4D Imaging Radar). In response to this challenge, we introduce an adaptable point-based single-stage 3D detector, AS-Det, engineered to excel on both LiDAR and 4D Radar point clouds. Specifically, we propose a novel active sampling strategy that actively mines object-related information to achieve efficient sampling and representation across different types of point clouds through end-to-end training. Additionally, we introduce a lightweight multi-scale center feature aggregation module to exploit multi-scale object context for precise and low-cost detection. By integrating the abovementioned modules, AS-Det achieves highly adaptive detection on various point clouds, encompassing different sensors and scales. Experimental results demonstrate the superior performance and adaptability of AS-Det on both LiDAR and 4D Radar point clouds.
Ziheng Ding, Xiaze Zhang, Ying Cheng 0005, Rui Feng 0001
AAAI5
2025 Fine-Grained Knowledge-Guided Alignment for Medical Vision-Language Pre-Training
abstract
Medical contrastive Vision-Language Pre-training (VLP) has emerged as a promising approach, enabling models to learn joint representations from paired medical images and radiology reports. Despite existing methods exploring local visual representation learning techniques, they often fall short in local alignment and knowledge infusion, e.g., uniform token treatment and isolated knowledge assignment. To address these issues, we propose a novel Fine-grained Knowledge-Guided Alignment (FKGA) framework for medical VLP. Specifically, we propose a Fine-grained Disease Knowledge Integration (FDKI) module to inject detailed disease descriptions into corresponding disease tokens in reports. Based on these semantic-enriched tokens, we introduce global instance-wise and local token-wise contrastive learning to further align the semantically related visual and textual modalities. In contrast to previous local visual representation learning methods, our design of semantic-enriched token alignment and context-preserved knowledge infusion enhances the semantic understanding of diseases. Extensive experimental results on five downstream tasks demonstrate that our proposed method outperforms other state-of-the-art methods across seven datasets.
Yaning Pan, Ying Cheng 0005, Qingqiu Li, Runtian Yuan, Rui Feng 0001
BIBM6
2025 RoBGuard: Enhancing LLMs to Assess Risk of Bias in Clinical Trial Documents
abstract
Randomized Controlled Trials (RCTs) are rigorous clinical studies crucial for reliable decision-making, but their credibility can be compromised by bias. The Cochrane Risk of Bias tool (RoB 2) assesses this risk, yet manual assessments are time-consuming and labor-intensive. Previous approaches have employed Large Language Models (LLMs) to automate this process. However, they typically focus on manually crafted prompts and a restricted set of simple questions, limiting their accuracy and generalizability. Inspired by the human bias assessment process, we propose RoBGuard, a novel framework for enhancing LLMs to assess the risk of bias in RCTs. Specifically, RoBGuard integrates medical knowledge-enhanced question reformulation, multimodal document parsing, and multi-expert collaboration to ensure both completeness and accuracy. Additionally, to address the lack of suitable datasets, we introduce two new datasets: RoB-Item and RoB-Domain. Experimental results demonstrate RoBGuard’s effectiveness on the RoB-Item dataset, outperforming existing methods.
Changkai Ji, Yingwen Wang, Yuejie Zhang, Ying Cheng 0005, Rui Feng 0001
COLING7
2025 SplitOcc: Multi-Resolution Sparse Voxel for Efficient LiDAR-Based Semantic Scene Completion
abstract
LiDAR-based Semantic Scene Completion (SSC) is crucial for enhancing environmental perception and ensuring safety in autonomous driving. However, current methods face challenges in balancing accuracy and computational efficiency. On the one hand, projection-based methods reduce complexity but often suffer from spatial information loss. On the other hand, voxel-based methods preserve 3D structures but are computationally expensive. To address these limitations, we introduce SplitOcc, a novel multi-resolution approach that utilizes low-resolution voxels to represent large structures (e.g., road) and high-resolution voxels for detailed objects (e.g., bicycle). By employing multi-resolution sparse semantic voxels, SplitOcc can understand and represent the environment efficiently and accurately. Furthermore, the proposed Multi-Label Loss and Delayed-Drop strategies improve accuracy by preserving key semantic details during reconstruction. Extensive experiments demonstrate that our SplitOcc outperforms existing state-of-the-art methods across multiple evaluation metrics, showing notable improvements in both perception accuracy and detail preservation.
Chaoyi Sun, Xiaze Zhang, Ziheng Ding, Ying Cheng 0005, Rui Feng 0001
ECAI6
2025 Sketch-based Point Cloud Generation with Diffusion Model and Pre-training Enhancement
abstract
Diffusion models, known for their success in various generative tasks like image generation and super-resolution, are applied in this study for point cloud generation, a field that has not been extensively explored due to the complexity of point clouds. We propose a novel method using a diffusion model to generate high-quality 3D point clouds from 2D sketches. This method employs a self-supervised contrastive learning scheme to align sketch and point cloud modalities. Additionally, it incorporates a specific partition mixing strategy to integrate edge information during pre-training. Evaluated on two benchmark datasets, our method outperforms existing state-of-the-art approaches, showcasing the potential of diffusion models in point cloud generation and setting a new direction for future research.
Yangdong Chen, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP4
2025 Uncertainty-Aware Dynamic Fusion for Multimodal Clinical Prediction Tasks
abstract
Multimodal fusion offers significant potential for enhancing medical diagnosis, particularly in the Intensive Care Unit (ICU), where integrating diverse data sources is crucial. Traditional static fusion models often fail to account for sample-wise variations in modality importance, which can impact prediction accuracy. To address this issue, we propose a dynamic Uncertainty-Aware Weighting (UAW) strategy that adaptively adjusts the importance of different modalities based on their reliability. This strategy is coupled with an Expert Ensemble Fusion (EEF) module, which leverages self-attention mechanisms and modality-specific FeedForward Networks (FFNs) to preserve and integrate critical information from various modalities. The proposed method demonstrates its efficacy through extensive experiments on phenotype classification and mortality prediction tasks, showing improved accuracy and robustness in handling diverse clinical data.
Ying Cheng 0005, Yuejie Zhang, Rui Feng 0001
ICASSP5
2025 Minimizing Disparities between Real and Pseudo Queries for Unsupervised Visual Grounding
abstract
Visual grounding involves the identification and localization of image regions given textual descriptions. To reduce the manual labeling effort on region-text pairs, unsupervised visual grounding aims to generate pseudo bounding box and query pairs for training grounding models. However, there exists significant disparities between real and pseudo queries in terms of object, attribute distributions, and textual formats, limiting the generalization performance of unsupervised grounding methods. To address this challenge, we propose a novel unsupervised visual grounding framework. During training, we prompt Multimodal Large Language Models to generate pseudo queries, in which the entities are beyond the object detector’s pre-defined limited categories, and are associated with richer attributes. We further devise a Modifier Tree structure to bridge the gap of textual format between real and pseudo queries. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art unsupervised approaches on public benchmark datasets, particularly when dealing with complex queries.
Changkai Ji, Jilan Xu, Yanhao Zhu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP6
2025 Human Simulacra: Benchmarking the Personification of Large Language Models
abstract
Large Language Models (LLMs) are recognized as systems that closely mimic aspects of human intelligence. This capability has attracted the attention of the social science community, who see the potential in leveraging LLMs to replace human participants in experiments, thereby reducing research costs and complexity. In this paper, we introduce a benchmark for LLMs personification, including a strategy for constructing virtual characters' life stories from the ground up, a Multi-Agent Cognitive Mechanism capable of simulating human cognitive processes, and a psychology-guided evaluation method to assess human simulations from both self and observational perspectives. Experimental results demonstrate that our constructed simulacra can produce personified responses that align with their target characters. We hope this work will serve as a benchmark in the field of human simulation, paving the way for future research.
Qiujie Xie, Qiming Feng, Qingqiu Li, Linyi Yang, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003, Yue Zhang 0004
ICLR7
2025 EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
abstract
Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and textual instructions, the goal is to generate future frames of the ego-centric video. Inspired by the notion that hand-object interactions (HOI) in ego-centric videos represent the primary intentions and actions of the current actor, we present EgoExo-Gen that explicitly models the hand-object dynamics for cross-view video prediction. EgoExo-Gen consists of two stages. First, we design a cross-view HOI mask prediction model that anticipates the HOI masks in future ego-frames by modeling the spatio-temporal ego-exo correspondence. Next, we employ a video diffusion model to predict future ego-frames using the first ego-frame and textual instructions, while incorporating the HOI masks as structural guidance to enhance prediction quality. To facilitate training, we develop a fully automated pipeline to generate pseudo HOI masks for both ego- and exo-videos by exploiting vision foundation models. Extensive experiments demonstrate that our proposed EgoExo-Gen achieves better prediction performance compared to previous video prediction models on the public Ego-Exo4D and H2O benchmark datasets, with the HOI masks significantly improving the generation of hands and interactive objects in the ego-centric videos.
Jilan Xu, Yifei Huang 0002, Baoqi Pei, Junlin Hou, Qingqiu Li, Guo Chen 0006, Yuejie Zhang, Rui Feng 0001, Weidi Xie
ICLR8
2025 ConTrack3D: Contrastive Learning Contributes Concise 3D Multi-Object Tracking
abstract
Online object detection and tracking are crucial for embodied intelligence systems, including autonomous vehicles and robotics. Traditional approaches employ a pipeline structure to perform detection and tracking separately, which can not fully leverage information from the detector. Moreover, most prior tracking methods rely on motion models such as constant velocity for state updates, which can lead to incorrect associations when the velocity estimates are inaccurate. To address these limitations, we propose ConTrack3D, an online tracking approach that jointly performs detection and tracking in an end-to-end manner. Specifically, ConTrack3D incorporates a Joint Encoder module to capture detection embeddings and a Temporal Extender module for data-driven state updates. By employing contrastive learning, ConTrack3D learns discriminative tracking representation for more accurate association. ConTrack3D is evaluated on the nuScenes benchmark, and the experimental results demonstrate its significant improvements in tracking performance.
Ruibin Du, Ziheng Ding, Xiaze Zhang, Ying Cheng 0005, Rui Feng 0001
ICRA6
2025 TGSAM-2: Text-Guided Medical Image Segmentation Using Segment Anything Model 2
Runtian Yuan, Ling Zhou 0002, Jilan Xu, Qingqiu Li, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
MICCAI (10)7
2025 Open-Set Image Tagging with Multi-Grained Text Supervision
abstract
This paper introduces the Recognize Anything Plus Model (RAM++), an open-set image tagging model effectively leveraging multi-grained text supervision. Previous approaches (e.g., CLIP) primarily utilize global text supervision paired with images, leading to sub-optimal performance in recognizing multiple individual semantic tags. In contrast, RAM++ seamlessly integrates individual tag supervision with global text supervision, all within a unified alignment framework. This integration not only ensures efficient recognition of predefined tag categories, but also enhances generalization capabilities for diverse open-set categories. Furthermore, RAM++ employs large language models (LLMs) to convert semantically constrained tag supervision into more expansive tag description supervision, thereby enriching the scope of open-set visual description concepts. Comprehensive evaluations on various image recognition benchmarks demonstrate RAM++ exceeds existing state-of-the-art (SOTA) open-set image tagging models on most aspects. Specifically, for predefined commonly used tag categories, RAM++ showcases 10.2 mAP and 15.4 mAP enhancements over CLIP on OpenImages and ImageNet. For open-set categories beyond predefined, RAM++ records improvements of 5.0 mAP and 6.4 mAP over CLIP and RAM respectively on OpenImages. For diverse human-object interaction phrases, RAM++ achieves 7.8 mAP and 4.7 mAP improvements on the HICO benchmark.
Yi-Jie Huang, Youcai Zhang, Rui Feng 0001, Yuejie Zhang, Yanchun Xie, Lei Zhang 0001
ACM Multimedia5
2025 Semantic-Aware Hard Negative Mining for Medical Vision-Language Contrastive Pretraining
abstract
Existing medical vision-language contrastive pretraining methods aim to bring the paired image-report embeddings close together while pushing the unpaired ones apart. However, medical images often exhibit high inter-class visual similarity with only subtle differences, leading to the presence of hard negative samples that are semantically distinct from the anchor but incorrectly close to it in the embedding space, making it challenging to distinguish semantically dissimilar samples. Previous methods consider only the embedding similarity between samples to identify hard negatives, often wrongly treating false negatives as hard negatives. To address this issue, we design a simple yet effective approach called Semantic-Aware Hard Negative mining (SAHN), distinguishing hard negatives from false negatives and encouraging the model to pay greater attention to hard negatives. Specifically, hard negatives are identified as samples with high embedding similarity but low semantic similarity to the anchor and assigned greater importance weights. By integrating these importance weights into the InfoNCE loss, SAHN enhances the model's ability to separate semantically dissimilar samples while clustering semantically similar ones. We further conduct a gradient-based theoretical analysis to validate the effectiveness of SAHN. Extensive experimental results on four downstream medical tasks covering image classification, object detection, semantic segmentation, and cross-modal retrieval demonstrate the superiority of our approach.
Ying Cheng 0005, Yaning Pan, Rui Feng 0001
ACM Multimedia6
2025 Text-Promptable Propagation for Referring Medical Image Sequence Segmentation
abstract
Referring Medical Image Sequence Segmentation (Ref-MISS) is a novel and challenging task that aims to segment anatomical structures in medical image sequences (e.g., endoscopy, ultrasound, CT, and MRI) based on natural language descriptions. Existing 2D and 3D segmentation models struggle to explicitly track objects of interest across medical image sequences, and lack support for interactive, text-driven guidance. To address these limitations, we propose Text-Promptable Propagation (TPP), which enables the recognition of referred objects through cross-modal referring interaction, and maintains continuous tracking across the sequence via Transformer-based triple propagation, using text embeddings as queries. To support this task, we curate a large-scale benchmark, Ref-MISS-Bench, which covers 4 imaging modalities and 20 different organs and lesions. Experimental results on this benchmark demonstrate that TPP consistently outperforms state-of-the-art methods in both medical segmentation and referring video object segmentation. Code and data are available at https://github.com/yuanruntian/TPP.
Runtian Yuan, Mohan Chen 0001, Jilan Xu, Ling Zhou 0002, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ACM Multimedia7
2025 EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues
abstract
Qiming Feng, Qiujie Xie, Xiaolong Wang, Qingqiu Li, Yuejie Zhang, Rui Feng, Tao Zhang, Shang Gao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qiming Feng, Qiujie Xie, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
NAACL (Long Papers)6
2025 AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation
abstract
Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding, current MLMMs still face two major challenges: (1) insufficient region-level understanding and interaction, and (2) limited accuracy and interpretability due to single-step prediction. In this paper, we address these challenges by empowering MLMMs with anatomy-centric reasoning capabilities to enhance their interactivity and explainability. Specifically, we propose an Anatomical Ontology-Guided Reasoning (AOR) framework that accommodates both textual and optional visual prompts, centered on region-level information to enable multimodal multi-step reasoning. We also develop AOR-Instruction, a large instruction dataset for MLMMs training, under the guidance of expert physicians. Our experiments demonstrate AOR's superior performance in both Visual Question Answering (VQA) and report generation tasks. Code and data are available at: https://github.com/Liqq1/AOR.
Qingqiu Li, Zihang Cui, Seongsu Bae, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen, Shang Gao 0003, Junjun He
NeurIPS7
2025 A Medical Multimodal Large Language Model for Pediatric Pneumonia
abstract
Pediatric pneumonia is the leading cause of death among children under five years worldwide, imposing a substantial burden on affected families. Currently, there are three significant hurdles in diagnosing and treating pediatric pneumonia. Firstly, pediatric pneumonia shares similar symptoms with other respiratory diseases, making rapid and accurate differential diagnosis challenging. Secondly, primary hospitals often lack sufficient medical resources and experienced doctors. Lastly, providing personalized diagnostic reports and treatment recommendations is labor-intensive and time-consuming. To tackle these challenges, we proposed a Medical Multimodal Large Language Model for Pediatric Pneumonia (P2Med-MLLM). It was capable of handling diverse clinical tasks-such as generating free-text medical records and radiology reports-within a unified framework. Specifically, P2Med-MLLM was trained on a large-scale dataset, including real clinical information from 163,999 outpatient and 8,684 inpatient cases. It can process both plain text data (e.g., outpatient and inpatient records) and interleaved image-text pairs (e.g., 2D chest X-ray images, 3D chest Computed Tomography images, and corresponding radiology reports). We designed a three-stage training strategy to enable P2Med-MLLM to comprehend medical knowledge and follow instructions for various clinical decision-support tasks. To rigorously evaluate P2Med-MLLM's performance, we conducted automatic scoring by the large language model and manual scoring by the specialist on the test set of 642 samples, meticulously verified by pediatric pulmonology specialists. The results demonstrated the reliability of automated scoring and the superiority of P2Med-MLLM. This work plays a crucial role in assisting doctors with prompt diagnosis and treatment planning, reducing severe symptom mortality rates, and optimizing the allocation of medical resources.
Tianhao Cheng, Jinwu Fang, Rui Feng 0001, Daoying Geng
IEEE J. Biomed. Health Informatics6
2025 QMix: Quality-Aware Learning With Mixed Noise for Robust Retinal Disease Diagnosis
abstract
Due to the complex nature of medical image acquisition and annotation, medical datasets inevitably contain noise. This adversely affects the robustness and generalization of deep neural networks. Previous noise learning methods mainly considered noise arising from images being mislabeled, i.e., label noise, assuming all mislabeled images were of high quality. However, medical images can also suffer from severe data quality issues, i.e., data noise, where discriminative visual features for disease diagnosis are missing. In this paper, we propose QMix, a noise learning framework that learns a robust disease diagnosis model under mixed noise scenarios. QMix alternates between sample separation and quality-aware semi-supervised training in each epoch. The sample separation phase uses a joint uncertainty-loss criterion to effectively separate (1) correctly labeled images, (2) mislabeled high-quality images, and (3) mislabeled low-quality images. The semi-supervised training phase then learns a robust disease diagnosis model from the separated samples. Specifically, we propose a sample-reweighing loss to mitigate the effect of mislabeled low-quality images during training, and a contrastive enhancement loss to further distinguish them from correctly labeled images. QMix achieved state-of-the-art performance on six public retinal image datasets and exhibited significant improvements in robustness against mixed noise. Code will be available upon acceptance.
Junlin Hou, Jilan Xu, Rui Feng 0001, Hao Chen 0011
IEEE Trans. Medical Imaging3
2024 Towards Evidential and Class Separable Open Set Object Detection
abstract
Detecting in open-world scenarios poses a formidable challenge for models intended for real-world deployment. The advanced closed set object detectors achieve impressive performance under the closed set setting, but often produce overconfident misprediction on unknown objects due to the lack of supervision. In this paper, we propose a novel Evidential Object Detector (EOD) to formulate the Open Set Object Detection (OSOD) problem from the perspective of Evidential Deep Learning (EDL) theory, which quantifies classification uncertainty by placing the Dirichlet Prior over the categorical distribution parameters. The task-specific customized evidential framework, equipped with meticulously designed model architecture and loss function, effectively bridges the gap between EDL theory and detection tasks. Moreover, we utilize contrastive learning as an implicit means of evidential regularization and to encourage the class separation in the latent space. Alongside, we innovatively model the background uncertainty to further improve the unknown discovery ability. Extensive experiments on benchmark datasets demonstrate the outperformance of the proposed method over existing ones.
Rui Feng 0001
AAAI4
2024 DeepPointMap: Advancing LiDAR SLAM with Unified Neural Descriptors
abstract
Point clouds have shown significant potential in various domains, including Simultaneous Localization and Mapping (SLAM). However, existing approaches either rely on dense point clouds to achieve high localization accuracy or use generalized descriptors to reduce map size. Unfortunately, these two aspects seem to conflict with each other. To address this limitation, we propose an unified architecture, DeepPointMap, achieving excellent preference on both aspects. We utilize neural network to extract highly representative and sparse neural descriptors from point clouds, enabling memory-efficient map representation and accurate multi-scale localization tasks (e.g., odometry and loop-closure). Moreover, we showcase the versatility of our framework by extending it to more challenging multi-agent collaborative SLAM. The promising results obtained in these scenarios further emphasize the effectiveness and potential of our approach.
Xiaze Zhang, Ziheng Ding, Yuejie Zhang, Wenchao Ding 0001, Rui Feng 0001
AAAI6
2024 PRIDE: Pediatric Radiological Image Diagnosis Engine Guided by Medical Domain Knowledge
abstract
Pediatric chest X-rays (CXRs) are crucial for diagnosing respiratory diseases in children. However, most deep learning models perform poorly on pediatric data due to domain gaps, as they are primarily trained on adult datasets with limited pediatric samples. Recent large models, though effective, still struggle with domain-specific terminology and specialized medical reasoning essential for pediatric diagnoses. To address these challenges, we propose PRIDE: a two-stage Pediatric Radiological Image Diagnosis Engine guided by medical knowledge. PRIDE works consistently with real clinical workflows by first gathering medical evidence from patient data, followed by applying clinical knowledge to enhance diagnostic accuracy. In the Multi-source Radiological Findings Recognition Stage, PRIDE integrates insights from generalized medical models and fine-tuned adult and pediatric radiological models to provide comprehensive findings results. In the Knowledge and Evidence-guided Diagnosis Stage, a Multimodal Large Language Model (MLLM) acts as a pediatric clinician, making diagnostic decisions using CXRs, radiological findings, demographic data, and clinical knowledge. Our evaluation on the VinDr-PCXR pediatric dataset demonstrates that PRIDE outperforms existing methods. Ablation studies further confirm the importance and effectiveness of its key components.
Yuze Zhao, Shijie Pang, Yingwen Wang, Rui Feng 0001
BIBM7
2024 Evidential Open Set Recognition for Imbalanced Medical Images via Multi-level Data Augmentation
abstract
Due to the exstence of common and rare diseases, the complex clinical scenario often poses the challenging class imbalanced open set recognition problem. Unfortunately, most existing approaches are ill-suited for such situations with limited and imbalanced data during training and the possibility of encountering unseen classes during test. In this work, we propose a novel Multi-level mixup-based Evidential Open Set Recognition (ME-OSR) approach to more explicitly and effectively address the open set recognition for class-imbalanced medical images. Briefly, we first extract disentangled discriminative and background image features. Then, based on the original images and extracted features, we propose to sample and conduct Multi-level Open Mixup (MOM) for a more balanced open set data augmentation. It includes extensive intra- and interclass mixup operations in both image and feature spaces, which can augment rare classes with different feature combinations and generate potential pseudo-unknown class examples in the open set to boost the model training. Based on the augmented data and their extracted discriminative features, we propose a Regularized Evidential Deep Learning (REDL) classifier to work with the augmented data to achieve open set recognition with prediction uncertainty estimation and unknown example rejection. Through comparative experiments and ablation studies on several representative medical datasets, we showed that our proposed method outperforms other state-of-the-arts on four popular medical OSR datasets.
Yiqian Xu, Ying Zhang 0005, Rui Feng 0001
BIBM4
2024 Retrieval-Augmented Egocentric Video Captioning
abstract
Understanding human actions from videos offirst-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of exploiting existing large-scale third-person videos. In this paper, (1) we develop EgoInstructor, a retrieval-augmented multimodal captioning model that automatically retrieves semantically relevant third-person instructional videos to enhance the video captioning of egocentric videos, (2) for training the cross-view retrieval module, we devise an au-tomatic pipeline to discover ego-exo video pairs from distinct large-scale egocentric and exocentric datasets, (3) we train the cross-view retrieval module with a novel EgoEx-oNCE loss that pulls egocentric and exocentric video features closer, by aligning them to shared text features that describe similar actions, (4) through extensive experiments, our cross-view retrieval module demonstrates superior performance across seven benchmarks. Regarding egocen-tric video captioning, EgoInstructor exhibits significant improvements by leveraging third-person videos as references.
Jilan Xu, Yifei Huang 0002, Junlin Hou, Guo Chen 0006, Yuejie Zhang, Rui Feng 0001, Weidi Xie
CVPR6
2024 Fine-Granularity Face Sketch Synthesis
abstract
Generative Adversarial Networks (GANs) are often used in face sketch synthesis due to their powerful ability in image generation. However, most GAN based synthesis methods took the entire face as the minimum unit. Differently, we propose a novel fine-granularity face sketch synthesis framework in this paper. The core idea is to first capture local information at a fine granularity (i.e., facial component), and then generate a complete face sketch based on the fine-grained information. Specifically, we partition the face sketch into multiple components, and then train a parallel network for each component. A condition enhanced detail repair network is further designed to correct the mismatches and deformations produced during parallel generation. Extensive experiments show that our approach outperforms state-of-the-art methods from both the qualitative and quantitative perspectives.
Yangdong Chen, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICASSP4
2024 ControlCap: Controllable Captioning via No-Fuss Lexicon
abstract
Controllable captioning has received much attention in recent years. Although substantial progress has been made, existing methods still face challenges such as high training costs, intricate control signals and limited control capabilities. To address these issues, we propose a straightforward and unified framework called ControlCap. It uses a no-fuss lexicon as control signal and controls the style and content of visual descriptions through Soft Guidance (a global guide to the caption distribution) and Hard Force (integrating signals without additional training). Extensive experiments, both quantitative and qualitative, have been conducted on three benchmark captioning tasks. Results demonstrate the control ability of ControlCap: it can produce controlled captions that are coherent and diverse while keeping the core content intact.
Qiujie Xie, Qiming Feng, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP4
2024 Lesion-Aware Open Set Medical Image Recognition with Domain Shift
abstract
Medical image diagnosis in real clinical scenario faces the challenging open set recognition problem with domain shift. Based on the observation that it is often merely the local lesion areas in the whole image decide or cause the true seen or potential unseen disease labels, we propose a unified deep learning framework to address it by establishing a lesion-aware mask generator to guide the model towards focusing on the lesion areas for both seen and unseen diseases, followed by a multi-task feature learning module to extract the lesion image features; and an open set recognizer with improved capability to reject unseen classes. Last but not least, we propose tailored training strategies with mixed up data and mask augmentation for improved OSR and multi-clue cross-domain feature alignment for domain generalization. Through comparative experiments and ablation studies on the representative skin lesion benchmarks, we successfully validate the effectiveness of our proposed framework.
Yiqian Xu, Rui Feng 0001
ICASSP3
2024 Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning
abstract
Video paragraph captioning task aims to generate a detailed, fluent and relevant paragraph for a given video. Prior studies often focus on isolating visual objects (potential main components in a sentence) from the overall video content. They rarely explore the latent semantic relations between objects and high-level video concepts, resulting in dull or even incorrect descriptions. To create fine-grained and contextually relevant paragraph captions, we propose a novel framework that constructs a concept graph from a commonsense knowledge base and infers richer semantic meaning from the visual objects. Moreover, we employ a Vision-Guided Concept Selection Network that incorporates an under-sentence supervision mechanism to align the external knowledge with the visual information. Through extensive experiments on ActivityNet captions and YouCook2, the effectiveness of our method is demonstrated compared to state-of-the-art methods.
Guorui Yu, Yimin Hu, Yiqian Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP5
2024 Cross-Image Distillation for Semi-Supervised Semantic Segmentation
abstract
Semi-supervised semantic segmentation approaches have drawn much more attention in recent years, which aim to exploit a large amount of unlabeled data together with a small number of labeled data. However, existing models usually regarded segmentation as pixel-wise classification, neglecting global semantic relations among pixels across various images. Moreover, scarce annotated data usually exhibits a biased distribution against the desired one, hindering performance improvement. To address these challenging problems, we propose a novel cross-image distillation framework for semi-supervised semantic segmentation. Specifically, we introduce a relation distillation module to model inter-channel correlations between features of labeled samples and unlabeled samples. In addition, we propose a style distillation strategy to explicitly calibrate the learned feature distributions of labeled and unlabeled data to be aligned. Experimental results on two popular benchmarks demonstrate that our proposed approach achieves superior performance over other state-of-the-art methods. We will release the code soon.
Junlin Hou, Rui Feng 0001
ICASSP6
2024 Tag2Text: Guiding Vision-Language Model via Image Tagging
abstract
This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a limited detector, our approach utilizes tags parsed from its paired text to learn an image tagger and meanwhile provides guidance to vision-language models. Given that, Tag2Text can utilize large-scale annotation-free image tags in accordance with image-text pairs, and provides more diverse tag categories beyond objects. Strikingly, Tag2Text showcases the ability of a foundational image tagging model, with superior zero-shot performance even comparable to full supervision manner. Moreover, by leveraging tagging guidance, Tag2Text effectively enhances the performance of vision-language models on both generation-based and alignment-based tasks. Across a wide range of downstream benchmarks, Tag2Text achieves state-of-the-art results with similar model sizes and data scales, demonstrating the efficacy of the proposed tagging guidance.
Youcai Zhang, Jinyu Ma, Rui Feng 0001, Yuejie Zhang, Yandong Guo, Lei Zhang 0001
ICLR5
2024 Temporal Feature Aggregation for Efficient 2D Video Grounding
abstract
Video grounding aims to locate the target video moment in an untrimmed video based on a text query. Most existing methods employ 3D CNNs as the video feature extractor, incurring substantial computational costs. Only a few methods use 2D backbones for video feature extraction, and they suffer from diminished accuracy due to the inherent lack of temporal information within 2D features. To address this problem, we propose a novel 2D video grounding method called TFA that improves accuracy while minimizing computational costs. Our approach involves a query-guided temporal feature aggregation module designed to explicitly capture temporal information. We disentangle time intervals of input video frames and prediction spans to reduce computational overhead. Additionally, we introduce deformable attention into the multi-modal encoder for further enhancement. Extensive experiments on two public datasets demonstrate that our method outperforms previous 2D video grounding methods and achieves competitive results with most 3D methods at significantly reduced costs.
Mohan Chen 0001, Yiren Zhang, Jueqi Wei, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICME5
2024 Memory-Augmented Transformer for Efficient End-to-End Video Grounding
abstract
Video grounding aims to localize a specific segment corresponding to a text query in an untrimmed video. Due to the tremendous computational cost required to process the video frames, the de facto paradigm of video grounding is to extract video features using pretrained video encoders. The parameters of the video encoders are fixed during training, which limits the performance of the localization model. To solve this problem, we propose a Memory-Augmented Transformer (MAT) model. Specifically, each video is split into non-overlapping clips, and our MAT processes videos in a clip-by-clip manner while caching video features into FIFO cached memory queues. By enabling early return, our MAT outperforms previous methods with only less than 60% frames seen. Extensive experimental results on three public benchmark datasets demonstrate that our MAT can achieve competitive performance while being much more efficient than currently prevailing two-stage methods. Code is available at https://github.com/xuyw1997/MAT.
Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICME4
2024 ADSNet: Cross-Domain LTV Prediction with an Adaptive Siamese Network in Advertising
abstract
Advertising platforms have evolved in estimating Lifetime Value (LTV) to better align with advertisers' true performance metric which considers cumulative sum of purchases a customer contributes over a period. Accurate LTV estimation is crucial for the precision of the advertising system and the effectiveness of advertisements. However, the sparsity of real-world LTV data presents a significant challenge to LTV predictive model(i.e., pLTV), severely limiting the their capabilities. Therefore, we propose to utilize external data, in addition to the internal data of advertising platform, to expand the size of purchase samples and enhance the LTV prediction model of the advertising platform. To tackle the issue of data distribution shift between internal and external platforms, we introduce an Adaptive Difference Siamese Network (ADSNet), which employs cross-domain transfer learning to prevent negative transfer. Specifically, ADSNet is designed to learn information that is beneficial to the target domain. We introduce a gain evaluation strategy to calculate information gain, aiding the model in learning helpful information for the target domain and providing the ability to reject noisy samples, thus avoiding negative transfer. Additionally, we also design a Domain Adaptation Module as a bridge to connect different domains, reduce the distribution distance between them, and enhance the consistency of representation space distribution. We conduct extensive offline experiments and online A/B tests on a real advertising platform. Our proposed ADSNet method outperforms other methods, improving GINI by 2%. The ablation study highlights the importance of the gain evaluation strategy in negative gain sample rejection and improving model performance. Additionally, ADSNet significantly improves long-tail prediction. The online A/B tests confirm ADSNet's efficacy, increasing online LTV by 3.47% and GMV by 3.89%.
Ying Cheng 0005, Qi He 0011, Xing Zhou 0003, Rui Feng 0001, Jie Jiang 0008
KDD6
2024 Anatomical Structure-Guided Medical Vision-Language Pre-training
Qingqiu Li, Xiaohan Yan, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen
MICCAI (11)6
2024 DeepPointMap2: Accurate and Robust LiDAR-Visual SLAM with Neural Descriptors
abstract
Simultaneous Localization and Mapping (SLAM) plays a pivotal role in autonomous driving and robotics. Existing methods often rely on hand-craft feature extraction and cross-modal fusion techniques, resulting in limited feature representation capability and reduced robustness. To address this challenge, we introduce DeepPointMap2, a novel learning-based LiDAR-Visual SLAM architecture that leverages neural descriptors to tackle multiple SLAM sub-tasks in a unified manner. Our approach employs neural networks to extract multi-modal tokens, which are then adaptively fused by the Visual-Point Fusion Module to generate sparse 3D neural descriptors, ensuring precise and robust performance. As a pioneering work, our method achieves state-of-the-art localization performance among various Visual-, LiDAR-, and Visual-LiDAR-based methods in widely-used benchmarks, as shown in the experiment results. Furthermore, the approach proves to be robust in scenarios involving camera failure and LiDAR obstruction.
Xiaze Zhang, Ziheng Ding, Ying Cheng 0005, Wenchao Ding 0001, Rui Feng 0001
ACM Multimedia6
2024 CT2C-QA: Multimodal Question Answering over Chinese Text, Table and Chart
abstract
Multimodal Question Answering (MMQA) is crucial as it enables comprehensive understanding and accurate responses by integrating insights from diverse data representations such as tables, charts, and text. Most existing researches in MMQA only focus on two modalities such as image-text QA, table-text QA and chart-text QA, and there remains a notable scarcity in studies that investigate the joint analysis of text, tables, and charts. In this paper, we present CT2C-QA, a pioneering Chinese reasoning-based QA dataset that includes an extensive collection of text, tables, and charts, meticulously compiled from 200 selectively sourced webpages. Our dataset simulates real webpages and serves as a great test for the capability of the model to analyze and reason with multimodal data, because the answer to a question could appear in various modalities, or even potentially not exist at all. Additionally, we present AED (Allocating, Expert and Decision), a multi-agent system implemented through collaborative deployment, information interaction, and collective decision-making among different agents. Specifically, the Assignment Agent is in charge of selecting and activating expert agents, including those proficient in text, tables, and charts. The Decision Agent bears the responsibility of delivering the final verdict, drawing upon the analytical insights provided by these expert agents. We execute a comprehensive analysis, comparing AED with various state-of-the-art models in MMQA, including GPT-4. The experimental outcomes demonstrate that current methodologies, including GPT-4, are yet to meet the benchmarks set by our dataset.
Tianhao Cheng, Yuejie Zhang, Ying Cheng 0005, Rui Feng 0001
ACM Multimedia5
2024 Mixtures of Experts for Audio-Visual Learning
abstract
With the rapid development of multimedia technology, audio-visual learning has emerged as a promising research topic within the field of multimodal analysis. In this paper, we explore parameter-efficient transfer learning for audio-visual learning and propose the Audio-Visual Mixture of Experts (\ourmethodname) to inject adapters into pre-trained models flexibly. Specifically, we introduce unimodal and cross-modal adapters as multiple experts to specialize in intra-modal and inter-modal information, respectively, and employ a lightweight router to dynamically allocate the weights of each expert according to the specific demands of each task. Extensive experiments demonstrate that our proposed approach \ourmethodname achieves superior performance across multiple audio-visual tasks, including AVE, AVVP, AVS, and AVQA. Furthermore, visual-only experimental results also indicate that our approach can tackle challenging scenes where modality information is missing. The source code is available at \url{https://github.com/yingchengy/AVMOE}.
Ying Cheng 0005, Rui Feng 0001
NeurIPS4
2024 Learning Music-Dance Representations Through Explicit-Implicit Rhythm Synchronization
abstract
Although audio-visual representation has been proven to be applicable in many downstream tasks, the representation of dancing videos, which is more specific and always accompanied by music with complex auditory contents, remains challenging and uninvestigated. Considering the intrinsic alignment between the cadent movement of the dancer and music rhythm, we introduceMuDaR, a novelMusic-DanceRepresentation learning framework to perform the synchronization of music and dance rhythms both in explicit and implicit ways. Specifically, we derive the dance rhythms based on visual appearance and motion cues inspired by the music rhythm analysis. Then the visual rhythms are temporally aligned with the music counterparts, which are extracted by the amplitude of sound intensity. Meanwhile, we exploit the implicit coherence of rhythms implied in audio and visual streams by contrastive learning. The model learns the joint embedding by predicting the temporal consistency between audio-visual pairs. The music-dance representation, together with the capability of detecting audio and visual rhythms, can further be applied to three downstream tasks: (a) dance classification, (b) music-dance retrieval, and (c) music-dance retargeting. Extensive experiments demonstrate that our proposed framework outperforms other self-supervised methods by a large margin.
Jiashuo Yu, Junfu Pu, Ying Cheng 0005, Rui Feng 0001, Ying Shan
IEEE Trans. Multim.4
2023 Enhanced Knowledge Injection for Radiology Report Generation
abstract
Automatic generation of radiology reports holds crucial clinical value, as it can alleviate substantial workload on radiologists and remind less experienced ones of potential anomalies. Despite the remarkable performance of various image captioning methods in the natural image field, generating accurate reports for medical images still faces challenges, i.e., disparities in visual and textual data, and lack of accurate domain knowledge. To address these issues, we propose an enhanced knowledge injection framework, which utilizes two branches to extract different types of knowledge. The Weighted Concept Knowledge (WCK) branch is responsible for introducing clinical medical concepts weighted by TF-IDF scores. The Multimodal Retrieval Knowledge (MRK) branch extracts triplets from similar reports, emphasizing crucial clinical information related to entity positions and existence. By integrating this finer-grained and well-structured knowledge with the current image, we are able to leverage the multi-source knowledge gain to ultimately facilitate more accurate report generation. Extensive experiments have been conducted on two public benchmarks, demonstrating that our method achieves superior performance over other state-of-the-art methods. Ablation studies further validate the effectiveness of two extracted knowledge sources.
Qingqiu Li, Jilan Xu, Runtian Yuan, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
BIBM6
2023 Semi-MedSeq: Semi-supervised Semantic Segmentation for Medical Image Sequences
abstract
In clinical practice, medical imaging techniques include 2D video-based examinations that capture sequential scans, and 3D volumetric imaging that forms a comprehensive 3D representation from a stack of 2D slices. The medical image sequences produced by the above techniques provide valuable spatio-temporal characteristics for analysis and segmentation, but the annotation of image sequences is extremely time-consuming and labor-intensive. To exploit the coherence and address the scarcity of labeled data, we propose a novel semi-supervised semantic segmentation framework for medical image sequences, which consists of a conditional network and a denoising network. Specifically, we embed a Sequential Feature Reconstruction module into both networks. This module reconstructs the target frame from contiguous frames and captures their shared visual features. Guided by the context-enhancing information from the conditioning network, the denoising network suppresses background noise via a Diffusion-based Noise Elimination module. Extensive experiments are conducted on 2D and 3D tasks, including cardiac segmentation, polyp segmentation, placenta vessel segmentation and abdomen multi-organ segmentation. The results show our method is superior to existing semi-supervised methods and exhibits advantages over fully-supervised medical image segmentation methods with only 1/2 labeled data, validating its effectiveness and generalization ability.
Runtian Yuan, Jilan Xu, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM5
2023 Learning Open-Vocabulary Semantic Segmentation Models From Natural Language Supervision
abstract
This paper considers the problem of open-vocabulary semantic segmentation (OVS), that aims to segment objects of arbitrary classes beyond a pre-defined, closed-set categories. The main contributions are as follows: First, we propose a transformer-based model for OVS, termed as OVSegmentor, which only exploits web-crawled imagetext pairs for pre-training without using any mask annotations. OVSegmentor assembles the image pixels into a set of learnable group tokens via a slotattention based binding module, then aligns the group tokens to corresponding caption embeddings. Second, we propose two proxy tasks for training, namely masked entity completion and cross-image mask consistency. The former aims to infer all masked entities in the caption given group tokens, that enables the model to learn fine-grained alignment between visual groups and text entities. The latter enforces consistent mask predictions between images that contain shared entities, encouraging the model to learn visual invariance. Third, we construct CC4M dataset for pre-training by filtering CC12M with frequently appeared entities, which significantly improves training efficiency. Fourth, we perform zero-shot transfer on four benchmark datasets, PASCAL VOC, PASCAL Context, COCO Object, and ADE20K. OVSegmentor achieves superior results over state-of-the-art approaches on PASCAL VOC using only 3% data (4M vs 134M) for pre-training.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Yi Wang 0074, Yu Qiao 0001, Weidi Xie
CVPR4
2023 Large Language Models are Complex Table Parsers
abstract
With the Generative Pre-trained Transformer 3.5 (GPT-3.5)exhibiting remarkable reasoning and comprehension abilities in Natural Language Processing (NLP), most Question Answering (QA) research has primarily centered around general QA tasks based on GPT, neglecting the specific challenges posed by Complex Table QA.In this paper, we propose to incorporate GPT-3.5 to address such challenges, in which complex tables are reconstructed into tuples and specific prompt designs are employed for dialogues.Specifically, we encode each cell's hierarchical structure, position information, and content as a tuple.By enhancing the prompt template with an explanatory description of the meaning of each tuple and the logical reasoning process of the task, we effectively improve the hierarchical structure awareness capability of GPT-3.5 to better parse the complex tables.Extensive experiments and results on Complex Table QA datasets, i.e., the open-domain dataset HiTAB and the aviation domain dataset AIT-QA show that our approach significantly outperforms previous work on both datasets, leading to state-of-theart (SOTA) performance.
Changkai Ji, Yuejie Zhang, Yingwen Wang, Rui Feng 0001
EMNLP7
2023 Diabetic Retinopathy Grading with Weakly-Supervised Lesion Priors
abstract
Explicit information of lesions can provide visual instructions for diabetic retinopathy (DR) grading on fundus images. However, pixel-level lesion annotations are extremely difficult and time-consuming to acquire. In this work, we propose a novel weakly-supervised lesion-aware network for DR grading, which enhances the discriminative features with lesion priors by only image-level supervision. Specifically, we design a lesion attention module that generates lesion activation maps by introducing an auxiliary task of binary DR identification. Lesion activation maps are utilized to assist the network to focus on the most relevant regions for boosting DR grading performance. Besides, we particularly devise an adaptive joint loss to balance the DR identification and DR grading tasks dynamically. Extensive results on the public DR dataset demonstrate the superiority and generality of our proposed lesion-aware network. The interpretability of generated lesion activation maps is also verified by the comparison with ground truth segmentation masks.
Junlin Hou, Jilan Xu, Rui Feng 0001, Yue Zhang 0004, Haidong Zou, Lina Lu, Wenwen Xue
ICASSP4
2023 Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal Knowledge
abstract
Video paragraph captioning task aims at generating a fine-grained, coherent and relevant paragraph for a video. Different from the images where objects are static, the temporal states of objects are changing in videos. The dynamic information could be contributed to understanding the whole video content. Existing works rarely put focus on modeling the dynamic changing state of the objects in the videos, causing the activities occurred in videos are poorly or wrongly depicted in paragraphs. To address this problem, we propose a novel Object State Tracking Network, which can capture the temporal state change of objects. However, due to the similarity of the consecutive frames in the videos, the information of the video is redundant and noisy. We further propose a semantic alignment mechanism, and enable the sentence information to refine the visual information. Extensive experiments on ActivityNet Captions demonstrate the effectiveness of our method.
Yimin Hu, Guorui Yu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICASSP4
2023 SCSGNet: Spatial-Correlated and Shape-Guided Network for Breast Mass Segmentation
abstract
Automatic and accurate breast mass segmentation plays a crucial role in the early diagnosis of breast cancer. However, it has been a challenging task for two main reasons: (1) Breast masses are diverse; and (2) The boundaries of masses are ambiguous. To address these problems, we propose a Spatial-Correlated and Shape-Guided Network (SCSGNet), which combines global context extraction with local boundary refinement. Specifically, the high-level features are aggregated to produce a global map as the initial guidance area, and a Series-Parallel Feature Fusion (SPFF) module is added to capture masses of different shapes and sizes. Besides, we design a Dynamic Long-range Correlation Capture (DLCC) module to capture the spatial correlation of masses at different positions. Finally, we devise a Triplet Attention Guide (TAG) module to iteratively update the feature map and refine the boundary. Experiments on two public datasets demonstrate that our method achieves superior performance over other state-of-the-art methods.
Qingqiu Li, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001
ICASSP5
2023 Boosting Fine-Grained Sketch-Based Image Retrieval with Self-Supervised Learning
abstract
Fine-grained sketch-based image retrieval (FG-SBIR) aims at aligning images and sketches at the instance level. It is a challenging task as there are significant differences between sketch and image. Existing methods usually produce less desired performance due to the lack of large-scale fine-grained image-sketch datasets and the strong dependence on the classification models pretrained on ImageNet. In this paper, we propose a better self-supervised pre-trained FG-SBIR model which does not depend on large-scale annotated datasets. Only images and their corresponding edge maps are used at the pre-training stage. Mixed modal transformation is designed to generate different mixed-up views. The FG-SBIR model is pre-trained by minimizing the distance between the views of the same instance and then fine-tuned by a simple triplet loss. With a plain downstream network, it achieves generally better performance than state-of-the-art models on three widely used FG-SBIR datasets.
Zhaolong Zhang, Yangdong Chen, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022
ICASSP4
2023 Video Captioning via Relation-Aware Graph Learning
abstract
Recent neural models for video captioning usually employed an encoder-decoder framework. However, most approaches either neglected the spatial and temporal interactions between objects in a video or implicitly modelled the interactions, resulting in less desired performance. In this paper, we propose a novel relation-aware graph learning framework. It explicitly models both spatial and temporal relations for objects. In particular, a relation-aware graph is designed to depict the spatial relations between different objects in a scene. Parallelly, a temporal graph network is designed to perform relational reasoning for the same objects in adjacent frames. Features of both types of relations are learned and fused for the follow-up language decoder. Experiments on two bench-mark datasets show the effectiveness of our framework. It achieves state-of-the-art performance with CIDEr scores on MSVD and MSR-VTT.
Heming Jing, Qiujie Xie, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP5
2023 Class-aware Variational Auto-encoder for Open Set Recognition
abstract
Compared with traditional classification models trained under the closed world assumption, Open Set Recognition (OSR) requires accurate classification for known classes as well as rejection for unknown ones. By modeling the distribution of each known class, Conditional Variational Auto-encoder (CVAE) has achieved great success in OSR, even though it was originally proposed for image generation. In this paper, we propose a novel two-stage learning framework, Class-aware Variational Auto-encoder (CA-VAE) to better adapt CVAE to the OSR task. Pre-derived attention images are taken as the objective target for reconstruction, thus model is implicitly directed to focus on the class-discriminative regions of the image. In this way, the learned latent representation is de-biased towards class-aware. Experiments on standard image datasets demonstrate the outperformance of the proposed method over existing ones, which achieves new state-of-the-art results. Codes are available at https://github.com/roywang021/CA-VAE.
Ling Su, Yingzi Ye, Yuejie Zhang, Rui Feng 0001
ICME8
2023 Conditional Video-Text Reconstruction Network with Cauchy Mask for Weakly Supervised Temporal Sentence Grounding
abstract
Temporal sentence grounding aims to detect the target segment most related to a given query in an untrimmed video. To alleviate the expensive annotation cost for temporal labels, researchers paid more attention to weakly supervised setting. Prior studies neglected the utilization of video representation reconstruction, which led to an unbalanced alignment learning. Moreover, they used different strategies to generate proposals which ignored the temporal structure in a query. In this paper, we propose a novel Conditional Video-Text Reconstruction Network (CVTRN). It supports conditional reconstruction of video and text representation. Specifically, video and text features are fused to compute semantic alignment, which is the condition of reconstruction. A new mask strategy for mask conditioned sentence reconstruction is also devised. This strategy focuses more on boundary regions than the widely used Gaussian mask in previous methods. Experimental results on two public benchmark datasets show that our CVTRN outperforms the state-of-the-art methods.
Jueqi Wei, Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
ICME5
2023 SPTNET: Span-based Prompt Tuning for Video Grounding
abstract
When a Pre-trained Language Model (PLM) is adopted in video grounding task, it usually acts as a text encoder without having its knowledge fully utilized. Also, there exists an inconsistency problem between the pre-training and downstream objectives. To solve the issues, we propose a new paradigm, named Span-based Prompt Tuning (SPTNet). It can convert the video grounding task into a cloze form. Specifically, a query is first changed into a form with mask token by a template, then the video and the query embeddings are integrated through a cross-modal transformer. The start and end points of the query matching time span are predicted with the embedding of the mask token. Experimental results on two public benchmarks ActivityNet Captions and Charades-STA show that our SPTNet achieves surpassing performance compared with state-of-the-art methods.
Yiren Zhang, Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
ICME5
2023 CAMG: Context-Aware Moment Graph Network for Multimodal Temporal Activity Localization via Language
Yuelin Hu, Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
NLPCC (1)4
2023 Enhancing Open-Set Object Detection via Uncertainty-Boxes Identification
Wei Ji 0009, Dongqin Wu, Weijia Fu, Yingwen Wang, Yuejie Zhang, Rui Feng 0001
PRCV (8)7
2023 Deep cross-modal hashing with fine-grained similarity
Yangdong Chen, Jiaqi Quan, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022
Appl. Intell.4
2023 Class-Incremental Generalized Zero-Shot Learning
Zhenfeng Sun, Rui Feng 0001, Yanwei Fu 0001
Multim. Tools Appl.2
2023 Rethinking Local and Global Feature Representation for Dense Prediction
Mohan Chen 0001, Li Zhang 0040, Rui Feng 0001, Xiangyang Xue 0001, Jianfeng Feng
Pattern Recognit.3
2022 Single-Modality Endoscopic Polyp Segmentation via Random Color Reversal Synthesis and Two-Branched Learning
abstract
Endoscopic polyp segmentation plays a fundamental role in the diagnosis and treatment of colorectal cancer. However, polyp segmentation often suffers from limited accuracy due to its large variations in appearance, blurry boundary and severe imbalanced illumination. In this paper, we propose a novel Translation Assisted Segmentation Network (TASNet) for polyp segmentation of single-modality endoscopic images. It consists of two branches, i.e. an image-to-image translation branch and an image segmentation branch. These two branches communicate via a shared encoder. For the image-to-image translation branch, a Color Reversal Strategy is established to treat the original image as source image and synthesize target images. Moreover, we introduce a Random Color Reversal Synthesis module for progressive segmentation. Extensive experiments show that our framework achieves superior performance than state-of-the-art methods on five widely-used endoscopic image datasets.
Mingzhu Chen, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM6
2022 Cross-Field Transformer for Diabetic Retinopathy Grading on Two-field Fundus Images
abstract
Automatic diabetic retinopathy (DR) grading based on fundus photography has been widely explored to benefit the routine screening and early treatment. Existing researches generally focus on single-field fundus images, which have limited field of view for precise eye examinations. In clinical applications, ophthalmologists adopt two-field fundus photography as the dominating tool, where the information from each field (i.e., macula-centric and optic disc-centric) is highly correlated and complementary, and benefits comprehensive decisions. However, automatic DR grading based on two-field fundus photography remains a challenging task due to the lack of publicly available datasets and effective fusion strategies. In this work, we first construct a new benchmark dataset (DRTiD) for DR grading, consisting of 3,100 two-field fundus images. To the best of our knowledge, it is the largest public DR dataset with diverse and high-quality two-field images. Then, we propose a novel DR grading approach, namely Cross-Field Transformer (CrossFiT), to capture the correspondence between two fields as well as the long-range spatial correlations within each field. Considering the inherent two-field geometric constraints, we particularly define aligned position embeddings to preserve relative consistent position in fundus. Besides, we perform masked cross-field attention during interaction to filter the noisy relations between fields. Extensive experiments on our DRTiD dataset and a public DeepDRiD dataset demonstrate the effectiveness of our CrossFiT network. The new dataset and the source code of CrossFiT will be publicly available at https://github.com/DU-VTS/DRTiD.
Junlin Hou, Jilan Xu, Yuejie Zhang, Haidong Zou, Lina Lu, Wenwen Xue, Rui Feng 0001
BIBM9
2022 An Interpretable Causal Approach for Bronchopulmonary Dysplasia Prediction
abstract
In this paper, we focus on the Bronchopulmonary Dysplasia (BPD) prediction task, which aims to identify the BPD in premature infants based on the given medical images. The existing methods for this task sometimes learn the spurious relations (confounders) between the image and label while ignore the causal features due to the limited size of the dataset, which largely influences the interpretability and robustness of the model. To address this challenge problem, we propose an interpretable causal method for BPD prediction, which can eliminate the irrelevant features and capture the causal features. We term our method as Causal Intervention by semantic instrumental Variable (CisiV). First, we design a causal structure graph modeling module to learn the representations of instrumental variables and confounders automatically. The constraints of confounders further guarantee the instrumental variables validity. Meanwhile, we add congenital attributes of premature infants to enable the learned instrumental variables to have medical semantics via a semantic matching module. We apply our model to BPD prediction in premature infants and achieve promising results. Extensive experiments indicate that CisiV could facilitate early intervention and provide support for clinical decision-making. Codes will be released to the community.
Wei Ji 0009, Liangfeng Tang, Yuejie Zhang, Rui Feng 0001
BIBM6
2022 Multi-contrast High Quality MR Image Super-Resolution with Dual Domain Knowledge Fusion
abstract
Multi-contrast high quality high-resolution (HR) Magnetic Resonance (MR) images enrich available information for diagnosis and analysis. Deep convolutional neural network methods have shown promising ability for MR image super-resolution (SR) given low-resolution (LR) MR images. Methods taking HR images as references (Ref) have made progress to enhance the effect of MR images SR. However, existing multi-contrast MR image SR approaches are based on contrasting-expanding backbones, which lose high frequency information of Ref image during downsampling. They also failed to transfer textures of Ref image into target domain. In this paper, we propose Edge Mask Transformer UNet (EMFU) for accelerating MR images SR. We propose Edge Mask Transformer (EMF) to generate global details and texture representation of target domain. Dual domain fusion module in UNet aggregates semantic information of the representation and LR image of target domain. Specifically, we extract and encode edge masks to guide the attention in EMF by re-distributing the embedding tensors, so that the network allocates more attention to image edge area. We also design a dual domain fusion module with self-attention and cross-attention to deeply fuse semantic information of multiple protocols for MRI. Extensive experiments show the effectiveness of our proposed EMFU, which surpasses state-of-the-art methods on benchmarks quantitatively and visually. Codes will be released to the community.
Runhan Wang, Weijia Fu, Yuejie Zhang, Rui Feng 0001
BIBM6
2022 MedSeq: Semantic Segmentation for Medical Image Sequences
abstract
Medical image segmentation plays a critical role in computer-aided diagnosis, while the diversity and complexity of medical images make it difficult to segment precisely. In practice, medical images of specific modalities (e.g. Magnetic Resonance Imaging, Colonoscopy and Ultrasonography) are collected as sequences independently for every patient. However, 1) there exists few works exploiting sequence information among successive frames, neglecting inter-frame relationships that are useful to locate target objects; 2) the performance of medical image segmentation is limited to the low contrast or blurry boundary of medical images, and intra-frame dependencies are not fully explored. Thus in this paper, we propose MedSeq for segmenting objects of interest in medical image sequences. Following the “locate-then-refine” paradigm, we locate target regions by modeling cross-frame relationships and then perform refinement on coarse masks. More specifically, we design a Cross-frame Attention module to learn correlations among frames, taking advantages of their similar appearances. For refinement, we propose a novel Boundary-aware Transformer to improve the segmentation of boundary patches. Extensive experiments are conducted on benchmark datasets of Cardiac Segmentation and Video Polyp Segmentation. Our method achieves superior performance over the state-of-the-art methods.
Runtian Yuan, Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM5
2022 Pyramid Region-based Slot Attention Network for Temporal Action Proposal Generation
Lingbo Liu, Rui Feng 0001
BMVC6
2022 CREAM: Weakly Supervised Object Localization via Class RE-Activation Mapping
abstract
Weakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) de-rived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomplete localization problem). In this paper, we empirically prove that this problem is associated with the mixup of the activation values between less discrimi-native foreground regions and the background. To address it, we propose Class RE-Activation Mapping (CREAM), a novel clustering-based approach to boost the activation values of the integral object regions. To this end, we in-troduce class-specific foreground and background context embeddings as cluster centroids. A CAM-guided momen-tum preservation strategy is developed to learn the context embeddings during training. At the inference stage, the re-activation mapping is formulated as a parameter es-timation problem under Gaussian Mixture Model, which can be solved by deriving an unsupervised Expectation- Maximization based soft-clustering algorithm. By simply integrating CREAM into various WSOL approaches, our method significantly improves their performance. CREAM achieves the state-of-the-art performance on CUB, ILSVRC and OpenImages benchmark datasets. Code will be avail-able at https://github.com/lazzcharles/CREAM.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
CVPR4
2022 ST2PE: Spatial and Temporal Transformer for Pose Estimation
Yuan Wu 0004, Yanlu Cai, Rui Feng 0001, Cheng Jin 0001
ICANN (2)3
2022 Semantic-Driven Saliency-Context Separation for Video Captioning
abstract
Video captioning aims at generating a natural language de-scription for a given video clip including not only salient sce-narios but also contextual scenarios. The former reveal the highlight of a video and are usually the focus of most existing captioning methods. The latter, however, are not well ex-plored and even ignored easily, though they may provide cer-tain detailed and latent information that can help with a better understanding of the video. To effectively exploit the infor-mation contained in both, a novel video captioning network is proposed. It has two key modules: Cross-Modality Selection (CMS) and Saliency-Context Adaptive Decoder (SCAD). Specifically, CMS mainly focuses on utilizing the semantic information to distinguish saliency and context. Meanwhile, SCAD adaptively identifies both the saliency and context to generate more detailed and precise captions. Experiments on two benchmark datasets, i.e., MSVD and MSR-VTT, demon-strate the effectiveness of our model through the comparison with state-of-the-art methods.
Heming Jing, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICME3
2022 Self-Supervised Video Representation Learning with Motion-Contrastive Perception
abstract
Visual-only self-supervised learning has achieved significant improvement in video representation learning. Existing related methods encourage models to learn video representations by utilizing contrastive learning or designing specific pretext tasks. However, some models are likely to focus on the background, which is unimportant for learning video representations. To alleviate this problem, we propose a new view called long-range residual frame to obtain more motion-specific information. Based on this, we propose the Motion-Contrastive Perception Network (MCPNet), which consists of two branches, namely, Motion Information Perception (MIP) and Contrastive Instance Perception (CIP), to learn generic video representations by focusing on the changing areas in videos. Specifically, the MIP branch aims to learn fine-grained motion features, and the CIP branch performs contrastive learning to learn overall semantics information for each instance. Experiments on two benchmark datasets UCF-101 and HMDB-51 show that our method outperforms current state-of-the-art visual-only self-supervised approaches.
Ying Cheng 0005, Yuejie Zhang, Rui Feng 0001
ICME5
2022 STDNet: Spatio-Temporal Decomposed Network for Video Grounding
abstract
Previous methods for video grounding treated either the query or the video as a whole, while neglecting their respective semantics in the orthogonal space and time dimensions. Since spatial semantics appears frequently in a video, temporal semantics is more discriminative and deserves more attention. Based on such considerations, we propose a novel Spatio-Temporal Decomposed Network (STDNet) which decomposes the query and the video into their spatial and temporal semantics, respectively. Specifically, spatial and temporal words are selected from the query, and the video is split into two pathways. Spatial cross-modal attention is computed first and serves as prior knowledge for temporal attention. A new localization strategy is also devised which regresses the segment's start conditioned on the end and essentially breaks the independence assumption made in previous methods. Experimental results on three public benchmark datasets show that our STDNet outperforms the state-of-the-art methods.
Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICME3
2022 Dynamically Connected Graph Representation for Object Detection
Shuyu Miao, Rui Feng 0001
ICONIP (3)4
2022 TCCNet: Temporally Consistent Context-Free Network for Semi-supervised Video Polyp Segmentation
abstract
Automatic video polyp segmentation (VPS) is highly valued for the early diagnosis of colorectal cancer. However, existing methods are limited in three respects: 1) most of them work on static images, while ignoring the temporal information in consecutive video frames; 2) all of them are fully supervised and easily overfit in presence of limited annotations; 3) the context of polyp (i.e., lumen, specularity and mucosa tissue) varies in an endoscopic clip, which may affect the predictions of adjacent frames. To resolve these challenges, we propose a novel Temporally Consistent Context-Free Network (TCCNet) for semi-supervised VPS. It contains a segmentation branch and a propagation branch with a co-training scheme to supervise the predictions of unlabeled image. To maintain the temporal consistency of predictions, we design a Sequence-Corrected Reverse Attention module and a Propagation-Corrected Reverse Attention module. A Context-Free Loss is also proposed to mitigate the impact of varying contexts. Extensive experiments show that even trained under 1/15 label ratio, TCCNet is comparable to the state-of-the-art fully supervised methods for VPS. Also, TCCNet surpasses existing semi-supervised methods for natural image and other medical image segmentation tasks.
Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
IJCAI4
2022 IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training
abstract
Vision-Language Pre-training (VLP) with large-scale image-text pairs has demonstrated superior performance in various fields. However, the image-text pairs co-occurrent on the Internet typically lack explicit alignment information, which is suboptimal for VLP. Existing methods proposed to adopt an off-the-shelf object detector to utilize additional image tag information. However, the object detector is time-consuming and can only identify the pre-defined object categories, limiting the model capacity. Inspired by the observation that the texts incorporate incomplete fine-grained image information, we introduce IDEA, which stands for increasing text diversity via online multi-label recognition for VLP. IDEA shows that multi-label learning with image tags extracted from the texts can be jointly optimized during VLP. Moreover, IDEA can identify valuable image tags online to provide more explicit textual supervision. Comprehensive experiments demonstrate that IDEA can significantly boost the performance on multiple downstream datasets with a small extra computational cost.
Youcai Zhang, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang, Yandong Guo
ACM Multimedia6
2022 MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video Parsing
abstract
Recognizing and localizing events in videos is a fundamental task for video understanding. Since events may occur in auditory and visual modalities, multimodal detailed perception is essential for complete scene comprehension. Most previous works attempted to analyze videos from a holistic perspective. However, they do not consider semantic information at multiple scales, which makes the model difficult to localize events in different lengths. In this paper, we present a Multimodal Pyramid Attentional Network (MM-Pyramid ) for event localization. Specifically, we first propose the attentive feature pyramid module. This module captures temporal pyramid features via several stacking pyramid units, each of them is composed of a fixed-size attention block and dilated convolution block. We also design an adaptive semantic fusion module, which leverages a unit-level attention block and a selective fusion block to integrate pyramid features interactively. Extensive experiments on audio-visual event localization and weakly-supervised audio-visual video parsing tasks verify the effectiveness of our approach.
Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang
ACM Multimedia4
2022 Modality-aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence Detection
abstract
Weakly-supervised audio-visual violence detection aims to distinguish snippets containing multimodal violence events with video-level labels. Many prior works perform audio-visual integration and interaction in an early or intermediate manner, yet overlooking the modality heterogeneousness over the weakly-supervised setting. In this paper, we analyze the modality asynchrony and undifferentiated instances phenomena of the multiple instance learning (MIL) procedure, and further investigate its negative impact on weakly-supervised audio-visual learning. To address these issues, we propose a modality-aware contrastive instance learning with self-distillation (MACIL-SD) strategy . Specifically, we leverage a lightweight two-stream network to generate audio and visual bags, in which unimodal background, violent, and normal instances are clustered into semi-bags in an unsupervised way. Then audio and visual violent semi-bag representations are assembled as positive pairs, and violent semi-bags are combined with background and normal instances in the opposite modality as contrastive negative pairs. Furthermore, a self-distillation module is applied to transfer unimodal visual knowledge to the audio-visual model, which alleviates noises and closes the semantic gap between unimodal and multimodal features. Experiments show that our framework outperforms previous methods with lower complexity on the large-scale XD-Violence dataset. Results also demonstrate that our proposed approach can be used as plug-in modules to enhance other networks. Codes are available at https://github.com/JustinYuu/MACIL_SD.
Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang
ACM Multimedia4
2022 AE-Net: Fine-grained sketch-based image retrieval via attention-enhanced network
Yangdong Chen, Zhaolong Zhang, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.5
2022 Balanced single-shot object detection using cross-context attention-guided network
Shuyu Miao, Shanshan Du, Rui Feng 0001, Yuejie Zhang, Tianbi Liu, Weiguo Fan
Pattern Recognit.3
2022 Stacked Multimodal Attention Network for Context-Aware Video Captioning
abstract
Recent neural models for video captioning usually employ an attention-based encoder-decoder framework. However, current approaches mainly attend to the motion features and object features of the video when generating the caption, but ignore the potential but useful historical information. Besides, exposure bias and vanishing gradients problems always exist in current caption generation models. In this paper, we propose a novel video captioning framework, named Stacked Multimodal Attention Network (SMAN). It adopts additional visual and textual historical information during caption generation as context features, employs a stacked architecture to process different features gradually, and utilizes the Reinforcement Learning method and coarse-to-fine training strategy to further improve the generated results. Both quantitative and qualitative experiments on the benchmark datasets ofMSVDandMSR-VTTshow the effectiveness and feasibility of our framework. The codes are available onhttps://github.com/zhengyi123456/SMAN.
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
IEEE Trans. Circuits Syst. Video Technol.3
2021 CellDet: Dual-Task Cell Detection Network for IHC-Stained Image Analysis
abstract
Cell detection on immunohistochemistry stained (IHC-stained) images plays an essential role in computer assisted prediction of tumor progression and treatment response. Currently available cell detection datasets provide either point level or bounding box level annotations for deep object detection network training. And these widely used networks usually employ standard pyramid structured multi-scale feature fusion. However, we find that these methods have obvious limitation when facing large amounts of cells in similar scale with severe overlapping. To address this problem, we propose a novel CellDet network with (1) Scale Consistency Feature Fusion Module (SCFFM) and (2) Dual Task Detection Module to simultaneously exploit the complementary information from both point and bounding box annotations. In order to verify the effectiveness our proposed method, We make efforts to relabel the public SHIDC-B-Ki-67 dataset with bounding box annotations. Extensive experimental results show that the proposed CellDet outperforms other state-of-the-art cell detection methods with a remarkable margin. We will release our source code and dataset in https://github.com/JiweiMaster/celldet.
Wei Ji 0009, Wenbin Pan, Rui Feng 0001, Yuejie Zhang, Gang Jin
BIBM6
2021 Semi-supervised Medical Image Segmentation with Distribution Calibration and Non-local Semantic Constraint
abstract
The available medical images with accurate segmentation masks are usually limited due to the expensive and time-consuming annotation cost. Many semi-supervised approaches tried to exploit a large amount of unlabeled data together with the small number of labeled data. However, their learned segmentation models can easily become overfitted on the biased labeled data, mainly because of the misalignment between the labeled and unlabeled data distribution. To address this challenging problem, we propose a novel semi-supervised model with distribution calibration and non-local semantic constraint for medical image segmentation. In specific, we explicitly calibrate the learned feature distributions of the labeled and unlabeled data to make them aligned. Meanwhile, we add a special nonlocal semantic loss to encourage the learned features to be more discriminative for the segmentation task at the same time. Consequently, our final segmentation networks have the advantage to better generalize on the unlabeled data in both training and test set. Experimental results on three popular medical image segmentation benchmarks demonstrate that our proposed model achieves superior performance over other state-of-the-art methods. We will release our source code in this URL.
Junlin Hou, Rui Feng 0001, Yuejie Zhang
BIBM4
2021 Improving Multimodal Speech Enhancement by Incorporating Self-Supervised and Curriculum Learning
abstract
Speech enhancement in realistic scenarios still remains many challenges, such as complex background signals and data limitations. In this paper, we present a co-attention based framework that incorporates self-supervised and curriculum learning to derive the target speech in noisy environments. Specifically, we first leverage self-supervision to pre-train the co-attention model on the task of audio-visual synchronization. The pre-trained model can focus on the lip of speakers automatically, and then the self-supervised features from the model are combined with a u-net regression network to separate the spectrograms of sound mixtures. To make the training process easier and further improve the performance, we introduce the curriculum learning scheme for the training stage of speech enhancement. Extensive experiments show that our model achieves superior performance over previous self-supervised method for speech enhancement, and demonstrate the generalizability of our approach to the transferred dataset.
Ying Cheng 0005, Mengyu He, Jiashuo Yu, Rui Feng 0001
ICASSP4
2021 Object-Oriented Relational Distillation for Object Detection
abstract
Object detection models have achieved increasingly better performance based on more complex architecture designs, but the heavy computation limits their further widespread application on the devices with insufficient computational power. To this end, we propose a novel Object-Oriented Relational Distillation (OORD) method that drives small detection models to have an effective performance like large detection models with constant efficiency. Here, we introduce to distill relative relation knowledge from teacher/large models to student/small models, which promotes the small models to learn better soft feature representation by the guiding of large models. OORD consists of two parts, i.e., Object Extraction (OE) and Relation Distillation (RD). OE extracts foreground features to avoid background feature interference, and RD distills the relative relations between the foreground features through graph convolution. Related experiments conducted on various kinds of detection models show the effectiveness of OORD, which improves the performance of the small model by nearly 10% without additional inference time cost.
Shuyu Miao, Rui Feng 0001
ICASSP2
2021 Exploiting Deep Cross-Slice Features From CT Images For Multi-Class Pneumonia Classification
abstract
Computed Tomography (CT) scanning is widely used for chest diseases detection including pneumonia due to its diagnostic efficacy and efficiency. Recent studies have shown that the distribution of infection regions caused by COVID19 in CT images is different from other pneumonia, and the COVID-19 cases are more likely to suffer severe and large-area infections. However, most deep learning methods only focus on intra-slice features and ignore cross-slice features. In this paper, we propose a novel two-stage method to fully exploit deep cross-slice features from volumetric CT data, including a dual-task supervised CNN and a context-aware Bi-LSTM. To further demonstrate the effectiveness of our model, we conduct extensive experiments on a chest CT imaging dataset with a total of 801 patients (250 healthy people, 238 COVID-19 patients, 191 H1N1 patients, and 122 CAP patients). The experimental results indicate the superiority of our proposed model on the multi-class pneumonia classification task.
Jiawang Cao, Junlin Hou, Longquan Jiang 0003, Weiya Shi, Rui Feng 0001
ICIP8
2021 Temporally Coarse to Fine Snippets Relationship Learning with Graph Convolution for Temporal Action Proposal Generation
abstract
Previous works have shown that explicit snippets relationship modeling can be helpful for feature learning on untrimmed action videos. However, the snippets relationship learning in these methods are far from optimal in that they failed to consider the valuable temporally coarse-grained features, learnable soft relationship weights, and separate relationship learning in different temporal orders. To address this issue, we proposed a novel SGC-Block for improved snippet relationship learning, which enables the temporally coarse-to-fine soft valued snippet-wise relationship learning in different temporal directions. The SGC-Block constructs the snippets graph and explicitly models the (1) temporal relations (TPR); (2) coarsegrained snippet-wise relations (CSR); (3) fine-grained snippet-wise relations (FSR); and an additional (4) adaptive relations (ADR). Especially, the novel CSR is inspired by the feature pyramid pooling structure to obtain the coarse feature presentations in the temporal dimension. Experimental results showed that our proposed approach outperforms most state-of-the-art methods on the THUMOS14 and ActivityNet-1.3 benchmarks.
Shuyu Miao, Rui Feng 0001
ICME4
2021 MPN: Multimodal Parallel Network for Audio-Visual Event Localization
abstract
Audio-visual event localization aims to localize an event that is both audible and visible in the wild, which is a widespread audio-visual scene analysis task for unconstrained videos. To address this task, we propose a Multimodal Parallel Network (MPN), which can perceive global semantics and unmixed local information parallelly. Specifically, our MPN framework consists of a classification subnetwork to predict event categories and a localization subnetwork to predict event boundaries. The classification subnetwork is constructed by the Multimodal Co-attention Module (MCM) and obtains global contexts. The localization subnetwork consists of Multimodal Bottleneck Attention Module (MBAM), which is designed to extract fine-grained segment-level contents. Extensive experiments demonstrate that our framework achieves the state-of-the-art performance both in fully supervised and weakly supervised settings on the Audio-Visual Event (AVE) dataset.
Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001
ICME3
2021 Disentangled Feature Network for Fine-Grained Recognition
Shuyu Miao, Mingming Gong, Rui Feng 0001
ICONIP (2)7
2021 Exploring Logical Reasoning for Referring Expression Comprehension
abstract
Referring expression comprehension aims to localize the target object in an image referred by a natural language expression. Most existing approaches neglect the implicit logical correlations among fine-grained cues, e.g., categories, attributes, which are beneficial for distinguishing objects. In this paper, we propose a logic-guided approach to explore logical knowledge for referring expression comprehension in a hierarchical modular-based framework. Specifically, we propose to extract fine-grained cues in visual and textual domains and perform logical reasoning over them with explicit logical expressions to regularize the matching process without extra parameters. Besides, we propose to improve existing modular-based methods by introducing context information of objects in the relationship module. Extensive experiments are conducted on three referring expression datasets, and the results demonstrate that our model can produce more consistent predictions and further achieve superior performance compared with previous methods.
Ying Cheng 0005, Jiashuo Yu, Yuejie Zhang, Rui Feng 0001
ACM Multimedia6
2021 Cross-modal retrieval with dual multi-angle self-attention
abstract
Abstract In recent years, cross‐modal retrieval has been a popular research topic in both fields of computer vision and natural language processing. There is a huge semantic gap between different modalities on account of heterogeneous properties. How to establish the correlation among different modality data faces enormous challenges. In this work, we propose a novel end‐to‐end framework named Dual Multi‐Angle Self‐Attention (DMASA) for cross‐modal retrieval. Multiple self‐attention mechanisms are applied to extract fine‐grained features for both images and texts from different angles. We then integrate coarse‐grained and fine‐grained features into a multimodal embedding space, in which the similarity degrees between images and texts can be directly compared. Moreover, we propose a special multistage training strategy, in which the preceding stage can provide a good initial value for the succeeding stage and make our framework work better. Very promising experimental results over the state‐of‐the‐art methods can be achieved on three benchmark datasets of Flickr8k, Flickr30k, and MSCOCO.
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
J. Assoc. Inf. Sci. Technol.4
2021 Periphery-aware COVID-19 diagnosis with contrastive representation enhancement
Junlin Hou, Jilan Xu, Longquan Jiang 0003, Shanshan Du, Rui Feng 0001, Yuejie Zhang, Xiangyang Xue 0001
Pattern Recognit.5
2020 Zero-Shot Sketch-Based Image Retrieval via Graph Convolution Network
abstract
Zero-Shot Sketch-based Image Retrieval (ZS-SBIR) has been proposed recently, putting the traditional Sketch-based Image Retrieval (SBIR) under the setting of zero-shot learning. Dealing with both the challenges in SBIR and zero-shot learning makes it become a more difficult task. Previous works mainly focus on utilizing one kind of information, i.e., the visual information or the semantic information. In this paper, we propose a SketchGCN model utilizing the graph convolution network, which simultaneously considers both the visual information and the semantic information. Thus, our model can effectively narrow the domain gap and transfer the knowledge. Furthermore, we generate the semantic information from the visual information using a Conditional Variational Autoencoder rather than only map them back from the visual space to the semantic space, which enhances the generalization ability of our model. Besides, feature loss, classification loss, and semantic loss are introduced to optimize our proposed SketchGCN model. Our model gets a good performance on the challenging Sketchy and TU-Berlin datasets.
Zhaolong Zhang, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
AAAI3
2020 Data-Efficient Histopathology Image Analysis with Deformation Representation Learning
abstract
Histopathological examination of tissue biopsies plays a fundamental role in disease assessment. Automatic histopathology image analysis requires substantial task-specific annotations, which are often expensive and laborious in realworld scenarios. This insufficient annotation of data limits the generalization ability of supervised learning models. To address this challenge, we propose a self-supervised Deformation Representation Learning (DRL) framework to learn semantic features from unlabeled data. As a novel paradigm, our approach utilizes deformation as supervisory signals based on two critical features, i.e., local structure heterogeneity and global context homogeneity. Given an original histopathology image and its deformed counterpart, there exists a moderate difference in local structures. In contrast, due to the transformation-invariance, both images share a similar global context compared with other images. Specifically, an encoder network is trained to distinguish the local inconsistency by measuring the mutual information and maintain the global consistency with noise contrastive estimation. Extensive experiments on public histopathology image datasets show that the learned representations are generalizable for various downstream tasks, such as transfer learning on segmentation and semi-supervised classification. Our approach achieves superior results over other self-supervised methods and the ImageNet pre-trained model, and it reveals the ability as a novel pre-training scheme in histopathology image analysis.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Chunyang Ruan, Tao Zhang 0022, Weiguo Fan
BIBM4
2020 OSD: An Occlusion Skeleton Dataset for Action Recognition
abstract
Currently available 2D skeleton datasets for action recognition mostly contain nonoccluded skeleton samples. Models trained on such datasets lack generalization ability in occlusion situations. In this paper we propose an occlusion projection method, which projects a 3D occlusion object into 2D plane to generate a 2D occluded area. Based on this method, we build a 2D occlusion skeleton dataset named OSD with 56,800 occluded skeleton samples and 60 distinct classes. Experimental results show that the model trained on OSD has better generalization ability in occlusion situations compared with the model trained on datasets with nonoccluded samples, which proves the effectiveness of OSD.
Yuan Wu 0004, Haoyue Qiu, Rui Feng 0001
IEEE BigData4
2020 Learning Class-Based Graph Representation for Object Detection
Shuyu Miao, Rui Feng 0001, Yuejie Zhang, Weiguo Fan
ECAI2
2020 Representation Reconstruction Head for Object Detection
abstract
There are two kinds of detection heads in object detection frameworks. Between them, the heads based on full connection contribute to mapping the learned feature representation to the sample label space, while the heads based on full convolution facilitate preserving location sensitivity information. However, to enjoy the benefits from both detection heads is still underexplored. In this paper, we propose a generalized Representation Reconstruction Head (RRHead) to break through the limitation that most detection heads focus on unilateral self-advantage while ignoring another one. RRHead enhances multi scale feature representation for better feature mapping, and employs location sensitivity representation for better location preservation. These optimize fully-convolutional-based heads and fully-connected-based heads separately. RRHead can be embedded in existing detection frameworks to heighten the rationality and reliability of the detection head representation without any additional modification. Extensive experiments show that our proposed RRHead improves the detection performance of the existing frameworks by a large margin on several challenging benchmarks, and achieves new state-of-the-art performance.
Shuyu Miao, Rui Feng 0001, Yuejie Zhang
ICIP2
2020 DG-FPN: Learning Dynamic Feature Fusion Based on Graph Convolution Network For Object Detection
abstract
Feature Pyramid Network (FPN) is one of the most popular feature fusion methods to address the multi-scale issue in object detection. Current FPN-based methods are mostly designed manually, which cannot guarantee the optimal feature fusion. Besides, the predetermined methods generally provide the same strategy to various targets, which are not distinctive among targets with different scales. In this paper, we present a novel dynamic feature fusion method based on the graph convolution network (GCN), called DG-FPN. The proposed GCN-based method can dynamically transfer knowledge with learnable weights across all nodes, making it possible to learn the optimal feature fusion for detectors. Furthermore, the pixel-based adjacency matrix is proposed to offer customized fusion strategy for each target, achieving dynamic feature fusion. To optimize matrix-driven learning, semantic information is introduced to guide the process of fusion. Experiments show that DG-FPN significantly improves the performance of baseline networks on the challenging MS-COCO object benchmark, especially in small objects.
Shuyu Miao, Rui Feng 0001
ICME3
2020 Video Captioning With Temporal And Region Graph Convolution Network
abstract
Video captioning aims to generate a natural language description for a given video clip that includes not only spatial information but also temporal information. To better exploit such spatial-temporal information attached to videos, we propose a novel video captioning framework with Temporal Graph Network (TGN) and Region Graph Network (RGN). TGN mainly focuses on utilizing the sequential information of frames that most of existing methods ignore. RGN is designed to explore the relationships among salient objects. Different from previous work, we introduce Graph Convolution Network (GCN) to encode frames with their sequential information and build a region graph for utilizing object information. We also particularly adopt a stack GRU decoder with a coarse-to-fine structure for caption generation. Very promising experimental results on two benchmark datasets (MSVD and MSR-VTT) show the effectiveness of our model.
Xinlong Xiao, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003, Weiguo Fan
ICME3
2020 Learning Error-Driven Curriculum for Crowd Counting
abstract
Density regression has been widely employed in crowd counting. However, the frequency imbalance of pixel values in the density map is still an obstacle to improve the performance. In this paper, we propose a novel learning strategy for learning error-driven curriculum, which uses an additional network to supervise the training of the main network. A tutoring network called TutorNet is proposed to repetitively indicate the critical errors of the main network. TutorNet generates pixel-level weights to formulate the curriculum for the main network during training, so that the main network will assign a higher weight to those hard examples than easy examples. Furthermore, we scale the density map by a factor to enlarge the distance among inter-examples, which is well known to improve the performance. Extensive experiments on two challenging benchmark datasets show that our method has achieved state-of-the-art performance.
Wenxi Li, Zhuoqun Cao, Songjian Chen, Rui Feng 0001
ICPR5
2020 A Dynamic-Attention on Crowd Region with Physical Optical Flow Features for Crowd Counting
abstract
Crowd counting is widely used in various video surveillance applications. However, most of the existing approaches treat videos as a single frame, which increase redundant information and have low efficiency, due to ignoring the context history information of neighboring frames. In this paper, we propose a novel two-stream dynamic-attention network (DANet) to associate the temporal and spatial information. Specifically, the DANet includes two stages, one of which is to generate the region-attention map and the second is to refine the high-quality density map. In each stage, we develop a hierarchical fusion strategy to guide spatial attention, which can iteratively refine the region of crowds. Besides, the dynamic-attention module guided by the physical optical flow can be dynamically integrated into any network module to optimize the generation of features for improving the effect. Therefore, it can be plugged into many computer vision architectures. Finally, experimental results on three challenging benchmark datasets show that DANet outperforms most of the previous methods. Incorporating such dynamic-attention into a framework could boost the performance of end-to-end CNN-based methods.
Wenxi Li, Songjian Chen, Rui Feng 0001
IJCNN4
2020 Fuzzy Graph Neural Network for Few-Shot Learning
abstract
Recent works have shown that graph neural net-works (GNNs) can substantially improve the performance of few-shot learning benefitting from their natural ability to learn inter-class uniqueness and intra-class commonality. However, previous GNN methods have not achieved satisfactory performance due to the absence of a strong relational inductive bias which determines how entities interact and are isolated. In this paper, inspired by the fuzzy theory, we propose a novel meta-learning method called Fuzzy GNN (FGNN), which obtains superior relational inductive biases in each episode, for few-shot learning. Specifically, we employ an edge-focused GNN to perform the edge prediction by iteratively updating the edge-labels. According to the output of edge prediction, we design a fuzzy membership function to achieve more exact relationship representations for node classification. The parameters of the FGNN are learned by episodic training with mixed loss including node-label and edge-label. Extensive experimental evaluation clearly demonstrates the effectiveness of FGNN. The results show that our method achieves state-of-the-art performance and a significant improvement over other GNN methods on two few-shot learning benchmarks.
Junlin Hou, Rui Feng 0001
IJCNN3
2020 Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation Learning
abstract
When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be utilized as free supervised information to train a neural network by solving the pretext task of audio-visual synchronization. In this paper, we propose a novel self-supervised framework with co-attention mechanism to learn generic cross-modal representations from unlabelled videos in the wild, and further benefit downstream tasks. Specifically, we explore three different co-attention modules to focus on discriminative visual regions correlated to the sounds and introduce the interactions between them. Experiments show that our model achieves state-of-the-art performance on the pretext task while having fewer parameters compared with existing methods. To further evaluate the generalizability and transferability of our approach, we apply the pre-trained model on two downstream tasks, i.e., sound source localization and action recognition. Extensive experiments demonstrate that our model provides competitive results with other self-supervised methods, and also indicate that our approach can tackle the challenging scenes which contain multiple sound sources.
Ying Cheng 0005, Zhihao Pan, Rui Feng 0001, Yuejie Zhang
ACM Multimedia4
2020 Automated diabetic retinopathy grading and lesion detection based on the modified R-FCN object-detection algorithm
abstract
In this work, we develop a computer‐aided retinal image screening system that can perform automated diabetic retinopathy (DR) grading and DR lesion detection in retinal fundus images. We propose a modified object‐detection method for this task via a region‐based fully convolutional network (R‐FCN). A feature pyramid network and a modified region proposal network are applied to enhance the detection of small objects. The DR‐grading model based on the modified R‐FCN is evaluated on the Messidor data set and images provided by the Shanghai Eye Hospital. High sensitivity of 99.39% and specificity of 99.93% are obtained on the hospital data. Moreover, high sensitivity of 92.59% and specificity of 96.20% are obtained on the Messidor data set. The modified R‐FCN lesion‐detection model is validated on the hospital data set and achieves a 92.15% mean average precision. The proposed R‐FCN can efficiently accomplish DR grading and lesion detection with high accuracy.
Jianxu Luo, Rui Feng 0001, Lina Lu, Haidong Zou
IET Comput. Vis.4
2020 Deep cascaded cross-modal correlation learning for fine-grained sketch-based image retrieval
Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.4
2020 Deep reinforcement hashing with redundancy elimination for effective image retrieval
Juexu Yang, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Weiguo Fan
Pattern Recognit.3
2017 Multi-task Deep Neural Network for Joint Face Recognition and Facial Attribute Prediction
abstract
Deep neural networks have significantly improved the performance of face recognition and facial attribute prediction, which however are still very challenging on the million scale dataset, i.e. MegaFace. In this paper, we for the first time, advocate a multi-task deep neural network for jointly learning face recognition and facial attribute prediction tasks. Extensive experimental evaluation clearly demonstrates the effectiveness of our architecture. Remarkably, on the largest face recognition benchmark -- MegaFace dataset, our networks can achieve the Rank-1 identication accuracy of 77.74% and face verication accuracy 79.24% TAR at 10-6 FAR, which are the best performance on the small protocol among all the publicly released methods.
Zhanxiong Wang, Keke He, Yanwei Fu 0001, Rui Feng 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001
ICMR4
2017 Adaptively Weighted Multi-task Deep Network for Person Attribute Classification
abstract
Multi-task learning aims to boost the performance of multiple prediction tasks by appropriately sharing relevant information among them. However, it always suffers from the negative transfer problem. And due to the diverse learning difficulties and convergence rates of different tasks, jointly optimizing multiple tasks is very challenging. To solve these problems, we present a weighted multi-task deep convolutional neural network for person attribute analysis. A novel validation loss trend algorithm is, for the first time proposed to dynamically and adaptively update the weight for learning each task in the training process. Extensive experiments on CelebA, Market-1501 attribute and Duke attribute datasets clearly show that state-of-the-art performance is obtained; and this validates the effectiveness of our proposed framework.
Keke He, Zhanxiong Wang, Yanwei Fu 0001, Rui Feng 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001
ACM Multimedia4
2016 Flexible multi-task learning with latent task grouping
Jian Pu, Yu-Gang Jiang 0001, Rui Feng 0001, Xiangyang Xue 0001
Neurocomputing4
2015 Cross-Modal Image-Tag Relevance Learning for Social Images
abstract
A new algorithm is developed in this paper to support more effective cross-modal image-tag relevance learning for large-scale social images, which integrates the multimodal feature representation, multimodal relevance measurement, and cross- modal relevance fusion. The main contribution of our work is that we provide a more reasonable base to learn cross-modal relevance among social images, which can be acquired from integrating multimodal image and tag relevance with multiple features in different modalities. Very positive results were obtained in our experiments using a large quantity of public social image data.
Zhengxiang Cai, Rui Feng 0001, Cheng Jin 0001, Yuejie Zhang, Tao Zhang 0022
ACM Multimedia3
2013 Understanding and Predicting Interestingness of Videos
abstract
The amount of videos available on the Web is growing explosively. While some videos are very interesting and receive high rating from viewers, many of them are less interesting or even boring. This paper conducts a pilot study on the understanding of human perception of video interestingness, and demonstrates a simple computational method to identify more interesting videos. To this end we first construct two datasets of Flickr and YouTube videos respectively. Human judgements of interestingness are collected and used as the ground-truth for training computational models. We evaluate several off-the-shelf visual and audio features that are potentially useful for predicting interestingness on both datasets. Results indicate that audio and visual features are equally important and the combination of both modalities shows very promising results.
Yu-Gang Jiang 0001, Rui Feng 0001, Xiangyang Xue 0001, Yingbin Zheng, Hanfang Yang
AAAI3
2013 Beauty is here: evaluating aesthetics in videos using multimodal features and free training data
abstract
The aesthetics of videos can be used as a useful clue to improve user satisfaction in many applications such as search and recommendation. In this paper, we demonstrate a computational approach to automatically evaluate the aesthetics of videos, with particular emphasis on identifying beautiful scenes. Using a standard classification pipeline, we analyze the effectiveness of a comprehensive set of features, ranging from low-level visual features, mid-level semantic attributes, to style descriptors. In addition, since there is limited public training data with manual labels of video aesthetics, we explore freely available resources with a simple assumption that people tend to share more aesthetically appealing works than unappealing ones. Specifically, we use images from DPChallenge and videos from Flickr as positive training data and the Dutch documentary videos as negative data, where the latter contain mostly old materials of low visual quality. Our extensive evaluations show that combining multiple features is helpful, and very promising results can be obtained using the noisy but annotation-free training data. On the NHK Multimedia Challenge dataset, we attain a Spearman's rank correlation coefficient of 0.41.
Qi Dai 0001, Rui Feng 0001, Yu-Gang Jiang 0001
ACM Multimedia3
2009 Web image retrieval reranking with multi-view clustering
abstract
General image retrieval is often carried out by a text-based search engine, such as Google Image Search. In this case, natural language queries are used as input to the search engine. Usually, the user queries are quite ambiguous and the returned results are not well-organized as the ranking often done by the popularity of an image. In order to address these problems, we propose to use both textual and visual contents of retrieved images to reRank web retrieved results. In particular, a machine learning technique, a multi-view clustering algorithm is proposed to reorganize the original results provided by the text-based search engine. Preliminary results validate the effectiveness of the proposed framework.
Mingmin Chi, Peiwu Zhang, Yingbin Zhao, Rui Feng 0001, Xiangyang Xue 0001
WWW4