Yanming Guo

dblp:46/3423 · DBLP profile ↗
← Back
44ranked-venue papers
5as first author
29since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021
YearPublicationVenuePosition
2027 JAIL: Adaptive multi-turn jailbreak attacks reveal limitations of LLM safety alignment
Yunhao Feng, Mingrui Lao, Yishan Li, Yuxiang Xie, Yanming Guo
Expert Syst. Appl.7
2026 Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative Learning
abstract
In real-world applications, video action recognition models must continuously learn new action categories while retaining previously acquired knowledge. However, most existing approaches rely on storing historical data for replay, which introduces storage burdens and raises data privacy concerns. To address these challenges, we investigate the problem of Exemplar-Free Continual Video Action Recognition (EF-CVAR) and propose a novel framework named Slow-Fast Collaborative Learning (SFCL). SFCL integrates two complementary learning paradigms: a slow branch based on gradient-driven deep learning, which provides strong adaptability to new tasks, and a fast branch based on analytic learning (e.g., Recursive Least Squares), which efficiently preserves old knowledge without requiring access to past samples. To enable effective collaboration between the two branches, we design the Slow-Fast Dynamic Re-parameterization (SFDR) mechanism for adaptive fusion, and the Knowledge Reflection Mechanism (KRM), which mitigates forgetting and task-recency bias via pseudo-feature generation and dual-level knowledge distillation. Extensive experiments on UCF101, HMDB51, and Something-Something V2 demonstrate that SFCL achieves superior performance compared to existing replay-based methods, despite being exemplar-free. Notably, in long-duration continual learning scenarios, SFCL exhibits remarkable robustness, achieving up to a 30.39\% improvement in accuracy over baselines while maintaining a low forgetting rate, highlighting its scalability and effectiveness in real-world video recognition tasks.
Xueyi Zhang 0001, Siqi Cai 0002, Mingrui Lao, Yanming Guo, Huiping Zhuang
AAAI7
2026 Overcoming semantic manifold deviation for robust multimodal violence detection with incomplete modality
Jianan Zhu, Yanming Guo, Lai Kang, Yirun Ruan, Mingrui Lao
Expert Syst. Appl.2
2026 GRAIN: Gravity-resistance adaptive framework for identifying influential nodes using multi-order structural diversity
Yirun Ruan, Xinghua Qin, Sizheng Liu, Jun Tang 0001, Yanming Guo
Inf. Process. Manag.6
2026 Unraveling domain styles for enhanced cross-domain generalization
Zhonghua Yao, Juncheng Lian, Yanming Guo
Knowl. Based Syst.4
2026 Towards few-shot deepfake detection with an enhanced CLIP model
Yumin Yang, Xueyi Zhang 0001, Yanming Guo
Neural Networks4
2026 VMAD: Visual-Enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
abstract
Zero-shot anomaly detection (ZSAD) enables the inspection of unseen objects by bridging textual prompts and visual features, showing great potential in flexible manufacturing. While existing ZSAD methods rely on predefined prompts and struggle with unseen defects, Multimodal Large Language Models (MLLMs) offer promising solutions through their generative and interpretative capabilities. However, adapting MLLMs to Industrial Anomaly Detection (IAD) remains challenging due to fine-grained anomaly patterns and subtle visual distinctions. We propose VMAD (Visual-enhanced MLLM Anomaly Detection), a framework that enriches MLLM with visual IAD knowledge through two key components: a Defect-Sensitive Structure Learning scheme that transfers patch-similarities for improved discrimination, and a Locality-enhanced Token Compression that leverages multi-level local features for fine-grained detection. We also introduce RIAD, a comprehensive IAD dataset with detailed anomaly annotations. Extensive experiments on MVTec-AD, Visa, WFDD, and RIAD demonstrate VMAD’s superior performance. The dataset and code will be publicly available at https://github.com/denghuilin-cyber/VMAD.
Huilin Deng, Hongchen Luo, Wei Zhai, Yanming Guo, Yang Cao 0010, Yu Kang 0001
IEEE Trans Autom. Sci. Eng.4
2026 Corrections to "VMAD: Visual-Enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection"
abstract
In the above article [1], an earlier draft of Fig. 6 was inadvertently included. The correct Fig. 6 is presented on the next page.Fig. 6.Zero-shot anomaly segmentation on MVTec-AD, WFDD, and ViSA datasets.
Huilin Deng, Hongchen Luo, Wei Zhai, Yanming Guo, Yang Cao 0010, Yu Kang 0001
IEEE Trans Autom. Sci. Eng.4
2025 Generalization-Preserved Learning: Closing the Backdoor to Catastrophic Forgetting in Continual Deepfake Detection
Xueyi Zhang 0001, Peiyin Zhu, Zhiyuan Yan 0002, Jikang Cheng, Mingrui Lao, Siqi Cai 0002, Yanming Guo
ICCV8
2025 Boosting Adversarial Robustness Through Structure-Guided Adversarial Distillation
Yanming Guo, Chengsi Du, Dengjin Li, Mingrui Lao
ICIC (2)2
2025 Exploiting Event Temporal Dynamics and Sparsity Characteristics for RGB-Event Fusion Semantic Segmentation
abstract
Frame-based semantic segmentation faces information loss due to the limited dynamic range of conventional cameras. Event cameras, with their high dynamic range and temporal resolution, offer a promising solution. Unlike RGB images, event cameras produce sparse, asynchronous event streams, prompting a reconsideration of the event generation mechanism and a comprehensive exploration of their characteristics. Previous event representation methods have been limited to fixed time windows, neglecting the rich temporal dynamics inherent in event streams. Furthermore, the inherent noise and calibration deficiencies in event data present significant challenges for RGB-event fusion. We propose an Event-driven Fusion Network (EFNet) to improve semantic segmentation by leveraging event camera characteristics. To address the first challenge, we introduce a Dual-Temporal Event Integration (DT-EI) module, which leverages temporal dynamics to generate and integrate dual-temporal event representations, capturing objects with varying motion speeds. For the second challenge, we propose an Event-Count Guided Recalibration and Fusion (EGRF) module, which utilizes event sparsity to guide feature recalibration through Event-Count Attention Maps, combined with a bidirectional cross-attention mechanism for adaptive fusion of event and image features. Experimental results demonstrate that EFNet outperforms state-of-the-art methods in event-based semantic segmentation.
Yingmei Wei, Yanming Guo, Jiangming Chen
ICMR3
2025 Choose Your Expert: Uncertainty-Guided Expert Selection for Continual Deepfake Detection
abstract
The rapid evolution of deepfake techniques presents dual challenges for detection models: adapting to continuously shifting attack distributions while retaining previously learned knowledge. Although recent continual deepfake detection methods have made progress, they often rely on replay-based training, which limits scalability and deployment. Meanwhile, the task structure of deepfake detection offers a unique opportunity that remains under-explored: it is inherently a binary classification problem with a fixed label space, where the main difficulty lies in distributional drift rather than class expansion. This insight enables the modeling of each incremental distribution shift as a dedicated expert, focusing on specific forgery patterns. To this end, we propose a novel analytically driven, replay-free continual detection framework that eliminates the need for iterative gradient updates. In this framework, task-specific experts are constructed via closed-form ridge regression, requiring only a single forward pass and ensuring non-interference with previous tasks. To enhance the model's capacity for fine-grained forgery recognition, we introduce a lightweight Forgery-Aware Residual Enhancer (FARE). At inference, an Uncertainty-Guided Expert Selection module (UGES) dynamically routes each sample to the most confident expert, which does not require prior knowledge of the attack type. The proposed framework achieves a favorable trade-off between efficiency, privacy, and generalization. It achieves state-of-the-art performance across four benchmark datasets, with an average accuracy of 91.82% and only 1.78% forgetting. Notably, it improves cross-forgery generalization by 9.28% on unseen forgery types, demonstrating strong generalization.
Xueyi Zhang 0001, Peiyin Zhu, Jinping Sui, Xiaoda Yang, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Jun Tang 0001
ACM Multimedia8
2025 TrustCLIP: Learning from Noisy Labels via Semantic Label Verification and Trust-aligned Gradient Projection
abstract
Prompt learning has emerged as an efficient adaptation paradigm for vision-language models (VLMs), yet it remains highly vulnerable to label noise, which limits its real-world applicability. We propose TrustCLIP, a noise-robust prompt tuning framework that leverages the inherent semantic structure of CLIP through two key components: Semantic Label Verification (SLV) and Trust-aligned Gradient Projection (TGP). SLV defines a semantic trust boundary based on CLIP's zero-shot predictions to identify reliable samples for standard supervised training. For uncertain samples, TGP projects their gradients into a trust-aligned subspace constructed from the gradients of clean samples, thereby preserving semantically aligned learning signals while suppressing noise-induced optimization drift. Unlike prior approaches, TrustCLIP doesn't require additional parameters, loss reweighting, or uncertainty estimation. Extensive experiments on 7 benchmark datasets with both synthetic and real-world noisy labels demonstrate that TrustCLIP consistently outperforms state-of-the-art methods in terms of both robustness and transferability.
Xueyi Zhang 0001, Peiyin Zhu, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Haizhou Li 0001
ACM Multimedia7
2025 Learning from Peers: Collaborative Ensemble Adversarial Training
Dengjin Li, Yanming Guo, Yuxiang Xie, Jiangming Chen, Mingrui Lao
PRCV (1)2
2025 Boosting Discriminability for Robust Multimodal Entity Linking with Visual Modality Missing
abstract
Multimodal Entity Linking (MEL) aims to retrieve ambiguous mentions within multimodal contexts to the referent entities in a multimodal knowledge base, typically based on the assumption of modality completeness. However, when deployed in open-world applications, MEL systems may encounter uncertainly missing of visual modalities from user-proposed mentions. In this paper, we propose a novel setting dubbed MEL-MM to simulate the practical challenge, and reveal that the semantic discriminability is a crucial factor to enhance the anti-missingness resilience. To this end, we introduce an innovative yet efficient approach termed Cross-View Introspective Ranking Distillation (CVIRD), which seeks to sufficiently align the linking similarities between teacher and student models trained from modality-complete and incomplete data. To be specific, as the first concept in CVIRD, Missing-Aware Ranking Distillation (MARD) focuses on modeling the discriminability by formulating the similarity rankings between mention and entities in a missing-sensitive and differentiable manner. Moreover, the second concept of Cross-View Distillation with Introspection (CVDI) aims to improve discriminability extraction in MARD through multi-level distillation, considering both cross-view retrieval and self-consistency. Experiments verify the effectiveness and model-agnostic ability of our method, which achieves superior performance in contrast to competitive missingness-resilient strategies.
Mingrui Lao, Yanming Guo, Xueyi Zhang 0001, Siqi Cai 0002, Zhaoyun Ding, Haizhou Li 0001
SIGIR3
2025 GLC: A dual-perspective approach for identifying influential nodes in complex networks
abstract
Identifying influential spreaders is crucial for understanding the dynamics of information diffusion within complex networks. Several centrality methods have been proposed to address this, but these studies often concentrate on only one aspect. To solve this problem, we introduce a dual-perspective approach which considers both global and local perspectives for identifying influential nodes in complex networks. From a global perspective, if a node has the capability to efficiently transmit information to various clusters within a network, then the information originating from that node will quickly spread across a large area. From a local perspective, when a node has a greater number of neighbors—especially those that are significant within the network—the information emanating from that node is less likely to be confined to a localized region. Based on this understanding, we first design a novel clustering method to detect groups in which the connections among nodes are denser than those with the rest of the network. The most influential nodes in each group are identified as global critical nodes. Subsequently, the local influence of a node is defined by the number and significance of its neighboring nodes. Ultimately, nodes are ranked according to their local influence, their proximity to the global critical nodes using the shortest paths, and the importance of these global critical nodes. To evaluate the performance of the proposed method, the susceptible-infected-removed (SIR) diffusion model is used. Results of the investigation on real networks and realistic synthetic benchmarks show that the proposed method can identify nodes with high influence better than other centrality methods.
Yirun Ruan, Sizheng Liu, Jun Tang 0001, Yanming Guo
Expert Syst. Appl.4
2025 FCAT: Federated causal adversarial training
Yunhao Feng, Yanming Guo, Mingrui Lao, Yishan Li, Yuxiang Xie
Knowl. Based Syst.2
2025 Fed-GCC: Global classifier consensus for conventional/task-free federated class-incremental learning
Dianqi Liu, Yanming Guo, Jun Tang 0001, Yirun Ruan
Knowl. Based Syst.3
2025 Few-shot event-based action recognition
abstract
Despite the evident superiority of event cameras in practical vision applications (e.g., action recognition), owing to their distinctive sensing mechanism, existing event-based action recognition methods rely heavily on large-scale training data. However, the expensive cost of camera deployment and the requirement of data privacy protection make it challenging to collect substantial data in real-world scenarios. To address this limitation, we explore a novel yet practical task, Few-Shot Event-Based Action Recognition (FSEAR), which aims at leveraging a minimal number of intractable event action data for model training and accurately classifying unlabeled data into a specific category. Accordingly, we design a new framework for FSEAR, including a Noise-Aware Event Encoder (NAE) and a Distilled Prototypical Distance Fusion (DPDF). The former efficiently filters noise within the spatiotemporal domain while retaining vital information related to action timing. The latter conducts multi-scale measurements across geometric, directional, and distributional dimensions. These two modules benefit mutually and thus effectively exploit the potential characteristics of event data. Extensive experiments on four distinct event action recognition datasets have demonstrated the significant advantages of our model over other few-shot learning methods. Our code and models will be publicly released.
Zanxi Ruan, Nan Pu, Jiangming Chen, Songqun Gao, Yanming Guo, Qiuyu Kong, Yuxiang Xie, Yingmei Wei
Neural Networks5
2025 Polarization Calibration Verification of the Directional Polarimetric Camera on the Terrestrial Ecosystem Carbon Inventory Satellite
abstract
Space-borne multi-angle polarization remote sensing is considered to be one of the most important tools to obtain global aerosol parameters in assessment of climate change. Accurate calibration is a prerequisite for quantitative polarization remote sensing. However, most research focuses on the polarization calibration methods and the monitoring of the calibration coefficient stability. Few studies investigate the polarization calibration verification of the subsequent new satellite sensors. The Directional Polarimetric Camera (DPC) onboard the Chinese Terrestrial Ecosystem Carbon Inventory Satellite (abbreviated as TECIS, with the Chinese name ”Gou Mang”) is a brand-new polarization sensor. To evaluate the polarization calibration of this sensor, we propose a verification scheme for polarization calibration, which can verify both the degree of polarization (DoP) and polarized reflectance. The accuracy of DoP is verified by using sunglint on the ocean. When detecting the sunglint, constraints such as observational geometry, cloud identification, and wind speed are introduced, and the polarized reflectance is atmospherically corrected according to the marine aerosol model and aerosol optical depth from MODIS. The accuracy of polarized reflectance at the Top of Atmosphere (TOA) is verified based on AERONET inversion products and the Bidirectional Polarization Distribution Functions (BPDF). Experiments show that the accuracies of DoP of the three polarization channels (490, 670, 865 nm) of DPC are 5.57%, 2.07%, and 1.97% respectively, and the average accuracy of the multi-angle polarized reflectance at the TOA of the 865nm channel is 0.209%.
Donghai Xie, Liuyan Guo, Yu Wu 0002, Tongyuan Zou, Hailiang Gao, Yutang Yu, Yinzhen Wang, Chang Yi, Shi Jin 0002, Yanming Guo
IEEE Trans. Geosci. Remote. Sens.16
2025 COLA: Context-Aware Language-Driven Test-Time Adaptation
abstract
Test-time adaptation (TTA) has gained increasing popularity due to its efficacy in addressing "distribution shift" issue while simultaneously protecting data privacy. However, most prior methods assume that a paired source domain model and target domain sharing the same label space coexist, heavily limiting their applicability. In this paper, we investigate a more general source model capable of adaptation to multiple target domains without needing shared labels. This is achieved by using a pre-trained vision-language model (VLM), e.g., CLIP, that can recognize images through matching with class descriptions. While the zero-shot performance of VLMs is impressive, they struggle to effectively capture the distinctive attributes of a target domain. To that end, we propose a novel method - Context-aware Language-driven TTA (COLA). The proposed method incorporates a lightweight context-aware module that consists of three key components: a task-aware adapter, a context-aware unit, and a residual connection unit for exploring task-specific knowledge, domain-specific knowledge from the VLM and prior knowledge of the VLM, respectively. It is worth noting that the context-aware module can be seamlessly integrated into a frozen VLM, ensuring both minimal effort and parameter efficiency. Additionally, we introduce a Class-Balanced Pseudo-labeling (CBPL) strategy to mitigate the adverse effects caused by class imbalance. We demonstrate the effectiveness of our method not only in TTA scenarios but also in class generalisation tasks. The source code is available at https://github.com/NUDT-Bai-Group/COLA-TTA.
Aiming Zhang, Liang Bai 0003, Jun Tang 0001, Yanming Guo, Yirun Ruan, Yun Zhou 0001, Zhihe Lu
IEEE Trans. Image Process.5
2024 Boosting Adversarial Robustness Distillation Via Hybrid Decomposed Knowledge
abstract
Adversarial Robust Distillation (ARD) has emerged as a potent defense mechanism tailored to small models against adversarial threats. However, mainstream ARD methods typically exploit teachers’ response as the transferred knowledge, while neglecting the analysis of involved target-related knowledge to mitigate adversarial attacks. Furthermore, these methods primarily focus on logits-level distillation, which overlook the features-level knowledge in teacher models. In this paper, we introduce a novel Hybrid Decomposed Distillation (HDD) approach, which attempts to identify the vital knowledge against adversarial threats through dual-level distillation. Specifically, we first seek to separate the predictions of teacher model into target-related and target-unrelated knowledge for flexible yet efficient logits-level distillation. Besides, to further boost the distillation efficacy, HDD leverages the channel correlations to decompose intermediate features into highly and less relevant components. Extensive experiments on two benchmarks demonstrate that our HDD achieves superior performance in both clean accuracy and robustness, in contrast to current state-of-the-art methods.
Mingrui Lao, Yanming Guo
ICASSP3
2024 Language Without Borders: A Dataset and Benchmark for Code-Switching Lip Reading
abstract
Lip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy environments. With the increasing interpersonal communications in social media owing to globalization, the existing monolingual datasets for lip reading may not be sufficient to meet the exponential proliferation of bilingual and even multilingual users. However, to our best knowledge, research on code-switching is only explored in speech recognition, while the attempts in lip reading are seriously neglected. To bridge this gap, we have collected a bilingual code-switching lip reading benchmark composed of Chinese and English, dubbed CSLR. As the pioneering work, we recruited 62 speakers with proficient foundations in bothspoken Chinese and English to express sentences containing both involved languages. Through rigorous criteria in data selection, CSLR benchmark has accumulated 85,560 video samples with a resolution of 1080x1920, totaling over 71.3 hours of high-quality code-switching lip movement data. To systematically evaluate the technical challenges in CSLR, we implement commonly-used lip reading backbones, as well as competitive solutions in code-switching speech for benchmark testing. Experiments show CSLR to be a challenging and under-explored lip reading task. We hope our proposed benchmark will extend the applicability of code-switching lip reading, and further contribute to the communities of cross-lingual communication and collaboration. Our dataset and benchmark are accessible at https://github.com/cslr-lipreading/CSLR.
Xueyi Zhang 0001, Mingrui Lao, Jun Tang 0001, Yanming Guo, Siqi Cai 0002, Xianghu Yue, Haizhou Li 0001
NeurIPS5
2022 VQA-BC: Robust Visual Question Answering Via Bidirectional Chaining
abstract
Current VQA models are suffering from the problem of overdependence on language bias, which severely reduces their robustness in real-world scenarios. In this paper, we analyze VQA models from the view of forward/backward chaining in the inference engine, and propose to enhance their robustness via a novel Bidirectional Chaining (VQA-BC) framework. Specifically, we introduce a backward chaining with hardnegative contrastive learning to reason from the consequence (answers) to generate crucial known facts (question-related visual region features). Furthermore, to alleviate the overconfident problem in answer prediction (forward chaining), we present a novel introspective regularization to connect forward and backward chaining with label smoothing. Extensive experiments verify that VQA-BC not only effectively overcomes language bias on out-of-distribution dataset, but also alleviates the over-correct problem caused by ensemble-based method on in-distribution dataset. Compared with competitive debiasing strategies, our method achieves state-of-the-art performance to reduce language bias on VQA-CP v2 dataset.
Mingrui Lao, Yanming Guo, Wei Chen 0072, Nan Pu, Michael S. Lew
ICASSP2
2022 Chinese named entity recognition: The state of the art
abstract
Named Entity Recognition(NER), one of the most fundamental problems in natural language processing, seeks to identify the boundaries and types of entities with specific meanings in natural language text. As an important international language, Chinese has uniqueness in many aspects, and Chinese NER (CNER) is receiving increasing attention. In this paper, we give a comprehensive survey of recent advances in CNER. We first introduce some preliminary knowledge, including the common datasets, tag schemes, evaluation metrics and difficulties of CNER. Then, we separately describe recent advances in traditional research and deep learning research of CNER, in which the CNER with deep learning is our focus. We summarize related works in a basic three-layer architecture, including character representation, context encoder, and context encoder and tag decoder. Meanwhile, the attention mechanism and adversarial-transfer learning methods based on this architecture are introduced. Finally, we present the future research trends and challenges of CNER.
Pan Liu 0006, Yanming Guo, Fenglei Wang
Neurocomputing2
2021 A Language Prior Based Focal Loss for Visual Question Answering
abstract
According to current research, one of the major challenges in Visual Question Answering (VQA) models is the overdependence on language priors (and neglect of the visual modality). VQA models tend to predict answers only based on superficial correlations between the first few words in question and frequency of related answer candidates. To address this issue, we propose a novel Language Prior based Focal Loss (LP-Focal Loss) by rescaling the standard cross entropy loss. Specifically, we employ a question-only branch to capture the language biases for each answer candidate based on the corresponding question input. Then, the LP-Focal Loss dynamically assigns lower weights to biased answers when computing the training loss, thereby reducing the contribution of more-biased instances in the train split. Extensive experiments show that the LP-Focal Loss can be generally applied to common baseline VQA models, and achieves significantly better performance on the VQA-CP v2 dataset, with an overall 18% accuracy boost over benchmark models.
Mingrui Lao, Yanming Guo, Yu Liu 0012, Michael S. Lew
ICME2
2021 DCA-CLA: A scRNA-seq Classification Framework based on Deep Count Autoencoder
abstract
Identifying cell types is crucial for single-cell RNA sequencing (scRNA-seq) analysis and can be potentially utilized to understand high-level biological processes. Supervised models based on neural networks have recently been successfully applied in the scRNA-seq cell type classification problem and achieved promising results. While most existing works directly use the raw or transformed data, we argue that the original data are too sparse and high-dimensional, and extracting their effective low-dimensional features can better train downstream classifiers, thereby improving the cell type classification performance. In this paper, we propose a novel framework, named Deep Count Autoencoder-based Classifier (DCA-CLA), to leverage the discriminative low-dimensional features for classification. Specifically, DCA-CLA first denoises the original count matrix and extracts the data features from the hidden layer using a deep count autoencoder module, then it feeds these bottleneck features into the classifier network to train the learnable parameters and test the performance. Experimental results on eight separate datasets and four pairs of datasets demonstrate that the proposed DCA-CLA framework achieves competitive performance over the state-of-the-art frameworks.
Yanming Guo, Songyang Lao, Jinlin Guo
IJCNN2
2021 From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question Answering
abstract
Most Visual Question Answering (VQA) models are faced with language bias when learning to answer a given question, thereby failing to understand multimodal knowledge simultaneously. Based on the fact that VQA samples with different levels of language bias contribute differently for answer prediction, in this paper, we overcome the language prior problem by proposing a novel Language Bias driven Curriculum Learning (LBCL) approach, which employs an easy-to-hard learning strategy with a novel difficulty metric Visual Sensitive Coefficient (VSC). Specifically, in the initial training stage, the VQA model mainly learns the superficial textual correlations between questions and answers (easy concept) from more-biased examples, and then progressively focuses on learning the multimodal reasoning (hard concept) from less-biased examples in the following stages. The curriculum selection of examples on different stages is according to our proposed difficulty metric VSC, which is to evaluate the difficulty driven by the language bias of each VQA sample. Furthermore, to avoid the catastrophic forgetting of the learned concept during the multi-stage learning procedure, we propose to integrate knowledge distillation into the curriculum learning framework. Extensive experiments show that our LBCL can be generally applied to common VQA baseline models, and achieves remarkably better performance on the VQA-CP v1 and v2 datasets, with an overall 20% accuracy boost over baseline models.
Mingrui Lao, Yanming Guo, Yu Liu 0012, Wei Chen 0072, Nan Pu, Michael S. Lew
ACM Multimedia2
2021 Multi-stage hybrid embedding fusion network for visual question answering
Mingrui Lao, Yanming Guo, Nan Pu, Wei Chen 0072, Yu Liu 0012, Michael S. Lew
Neurocomputing2
2020 Development and Application of an Intensive Care Medical Data Set for Deep Learning
abstract
A large number of patient healthcare data have been collected in the process of diagnosis and treatment of intensive care medicine, which provides major benefits for patient safety and quality. Unfortunately, the application of medical data is greatly limited. Key barriers to the use of the data include difficulties in data extraction and cleaning, and the construction of high-quality data sets promotes the research of medical big data analysis. In China, there is few intensive care data set built by clinicians has been used for clinical outcome prediction. This study developed and evaluated an Intensive Care Medical (ICM) data set for critically care patients that can be used for deep learning. The ICM data set contained four types of data collected routinely in Chinese hospitals, including all-cause characteristics of administrative information, vital signs, laboratory tests, and intravenous medication records. A total of 17,291 ICU admissions involving 12,815 patients aged 14 years and older were extracted from the data set. Deep learning model achieved high accuracy for tasks in hospital mortality predicting (AUROC[area under the receiver operator curve] reach 0.8941). We believe that the ICM data set can be used to create accurate predictions for a variety of clinical scenarios.
Shangping Zhao, Pan Liu 0006, Guanxiu Tang, Yanming Guo
IEEE BigData4
2020 PFNet: a novel part fusion network for fine-grained visual categorization
Jingyun Liang, Jinlin Guo, Yanming Guo, Songyang Lao
Multim. Tools Appl.3
2019 CycleMatch: A cycle-consistent embedding network for image-text matching
Yu Liu 0012, Yanming Guo, Li Liu 0002, Erwin M. Bakker, Michael S. Lew
Pattern Recognit.2
2018 A Dual Prediction Network for Image Captioning
abstract
General captioning practice involves a single forward prediction, with the aim of predicting the word in the next timestep given the word in the current timestep. In this paper, we present a novel captioning framework, namely Dual Prediction Network (DPN), which is end-to-end trainable and addresses the captioning problem with dual predictions. Specifically, the dual predictions consist of a forward prediction to generate the next word from the current input word, as well as a backward prediction to reconstruct the input word using the predicted word. DPN has two appealing properties: 1) By introducing an extra supervision signal on the prediction, DPN can better capture the interplay between the input and the target; 2) Utilizing the reconstructed input, DPN can make another new prediction. During the test phase, we average both predictions to formulate the final target sentence. Experimental results on the MS COCO dataset demonstrate that, benefiting from the reconstruction step, both generated predictions in DPN outperform the predictions of methods based on the general captioning practice (single forward prediction), and averaging them can bring a further accuracy boost. Overall, DPN achieves competitive results with state-of-the-art approaches, across multiple evaluation metrics.
Yanming Guo, Yu Liu 0012, Maaike de Boer, Li Liu 0002, Michael S. Lew
ICME1
2018 An Extensive Study of Cycle-Consistent Generative Networks for Image-to-Image Translation
abstract
Image-to-image translation between different domains has been an important research direction, with the aim of arbitrarily manipulating the source image content to become similar to a target image. Recently, cycle-consistent generative network (CycleGAN) has become a fundamental approach for general-purpose image-to-image translation, while almost no work has examined what factors may influence its performance. To provide more insights, we propose two new models roughly based on CycleGAN, namely Long CycleGAN and Nest CycleGAN. First, Long CycleGAN cascades several generators to perform the domain translation in a long cycle. It shows the benefit of stacking more generators on the generation quality. In addition to the long cycle, Nest CycleGAN develops new inner cycles to bridge intermediate generators directly, which can help constrain the unsupervised mappings. In the experiments, we conduct qualitative and quantitative comparisons for tasks including photo↔label, photo↔sketch, and photo colorization. The quantitative and qualitative results demonstrate the effectiveness of our two proposed models.
Yu Liu 0012, Yanming Guo, Wei Chen 0072, Michael S. Lew
ICPR2
2018 CNN-RNN: a large-scale hierarchical image classification framework
abstract
Objects are often organized in a semantic hierarchy of categories, where fine-level categories are grouped into coarse-level categories according to their semantic relations. While previous works usually only classify objects into the leaf categories, we argue that generating hierarchical labels can actually describe how the leaf categories evolved from higher level coarse-grained categories, thus can provide a better understanding of the objects. In this paper, we propose to utilize the CNN-RNN framework to address the hierarchical image classification task. CNN allows us to obtain discriminative features for the input images, and RNN enables us to jointly optimize the classification of coarse and fine labels. This framework can not only generate hierarchical labels for images, but also improve the traditional leaf-level classification performance due to incorporating the hierarchical information. Moreover, this framework can be built on top of any CNN architecture which is primarily designed for leaf-level classification. Accordingly, we build a high performance network based on the CNN-RNN paradigm which outperforms the original CNN (wider-ResNet) and also the current state-of-the-art. In addition, we investigate how to utilize the CNN-RNN framework to improve the fine category classification when a fraction of the training data is only annotated with coarse labels. Experimental results demonstrate that CNN-RNN can use the coarse-labeled training data to improve the classification of fine categories, and in some cases it even surpasses the performance achieved by fully annotated training data. This reveals that, CNN-RNN can alleviate the challenge of specialized and expensive annotation of fine labels.
Yanming Guo, Yu Liu 0012, Erwin M. Bakker, Yuanhao Guo, Michael S. Lew
Multim. Tools Appl.1
2018 Fusion that matters: convolutional fusion networks for visual recognition
abstract
In recent years, deep learning has been successfully applied to diverse multimedia research areas, with the aim of learning powerful and informative representations for a variety of visual recognition tasks. In this work, we propose convolutional fusion networks (CFN) to integrate multi-level deep features and fuse a richer visual representation. Despite recent advances in deep fusion networks, they still have limitations due to expensive parameters and weak fusion modules. Instead, CFN uses 1 × 1 convolutional layers and global average pooling to generate side branches with few parameters, and employs a locally-connected fusion module, which can learn adaptive weights for different side branches and form a better fused feature. Specifically, we introduce three key components of the proposed CFN, and discuss its differences from other deep models. Moreover, we propose fully convolutional fusion networks (FCFN) that are an extension of CFN for pixel-level classification applied to several tasks, such as semantic segmentation and edge detection. Our experiments demonstrate that CFN (and FCFN) can achieve promising performance by consistent improvements for both image-level and pixel-level classification tasks, compared to a plain CNN. We release our codes on https://github.com/yuLiu24/CFN . Also, we make a live demo ( goliath.liacs.nl ) using a CFN model trained on the ImageNet dataset.
Yu Liu 0012, Yanming Guo, Theodoros Georgiou 0001, Michael S. Lew
Multim. Tools Appl.2
2018 Learning visual and textual representations for multimodal matching and classification
Yu Liu 0012, Li Liu 0002, Yanming Guo, Michael S. Lew
Pattern Recognit.3
2018 Bag of Surrogate Parts Feature for Visual Recognition
abstract
Convolutional neural networks (CNNs) have attracted significant attention in visual recognition. Several recent studies have shown that, in addition to the fully connected layers, the features derived from the convolutional layers of CNNs can also achieve promising performance in image classification tasks. In this paper, we propose a new feature from the convolutional layers, called Bag of Surrogate Parts (BoSP), and its spatial variant, Spatial-BoSP (S-BoSP). The main idea is, we assume the feature maps in the convolutional layers as surrogate parts, and densely sample and assign image regions to these surrogate parts by observing the activation values. Together with BoSP/S-BoSP, we further propose another two schemes to enhance the performance: scale pooling and global-part prediction. Scale pooling aims to handle the objects with different scales and deformations, and global-part prediction combines the predictions of global and part features. By conducting extensive experiments on generic object, fine-grained object and scene datasets, we find the proposed scheme can not only achieve superior performance to the fully connected feature, but also produces competitive or, in some cases, remarkably better performance than the state of the art.
Yanming Guo, Yu Liu 0012, Songyang Lao, Erwin M. Bakker, Michael S. Lew
IEEE Trans. Multim.1
2017 Learning a Recurrent Residual Fusion Network for Multimodal Matching
Yu Liu 0012, Yanming Guo, Erwin M. Bakker, Michael S. Lew
ICCV2
2017 On the Exploration of Convolutional Fusion Networks for Visual Recognition
Yu Liu 0012, Yanming Guo, Michael S. Lew
MMM (1)2
2017 What Convnets Make for Image Captioning?
Yu Liu 0012, Yanming Guo, Michael S. Lew
MMM (1)2
2016 Bag of Surrogate Parts: one inherent feature of deep CNNs
Yanming Guo, Michael S. Lew
BMVC1
2016 Deep learning for visual understanding: A review
Yanming Guo, Yu Liu 0012, Ard Oerlemans, Songyang Lao, Song Wu 0003, Michael S. Lew
Neurocomputing1
2015 DeepIndex for Accurate and Efficient Image Retrieval
abstract
In the well-known Bag-of-Words model, local features, such as the SIFT descriptor, are extracted and quantized into visual words. Then, an index is created to reduce computational burden. However, local clues serve as low-level representations that can not represent high-level semantic concepts. Recently, the success of deep features extracted from convolutional neural networks(CNN) has shown promising results toward bridging the semantic gap. Inspired by this, we attempt to introduce deep features into inverted index based image retrieval and thus propose the DeepIndex framework. Moreover, considering the compensation of different deep features, we incorporate multiple deep features from different fully connected layers, resulting in the multiple DeepIndex. We find the optimal integration of one midlevel deep feature and one high-level deep feature, from two different CNN architectures separately. This can be treated as an attempt to further reduce the semantic gap. Extensive experiments on three benchmark datasets demonstrate that, the proposed DeepIndex method is competitive with the state-of-the-art on Holidays(85:65% mAP), Paris(81:24% mAP), and UKB(3:76 score). In addition, our method is efficient in terms of both memory and time cost.
Yu Liu 0012, Yanming Guo, Song Wu 0003, Michael S. Lew
ICMR2