Chuanyi Zhang

dblp:87/6424 · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow
abstract
Remote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships beyond simple recognition tasks. In this paper, we aim to construct an Earth observation workflow to handle complex queries by reasoning about spatial context and user intent. As a reasoning workflow, it should autonomously explore and construct its own inference paths, rather than being confined to predefined ground‑truth sequences. Ideally, its architecture ought to be unified yet generalized, possessing capabilities to perform diverse reasoning tasks through one model without requiring additional fine-tuning. Existing remote sensing approaches rely on supervised fine-tuning paradigms and task‑specific heads, limiting both autonomous reasoning and unified generalization. To this end, we propose RemoteReasoner, a unified workflow for geospatial reasoning. The design of RemoteReasoner integrates a multi-modal large language model (MLLM) for interpreting user instructions and localizing targets, together with task transformation strategies that enable multi-granularity tasks, including object-, region-, and pixel-level. In contrast to existing methods, our framework is trained with reinforcement learning (RL) to endow the MLLM sufficient reasoning autonomy. At the inference stage, our transformation strategies enable diverse task output formats without requiring task-specific decoders or further fine-tuning. Experiments demonstrated that RemoteReasoner achieves state-of-the-art performance across multi-granularity reasoning tasks. Furthermore, it retains the MLLM's inherent generalization capability, demonstrating robust performance on unseen tasks and categories.
Liang Yao 0001, Fan Liu 0003, Hongbo Lu, Chuanyi Zhang, Shengxiang Xu, Shimin Di
AAAI4
2026 SGPVT: Self-Generated Proximal Visual Tokens for Mitigating Proximal Collateral Damage in MLLM Unlearning
abstract
Jiaqi Li, Zhijing Zhang, Jiahui Geng, Sheng Bi, Chuanyi Zhang, Fan Liu, Guilin Qi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiaqi Li 0031, Zhijing Zhang, Jiahui Geng, Chuanyi Zhang, Fan Liu 0003, Guilin Qi
ACL (1)5
2026 Heterogeneous Knowledge Distillation Fostered Pre-training for remote sensing object detection
Fan Liu 0003, Chuanyi Zhang, Liang Yao 0001
Pattern Recognit.3
2025 Making Large Vision Language Models to Be Good Few-Shot Learners
abstract
Few-shot classification (FSC) is a fundamental yet challenging task in computer vision that involves recognizing novel classes from limited data. While previous methods have focused on enhancing visual features or incorporating additional modalities, Large Vision Language Models (LVLMs) offer a promising alternative due to their rich knowledge and strong visual perception. However, LVLMs risk learning specific response formats rather than effectively extracting useful information from support data in FSC. In this paper, we investigate LVLMs' performance in FSC and identify key issues such as insufficient learning and the presence of severe position biases. To tackle above challenges, we adopt the meta-learning strategy to teach models ``learn to learn". By constructing a rich set of meta-tasks for instruction fine-tuning, LVLMs enhance the ability to extract information from few-shot support data for classification. Additionally, we further boost LVLM's few-shot learning capabilities through label augmentation (LA) and candidate selection (CS) in the fine-tuning and inference stages, respectively. LA is implemented via a character perturbation strategy to ensure the model focuses on support information. CS leverages attribute descriptions to filter out unreliable candidates and simplify the task. Extensive experiments demonstrate that our approach achieves superior performance on both general and fine-grained datasets. Furthermore, our candidate selection strategy has been proven beneficial for training-free LVLMs.
Fan Liu 0003, Wenwen Cai, Jian Huo, Chuanyi Zhang, Delong Chen
AAAI4
2025 Open-World Attribute Mining for E-Commerce Products with Multimodal Self-Correction Instruction Tuning
abstract
In e-commerce, effective product Attribute Mining (AM) is essential for enhancing product features and aiding consumer decisions.However, current AM methods often focus on extracting attributes from unimodal text, underutilizing multimodal data.In this paper, we propose a novel framework called Multimodal Self-Correction Instruction Tuning (MSIT) to mine new potential attributes from images and texts with Multimodal Large Language Models (MLLMs).The tuning process involves two datasets: Attribute Generation Tuning Data (AGTD) and Chain-of-Thought Tuning Data (CTTD).AGTD is constructed utilizing incontext learning with a small set of seed attributes, aiding the MLLMs in accurately extracting attribute-value pairs from multimodal information.To introduce explicit reasoning and improve the extraction accuracy, we construct CTTD, which incorporates a structured 5-step reasoning process for self-correction.Finally, we employ a 3-stage inference process to filter out redundant attributes and sequentially validate each generated attribute.Comprehensive experimental results on two datasets show that MSIT outperforms state-of-the-art methods.We will release our code and data in the near future.
Jiaqi Li 0031, Xiaoli Shen, Chuanyi Zhang, Guilin Qi
ACL (1)4
2025 Cross Hemisphere-Aware Hybrid Neural Network for AD Diagnosis Based on PET Imaging
abstract
Positron Emission Tomography (PET) plays a vital role in the diagnosis of Alzheimer's Disease (AD) by revealing the brain's metabolic activity and detecting metabolic abnormalities in the early stages of AD. However, due to significant inter-subject variability in AD imaging presentations, metabolic patterns often differ across individuals, making it challenging for traditional deep learning models to effectively capture such complex and heterogeneous features. Considering that the brain exhibits structural symmetry between hemispheres but often presents pathological asymmetry and regional differences in disease progression, we propose an end-to-end hybrid framework to better model this nonuniform distribution. Our framework integrates 3D patch CNN for local feature encoding and incorporates Transformer Encoders to enhance global semantic representation. Subsequently, a Hemisphere-aware Cross Transformer is employed to fuse inter-hemispheric information, further improving feature discriminability. To enhance robustness under complex feature distributions, we employed a hybrid cross-entropy loss function to optimize the classification task. Experimental results based on the ADNI dataset demonstrate that our method achieves significant improvement in AD diagnosis and the mild cognitive impairment (MCI) conversion task. Importantly, our visualization of key brain regions further confirms the model's attention to ADrelated biomarkers, highlighting its potential to improve early AD diagnosis and clinical application.
Yanteng Zhang, Chuanyi Zhang, Congyu Zou, Qiang Liu 0021, Vince D. Calhoun
BIBM3
2025 Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
abstract
While densely annotated image captions significantly facilitate the learning of robust visionlanguage alignment, methodologies for systematically optimizing human annotation efforts remain underexplored.We introduce CHAIN-OF-TALKERS (COTALK), an AI-in-the-loop methodology designed to maximize the number of annotated samples and improve their comprehensiveness under fixed budget constraints (e.g., total human annotation time).The framework is built upon two key insights.First, sequential annotation reduces redundant workload compared to conventional parallel annotation, as subsequent annotators only need to annotate the "residual"-the missing visual information that previous annotations have not covered.Second, humans process textual input faster by reading while outputting annotations with much higher throughput via talking; thus a multimodal interface enables optimized efficiency.We evaluate our framework from two aspects: intrinsic evaluations that assess the comprehensiveness of semantic units, obtained by parsing detailed captions into object-attribute trees and analyzing their effective connections; extrinsic evaluation measures the practical usage of the annotated captions in facilitating vision-language alignment.Experiments with eight participants show our CHAIN-OF-TALKERS (CoTalk) improves annotation speed (0.42 vs. 0.30 units/sec) and retrieval performance (41.13% vs. 40.52%)over the parallel method.per minute? a review and meta-analysis of reading rate.
Delong Chen, Fan Liu 0003, Chuanyi Zhang, Liang Yao 0001, Yuhui Zheng
EMNLP5
2025 Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images
abstract
The Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where semantic contour extraction—such as identifying buildings, road networks, and coastlines-holds significant practical value. Those applications are currently handled via training specialized small models separately on small datasets in each domain. We introduce a foundation model derived from DirectSAM, termed DirectSAM-RS, which not only inherits the strong segmentation capability acquired from natural images, but also benefits from a large-scale dataset we created for remote sensing semantic contour extraction. This dataset comprises over 34k image-text-contour triplets, making it at least 30 times larger than individual dataset. DirectSAM-RS integrates a prompter module: a text encoder and cross-attention layers attached to the DirectSAM architecture, which allows flexible conditioning on target class labels or referring expressions. We evaluate the DirectSAM-RS in both zero-shot and fine-tuning setting, and demonstrate that it achieves state-of-the-art performance across several downstream benchmarks.
Shiyu Miao, Delong Chen, Fan Liu 0003, Chuanyi Zhang, Yanhui Gu, Shengjie Guo, Jun Zhou 0011
ICASSP4
2025 RemoteTrimmer: Adaptive Structural Pruning for Remote Sensing Image Classification
abstract
Since high resolution remote sensing image classifi-cation often requires a relatively high computation complexity, lightweight models tend to be practical and efficient. Model pruning is an effective method for model compression. However, existing methods rarely take into account the specificity of remote sensing images, resulting in significant accuracy loss after pruning. To this end, we propose an effective structural pruning approach for remote sensing image classification. Specifically, a pruning strategy that amplifies the differences in channel importance of the model is introduced. Then an adaptive mining loss function is designed for the fine-tuning process of the pruned model. Finally, we conducted experiments on two remote sensing classification datasets. The experimental results demonstrate that our method achieves minimal accuracy loss after compressing remote sensing classification models, achieving state-of-the-art (SoTA) performance.
Guangwenjie Zou, Liang Yao 0001, Fan Liu 0003, Chuanyi Zhang, Xin Li 0090, Shengxiang Xu, Jun Zhou 0001
ICASSP4
2025 RemoteSAM: Towards Segment Anything for Earth Observation
abstract
We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various input-output interfaces required across different task scenarios. Current systems cannot meet these requirements, as they typically utilize task-specific architecture trained on narrow data domains with limited semantic coverage. Our study addresses these limitations from two aspects: data and modeling. We first introduce an automatic data engine that enjoys significantly better scalability compared to previous human annotation or rule-based approaches. It has enabled us to create the largest dataset of its kind to date, comprising 270K image-text-mask triplets covering an unprecedented range of diverse semantic categories and attribute specifications. Based on this data foundation, we further propose a task unification paradigm that centers around referring expression segmentation. It effectively handles a wide range of vision-centric perception tasks, including classification, detection, segmentation, grounding, etc, using a single model without any task-specific heads. Combining these innovations on data and modeling, we present RemoteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as Falcon, GeoChat, and LHRS-Bot with significantly higher efficiency. Models and data are publicly available at https://github.com/1e12Leon/RemoteSAM.
Liang Yao 0001, Fan Liu 0003, Delong Chen, Chuanyi Zhang, Ziyun Chen 0004, Shimin Di, Yuhui Zheng
ACM Multimedia4
2025 UEMM-Air: Enable UAVs to Undertake More Multi-modal Tasks
abstract
The development of multi-modal Unmanned Aerial Vehicles (UAVs) environment perception systems is hindered by three critical gaps in existing datasets: (1) insufficient modalities and pixel misalignment, (2) noisy labels, and (3) limited task types. To address these gaps, we propose an automatic data construction approach and construct a multi-modal UAV-based environment perception dataset, UEMM-Air. Its synthetic nature ensures scalability, reproducibility, and rare-event coverage, making it suitable for large-scale model pre-training. Benefiting from our automated data collection and annotation pipeline, UEMM-Air encompasses 120k data pairs across 6 aligned modalities and supports 4 perception tasks, significantly exceeding existing datasets (max 60k data, 3 modalities, 2 tasks). Compared to existing synthetic datasets like SynDrone, UEMM-Air provides more accurate annotations by avoiding noisy labels from direct coordinate computation. Notably, models pre-trained on UEMM-Air achieve a 5.8% accuracy improvement compared to those utilizing other synthetic datasets, while requiring less than half the data. This benchmark establishes performance evaluation of UAV multi-modal environmental perception models, and hopefully encourages more research efforts towards enabling UAVs to undertake more multi-modal tasks. The dataset and its generation engine are openly accessible under a permissive license at https://github.com/1e12Leon/UEMM-Air.
Liang Yao 0001, Fan Liu 0003, Shengxiang Xu, Chuanyi Zhang, Shimin Di, Jianyu Jiang, Zequan Wang, Jun Zhou 0001
ACM Multimedia4
2025 Integrating Global and Local Information for Remote Sensing Image-Text Retrieval
abstract
Pre-trained Vision-Language Models (VLMs) have demonstrated promising performance in remote sensing image-text retrieval tasks. However, the scarcity of high-quality image-text datasets remains a challenge in fine-tuning VLMs for remote sensing. The captions in existing datasets tend to be uniform and lack details. To fully utilize rich detailed information from remote sensing images, we propose a method to fine-tune VLMs. We first construct a new visual-language dataset that balances both Global and Local information for Remote Sensing image-text retrieval (GLRS). Specifically, a Multi-modal Large Language Model (MLLM) is utilized to generate captions for local patches and global captions for the entire image. To effectively utilize local information, we propose a Global and Local image Captioning method (GLCap). With a Large Language Model (LLM), we further obtain higher-quality captions by merging both global and local captions. Finally, we fine-tune the weights of RS-M-CLIP with a progressive global-local fine-tuning strategy on GLRS. Experimental results demonstrate that our method outperforms state-of-the-art approaches on two common remote sensing image-text retrieval downstream tasks. The dataset will be publicly available once the paper is accepted.
Ziyun Chen 0004, Fan Liu 0003, Zhangqingyun Guan, Xiaocong Zhou, Chuanyi Zhang
IEEE Geosci. Remote. Sens. Lett.6
2025 Domain-Invariant Progressive Knowledge Distillation for UAV-Based Object Detection
abstract
Knowledge distillation (KD) is an effective method for compressing models in object detection tasks. Due to limited computational capability, unmanned aerial vehicle-based object detection (UAV-OD) widely adopt the KD technique to obtain lightweight detectors. Existing methods often overlook the significant differences in feature space caused by the large gap in scale between the teacher and student models. This limitation hampers the efficiency of knowledge transfer during the distillation process. Furthermore, the complex backgrounds in aerial images make it challenging for the student model to efficiently learn the object features. In this letter, we propose a novel KD framework for UAV-OD. Specifically, a progressive distillation approach is designed to alleviate the feature gap between teacher and student models. Then, a new feature alignment method is provided to extract object-related features for enhancing the student model’s knowledge reception efficiency. Finally, extensive experiments are conducted to validate the effectiveness of our proposed approach. The results demonstrate that our proposed method achieves state-of-the-art performance on two datasets.
Liang Yao 0001, Fan Liu 0003, Chuanyi Zhang, Zhiquan Ou
IEEE Geosci. Remote. Sens. Lett.3
2025 Multi-stage Bayesian Prototype Refinement with feature weighting for few-shot classification
Xiaocong Zhou, Shengxiang Xu, Fan Liu 0003, Chuanyi Zhang, Wenwen Cai, Jun Zhou 0011
Pattern Anal. Appl.5
2025 Boost UAV-Based Object Detection via Scale-Invariant Feature Disentanglement and Adversarial Learning
abstract
Detecting objects from Unmanned Aerial Vehicles (UAV) is often hindered by a large number of small objects, resulting in low detection accuracy. To address this issue, mainstream approaches typically utilize multi-stage inferences. Despite their remarkable detecting accuracies, real-time efficiency is sacrificed, making them less practical to handle real applications. To this end, we propose to improve the single-stage inference accuracy through learning scale-invariant features. Specifically, a Scale-Invariant Feature Disentangling module is designed to disentangle scale-related and scale-invariant features. Then an Adversarial Feature Learning scheme is employed to enhance disentanglement. Finally, scale-invariant features are leveraged for robust UAV-based object detection. Furthermore, we construct a multi-modal UAV object detection dataset, State-Air, which incorporates annotated UAV state parameters. We apply our approach to three lightweight detection frameworks on two benchmark datasets. Extensive experiments demonstrate that our approach can effectively improve model accuracy and achieve state-of-the-art (SoTA) performance on three datasets. Our code and dataset are publicly available at https://github.com/1e12Leon/SIFDAL.
Fan Liu 0003, Liang Yao 0001, Chuanyi Zhang, Xiruo Jiang, Jun Zhou 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Progressively Robust Loss for Deep Learning with Noisy Labels
abstract
Learning with noisy labels (LNL) plays a pivotal role in arming deep neural networks (DNNs) to combat label noise. Early noise-robust functions tend to promote robustness against noisy labels at the cost of sacrificing data-fitting ability. Recent robust loss methods typically try to balance noise-robustness and learning capability. However, most of them generally descend to partially robust losses, which are still exposed to the risk of overfitting noisy labels. To this end, we propose a novel paradigm named progressively robust loss framework to dynamically guide existing noise-robust losses from fast convergence to noise-tolerant, which is in accord with the deep models’ memorization effect. Furthermore, our theoretical analysis of the upper bounds of empirical risk errors illustrates the increasing noise-robustness of our approach. Experimental results on two synthetic benchmarks (CIFAR-100N and CIFAR-80N) and two real-world noisy datasets (WebFG-496 and Webvision) demonstrate the superiority of our approach over state-of-the-art robust loss methods in dealing with noisy labels. The code is available at https://github.com/ptcepgce/ptcepgce.
Zhenhuang Cai, Yuanbo Chen, Chuanyi Zhang, Zeren Sun, Yazhou Yao
IJCNN4
2024 Feature-weighted Multi-stage Bayesian Prototype for Few-shot Classification
abstract
Few-shot classification aims to recognize the query sample through a limited amount of support data, where a prototype classifier is commonly applied. However, although the prototype classifier is simple and non-parametric, it does not fully utilize the prior information of samples, leading to prototype bias. To this end, we propose a Feature-weighted Multi-stage Bayesian Prototype Classifier (FMBPC). Specifically, we utilize a feature weighting module to balance the effect of each support sample. Then, features of balanced support samples are utilized as prior information to construct the Bayesian prototype classifier, which can focus more on the important information. Ultimately, a multi-stage inferring strategy is adopted, where the support sample with the greatest distance is filtered in each stage. Prototypes and the corresponding classification score are updated after sample filtering. By integrating the multi-stage classification results, we successfully utilize multi-stage Bayesian inference to enhance the prototype classifier for more accurate few-shot classification results. Experimental results show the efficacy of our method, demonstrating notable advancements in few-shot classification accuracy.
Xiaocong Zhou, Fan Liu 0003, Chuanyi Zhang, Wenwen Cai, Jun Zhou 0001
MMAsia3
2024 Single Image Unlearning: Efficient Machine Unlearning in Multimodal Large Language Models
abstract
Machine unlearning (MU) empowers individuals with the `right to be forgotten' by removing their private or sensitive information encoded in machine learning models. However, it remains uncertain whether MU can be effectively applied to Multimodal Large Language Models (MLLMs), particularly in scenarios of forgetting the leaked visual data of concepts. To overcome the challenge, we propose an efficient method, Single Image Unlearning (SIU), to unlearn the visual recognition of a concept by fine-tuning a single associated image for few steps. SIU consists of two key aspects: (i) Constructing Multifaceted fine-tuning data. We introduce four targets, based on which we construct fine-tuning data for the concepts to be forgotten; (ii) Joint training loss. To synchronously forget the visual recognition of concepts and preserve the utility of MLLMs, we fine-tune MLLMs through a novel Dual Masked KL-divergence Loss combined with Cross Entropy loss. Alongside our method, we establish MMUBench, a new benchmark for MU in MLLMs and introduce a collection of metrics for its evaluation. Experimental results on MMUBench show that SIU completely surpasses the performance of existing methods. Furthermore, we surprisingly find that SIU can avoid invasive membership inference attacks and jailbreak attacks. To the best of our knowledge, we are the first to explore MU in MLLMs. We will release the code and benchmark in the near future.
Jiaqi Li 0031, Qianshan Wei, Chuanyi Zhang, Guilin Qi, Miaozeng Du, Yongrui Chen 0002, Fan Liu 0003
NeurIPS3
2023 Three Stream Based Multi-level Event Contrastive Learning for Text-Video Event Extraction
abstract
Text-video based multimodal event extraction refers to identifying event information from the given text-video pairs.Existing methods predominantly utilize video appearance features (VAF) and text sequence features (TSF) as input information.Some of them employ contrastive learning to align VAF with the event types extracted from TSF.However, they disregard the motion representations in videos and the optimization of contrastive objective could be misguided by the background noise from RGB frames.We observe that the same event triggers correspond to similar motion trajectories, which are hardly affected by the background noise.motivated by this, we propose a Three Stream Multimodal Event Extraction framework (TSEE) that simultaneously utilizes the features of text sequence and video appearance, as well as the motion representations to enhance the event extraction capacity.Firstly, we extract the optical flow features (OFF) as motion representations from videos to incorporate with VAF and TSF.Then we introduce a Multi-level Event Contrastive Learning module to align the embedding space between OFF and event triggers, as well as between event triggers and types.Finally, a Dual Querying Text module is proposed to enhance the interaction between modalities.Experimental results show that TSEE outperforms the state-of-theart methods, which demonstrates its superiority.
Jiaqi Li 0031, Chuanyi Zhang, Miaozeng Du, Dehai Min, Yongrui Chen 0002, Guilin Qi
EMNLP2
2023 Incorporating Domain Knowledge Graph into Multimodal Movie Genre Classification with Self-Supervised Attention and Contrastive Learning
abstract
Multimodal movie genre classification has always been regarded as a demanding multi-label classification task due to the diversity of multimodal data such as posters, plot summaries, trailers and metadata. Although existing works have made great progress in modeling and combining each modality, they still face three issues: 1) unutilized group relations in metadata, 2) unreliable attention allocation, and 3) indiscriminative fused features. Given that the knowledge graph has been proven to contain rich information, we present a novel framework that exploits the knowledge graph from various perspectives to address the above problems. As a preparation, the metadata is processed into a domain knowledge graph. A translate model for knowledge graph embedding is adopted to capture the relations between entities. Firstly we retrieve the relevant embedding from the knowledge graph by utilizing group relations in metadata and then integrate it with other modalities. Next, we introduce an Attention Teacher module for reliable attention allocation based on self-supervised learning. It learns the distribution of the knowledge graph and produces rational attention weights. Finally, a Genre-Centroid Anchored Contrastive Learning module is proposed to strengthen the discriminative ability of fused features. The embedding space of anchors is initialized from the genre entities in the knowledge graph. To verify the effectiveness of our framework, we collect a larger and more challenging dataset named MM-IMDb 2.0 compared with the MM-IMDb dataset. The experimental results on two datasets demonstrate that our model is superior to the state-of-the-art methods. Our code and dataset is available at https://github.com/aoluming/IDKG.git.
Jiaqi Li 0031, Guilin Qi, Chuanyi Zhang, Yongrui Chen 0002, Yiming Tan, Chenlong Xia
ACM Multimedia3
2023 Phertilizer: Growing a clonal tree from ultra-low coverage single-cell DNA sequencing of tumors
abstract
Emerging ultra-low coverage single-cell DNA sequencing (scDNA-seq) technologies have enabled high resolution evolutionary studies of copy number aberrations (CNAs) within tumors. While these sequencing technologies are well suited for identifying CNAs due to the uniformity of sequencing coverage, the sparsity of coverage poses challenges for the study of single-nucleotide variants (SNVs). In order to maximize the utility of increasingly available ultra-low coverage scDNA-seq data and obtain a comprehensive understanding of tumor evolution, it is important to also analyze the evolution of SNVs from the same set of tumor cells. We present Phertilizer, a method to infer a clonal tree from ultra-low coverage scDNA-seq data of a tumor. Based on a probabilistic model, our method recursively partitions the data by identifying key evolutionary events in the history of the tumor. We demonstrate the performance of Phertilizer on simulated data as well as on two real datasets, finding that Phertilizer effectively utilizes the copy-number signal inherent in the data to more accurately uncover clonal structure and genotypes compared to previous methods.
Leah L. Weber, Chuanyi Zhang, Idoia Ochoa, Mohammed El-Kebir
PLoS Comput. Biol.2
2023 Guided by Meta-Set: A Data-Driven Method for Fine-Grained Visual Recognition
abstract
The lack of sufficient training data has been one obstacle to fine-grained visual classification research because labeling subcategories generally requires specialist knowledge. As one optional approach to alleviating the data-hunger problem, leveraging web images as training data is drawing increasing attention. Nevertheless, web images potentially have false labels, which can misguide the training process. Although several works have been proposed to deal with label noise, it still can be difficult for the network to tackle complex real-world noisy labels without any prior knowledge. In the literature, we propose to leverage a small and clean meta-set to provide reliable prior knowledge for tackling noisy web images. Specifically, our method trains a network with two peer predicting heads, which learn from noisy web images (web head) and meta ones (meta head), respectively. The meta head produces pseudo soft labels for web images to revise their training loss, which can overcome the high noise ratio problem. Furthermore, a selection net is trained in a meta-learning strategy to identify in- and out-of-distribution noisy images. Then in-distribution ones are reused for training with pseudo soft labels produced by the meta head as supervision, while out-of-distribution ones are discarded. In this manner, the misguidance caused by label noise is remarkably alleviated and in-distribution noisy samples are properly exploited to boost model performance. The superiority of our proposed approach is demonstrated by mathematical theory with great interpretability as well as extensive experimental results on the real-world dataset WebFG-496.
Chuanyi Zhang, Guosheng Lin, Qiong Wang 0003, Fumin Shen, Yazhou Yao, Zhenmin Tang
IEEE Trans. Multim.1
2022 CORSID Enables de novo Identification of Transcription Regulatory Sequences and Genes in Coronaviruses
Chuanyi Zhang, Palash Sashittal, Mohammed El-Kebir
RECOMB1
2022 Exploiting Web Images for Fine-Grained Visual Recognition by Eliminating Open-Set Noise and Utilizing Hard Examples
abstract
Labeling objects at a subordinate level typically requires expert knowledge, which is not always available when using random annotators. As such, learning directly from web images for fine-grained recognition has attracted broad attention. However, the presence of label noise and hard examples in web images are two obstacles for training robust fine-grained recognition models. Therefore, in this paper, we propose a novel approach for removing irrelevant samples from real-world web images during training, while employing useful hard examples to update the network. Thus, our approach can alleviate the harmful effects of irrelevant noisy web images and hard examples to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Advanced-Softly-Update-Drop.
Huafeng Liu 0004, Chuanyi Zhang, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Zhenmin Tang, Jian Zhang 0002
IEEE Trans. Multim.2
2022 Robust Learning From Noisy Web Images Via Data Purification for Fine-Grained Recognition
abstract
Manually labeling fine-grained datasetsis laborious and typically requires domain-specific expert knowledge. Conversely, a vast amount of web data is relatively easy to obtain with nearly no human effort. Therefore, learning from noisy web data for fine-grained tasks is attracting increasing attention in recent years. However, the presence of noise in web images is a huge obstacle for training robust fine-grained recognition models. To this end, we propose a novel approach to identify noisy images as well as specifically distinguish in- and out-of-distribution samples. It can purify the noisy web training set by discarding out-of-distribution noise and relabeling in-distribution noisy samples. Then we can train the model on the purified dataset to alleviate the harmful effects of noise and make the most of web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Dataset-Purification.
Chuanyi Zhang, Qiong Wang 0003, Guosen Xie, Qi Wu 0001, Fumin Shen, Zhenmin Tang
IEEE Trans. Multim.1
2021 Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of pseudo labels within the image’s salient region. In this work, we propose a non-salient region object mining approach for weakly supervised semantic segmentation. We introduce a graph-based global reasoning unit to strengthen the classification network’s ability to capture global relations among disjoint and distant regions. This helps the network activate the object features outside the salient area. To further mine the non-salient region objects, we propose to exert the segmentation network’s self-correction ability. Specifically, a potential object mining module is proposed to reduce the false-negative rate in pseudo labels. Moreover, we propose a non-salient region masking module for complex images to generate masked pseudo labels. Our non-salient region masking module helps further discover the objects in the non-salient region. Extensive experiments on the PASCAL VOC dataset demonstrate state-of-the-art results compared to current methods. The source codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/nsrom.
Yazhou Yao, Tao Chen 0012, Guosen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Zhenmin Tang, Jian Zhang 0002
CVPR4
2021 Jo-SRC: A Contrastive Approach for Combating Noisy Labels
abstract
Due to the memorization effect in Deep Neural Networks (DNNs), training with noisy labels usually results in inferior model performance. Existing state-of-the-art methods primarily adopt a sample selection strategy, which selects small-loss samples for subsequent training. However, prior literature tends to perform sample selection within each mini-batch, neglecting the imbalance of noise ratios in different mini-batches. Moreover, valuable knowledge within high-loss samples is wasted. To this end, we propose a noise-robust approach named Jo-SRC (Joint Sample Selection and Model Regularization based on Consistency). Specifically, we train the network in a contrastive learning manner. Predictions from two different views of each sample are used to estimate its "likelihood" of being clean or out-of-distribution. Furthermore, we propose a joint loss to advance the model generalization performance by introducing consistency regularization. Extensive experiments have validated the superiority of our approach over existing state-of-the-art methods. The source code and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/Jo-SRC.
Yazhou Yao, Zeren Sun, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Jian Zhang 0002, Zhenmin Tang
CVPR3
2021 Extracting Useful Knowledge from Noisy Web Images via Data Purification for Fine-Grained Recognition
abstract
Fine-grained visual recognition tasks typically require training data with reliable acquisition and annotation processes. Acquiring such datasets with precise fine-grained annotations is very expensive and time-consuming. Conversely, a vast amount of web data is relatively easy to obtain with nearly no human effort. Nevertheless, the presence of label noise in web images becomes a huge obstacle for training robust fine-grained recognition models. In this work, we investigate the noisy label problem and propose a method that can specifically distinguish in- and out-of-distribution noisy samples. It can purify the web training data by discarding out-of-distribution noisy images and relabeling in-distribution ones. After purification, we can train the model on a less noisy web training set to achieve better robustness and performance. Extensive experiments on three real-world web datasets for fine-grained visual recognition demonstrate the superiority of our approach.
Chuanyi Zhang, Yazhou Yao, Xing Xu 0001, Jie Shao 0001, Jingkuan Song, Zechao Li, Zhenmin Tang
ACM Multimedia1
2020 Web-Supervised Network with Softly Update-Drop Training for Fine-Grained Visual Classification
abstract
Labeling objects at the subordinate level typically requires expert knowledge, which is not always available from a random annotator. Accordingly, learning directly from web images for fine-grained visual classification (FGVC) has attracted broad attention. However, the existence of noise in web images is a huge obstacle for training robust deep neural networks. In this paper, we propose a novel approach to remove irrelevant samples from the real-world web images during training, and only utilize useful images for updating the networks. Thus, our network can alleviate the harmful effects caused by irrelevant noisy web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art webly supervised methods. The data and source code of this work have been made anonymously available at: https://github.com/z337-408/WSNFGVC.
Chuanyi Zhang, Yazhou Yao, Huafeng Liu 0004, Guosen Xie, Xiangbo Shu, Tianfei Zhou, Zheng Zhang 0006, Fumin Shen, Zhenmin Tang
AAAI1
2020 Web-Supervised Network for Fine-Grained Visual Classification
abstract
Fine-grained visual classification (FGVC) is a tough task due to its high annotation cost of the fine-grained subcategories. To build a large-scale dataset at low manual cost, straightforwardly learning from web images for FGVC has attracted broad attention. However, there exist two characteristics in the need of concerning for the web dataset: 1) Noisy images; 2) A large proportion of hard examples. In this paper, we propose a simple yet effective approach to deal with noisy images and hard examples during training. Our method is a pure web-supervised method for FGVC. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to the state-of-the-art web-supervised methods. The data and source code of this work have been posted available at: https://github.com/NUST-Machine-Intelligence-Laboratory/WSNFG.
Chuanyi Zhang, Yazhou Yao, Jiachao Zhang, Jian Zhang 0002, Zhenmin Tang
ICME1
2020 Data-driven Meta-set Based Fine-Grained Visual Recognition
abstract
Constructing fine-grained image datasets typically requires domain-specific expert knowledge, which is not always available for crowd-sourcing platform annotators. Accordingly, learning directly from web images becomes an alternative method for fine-grained visual recognition. However, label noise in the web training set can severely degrade the model performance. To this end, we propose a data-driven meta-set based approach to deal with noisy web images for fine-grained recognition. Specifically, guided by a small amount of clean meta-set, we train a selection net in a meta-learning manner to distinguish in- and out-of-distribution noisy images. To further boost the robustness of the model, we also learn a labeling net to correct the labels of in-distribution noisy data. In this way, our proposed method can alleviate the harmful effects caused by out-of-distribution noise and properly exploit the in-distribution noisy samples for training. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art noise-robust methods.
Chuanyi Zhang, Yazhou Yao, Xiangbo Shu, Zechao Li, Zhenmin Tang, Qi Wu 0001
ACM Multimedia1
2020 VEF: a variant filtering tool based on ensemble methods
abstract
MOTIVATION: Variants identified by current genomic analysis pipelines contain many incorrectly called variants. These can be potentially eliminated by applying state-of-the-art filtering tools, such as Variant Quality Score Recalibration (VQSR) or Hard Filtering (HF). However, these methods are very user-dependent and fail to run in some cases. We propose VEF, a variant filtering tool based on decision tree ensemble methods that overcomes the main drawbacks of VQSR and HF. Contrary to these methods, we treat filtering as a supervised learning problem, using variant call data with known 'true' variants, i.e. gold standard, for training. Once trained, VEF can be directly applied to filter the variants contained in a given Variants Call Format (VCF) file (we consider training and testing VCF files generated with the same tools, as we assume they will share feature characteristics). RESULTS: For the analysis, we used whole genome sequencing (WGS) Human datasets for which the gold standards are available. We show on these data that the proposed filtering tool VEF consistently outperforms VQSR and HF. In addition, we show that VEF generalizes well even when some features have missing values, when the training and testing datasets differ in coverage, and when sequencing pipelines other than GATK are used. Finally, since the training needs to be performed only once, there is a significant saving in running time when compared with VQSR (4 versus 50 min approximately for filtering the single nucleotide polymorphisms of a WGS Human sample). AVAILABILITY AND IMPLEMENTATION: Code and scripts available at: github.com/ChuanyiZ/vef. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chuanyi Zhang, Idoia Ochoa
Bioinform.1