Zhongchao Shi

dblp:45/5323 · DBLP profile ↗
← Back
54ranked-venue papers
1as first author
33since 2021 · last 2026
0000-0002-5216-3827ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 1 first-author · 21 since 2021Artificial intelligence and machine learning · 24 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SlimNet: High-Quality and Efficient Object Removal by Eliciting Latent Capabilities of Diffusion Models
abstract
Object removal aims to eliminate undesired objects from images while plausibly restoring the underlying content with high visual fidelity. Existing diffusion-based methods often achieve strong results by relying on large-scale model finetuning, auxiliary control networks, or powerful diffusion backbones, which incur substantial computational overhead and limit practical efficiency. In this work, we propose SlimNet, a high-quality and efficient object removal framework that elicits the latent object removal capabilities of pretrained diffusion models. SlimNet injects lightweight adapters into a frozen diffusion backbone to selectively modulate intermediate representations, avoiding heavy architectural modifications. To further improve visual fidelity, we design a composite perceptual loss that enforces object removal accuracy, background preservation, and smooth boundary transitions, together with a semantic-aware data processing pipeline for automatic mask generation. Extensive experiments demonstrate that SlimNet achieves competitive removal quality compared to state-of-the-art methods, with favorable reductions in inference time and VRAM consumption, making it a practical solution for resource-constrained multimedia applications.
Yao Zhang 0010, Zhongchao Shi, Jianping Fan 0007, Guihua Zeng
ICMR5
2026 A Bootstrap Pipeline for Chat-Based Image Retrieval with Effective Question Generation
abstract
Chat-based image retrieval uses Large Language Models (LLM) to guide user input to enable more specific and precise search results, where LLM can enhance this process by asking user retrieval-oriented questions eliciting additional details about the target image. Despite the potential of this approach, no specialized Questioner model has been developed for this task due to the following significant challenges: (a) the difficulty of determining the optimal questions to ask; (b) the lack of a suitable protocol for fair model comparison; and (c) the notable scarcity of dialog-to-image retrieval data. To address these challenges, two fundamental principles are developed in this article to ensure the simplicity and effectiveness of the generated questions while enabling a fair comparison and accurate estimation of data quality and model performance. A bootstrap training methodology is introduced to collect retrieval-oriented dialog data and concurrently train the Questioner and the image Retriever. Under a fair comparison protocol, our extensive experiments have demonstrated that our proposed method can not only address the critical data gap but also achieve state-of-the-art results, which substantially surpass GPT-4o and GPT-4-Turbo through the fine-tuning of an 8B model.
Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui
ACM Trans. Multim. Comput. Commun. Appl.5
2025 DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with Attributes
abstract
Recent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capacity for fine-grained linguistic comprehension and leads to a significant decline in performance when faced with detailed descriptions or contextual information. To tackle these problems, we develop DoGA: Detect objects with Grouped Attributes, which employs commonly apparent attributes to bridge different granular semantics and uses specific attributes to identify the object discrepancy. Our DoGA incorporates three principle components: 1) Generation of attribute-based prompts, consisting of linguistic definitions enriched with common-sense visible attributes and hard negative notations deriving from the image-specific attribute features; 2) Paralleled entity fusion and optimization, designed to manage long attribute-based descriptions and negative concepts efficiently; and 3) Prompt-wise grouped training to accommodate model to perform many-to-many assignments, facilitating simultaneous training and inferring with multiple attribute-based synonyms. Extensive experiments demonstrate that training with synonymous attribute-based prompts allows DoGA to generalize multi-granular prompts and surpass previous state-of-the-art approaches, yielding 50.2 on the COCO and 38.0 on the LVIS benchmarks under the zero-short setting. We will make our code publicly available upon acceptance.
Yang Liu 0250, Feng Hou, Yunjie Peng, Gangjian Zhang, Yao Zhang 0010, Peng Wang 0095, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
AAAI10
2025 Omni-Query Active Learning for Source-Free Domain Adaptive Cross-Modality 3D Semantic Segmentation
abstract
Source-Free Domain Adaptation (SFDA) aims to transfer a pre-trained source model to the unlabeled target domain without accessing the source data, thereby effectively solving labeled data dependency and domain shift problems. However, the SFDA setting faces a bottleneck due to the absence of supervisory information. To mitigate this problem, Active Learning (AL) is introduced to combine with SFDA, endeavoring to actively label a small set of the most high-quality target points so that models with satisfactory performance can be obtained at an acceptable cost. Nevertheless, several issues remain unresolved, namely when to query new labels during training, what kind of samples deserve labeling to ensure rich information, and where the labels should be distributed to guarantee diversity. Thus we elaborate OmniQuery to omnibearing address the “When, What, and Where” problems about active points querying in source-free domain adaptation for cross-modal 3D semantic segmentation. The method consists of three main components: Query Decider, Point Ranker, and Budget Slicer. The Query Decider determines the optimal timing to query new points by fitting the validation curves during training. The Point Ranker nominates points for annotation by calculating the ambiguity of neighboring points in the feature space. The Budget Slicer allocates the annotation quota, i.e., labeling percentage of the point cloud, to different semantic regions by utilizing the advanced 2D semantic segmentation capabilities of the Segment Anything Model (SAM). Extensive experiments demonstrate the effectiveness of our proposed method, achieving up to 99.64% of fully supervised performance with only 3% of labels, and consistently outperforming comparison methods across various scenarios.
Jianxiang Xie, Yachao Zhang 0001, Zhongchao Shi, Jianping Fan 0007, Yuan Xie 0006, Yanyun Qu
AAAI4
2025 CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter
abstract
Yepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen, Li Liu, Jiang Tian, Zhongchao Shi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen, Jiang Tian, Zhongchao Shi
ACL (1)7
2025 Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter Levels
abstract
Junjie Ye, Yuming Yang, Yang Nan, Shuo Li, Qi Zhang, Tao Gui, Xuanjing Huang, Peng Wang, Zhongchao Shi, Jianping Fan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Junjie Ye 0005, Yuming Yang 0001, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001, Peng Wang 0095, Zhongchao Shi, Jianping Fan 0007
EMNLP9
2025 ARB-LLM: Alternating Refined Binarizations for Large Language Models
abstract
Large Language Models (LLMs) have greatly pushed forward advancements in natural language processing, yet their high memory and computational demands hinder practical deployment. Binarization, as an effective compression technique, can shrink model weights to just 1 bit, significantly reducing the high demands on computation and memory. However, current binarization methods struggle to narrow the distribution gap between binarized and full-precision weights, while also overlooking the column deviation in LLM weight distribution. To tackle these issues, we propose ARB-LLM, a novel 1-bit post-training quantization (PTQ) technique tailored for LLMs. To narrow the distribution shift between binarized and full-precision weights, we first design an alternating refined binarization (ARB) algorithm to progressively update the binarization parameters, which significantly reduces the quantization error. Moreover, considering the pivot role of calibration data and the column deviation in LLM weights, we further extend ARB to ARB-X and ARB-RC. In addition, we refine the weight partition strategy with column-group bitmap (CGB), which further enhance performance. Equipping ARB-X and ARB-RC with CGB, we obtain ARB-LLM$_{\text{X}}$ and ARB-LLM$ _{\text{RC}} $ respectively, which significantly outperform state-of-the-art (SOTA) binarization methods for LLMs. As a binary PTQ method, our ARB-LLM$ _{\text{RC}} $ is the first to surpass FP16 models of the same size. Code: https://github.com/ZHITENGLI/ARB-LLM.
Zhiteng Li, Xianglong Yan, Tianao Zhang, Haotong Qin, Jiang Tian, Zhongchao Shi, Linghe Kong, Yulun Zhang 0001, Xiaokang Yang 0001
ICLR7
2025 DenseGrounding: Improving Dense Language-Vision Semantics for Ego-centric 3D Visual Grounding
abstract
Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D visual grounding, where agents locate target objects in real-world 3D spaces based on verbal descriptions. However, this task faces two significant challenges: (1) loss of fine-grained visual semantics due to sparse fusion of point clouds with ego-centric multi-view images, (2) limited textual semantic context due to arbitrary language descriptions. We propose DenseGrounding, a novel approach designed to address these issues by enhancing both visual and textual semantics. For visual features, we introduce the Hierarchical Scene Semantic Enhancer, which retains dense semantics by capturing fine-grained global scene features and facilitating cross-modal alignment. For text descriptions, we propose a Language Semantic Enhancer that leverage large language models to provide rich context and diverse language descriptions with additional context during model training. Extensive experiments show that DenseGrounding significantly outperforms existing methods in overall accuracy, achieving improvements of **5.81%** and **7.56%** when trained on the comprehensive full training dataset and smaller mini subset, respectively, further advancing the SOTA in ego-centric 3D visual grounding. Our method also achieves **1st place** and receives **Innovation Award** in the 2024 Autonomous Grand Challenge Multi-view 3D Visual Grounding Track, validating its effectiveness and robustness.
Henry Zheng, Qihang Peng, Yong Xien Chng, Rui Huang 0012, Yepeng Weng, Zhongchao Shi, Gao Huang 0001
ICLR7
2025 Traversal Verification for Speculative Tree Decoding
abstract
Speculative decoding is a promising approach for accelerating large language models. The primary idea is to use a lightweight draft model to speculate the output of the target model for multiple subsequent timesteps, and then verify them in parallel to determine whether the drafted tokens should be accepted or rejected. To enhance acceptance rates, existing frameworks typically construct token trees containing multiple candidates in each timestep. However, their reliance on token-level verification mechanisms introduces two critical limitations: First, the probability distribution of a sequence differs from that of individual tokens, leading to suboptimal acceptance length. Second, current verification schemes begin from the root node and proceed layer by layer in a top-down manner. Once a parent node is rejected, all its child nodes should be discarded, resulting in inefficient utilization of speculative candidates. This paper introduces Traversal Verification, a novel speculative decoding algorithm that fundamentally rethinks the verification paradigm through leaf-to-root traversal. Our approach considers the acceptance of the entire token sequence from the current node to the root, and preserves potentially valid subsequences that would be prematurely discarded by existing methods. We theoretically prove that the probability distribution obtained through Traversal Verification is identical to that of the target model, guaranteeing lossless inference while achieving substantial acceleration gains. Experimental results on various models and multiple tasks demonstrate that our method consistently improves acceptance length and throughput over token-level verification.
Yepeng Weng, Qiao Hu 0001, Xujie Chen, Dianwen Mei, Huishi Qiu, Jiang Tian, Zhongchao Shi
NeurIPS8
2025 Learning Preference Distributions: A Label-Side Paradigm for Explainable Reward Models
abstract
The reward model is a critical component in training powerful large language models. However, current methods largely overlook the inherent subjectivity and variability in human preferences. Typically, this issue is indirectly addressed through model-side approaches, such as ensemble methods to estimate uncertainty from multiple predictions or uncertainty-aware regression to predict mean and variance. These indirect approaches fail to capture the intrinsic distributional characteristics and inter-rater disagreements present in human judgments. In contrast, we propose a direct, label-side solution by explicitly modeling human preference distributions. We recover missing information from scalar ratings to construct meaningful distribution labels. By employing Label Distribution Learning (LDL), each dimension of the resulting multidimensional discrete distribution explicitly corresponds to a specific preference score, naturally reflecting the subjective and multidimensional nature of human evaluations. Our approach improves explainability and confidence estimation, while also enabling more effective data selection and sample-efficient test-time adaptation. Empirical results demonstrate that our method not only achieves state-of-the-art performance but also provides a robust framework for uncertainty quantification and nuanced preference modeling.
Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui
SMC4
2025 Collective domain adversarial learning for unsupervised domain adaptation
Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui
Frontiers Comput. Sci.4
2025 Epistemic graph: A plug-and-play module for hybrid representation learning
abstract
In recent years, deep models have achieved remarkable success in various vision tasks. However, their performance heavily relies on large training datasets. In contrast, humans exhibit hybrid learning, seamlessly integrating structured knowledge for cross-domain recognition or relying on a smaller amount of data samples for few-shot learning. Motivated by this human-like epistemic process, we aim to extend hybrid learning to computer vision tasks by integrating structured knowledge with data samples for more effective representation learning. Nevertheless, this extension faces significant challenges due to the substantial gap between structured knowledge and deep features learned from data samples, encompassing both dimensions and knowledge granularity. In this paper, a novel Epistemic Graph Layer (EGLayer) is introduced to enable hybrid learning, enhancing the exchange of information between deep features and a structured knowledge graph. Our EGLayer is composed of three major parts, including a local graph module, a query aggregation model, and a novel correlation alignment loss function to emulate human epistemic ability. Serving as a plug-and-play module that can replace the standard linear classifier, EGLayer significantly improves the performance of deep models. Extensive experiments demonstrate that EGLayer can greatly enhance representation learning for the tasks of cross-domain recognition and few-shot learning, and the visualization of knowledge graphs can aid in model interpretation.
Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Yong Rui
Neurocomputing5
2025 Gradient-aware domain-invariant learning for domain generalization
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
Multim. Syst.7
2025 DomainVerse: A Benchmark Towards Real-World Distribution Shifts for Training-Free Adaptive Domain Generalization
abstract
Traditional cross-domain tasks, including unsupervised domain adaptation (UDA), domain generalization (DG) and test-time adaptation (TTA), rely heavily on the training model by source domain data whether for specific or arbitrary target domains. With the recent advance of vision-language models (VLMs), recognized as natural source models that can be transferred to various downstream tasks without any parameter training, we propose a novel cross-domain task directly combining the strengths of both UDA and DG, named Training-Free Adaptive Domain Generalization (TF-ADG). However, current cross-domain datasets have many limitations, such as unrealistic domains, unclear domain definitions, and the inability to fine-grained domain decomposition, which hinder the real-world application of current cross-domain models due to the lack of accurate and fair evaluation of fine-grained realistic domains. These insights motivate us to establish a novel realistic benchmark for TF-ADG. Benefiting from the introduced hierarchical definition of domain shifts, our proposed dataset DomainVerse addresses these issues by providing about 0.5 million images from 390 realistic, hierarchical, and balanced domains, allowing for decomposition across multiple domains within each image. With the help of the constructed DomainVerse and VLMs, we further propose two algorithms called Domain CLIP and Domain++ CLIP for training-free adaptive domain generalization. Extensive and comprehensive experiments demonstrate the significance of the dataset and the effectiveness of the proposed methods.
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui
IEEE Trans. Multim.8
2024 Trust it or not: Confidence-guided automatic radiology report generation
Yixin Wang 0003, Zihao Lin 0003, Zhe Xu 0012, Jie Luo 0003, Jiang Tian, Zhongchao Shi, Lifu Huang, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002
Neurocomputing7
2024 Domain-Aware Graph Network for Bridging Multi-Source Domain Adaptation
abstract
Domain adaptation (DA) addresses the challenge of distribution discrepancy between the training and test data, while multi-source domain adaptation (MSDA) is particularly appealing for realistic scenarios. With the emergence of extensive unlabeled datasets, self-supervised learning has gained significant popularity in deep learning. It is noteworthy that multi-source domain adaptation and self-supervised learning share a common objective: leveraging unlabeled data to acquire more informative representations. However, conventional self-supervised learning encounters two main limitations. Firstly, the traditional pretext task falls to transfer fine-grained knowledge to downstream task with general representation learning. Secondly, the scheme of the same feature extractor with distinct prediction heads makes the cross-task knowledge exchange and information sharing ineffective. In order to tackle these challenges, we introduce a novel approach called Domain-Aware Graph Network (DAGNet). DAGNet utilizes a graph neural network as a bridge to facilitate efficient cross-task knowledge exchange. By employing a mask token strategy, we enhance the robustness of representations by selectively masking certain domain or self-supervised information. In terms of datasets, the uneven and style-based domain shifts in current datasets make it challenging to measure the model's domain adaptation performance in real-world applications. To address this issue, we introduce a benchmark dataset DomainVerse with continuous spatio-temporal domain shifts encountered in the real world. Our extensive experiments demonstrate that DAGNet achieves state-of-the-art performance not only on mainstream multi-source domain adaptation datasets but also on different settings within DomainVerse. Code is available athttps://github.com/a791702141/SSG.
Feng Hou, Yang Zhang 0002, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui
IEEE Trans. Multim.5
2024 A Survey of Visual Transformers
abstract
Transformer, an attention-based encoder-decoder model, has already revolutionized the field of natural language processing (NLP). Inspired by such significant achievements, some pioneering works have recently been done on employing Transformer-liked architectures in the computer vision (CV) field, which have demonstrated their effectiveness on three fundamental CV tasks (classification, detection, and segmentation) as well as multiple sensory data stream (images, point clouds, and vision-language data). Because of their competitive modeling capabilities, the visual Transformers have achieved impressive performance improvements over multiple benchmarks as compared with modern convolution neural networks (CNNs). In this survey, we have reviewed over 100 of different visual Transformers comprehensively according to three fundamental CV tasks and different data stream types, where taxonomy is proposed to organize the representative methods according to their motivations, structures, and application scenarios. Because of their differences on training settings and dedicated vision tasks, we have also evaluated and compared all these existing visual Transformers under different configurations. Furthermore, we have revealed a series of essential but unexploited aspects that may empower such visual Transformers to stand out from numerous architectures, e.g., slack high-level semantic embeddings to bridge the gap between the visual Transformers and the sequential ones. Finally, two promising research directions are suggested for future investment. We will continue to update the latest articles and their released source codes at https://github.com/liuyang-ict/awesome-visual-transformers.
Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
IEEE Trans. Neural Networks Learn. Syst.8
2023 SAP-DETR: Bridging the Gap Between Salient Points and Queries-Based Transformer Detector for Fast Model Convergency
abstract
Recently, the dominant DETR-based approaches apply central-concept spatial prior to accelerating Transformer detector convergency. These methods gradually refine the reference points to the center of target objects and imbue object queries with the updated central reference information for spatially conditional attention. However, centralizing reference points may severely deteriorate queries' saliency and confuse detectors due to the indiscriminative spatial prior. To bridge the gap between the reference points of salient queries and Transformer detectors, we propose SAlient Point-based DETR (SAP-DETR) by treating object detection as a transformation from salient points to instance objects. Concretely, we explicitly initialize a query-specific reference point for each object query, gradually aggregate them into an instance object, and then predict the distance from each side of the bounding box to these points. By rapidly attending to query-specific reference regions and the conditional box edges, SAP-DETR can effectively bridge the gap between the salient point and the query-based Transformer detector with a significant convergency speed. Experimentally, SAP-DETR achieves 1.4× convergency speed with competitive performance and stably promotes the SoTA approaches by ∼1.0 AP. Based on ResNet-DC-101, SAP-DETR achieves 46.9 AP. The code will be released at https://github.com/liuyang-ict/SAP-DETR.
Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
CVPR6
2023 Learning How to Learn Domain-Invariant Parameters for Domain Generalization
abstract
Due to domain shift, deep neural networks (DNNs) usually fail to generalize well on unknown test data in practice. Domain generalization (DG) aims to overcome this issue by capturing domain-invariant representations from source domains. Motivated by the insight that only partial parameters of DNNs are optimized to extract domain-invariant representations, we expect a general model that is capable of well perceiving and emphatically updating such domain-invariant parameters. In this paper, we propose two modules of Domain Decoupling and Combination (DDC) and Domain-invariance-guided Backpropagation (DIGB), which can encourage such general model to focus on the parameters that have a unified optimization direction between pairs of contrastive samples. Our extensive experiments on two benchmarks have demonstrated that our proposed method has achieved state-of-the-art performance with strong generalization capability.
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
ICASSP7
2023 Dual Pseudo-Labels Interactive Self-Training for Semi-Supervised Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to match a specific person from a gallery of images captured from non-overlapping visible and infrared cameras. Most works focus on fully supervised VI-ReID, which requires substantial cross-modality annotation that is more expensive than the annotation in single-modality. To reduce the extensive cost of annotation, we explore two practical semi-supervised settings: uni-semi-supervised (annotating only visible images) and bi-semi-supervised (annotating partially in both modalities). These two semi-supervised settings face two challenges due to the large cross-modality discrepancies and the lack of correspondence supervision between visible and infrared images. Thus, it is diffi-cult to generate reliable pseudo-labels and learn modality-invariant features from noise pseudo-labels. In this paper, we propose a dual pseudo-label interactive self-training (DPIS) for these two semi-supervised VI-ReID. Our DPIS integrates two pseudo-labels generated by distinct models into a hybrid pseudo-label for unlabeled data. However, the hybrid pseudo-label still inevitably contains noise. To eliminate the negative effect of noise pseudo-labels, we introduce three modules: noise label penalty (NLP), noise correspondence calibration (NCC), and unreliable anchor learning (UAL). Specifically, NLP penalizes noise labels, NCC calibrates noisy correspondences, and UAL mines the hard-to-discriminate features. Extensive experimental results on SYSU-MM01 and RegDB demonstrate that our DPIS achieves impressive performance under these two semi-supervised settings.
Jiangming Shi, Yachao Zhang 0001, Xiangbo Yin, Yuan Xie 0006, Zhizhong Zhang 0001, Jianping Fan 0007, Zhongchao Shi, Yanyun Qu
ICCV7
2023 VS-Boost: Boosting Visual-Semantic Association for Generalized Zero-Shot Learning
abstract
Unlike conventional zero-shot learning (CZSL) which only focuses on the recognition of unseen classes by using the classifier trained on seen classes and semantic embeddings, generalized zero-shot learning (GZSL) aims at recognizing both the seen and unseen classes, so it is more challenging due to the extreme training imbalance. Recently, some feature generation methods introduce metric learning to enhance the discriminability of visual features. Although these methods achieve good results, they focus only on metric learning in the visual feature space to enhance features and ignore the association between the feature space and the semantic space. Since the GZSL method uses semantics as prior knowledge to migrate visual knowledge to unseen classes, the consistency between visual space and semantic space is critical. To this end, we propose relational metric learning which can relate the metrics in the two spaces and make the distribution of the two spaces more consistent. Based on the generation method and relational metric learning, we proposed a novel GZSL method, termed VS-Boost, which can effectively boost the association between vision and semantics. The experimental results demonstrate that our method is effective and achieves significant gains on five benchmark datasets compared with the state-of-the-art methods.
Xiaofan Li 0008, Yachao Zhang 0001, Shiran Bian, Yanyun Qu, Yuan Xie 0006, Zhongchao Shi, Jianping Fan 0007
IJCAI6
2023 Cross-modal Unsupervised Domain Adaptation for 3D Semantic Segmentation via Bidirectional Fusion-then-Distillation
abstract
Cross-modal Unsupervised Domain Adaptation (UDA) becomes a research hotspot because it reduces the laborious annotation of target domain samples. Existing methods only mutually mimic the outputs of cross-modality in each domain, which enforces the class probability distribution agreeable in different domains. However, these methods ignore the complementarity brought by the modality fusion representation in cross-modal learning. In this paper, we propose a cross-modal UDA method for 3D semantic segmentation via Bidirectional Fusion-then-Distillation, named BFtD-xMUDA, which explores cross-modal fusion in UDA and realizes distribution consistency between outputs of two domains not only for 2D image and 3D point cloud but also for 2D/3D and fusion. Our method contains three significant components: Model-agnostic Feature Fusion Module (MFFM), Bidirectional Distillation (B-Distill), and Cross-modal Debiased Pseudo-Labeling (xDPL). MFFM is employed to generate cross-modal fusion features for establishing a latent space, which enforces maximum correlation and complementarity between two heterogeneous modalities. B-Distill is introduced to exploit bidirectional knowledge distillation which includes cross-modality and cross-domain fusion distillation, and well-achieving domain-modality alignment. xDPL is designed to model the uncertainty of pseudo-labels by self-training scheme. Extensive experimental results demonstrate that our method outperforms state-of-the-art competitors in several adaptation scenarios.
Mingwei Xing, Yachao Zhang 0001, Yuan Xie 0006, Jianping Fan 0007, Zhongchao Shi, Yanyun Qu
ACM Multimedia6
2023 Hardware-friendly Scalable Image Super Resolution with Progressive Structured Sparsity
abstract
Single image super-resolution (SR) is an important low-level vision task, and the dynamic SR trading off performance and efficiency are increasingly in demand. The existing dynamic SR methods are divided into two classes: the structured pruning and non-structured compressing methods. The former removes redundant structures in the network, which often leads to significant performance degradation, and the latter searches for extremely sparse parameter masks, achieving promising performance, but they are not deployable in hardware platforms with irregular memory access. In order to solve the mentioned problems, we propose Hardware-friendly Scalable SR (HSSR) with progressively structured sparsity. The superiority of our method is that with only a single scalable model it covers multiple SR models with different sizes, without extra retraining or post-processing. HSSR contains the forward and backward processing. In the forward process, we gradually shrink the SR networks with structured iterative sparsity where grouping convolution together with knowledge distillation is conducted to reduce the amount of SR parameters and the computational complexity while keeping the performance, and in the backward process, we gradually expand the compressed SR networks with structured iterative recovery. Comprehensive experiments on benchmark datasets show that HSSR is perfectly compatible with common convolution baselines. Compared with the Slimmable method, our model is superior in performance, flops, and model size. Experimental results demonstrate that HSSR achieves significant compression, saving up to 1500K parameters and 100 GFlops calculation compared to the original model in real-world applications.
Fangchen Ye, Hongzhan Huang, Jianping Fan 0007, Zhongchao Shi, Yuan Xie 0006, Yanyun Qu
ACM Multimedia5
2023 Balanced masking strategy for multi-label image classification
Yao Zhang 0010, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui
Neurocomputing3
2023 Graph Attention Transformer Network for Multi-label Image Classification
abstract
Multi-label classification aims to recognize multiple objects or attributes from images. The key to solving this issue relies on effectively characterizing the inter-label correlations or dependencies, which bring the prevailing graph neural network. However, current methods often use the co-occurrence probability of labels based on the training set as the adjacency matrix to model this correlation, which is greatly limited by the dataset and affects the model’s generalization ability. This article proposes a Graph Attention Transformer Network, a general framework for multi-label image classification by mining rich and effective label correlation. First, we use the cosine similarity value of the pre-trained label word embedding as the initial correlation matrix, which can represent richer semantic information than the co-occurrence one. Subsequently, we propose the graph attention transformer layer to transfer this adjacency matrix to adapt to the current domain. Our extensive experiments have demonstrated that our proposed methods can achieve highly competitive performance on three datasets.
Shikai Chen, Yao Zhang 0010, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Cross-Domain Few-Shot Learning for Rare-Disease Skin Lesion Segmentation
abstract
Recently, deep learning (DL)-based skin lesion segmentation in dermoscopic images has advanced the efficient diagnosis of skin diseases. Commonly, most of the DL-based methods require a large amount of training data and can only perform accurate predictions on pre-defined classes. However, there exist some rare skin diseases with very limited labeled samples, which poses great challenges to typical DL-based methods. Few-shot learning (FSL) technique, which aims to train models with abundant seen classes and then generalizes to related unseen classes, is promising in addressing a similar problem. Unfortunately, simply borrowing the typical FSL is infeasible since collecting such abundant seen-class data (common skin diseases), is also difficult. In this paper, we propose a cross-domain few-shot segmentation (CD-FSS) framework, which enables the model to leverage the learning ability obtained from the natural domain, to facilitate rare-disease skin lesion segmentation with limited data of common diseases. Specifically, the framework consists of two processes, i.e., specific learning and generic learning, which are alternately optimized in a meta-training manner. A specific learner and a generic learner are tailored to build relationships between both processes. Experimental results demonstrate that our framework significantly improves the generalization ability from natural domain to unseen medical domain.
Yixin Wang 0003, Zhe Xu 0012, Jiang Tian, Jie Luo 0003, Zhongchao Shi, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002
ICASSP5
2022 Semi-Supervised 3D Medical Image Segmentation Via Boundary-Aware Consistent Hidden Representation Learning
abstract
This paper proposes a novel Boundary-aware Consistent Hidden Representation Learning Network (BA-CHRLN), which contains two branches for semi-supervised 3D medical image segmentation. Inspired by the contrastive learning, the two branches share the same encoder and each has its individual decoder, namely supervised decoder and unsupervised one. A stop-gradient operation is also utilized to prevent collapsing of solutions. Taking the unlabeled images as references, BA-CHRLN imposes the consistency by applying a perturbation on the high-level hidden feature representations, which significantly improves the encoder’s representation and the network’s robustness. A boundary-aware map is further introduced to capture the organ’s boundary without any prior knowledge and additional parameters. Experiments on the Left Atrium (LA) benchmark dataset demonstrate the effectiveness of the BA-CHRLN.
Linhu Liu, Jiang Tian, Xiangqian Cheng, Zhongchao Shi, Jianping Fan 0007, Yong Rui
ICIP4
2022 Disentangled Neural Architecture Search
abstract
Neural architecture search (NAS) has emerged as a hot topic recently which makes artificial intelligence techniques easier to apply and reduces the demand for experts knowledge by generating deep neural network architectures automatically. However, most existing methods for neural architecture search (NAS) heavily rely on underlying black-box controllers to generate potential candidates of network architectures, and suffer from serious problems of lacking interpretability and search efficiency. In this paper, we propose Disentangled Neural Architecture Search (DNAS), which addresses the two issues by adopting disentangled NAS controller and efficient dense-sampling strategy. Specifically, DNAS learns disentangled factors of network architecture by explicitly encouraging the latent factors to be independent. This approach could not only achieve semantic interpretability but also allow us to conveniently identify the promising regions of representations corresponding to high-performance architectures. We further propose a dense-sampling strategy that conducts targeted architecture search within the promising regions to accelerate the searching process. Our DNAS owns several attractive features: 1) it can successfully learn semantic representations of architectures, including operation selection, skip connections, and layer order; 2) it can speed up the process for neural architecture search more than 13× by using dense-sampling and disentangled factors; 3) it can achieve higher accuracy under less computational cost—DNAS achieves state-of-the-art performance of 94.16% on NASBench-101, and 22.7% top-1 test error on ImageNet with 1.6 GPU-days.
Xinyue Zheng, Peng Wang 0095, Qigang Wang, Zhongchao Shi, Jianping Fan 0007
IJCNN4
2022 Self-Supervised Graph Neural Network for Multi-Source Domain Adaptation
abstract
Domain adaptation (DA) tries to tackle the scenarios when the test data does not fully follow the same distribution of the training data, and multi-source domain adaptation (MSDA) is very attractive for real world applications. By learning from large-scale unlabeled samples, self-supervised learning has now become a new trend in deep learning. It is worth noting that both self-supervised learning and multi-source domain adaptation share a similar goal: they both aim to leverage unlabeled data to learn more expressive representations. Unfortunately, traditional multi-task self-supervised learning faces two challenges: (1) the pretext task may not strongly relate to the downstream task, thus it could be difficult to learn useful knowledge being shared from the pretext task to the target task; (2) when the same feature extractor is shared between the pretext task and the downstream one and only different prediction heads are used, it is ineffective to enable inter-task information exchange and knowledge sharing. To address these issues, we propose a novel Self-Supervised Graph Neural Network (SSG), where a graph neural network is used as the bridge to enable more effective inter-task information exchange and knowledge sharing. More expressive representation is learned by adopting a mask token strategy to mask some domain information. Our extensive experiments have demonstrated that our proposed SSG method has achieved state-of-the-art results over four multi-source domain adaptation datasets, which have shown the effectiveness of our proposed SSG method from different aspects.
Feng Hou, Yangzhou Du, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui
ACM Multimedia4
2022 Semi-supervised Medical Image Segmentation with Semantic Distance Distribution Consistency Learning
Linhu Liu, Jiang Tian, Zhongchao Shi, Jianping Fan 0007
PRCV (2)3
2021 AKFNET: An Anatomical Knowledge Embedded Few-Shot Network For Medical Image Segmentation
abstract
Automated organ segmentation in CTs is an essential prerequisite for many clinical applications, such as computer-aided diagnosis and intervention. As medical data annotation requires massive human labor from experienced radiologists, how to effectively improve the segmentation performance with limited annotated training data remains a challenging problem. Few-shot learning imitates the learning process of humans, which turns out to be a promising way to overcome the aforementioned challenge. In this paper, we propose a novel anatomical knowledge embedded few-shot network (AKFNet), where an anatomical knowledge embedded support unit (AKSU) is carefully designed to embed the anatomical priors from support images into our model. Moreover, a similarity guidance alignment unit (SGAU) is proposed to impose a mutual alignment between the support and query sets. As a result, AKFNet fully exploits anatomical knowledge and presents good learning capability. Without bells and whistles, AKFNet outperforms the state-of-the-art methods with 0.84-1.76% Dice increase. Transfer learning experiments further verify its learning capability.
Yanan Wei, Jiang Tian, Zhongchao Shi
ICIP4
2021 ACN: Adversarial Co-training Network for Brain Tumor Segmentation with Missing Modalities
Yixin Wang 0003, Yang Zhang 0002, Yang Liu 0250, Zihao Lin 0003, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002
MICCAI (7)7
2021 Modality-Aware Mutual Learning for Multi-modal Medical Image Segmentation
Yao Zhang 0010, Jiawei Yang 0002, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002
MICCAI (1)4
2020 Label Distribution Learning on Auxiliary Label Space Graphs for Facial Expression Recognition
abstract
Many existing studies reveal that annotation inconsistency widely exists among a variety of facial expression recognition (FER) datasets. The reason might be the subjectivity of human annotators and the ambiguous nature of the expression labels. One promising strategy tackling such a problem is a recently proposed learning paradigm called Label Distribution Learning (LDL), which allows multiple labels with different intensity to be linked to one expression. However, it is often impractical to directly apply label distribution learning because numerous existing datasets only contain one-hot labels rather than label distributions. To solve the problem, we propose a novel approach named Label Distribution Learning on Auxiliary Label Space Graphs(LDL-ALSG) that leverages the topological information of the labels from related but more distinct tasks, such as action unit recognition and facial landmark detection. The underlying assumption is that facial images should have similar expression distributions to their neighbours in the label space of action unit recognition and facial landmark detection. Our proposed method is evaluated on a variety of datasets and outperforms those state-of-the-art methods consistently with a huge margin.
Shikai Chen, Yuedong Chen, Zhongchao Shi, Xin Geng 0001, Yong Rui
CVPR4
2020 Selecting Useful Knowledge from Previous Tasks for Future Learning in a Single Network
abstract
Continual learning can learn new tasks incrementally while avoiding catastrophic forgetting. Recent work has shown that packing multiple tasks into a single network incrementally by iterative pruning and re-training network is a promising method. We build upon this idea and propose an improved version of PackNet. Specifically, we propose a novel gradient-based threshold method to reuse the knowledge of the previous tasks selectively when learning new tasks. Our experiments on a variety of classification tasks and different network architectures demonstrate that our method obtains competitive results when compared to PackNet.
Feifei Shi, Peng Wang 0095, Zhongchao Shi, Yong Rui
ICPR3
2020 DARN: Deep Attentive Refinement Network for Liver Tumor Segmentation from 3D CT volume
abstract
Automatic liver tumor segmentation from 3D Computed Tomography (CT) is a necessary prerequisite in the interventions of hepatic abnormalities and surgery planning. However, accurate liver tumor segmentation remains challenging due to the large variability of tumor sizes and inhomogeneous texture. Recent advances based on Fully Convolutional Network (FCN) in liver tumor segmentation draw on success of learning discriminative multi-level features. In this paper, we propose a Deep Attentive Refinement Network (DARN) for improved liver tumor segmentation from CT volumes by fully exploiting both low and high level features embedded in different layers of FCN. Different from existing works, we exploit attention mechanism to leverage the relation of different levels of features encoded in different layers of FCN. Specifically, we introduce a Semantic Attention Refinement (SemRef) module to selectively emphasize global semantic information in low level features with the guidance of high level ones, and a Spatial Attention Refinement (SpaRef) module to adaptively enhance spatial details in high level features with the guidance of low level ones. We evaluate our network on the public MICCAI 2017 Liver Tumor Segmentation Challenge dataset (LiTS dataset) and it achieves state-of-the-art performance. The proposed refinement modules are an effective strategy to exploit multi-level features and has great potential to generalize to other medical image segmentation tasks.
Yao Zhang 0010, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Zhiqiang He 0002
ICPR5
2020 Double-Uncertainty Weighted Method for Semi-supervised Learning
Yixin Wang 0003, Yao Zhang 0010, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002
MICCAI (1)5
2020 Algorithm Bias Detection and Mitigation in Lenovo Face Recognition Engine
Sheng Shi, Shanshan Wei, Zhongchao Shi, Yangzhou Du, Jianping Fan 0007, Yolanda Conyers
NLPCC (2)3
2020 Challenge Closed-Book Science Exam: A Meta-Learning Based Question Answering System
Xinyue Zheng, Peng Wang 0095, Qigang Wang, Zhongchao Shi
PKAW4
2019 Discrimination Assessment for Saliency Maps
Ruiyi Li, Yangzhou Du, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002
NLPCC (2)3
2019 Efficient Automatic Meta Optimization Search for Few-Shot Learning
Xinyue Zheng, Peng Wang 0095, Qigang Wang, Zhongchao Shi
PRCV (3)4
2019 Facial Motion Prior Networks for Facial Expression Recognition
abstract
Deep learning based facial expression recognition (FER) has received a lot of attention in the past few years. Most of the existing deep learning based FER methods do not consider domain knowledge well, which thereby fail to extract representative features. In this work, we propose a novel FER framework, named Facial Motion Prior Networks (FMPN). Particularly, we introduce an addition branch to generate a facial mask so as to focus on facial muscle moving regions. To guide the facial mask learning, we propose to incorporate prior domain knowledge by using the average differences between neutral faces and the corresponding expressive faces as the training guidance. Extensive experiments on three facial expression benchmark datasets demonstrate the effectiveness of the proposed method, compared with the state-of-the-art approaches.
Yuedong Chen, Shikai Chen, Zhongchao Shi, Jianfei Cai 0001
VCIP4
2018 SequentialSegNet: Combination with Sequential Feature for Multi-Organ Segmentation
abstract
Multi-organ segmentation from computed tomography (CT) images is essential for computer aided diagnosis (CAD), and recent advances in fully convolutional networks (FCNs) for volumetric image segmentation have demonstrated the importance of leveraging spatial information. In this paper, we propose a novel framework called SequentialSegNet, which efficiently combines features within a single CT image (intra-slice) and among multiple adjacent images (inter-slice) for a multi-organ segmentation. Experimental results show that our approach can effectively improve the segmentation performance on both large-size and small-size abdominal organs including liver, spleen and gallbladder.
Yao Zhang 0010, Yang Zhang 0002, Zhongchao Shi, Zhensheng Li, Zhiqiang He 0002
ICPR5
2018 A Diagnostic Report Generator from CT Volumes on Liver Tumor with Semi-supervised Attention Mechanism
Jiang Tian, Zhongchao Shi, Feiyu Xu 0001
MICCAI (2)3
2018 Dynamic Delay Based Cyclic Gradient Update Method for Distributed Training
Wenhui Hu, Peng Wang 0095, Qigang Wang, Zhengdong Zhou, Hui Xiang, Zhongchao Shi
PRCV (3)7
2017 Video Captioning with Listwise Supervision
abstract
Automatically describing video content with natural language is a fundamental challenging that has received increasing attention. However, existing techniques restrict the model learning on the pairs of each video and its own sentences, and thus fail to capture more holistically semantic relationships among all sentences. In this paper, we propose to model relative relationships of different video-sentence pairs and present a novel framework, named Long Short-Term Memory with Listwise Supervision (LSTM-LS), for video captioning. Given each video in training data, we obtain a ranking list of sentences w.r.t. a given sentence associated with the video using nearest-neighbor search. The ranking information is represented by a set of rank triplets that can be used to assess the quality of ranking list. The video captioning problem is then solved by learning LSTM model for sentence generation, through maximizing the ranking quality over all the sentences in the list. The experiments on MSVD dataset show that our proposed LSTM-LS produces better performance than the state of the art in generating natural sentences: 51.1% and 32.6% in terms of BLEU@4 and METEOR, respectively. Superior performances are also reported on the movie description M-VAD dataset.
Zhongchao Shi
AAAI3
2016 Boosting Video Description Generation by Explicitly Translating from Frame-Level Captions
abstract
Automatically describing video content with natural language is a fundamental challenge of computer vision. The recent advanced technique that approaches this problem is Recurrent Neural Networks (RNN). The need to train RNN on large-scale complex and diverse videos and their associated language, however, makes the task human-labeling intensive and computationally expensive. Moreover, the results can suffer from robustness problem, especially when there are rich of temporal dynamics in the sequence of video frames. We demonstrate in this paper that the above two limitations can be mitigated by jointly exploring the largely available data from image domain and representing each frame by high-level attributes rather than visual features. The former leverages the learnt models on image captioning benchmark to generate caption for each video frame, while the latter explicitly incorporates the obtained captions which are regarded as the attributes of each frame. Specifically, we propose a novel sequence to sequence architecture to generate descriptions for videos, in a sense that the inputs are the captions of sequential frames and it outputs words sequentially. On a widely used YouTube2Text dataset, our proposal is shown to be powerful with superior performance over several state-of-the-art methods including both architectures that are purely developed on video data and RNN-based models which translate directly from visual features to language.
Zhongchao Shi
ACM Multimedia2
2015 Click-through-based Deep Visual-Semantic Embedding for Image Search
abstract
The problem of image search is mostly considered from the perspectives of feature-based vector model and image ranker learning. A fundamental issue that underlies the success of these approaches is the similarity learning between query and image. The need of image surrounding texts in feature-based vector model, however, makes the similarity sensitive to the quality of text descriptions. On the other, the image ranker learning can suffer from robustness problem, originating from the fact that human labeled query-image pairs do not always predict user search intention precisely. We demonstrate in this paper that the above two issues can be well mitigated by jointly exploring visual-semantic embedding and the use of click-through data. Specifically, we propose a novel click-through-based deep visual-semantic embedding (C-DVSE) model for learning query and image similarity. The proposed model consists of two components: a deep convolutional neural networks followed by an image embedding layer for learning visual embedding, and a deep neural networks for generating query semantic embedding. The objective of our model is to maximize the correlation between semantic (query) and visual (clicked image) embedding. When the visual-semantic embedding is learnt, query-image similarity can be directly computed by cosine similarity on this embedding space. On a large-scale click-based image dataset with 11.7 million queries and one million images, our model is shown to be powerful for keyword-based image search with superior performance over several state-of-the-art methods.
Zhongchao Shi
ACM Multimedia2
2015 Learning query and image similarities with listwise supervision
abstract
One of the fundamental problems in image search is to learn the ranking functions, i.e., similarity between the textual query and visual image. A number of research paradigms, ranging from feature-based vector model to image ranker learning, have been applied to measure query-image similarity. However, most of the existing similarity learning methods either depend on surrounding texts for ranking images or learn image rankers to satisfy pairwise or tripletwise supervision. In this paper, we propose to leverage listwise supervision into a principled click-through-based query-image similarity learning framework. In particular, the algorithm utilizes click counts for each image in response to a query to get a ranking list. The ranking information is represented by a set of rank triplets that can be used to assess the quality of the ranking list. The image ranking problem is then solved efficiently by learning two linear projections for query and image space respectively, through maximizing the ranking quality over all the training data. When the two linear projections are learnt, query-image similarity can be directly computed by dot product on this projected subspace. On a large-scale click-through-based image dataset with 11.7 million queries and one million images, our learnt model via listwise supervision is shown to be powerful for keyword-based image search with superior performance over several state-of-the-art methods.
Zhongchao Shi
MMSP2
2015 Fusion of a panoramic camera and 2D laser scanner data for constrained bundle adjustment in GPS-denied environments
Shunping Ji, Xiaowei Shao, Peng Yang 0005, Zhongchao Shi, Ryosuke Shibasaki
Image Vis. Comput.6
2005 Remote sensing monitoring of grassland in key region of Northern China
Bin Xu 0008, Zhongchao Shi, Xiuchun Yang, Guixia Yang
IGARSS4
2004 A novel fingerprint matching method based on the Hough transform without quantization of the Hough space
abstract
In contrast to the conventional usage of the Hough transform for fingerprint alignment, our novel Hough transform-based fingerprint registration algorithm proposed in this paper doesn't need to discretize the Hough space when estimating the parameters of the rigid transformation between two fingerprint impressions. In our algorithm, the parameters (rotation, x translation and y translation) of the transform (rigid transform considered) can be estimated separately. In addition, the orientation fields from the fingerprints, instead of the minutiae, are used to estimate the most optimal rotation parameter. The experimental results obtained on the public domain collections of fingerprint images, FVC2002 db2 set A (800 fingerprints) and db3 set A (800 fingerprints), show that the new fingerprint registration algorithm presented in this paper is faster and more effective as compared to the conventional generalized Hough transform-based fingerprint alignment method.
Zhongchao Shi, Xuying Zhao, Yangsheng Wang
ICIG2
2004 A new segmentation algorithm for low quality fingerprint image
abstract
Aiming at the segmentation of low quality fingerprint images, a new framework, which is different from traditional methods that usually use some certain features to segment diversified images, is proposed. There are two contributions in this scheme: firstly, we introduce a quality estimation step before segmentation, which can remove a great many false traces effectively; secondly, a new feature eccentric moment is proposed to locate the blurry boundary. Then we segment the image using the new block feature of clarified image. Experimental results show that the proposed method can segment low quality fingerprint images properly.
Zhongchao Shi, Yangsheng Wang
ICIG1
2004 Remote sensing monitoring on dynamic status of grassland productivity and animal loading balance in Northern China
abstract
The study region includes 312 counties of 11 provinces and autonomous regions in Northern China. The methodology is mainly the combination of remote sensing, geographic information system and spatial database technology with field investigation of the grassland. The main conclusions are as follows. (1) the monitoring indicates that grass production of the major grasslands in Northern China is decreasing in recent decades, at a rate of 10-40%. (2) Compared with the grass production, over-grazing is a common phenomenon in grazing region. Among the 153 counties in the grazing region, over 50% is with over-loading of herd. Similar phenomenon is also observed in semi-grazing region where over 80% of the 159 counties have various degrees of herd over-loading. (3) Over-loading is severe in semi-grazing area than in grazing area. Comparison has been done to the loading balance between grazing and semi-grazing regions by the end of 1990's. In grazing region. there were 12 counties (banners) with extreme over-loading, accounting for 7.79% of total county numbers in this region. In semi-grazing region, there were 82 counties showing sign of overloading, accounting for 51.5% of total county numbers in this region. The proportion of extreme over-loading grazing region to the total regions occupies 4.82% in grazing region. This proportion is 34.85% in semi-grazing. Obviously, the severity of extreme overloading in semi-grazing highly surpasses that in grazing region. Understanding this animal loading difference among various types of region may help to facilitate proper administration of grazing for sustainable development of the grassland.
Bin Xu 0008, Xiaoping Xin, Zhongchao Shi, Haiqi Liu, Zhongxin Chen, Guixia Yang, Youqi Chen
IGARSS4