EDBT 2026 Demo / reviewers in the wild / expert
Zhongchao Shi
dblp:45/5323
· DBLP profile ↗
54ranked-venue papers
1as first author
33since 2021 · last 2026
0000-0002-5216-3827ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 1 first-author · 21 since 2021Artificial intelligence and machine learning · 24 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SlimNet: High-Quality and Efficient Object Removal by Eliciting Latent Capabilities of Diffusion ModelsabstractObject removal aims to eliminate undesired objects from images while plausibly restoring the underlying content with high visual fidelity. Existing diffusion-based methods often achieve strong results by relying on large-scale model finetuning, auxiliary control networks, or powerful diffusion backbones, which incur substantial computational overhead and limit practical efficiency. In this work, we propose SlimNet, a high-quality and efficient object removal framework that elicits the latent object removal capabilities of pretrained diffusion models. SlimNet injects lightweight adapters into a frozen diffusion backbone to selectively modulate intermediate representations, avoiding heavy architectural modifications. To further improve visual fidelity, we design a composite perceptual loss that enforces object removal accuracy, background preservation, and smooth boundary transitions, together with a semantic-aware data processing pipeline for automatic mask generation. Extensive experiments demonstrate that SlimNet achieves competitive removal quality compared to state-of-the-art methods, with favorable reductions in inference time and VRAM consumption, making it a practical solution for resource-constrained multimedia applications. Yao Zhang 0010, Zhongchao Shi, Jianping Fan 0007, Guihua Zeng |
ICMR | 5 |
| 2026 | A Bootstrap Pipeline for Chat-Based Image Retrieval with Effective Question GenerationabstractChat-based image retrieval uses Large Language Models (LLM) to guide user input to enable more specific and precise search results, where LLM can enhance this process by asking user retrieval-oriented questions eliciting additional details about the target image. Despite the potential of this approach, no specialized Questioner model has been developed for this task due to the following significant challenges: (a) the difficulty of determining the optimal questions to ask; (b) the lack of a suitable protocol for fair model comparison; and (c) the notable scarcity of dialog-to-image retrieval data. To address these challenges, two fundamental principles are developed in this article to ensure the simplicity and effectiveness of the generated questions while enabling a fair comparison and accurate estimation of data quality and model performance. A bootstrap training methodology is introduced to collect retrieval-oriented dialog data and concurrently train the Questioner and the image Retriever. Under a fair comparison protocol, our extensive experiments have demonstrated that our proposed method can not only address the critical data gap but also achieve state-of-the-art results, which substantially surpass GPT-4o and GPT-4-Turbo through the fine-tuning of an 8B model. Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with AttributesabstractRecent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capacity for fine-grained linguistic comprehension and leads to a significant decline in performance when faced with detailed descriptions or contextual information. To tackle these problems, we develop DoGA: Detect objects with Grouped Attributes, which employs commonly apparent attributes to bridge different granular semantics and uses specific attributes to identify the object discrepancy. Our DoGA incorporates three principle components: 1) Generation of attribute-based prompts, consisting of linguistic definitions enriched with common-sense visible attributes and hard negative notations deriving from the image-specific attribute features; 2) Paralleled entity fusion and optimization, designed to manage long attribute-based descriptions and negative concepts efficiently; and 3) Prompt-wise grouped training to accommodate model to perform many-to-many assignments, facilitating simultaneous training and inferring with multiple attribute-based synonyms. Extensive experiments demonstrate that training with synonymous attribute-based prompts allows DoGA to generalize multi-granular prompts and surpass previous state-of-the-art approaches, yielding 50.2 on the COCO and 38.0 on the LVIS benchmarks under the zero-short setting. We will make our code publicly available upon acceptance. Yang Liu 0250, Feng Hou, Yunjie Peng, Gangjian Zhang, Yao Zhang 0010, Peng Wang 0095, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
AAAI | 10 |
| 2025 | Omni-Query Active Learning for Source-Free Domain Adaptive Cross-Modality 3D Semantic SegmentationabstractSource-Free Domain Adaptation (SFDA) aims to transfer a pre-trained source model to the unlabeled target domain without accessing the source data, thereby effectively solving labeled data dependency and domain shift problems. However, the SFDA setting faces a bottleneck due to the absence of supervisory information. To mitigate this problem, Active Learning (AL) is introduced to combine with SFDA, endeavoring to actively label a small set of the most high-quality target points so that models with satisfactory performance can be obtained at an acceptable cost. Nevertheless, several issues remain unresolved, namely when to query new labels during training, what kind of samples deserve labeling to ensure rich information, and where the labels should be distributed to guarantee diversity. Thus we elaborate OmniQuery to omnibearing address the “When, What, and Where” problems about active points querying in source-free domain adaptation for cross-modal 3D semantic segmentation. The method consists of three main components: Query Decider, Point Ranker, and Budget Slicer. The Query Decider determines the optimal timing to query new points by fitting the validation curves during training. The Point Ranker nominates points for annotation by calculating the ambiguity of neighboring points in the feature space. The Budget Slicer allocates the annotation quota, i.e., labeling percentage of the point cloud, to different semantic regions by utilizing the advanced 2D semantic segmentation capabilities of the Segment Anything Model (SAM). Extensive experiments demonstrate the effectiveness of our proposed method, achieving up to 99.64% of fully supervised performance with only 3% of labels, and consistently outperforming comparison methods across various scenarios. Jianxiang Xie, Yachao Zhang 0001, Zhongchao Shi, Jianping Fan 0007, Yuan Xie 0006, Yanyun Qu |
AAAI | 4 |
| 2025 | CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative DrafterabstractYepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen, Li Liu, Jiang Tian, Zhongchao Shi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen, Jiang Tian, Zhongchao Shi |
ACL (1) | 7 |
| 2025 | Analyzing the Effects of Supervised Fine-Tuning on Model Knowledge from Token and Parameter LevelsabstractJunjie Ye, Yuming Yang, Yang Nan, Shuo Li, Qi Zhang, Tao Gui, Xuanjing Huang, Peng Wang, Zhongchao Shi, Jianping Fan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Junjie Ye 0005, Yuming Yang 0001, Qi Zhang 0001, Tao Gui, Xuanjing Huang 0001, Peng Wang 0095, Zhongchao Shi, Jianping Fan 0007 |
EMNLP | 9 |
| 2025 | ARB-LLM: Alternating Refined Binarizations for Large Language ModelsabstractLarge Language Models (LLMs) have greatly pushed forward advancements in natural language processing, yet their high memory and computational demands hinder practical deployment. Binarization, as an effective compression technique, can shrink model weights to just 1 bit, significantly reducing the high demands on computation and memory. However, current binarization methods struggle to narrow the distribution gap between binarized and full-precision weights, while also overlooking the column deviation in LLM weight distribution. To tackle these issues, we propose ARB-LLM, a novel 1-bit post-training quantization (PTQ) technique tailored for LLMs. To narrow the distribution shift between binarized and full-precision weights, we first design an alternating refined binarization (ARB) algorithm to progressively update the binarization parameters, which significantly reduces the quantization error. Moreover, considering the pivot role of calibration data and the column deviation in LLM weights, we further extend ARB to ARB-X and ARB-RC. In addition, we refine the weight partition strategy with column-group bitmap (CGB), which further enhance performance. Equipping ARB-X and ARB-RC with CGB, we obtain ARB-LLM$_{\text{X}}$ and ARB-LLM$ _{\text{RC}} $ respectively, which significantly outperform state-of-the-art (SOTA) binarization methods for LLMs.
As a binary PTQ method, our ARB-LLM$ _{\text{RC}} $ is the first to surpass FP16 models of the same size. Code: https://github.com/ZHITENGLI/ARB-LLM. Zhiteng Li, Xianglong Yan, Tianao Zhang, Haotong Qin, Jiang Tian, Zhongchao Shi, Linghe Kong, Yulun Zhang 0001, Xiaokang Yang 0001 |
ICLR | 7 |
| 2025 | DenseGrounding: Improving Dense Language-Vision Semantics for Ego-centric 3D Visual GroundingabstractEnabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D visual grounding, where agents locate target objects in real-world 3D spaces based on verbal descriptions. However, this task faces two significant challenges: (1) loss of fine-grained visual semantics due to sparse fusion of point clouds with ego-centric multi-view images, (2) limited textual semantic context due to arbitrary language descriptions. We propose DenseGrounding, a novel approach designed to address these issues by enhancing both visual and textual semantics. For visual features, we introduce the Hierarchical Scene Semantic Enhancer, which retains dense semantics by capturing fine-grained global scene features and facilitating cross-modal alignment. For text descriptions, we propose a Language Semantic Enhancer that leverage large language models to provide rich context and diverse language descriptions with additional context during model training. Extensive experiments show that DenseGrounding significantly outperforms existing methods in overall accuracy, achieving improvements of **5.81%** and **7.56%** when trained on the comprehensive full training dataset and smaller mini subset, respectively, further advancing the SOTA in ego-centric 3D visual grounding. Our method also achieves **1st place** and receives **Innovation Award** in the 2024 Autonomous Grand Challenge Multi-view 3D Visual Grounding Track, validating its effectiveness and robustness. Henry Zheng, Qihang Peng, Yong Xien Chng, Rui Huang 0012, Yepeng Weng, Zhongchao Shi, Gao Huang 0001 |
ICLR | 7 |
| 2025 | Traversal Verification for Speculative Tree DecodingabstractSpeculative decoding is a promising approach for accelerating large language models. The primary idea is to use a lightweight draft model to speculate the output of the target model for multiple subsequent timesteps, and then verify them in parallel to determine whether the drafted tokens should be accepted or rejected. To enhance acceptance rates, existing frameworks typically construct token trees containing multiple candidates in each timestep. However, their reliance on token-level verification mechanisms introduces two critical limitations: First, the probability distribution of a sequence differs from that of individual tokens, leading to suboptimal acceptance length. Second, current verification schemes begin from the root node and proceed layer by layer in a top-down manner. Once a parent node is rejected, all its child nodes should be discarded, resulting in inefficient utilization of speculative candidates. This paper introduces Traversal Verification, a novel speculative decoding algorithm that fundamentally rethinks the verification paradigm through leaf-to-root traversal. Our approach considers the acceptance of the entire token sequence from the current node to the root, and preserves potentially valid subsequences that would be prematurely discarded by existing methods. We theoretically prove that the probability distribution obtained through Traversal Verification is identical to that of the target model, guaranteeing lossless inference while achieving substantial acceleration gains. Experimental results on various models and multiple tasks demonstrate that our method consistently improves acceptance length and throughput over token-level verification. Yepeng Weng, Qiao Hu 0001, Xujie Chen, Dianwen Mei, Huishi Qiu, Jiang Tian, Zhongchao Shi |
NeurIPS | 8 |
| 2025 | Learning Preference Distributions: A Label-Side Paradigm for Explainable Reward ModelsabstractThe reward model is a critical component in training powerful large language models. However, current methods largely overlook the inherent subjectivity and variability in human preferences. Typically, this issue is indirectly addressed through model-side approaches, such as ensemble methods to estimate uncertainty from multiple predictions or uncertainty-aware regression to predict mean and variance. These indirect approaches fail to capture the intrinsic distributional characteristics and inter-rater disagreements present in human judgments. In contrast, we propose a direct, label-side solution by explicitly modeling human preference distributions. We recover missing information from scalar ratings to construct meaningful distribution labels. By employing Label Distribution Learning (LDL), each dimension of the resulting multidimensional discrete distribution explicitly corresponds to a specific preference score, naturally reflecting the subjective and multidimensional nature of human evaluations. Our approach improves explainability and confidence estimation, while also enabling more effective data selection and sample-efficient test-time adaptation. Empirical results demonstrate that our method not only achieves state-of-the-art performance but also provides a robust framework for uncertainty quantification and nuanced preference modeling. Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui |
SMC | 4 |
| 2025 | Collective domain adversarial learning for unsupervised domain adaptation
Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Xin Geng 0001, Yong Rui |
Frontiers Comput. Sci. | 4 |
| 2025 | Epistemic graph: A plug-and-play module for hybrid representation learningabstractIn recent years, deep models have achieved remarkable success in various vision tasks. However, their performance heavily relies on large training datasets. In contrast, humans exhibit hybrid learning, seamlessly integrating structured knowledge for cross-domain recognition or relying on a smaller amount of data samples for few-shot learning. Motivated by this human-like epistemic process, we aim to extend hybrid learning to computer vision tasks by integrating structured knowledge with data samples for more effective representation learning. Nevertheless, this extension faces significant challenges due to the substantial gap between structured knowledge and deep features learned from data samples, encompassing both dimensions and knowledge granularity. In this paper, a novel Epistemic Graph Layer (EGLayer) is introduced to enable hybrid learning, enhancing the exchange of information between deep features and a structured knowledge graph. Our EGLayer is composed of three major parts, including a local graph module, a query aggregation model, and a novel correlation alignment loss function to emulate human epistemic ability. Serving as a plug-and-play module that can replace the standard linear classifier, EGLayer significantly improves the performance of deep models. Extensive experiments demonstrate that EGLayer can greatly enhance representation learning for the tasks of cross-domain recognition and few-shot learning, and the visualization of knowledge graphs can aid in model interpretation. Shikai Chen, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Yong Rui |
Neurocomputing | 5 |
| 2025 | Gradient-aware domain-invariant learning for domain generalization
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
Multim. Syst. | 7 |
| 2025 | DomainVerse: A Benchmark Towards Real-World Distribution Shifts for Training-Free Adaptive Domain GeneralizationabstractTraditional cross-domain tasks, including unsupervised domain adaptation (UDA), domain generalization (DG) and test-time adaptation (TTA), rely heavily on the training model by source domain data whether for specific or arbitrary target domains. With the recent advance of vision-language models (VLMs), recognized as natural source models that can be transferred to various downstream tasks without any parameter training, we propose a novel cross-domain task directly combining the strengths of both UDA and DG, named Training-Free Adaptive Domain Generalization (TF-ADG). However, current cross-domain datasets have many limitations, such as unrealistic domains, unclear domain definitions, and the inability to fine-grained domain decomposition, which hinder the real-world application of current cross-domain models due to the lack of accurate and fair evaluation of fine-grained realistic domains. These insights motivate us to establish a novel realistic benchmark for TF-ADG. Benefiting from the introduced hierarchical definition of domain shifts, our proposed dataset DomainVerse addresses these issues by providing about 0.5 million images from 390 realistic, hierarchical, and balanced domains, allowing for decomposition across multiple domains within each image. With the help of the constructed DomainVerse and VLMs, we further propose two algorithms called Domain CLIP and Domain++ CLIP for training-free adaptive domain generalization. Extensive and comprehensive experiments demonstrate the significance of the dataset and the effectiveness of the proposed methods. Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui |
IEEE Trans. Multim. | 8 |
| 2024 | Trust it or not: Confidence-guided automatic radiology report generation
Yixin Wang 0003, Zihao Lin 0003, Zhe Xu 0012, Jie Luo 0003, Jiang Tian, Zhongchao Shi, Lifu Huang, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002 |
Neurocomputing | 7 |
| 2024 | Domain-Aware Graph Network for Bridging Multi-Source Domain AdaptationabstractDomain adaptation (DA) addresses the challenge of distribution discrepancy between the training and test data, while multi-source domain adaptation (MSDA) is particularly appealing for realistic scenarios. With the emergence of extensive unlabeled datasets, self-supervised learning has gained significant popularity in deep learning. It is noteworthy that multi-source domain adaptation and self-supervised learning share a common objective: leveraging unlabeled data to acquire more informative representations. However, conventional self-supervised learning encounters two main limitations. Firstly, the traditional pretext task falls to transfer fine-grained knowledge to downstream task with general representation learning. Secondly, the scheme of the same feature extractor with distinct prediction heads makes the cross-task knowledge exchange and information sharing ineffective. In order to tackle these challenges, we introduce a novel approach called Domain-Aware Graph Network (DAGNet). DAGNet utilizes a graph neural network as a bridge to facilitate efficient cross-task knowledge exchange. By employing a mask token strategy, we enhance the robustness of representations by selectively masking certain domain or self-supervised information. In terms of datasets, the uneven and style-based domain shifts in current datasets make it challenging to measure the model's domain adaptation performance in real-world applications. To address this issue, we introduce a benchmark dataset DomainVerse with continuous spatio-temporal domain shifts encountered in the real world. Our extensive experiments demonstrate that DAGNet achieves state-of-the-art performance not only on mainstream multi-source domain adaptation datasets but also on different settings within DomainVerse. Code is available athttps://github.com/a791702141/SSG. Feng Hou, Yang Zhang 0002, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui |
IEEE Trans. Multim. | 5 |
| 2024 | A Survey of Visual TransformersabstractTransformer, an attention-based encoder-decoder model, has already revolutionized the field of natural language processing (NLP). Inspired by such significant achievements, some pioneering works have recently been done on employing Transformer-liked architectures in the computer vision (CV) field, which have demonstrated their effectiveness on three fundamental CV tasks (classification, detection, and segmentation) as well as multiple sensory data stream (images, point clouds, and vision-language data). Because of their competitive modeling capabilities, the visual Transformers have achieved impressive performance improvements over multiple benchmarks as compared with modern convolution neural networks (CNNs). In this survey, we have reviewed over 100 of different visual Transformers comprehensively according to three fundamental CV tasks and different data stream types, where taxonomy is proposed to organize the representative methods according to their motivations, structures, and application scenarios. Because of their differences on training settings and dedicated vision tasks, we have also evaluated and compared all these existing visual Transformers under different configurations. Furthermore, we have revealed a series of essential but unexploited aspects that may empower such visual Transformers to stand out from numerous architectures, e.g., slack high-level semantic embeddings to bridge the gap between the visual Transformers and the sequential ones. Finally, two promising research directions are suggested for future investment. We will continue to update the latest articles and their released source codes at https://github.com/liuyang-ict/awesome-visual-transformers. Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | SAP-DETR: Bridging the Gap Between Salient Points and Queries-Based Transformer Detector for Fast Model ConvergencyabstractRecently, the dominant DETR-based approaches apply central-concept spatial prior to accelerating Transformer detector convergency. These methods gradually refine the reference points to the center of target objects and imbue object queries with the updated central reference information for spatially conditional attention. However, centralizing reference points may severely deteriorate queries' saliency and confuse detectors due to the indiscriminative spatial prior. To bridge the gap between the reference points of salient queries and Transformer detectors, we propose SAlient Point-based DETR (SAP-DETR) by treating object detection as a transformation from salient points to instance objects. Concretely, we explicitly initialize a query-specific reference point for each object query, gradually aggregate them into an instance object, and then predict the distance from each side of the bounding box to these points. By rapidly attending to query-specific reference regions and the conditional box edges, SAP-DETR can effectively bridge the gap between the salient point and the query-based Transformer detector with a significant convergency speed. Experimentally, SAP-DETR achieves 1.4× convergency speed with competitive performance and stably promotes the SoTA approaches by ∼1.0 AP. Based on ResNet-DC-101, SAP-DETR achieves 46.9 AP. The code will be released at https://github.com/liuyang-ict/SAP-DETR. Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
CVPR | 6 |
| 2023 | Learning How to Learn Domain-Invariant Parameters for Domain GeneralizationabstractDue to domain shift, deep neural networks (DNNs) usually fail to generalize well on unknown test data in practice. Domain generalization (DG) aims to overcome this issue by capturing domain-invariant representations from source domains. Motivated by the insight that only partial parameters of DNNs are optimized to extract domain-invariant representations, we expect a general model that is capable of well perceiving and emphatically updating such domain-invariant parameters. In this paper, we propose two modules of Domain Decoupling and Combination (DDC) and Domain-invariance-guided Backpropagation (DIGB), which can encourage such general model to focus on the parameters that have a unified optimization direction between pairs of contrastive samples. Our extensive experiments on two benchmarks have demonstrated that our proposed method has achieved state-of-the-art performance with strong generalization capability. Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
ICASSP | 7 |
| 2023 | Dual Pseudo-Labels Interactive Self-Training for Semi-Supervised Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to match a specific person from a gallery of images captured from non-overlapping visible and infrared cameras. Most works focus on fully supervised VI-ReID, which requires substantial cross-modality annotation that is more expensive than the annotation in single-modality. To reduce the extensive cost of annotation, we explore two practical semi-supervised settings: uni-semi-supervised (annotating only visible images) and bi-semi-supervised (annotating partially in both modalities). These two semi-supervised settings face two challenges due to the large cross-modality discrepancies and the lack of correspondence supervision between visible and infrared images. Thus, it is diffi-cult to generate reliable pseudo-labels and learn modality-invariant features from noise pseudo-labels. In this paper, we propose a dual pseudo-label interactive self-training (DPIS) for these two semi-supervised VI-ReID. Our DPIS integrates two pseudo-labels generated by distinct models into a hybrid pseudo-label for unlabeled data. However, the hybrid pseudo-label still inevitably contains noise. To eliminate the negative effect of noise pseudo-labels, we introduce three modules: noise label penalty (NLP), noise correspondence calibration (NCC), and unreliable anchor learning (UAL). Specifically, NLP penalizes noise labels, NCC calibrates noisy correspondences, and UAL mines the hard-to-discriminate features. Extensive experimental results on SYSU-MM01 and RegDB demonstrate that our DPIS achieves impressive performance under these two semi-supervised settings. Jiangming Shi, Yachao Zhang 0001, Xiangbo Yin, Yuan Xie 0006, Zhizhong Zhang 0001, Jianping Fan 0007, Zhongchao Shi, Yanyun Qu |
ICCV | 7 |
| 2023 | VS-Boost: Boosting Visual-Semantic Association for Generalized Zero-Shot LearningabstractUnlike conventional zero-shot learning (CZSL) which only focuses on the recognition of unseen classes by using the classifier trained on seen classes and semantic embeddings, generalized zero-shot learning (GZSL) aims at recognizing both the seen and unseen classes, so it is more challenging due to the extreme training imbalance. Recently, some feature generation methods introduce metric learning to enhance the discriminability of visual features. Although these methods achieve good results, they focus only on metric learning in the visual feature space to enhance features and ignore the association between the feature space and the semantic space. Since the GZSL method uses semantics as prior knowledge to migrate visual knowledge to unseen classes, the consistency between visual space and semantic space is critical. To this end, we propose relational metric learning which can relate the metrics in the two spaces and make the distribution of the two spaces more consistent. Based on the generation method and relational metric learning, we proposed a novel GZSL method, termed VS-Boost, which can effectively boost the association between vision and semantics. The experimental results demonstrate that our method is effective and achieves significant gains on five benchmark datasets compared with the state-of-the-art methods. Xiaofan Li 0008, Yachao Zhang 0001, Shiran Bian, Yanyun Qu, Yuan Xie 0006, Zhongchao Shi, Jianping Fan 0007 |
IJCAI | 6 |
| 2023 | Cross-modal Unsupervised Domain Adaptation for 3D Semantic Segmentation via Bidirectional Fusion-then-DistillationabstractCross-modal Unsupervised Domain Adaptation (UDA) becomes a research hotspot because it reduces the laborious annotation of target domain samples. Existing methods only mutually mimic the outputs of cross-modality in each domain, which enforces the class probability distribution agreeable in different domains. However, these methods ignore the complementarity brought by the modality fusion representation in cross-modal learning. In this paper, we propose a cross-modal UDA method for 3D semantic segmentation via Bidirectional Fusion-then-Distillation, named BFtD-xMUDA, which explores cross-modal fusion in UDA and realizes distribution consistency between outputs of two domains not only for 2D image and 3D point cloud but also for 2D/3D and fusion. Our method contains three significant components: Model-agnostic Feature Fusion Module (MFFM), Bidirectional Distillation (B-Distill), and Cross-modal Debiased Pseudo-Labeling (xDPL). MFFM is employed to generate cross-modal fusion features for establishing a latent space, which enforces maximum correlation and complementarity between two heterogeneous modalities. B-Distill is introduced to exploit bidirectional knowledge distillation which includes cross-modality and cross-domain fusion distillation, and well-achieving domain-modality alignment. xDPL is designed to model the uncertainty of pseudo-labels by self-training scheme. Extensive experimental results demonstrate that our method outperforms state-of-the-art competitors in several adaptation scenarios. Mingwei Xing, Yachao Zhang 0001, Yuan Xie 0006, Jianping Fan 0007, Zhongchao Shi, Yanyun Qu |
ACM Multimedia | 6 |
| 2023 | Hardware-friendly Scalable Image Super Resolution with Progressive Structured SparsityabstractSingle image super-resolution (SR) is an important low-level vision task, and the dynamic SR trading off performance and efficiency are increasingly in demand. The existing dynamic SR methods are divided into two classes: the structured pruning and non-structured compressing methods. The former removes redundant structures in the network, which often leads to significant performance degradation, and the latter searches for extremely sparse parameter masks, achieving promising performance, but they are not deployable in hardware platforms with irregular memory access. In order to solve the mentioned problems, we propose Hardware-friendly Scalable SR (HSSR) with progressively structured sparsity. The superiority of our method is that with only a single scalable model it covers multiple SR models with different sizes, without extra retraining or post-processing. HSSR contains the forward and backward processing. In the forward process, we gradually shrink the SR networks with structured iterative sparsity where grouping convolution together with knowledge distillation is conducted to reduce the amount of SR parameters and the computational complexity while keeping the performance, and in the backward process, we gradually expand the compressed SR networks with structured iterative recovery. Comprehensive experiments on benchmark datasets show that HSSR is perfectly compatible with common convolution baselines. Compared with the Slimmable method, our model is superior in performance, flops, and model size. Experimental results demonstrate that HSSR achieves significant compression, saving up to 1500K parameters and 100 GFlops calculation compared to the original model in real-world applications. Fangchen Ye, Hongzhan Huang, Jianping Fan 0007, Zhongchao Shi, Yuan Xie 0006, Yanyun Qu |
ACM Multimedia | 5 |
| 2023 | Balanced masking strategy for multi-label image classification
Yao Zhang 0010, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui |
Neurocomputing | 3 |
| 2023 | Graph Attention Transformer Network for Multi-label Image ClassificationabstractMulti-label classification aims to recognize multiple objects or attributes from images. The key to solving this issue relies on effectively characterizing the inter-label correlations or dependencies, which bring the prevailing graph neural network. However, current methods often use the co-occurrence probability of labels based on the training set as the adjacency matrix to model this correlation, which is greatly limited by the dataset and affects the model’s generalization ability. This article proposes a Graph Attention Transformer Network, a general framework for multi-label image classification by mining rich and effective label correlation. First, we use the cosine similarity value of the pre-trained label word embedding as the initial correlation matrix, which can represent richer semantic information than the co-occurrence one. Subsequently, we propose the graph attention transformer layer to transfer this adjacency matrix to adapt to the current domain. Our extensive experiments have demonstrated that our proposed methods can achieve highly competitive performance on three datasets. Shikai Chen, Yao Zhang 0010, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Cross-Domain Few-Shot Learning for Rare-Disease Skin Lesion SegmentationabstractRecently, deep learning (DL)-based skin lesion segmentation in dermoscopic images has advanced the efficient diagnosis of skin diseases. Commonly, most of the DL-based methods require a large amount of training data and can only perform accurate predictions on pre-defined classes. However, there exist some rare skin diseases with very limited labeled samples, which poses great challenges to typical DL-based methods. Few-shot learning (FSL) technique, which aims to train models with abundant seen classes and then generalizes to related unseen classes, is promising in addressing a similar problem. Unfortunately, simply borrowing the typical FSL is infeasible since collecting such abundant seen-class data (common skin diseases), is also difficult. In this paper, we propose a cross-domain few-shot segmentation (CD-FSS) framework, which enables the model to leverage the learning ability obtained from the natural domain, to facilitate rare-disease skin lesion segmentation with limited data of common diseases. Specifically, the framework consists of two processes, i.e., specific learning and generic learning, which are alternately optimized in a meta-training manner. A specific learner and a generic learner are tailored to build relationships between both processes. Experimental results demonstrate that our framework significantly improves the generalization ability from natural domain to unseen medical domain. Yixin Wang 0003, Zhe Xu 0012, Jiang Tian, Jie Luo 0003, Zhongchao Shi, Yang Zhang 0002, Jianping Fan 0007, Zhiqiang He 0002 |
ICASSP | 5 |
| 2022 | Semi-Supervised 3D Medical Image Segmentation Via Boundary-Aware Consistent Hidden Representation LearningabstractThis paper proposes a novel Boundary-aware Consistent Hidden Representation Learning Network (BA-CHRLN), which contains two branches for semi-supervised 3D medical image segmentation. Inspired by the contrastive learning, the two branches share the same encoder and each has its individual decoder, namely supervised decoder and unsupervised one. A stop-gradient operation is also utilized to prevent collapsing of solutions. Taking the unlabeled images as references, BA-CHRLN imposes the consistency by applying a perturbation on the high-level hidden feature representations, which significantly improves the encoder’s representation and the network’s robustness. A boundary-aware map is further introduced to capture the organ’s boundary without any prior knowledge and additional parameters. Experiments on the Left Atrium (LA) benchmark dataset demonstrate the effectiveness of the BA-CHRLN. Linhu Liu, Jiang Tian, Xiangqian Cheng, Zhongchao Shi, Jianping Fan 0007, Yong Rui |
ICIP | 4 |
| 2022 | Disentangled Neural Architecture SearchabstractNeural architecture search (NAS) has emerged as a hot topic recently which makes artificial intelligence techniques easier to apply and reduces the demand for experts knowledge by generating deep neural network architectures automatically. However, most existing methods for neural architecture search (NAS) heavily rely on underlying black-box controllers to generate potential candidates of network architectures, and suffer from serious problems of lacking interpretability and search efficiency. In this paper, we propose Disentangled Neural Architecture Search (DNAS), which addresses the two issues by adopting disentangled NAS controller and efficient dense-sampling strategy. Specifically, DNAS learns disentangled factors of network architecture by explicitly encouraging the latent factors to be independent. This approach could not only achieve semantic interpretability but also allow us to conveniently identify the promising regions of representations corresponding to high-performance architectures. We further propose a dense-sampling strategy that conducts targeted architecture search within the promising regions to accelerate the searching process. Our DNAS owns several attractive features: 1) it can successfully learn semantic representations of architectures, including operation selection, skip connections, and layer order; 2) it can speed up the process for neural architecture search more than 13× by using dense-sampling and disentangled factors; 3) it can achieve higher accuracy under less computational cost—DNAS achieves state-of-the-art performance of 94.16% on NASBench-101, and 22.7% top-1 test error on ImageNet with 1.6 GPU-days. Xinyue Zheng, Peng Wang 0095, Qigang Wang, Zhongchao Shi, Jianping Fan 0007 |
IJCNN | 4 |
| 2022 | Self-Supervised Graph Neural Network for Multi-Source Domain AdaptationabstractDomain adaptation (DA) tries to tackle the scenarios when the test data does not fully follow the same distribution of the training data, and multi-source domain adaptation (MSDA) is very attractive for real world applications. By learning from large-scale unlabeled samples, self-supervised learning has now become a new trend in deep learning. It is worth noting that both self-supervised learning and multi-source domain adaptation share a similar goal: they both aim to leverage unlabeled data to learn more expressive representations. Unfortunately, traditional multi-task self-supervised learning faces two challenges: (1) the pretext task may not strongly relate to the downstream task, thus it could be difficult to learn useful knowledge being shared from the pretext task to the target task; (2) when the same feature extractor is shared between the pretext task and the downstream one and only different prediction heads are used, it is ineffective to enable inter-task information exchange and knowledge sharing. To address these issues, we propose a novel Self-Supervised Graph Neural Network (SSG), where a graph neural network is used as the bridge to enable more effective inter-task information exchange and knowledge sharing. More expressive representation is learned by adopting a mask token strategy to mask some domain information. Our extensive experiments have demonstrated that our proposed SSG method has achieved state-of-the-art results over four multi-source domain adaptation datasets, which have shown the effectiveness of our proposed SSG method from different aspects. Feng Hou, Yangzhou Du, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui |
ACM Multimedia | 4 |
| 2022 | Semi-supervised Medical Image Segmentation with Semantic Distance Distribution Consistency Learning
Linhu Liu, Jiang Tian, Zhongchao Shi, Jianping Fan 0007 |
PRCV (2) | 3 |
| 2021 | AKFNET: An Anatomical Knowledge Embedded Few-Shot Network For Medical Image SegmentationabstractAutomated organ segmentation in CTs is an essential prerequisite for many clinical applications, such as computer-aided diagnosis and intervention. As medical data annotation requires massive human labor from experienced radiologists, how to effectively improve the segmentation performance with limited annotated training data remains a challenging problem. Few-shot learning imitates the learning process of humans, which turns out to be a promising way to overcome the aforementioned challenge. In this paper, we propose a novel anatomical knowledge embedded few-shot network (AKFNet), where an anatomical knowledge embedded support unit (AKSU) is carefully designed to embed the anatomical priors from support images into our model. Moreover, a similarity guidance alignment unit (SGAU) is proposed to impose a mutual alignment between the support and query sets. As a result, AKFNet fully exploits anatomical knowledge and presents good learning capability. Without bells and whistles, AKFNet outperforms the state-of-the-art methods with 0.84-1.76% Dice increase. Transfer learning experiments further verify its learning capability. Yanan Wei, Jiang Tian, Zhongchao Shi |
ICIP | 4 |
| 2021 | ACN: Adversarial Co-training Network for Brain Tumor Segmentation with Missing Modalities
Yixin Wang 0003, Yang Zhang 0002, Yang Liu 0250, Zihao Lin 0003, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
MICCAI (7) | 7 |
| 2021 | Modality-Aware Mutual Learning for Multi-modal Medical Image Segmentation
Yao Zhang 0010, Jiawei Yang 0002, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
MICCAI (1) | 4 |
| 2020 | Label Distribution Learning on Auxiliary Label Space Graphs for Facial Expression RecognitionabstractMany existing studies reveal that annotation inconsistency widely exists among a variety of facial expression recognition (FER) datasets. The reason might be the subjectivity of human annotators and the ambiguous nature of the expression labels. One promising strategy tackling such a problem is a recently proposed learning paradigm called Label Distribution Learning (LDL), which allows multiple labels with different intensity to be linked to one expression. However, it is often impractical to directly apply label distribution learning because numerous existing datasets only contain one-hot labels rather than label distributions. To solve the problem, we propose a novel approach named Label Distribution Learning on Auxiliary Label Space Graphs(LDL-ALSG) that leverages the topological information of the labels from related but more distinct tasks, such as action unit recognition and facial landmark detection. The underlying assumption is that facial images should have similar expression distributions to their neighbours in the label space of action unit recognition and facial landmark detection. Our proposed method is evaluated on a variety of datasets and outperforms those state-of-the-art methods consistently with a huge margin. Shikai Chen, Yuedong Chen, Zhongchao Shi, Xin Geng 0001, Yong Rui |
CVPR | 4 |
| 2020 | Selecting Useful Knowledge from Previous Tasks for Future Learning in a Single NetworkabstractContinual learning can learn new tasks incrementally while avoiding catastrophic forgetting. Recent work has shown that packing multiple tasks into a single network incrementally by iterative pruning and re-training network is a promising method. We build upon this idea and propose an improved version of PackNet. Specifically, we propose a novel gradient-based threshold method to reuse the knowledge of the previous tasks selectively when learning new tasks. Our experiments on a variety of classification tasks and different network architectures demonstrate that our method obtains competitive results when compared to PackNet. Feifei Shi, Peng Wang 0095, Zhongchao Shi, Yong Rui |
ICPR | 3 |
| 2020 | DARN: Deep Attentive Refinement Network for Liver Tumor Segmentation from 3D CT volumeabstractAutomatic liver tumor segmentation from 3D Computed Tomography (CT) is a necessary prerequisite in the interventions of hepatic abnormalities and surgery planning. However, accurate liver tumor segmentation remains challenging due to the large variability of tumor sizes and inhomogeneous texture. Recent advances based on Fully Convolutional Network (FCN) in liver tumor segmentation draw on success of learning discriminative multi-level features. In this paper, we propose a Deep Attentive Refinement Network (DARN) for improved liver tumor segmentation from CT volumes by fully exploiting both low and high level features embedded in different layers of FCN. Different from existing works, we exploit attention mechanism to leverage the relation of different levels of features encoded in different layers of FCN. Specifically, we introduce a Semantic Attention Refinement (SemRef) module to selectively emphasize global semantic information in low level features with the guidance of high level ones, and a Spatial Attention Refinement (SpaRef) module to adaptively enhance spatial details in high level features with the guidance of low level ones. We evaluate our network on the public MICCAI 2017 Liver Tumor Segmentation Challenge dataset (LiTS dataset) and it achieves state-of-the-art performance. The proposed refinement modules are an effective strategy to exploit multi-level features and has great potential to generalize to other medical image segmentation tasks. Yao Zhang 0010, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Zhiqiang He 0002 |
ICPR | 5 |
| 2020 | Double-Uncertainty Weighted Method for Semi-supervised Learning
Yixin Wang 0003, Yao Zhang 0010, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
MICCAI (1) | 5 |
| 2020 | Algorithm Bias Detection and Mitigation in Lenovo Face Recognition Engine
Sheng Shi, Shanshan Wei, Zhongchao Shi, Yangzhou Du, Jianping Fan 0007, Yolanda Conyers |
NLPCC (2) | 3 |
| 2020 | Challenge Closed-Book Science Exam: A Meta-Learning Based Question Answering System
Xinyue Zheng, Peng Wang 0095, Qigang Wang, Zhongchao Shi |
PKAW | 4 |
| 2019 | Discrimination Assessment for Saliency Maps
Ruiyi Li, Yangzhou Du, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
NLPCC (2) | 3 |
| 2019 | Efficient Automatic Meta Optimization Search for Few-Shot Learning
Xinyue Zheng, Peng Wang 0095, Qigang Wang, Zhongchao Shi |
PRCV (3) | 4 |
| 2019 | Facial Motion Prior Networks for Facial Expression RecognitionabstractDeep learning based facial expression recognition (FER) has received a lot of attention in the past few years. Most of the existing deep learning based FER methods do not consider domain knowledge well, which thereby fail to extract representative features. In this work, we propose a novel FER framework, named Facial Motion Prior Networks (FMPN). Particularly, we introduce an addition branch to generate a facial mask so as to focus on facial muscle moving regions. To guide the facial mask learning, we propose to incorporate prior domain knowledge by using the average differences between neutral faces and the corresponding expressive faces as the training guidance. Extensive experiments on three facial expression benchmark datasets demonstrate the effectiveness of the proposed method, compared with the state-of-the-art approaches. Yuedong Chen, Shikai Chen, Zhongchao Shi, Jianfei Cai 0001 |
VCIP | 4 |
| 2018 | SequentialSegNet: Combination with Sequential Feature for Multi-Organ SegmentationabstractMulti-organ segmentation from computed tomography (CT) images is essential for computer aided diagnosis (CAD), and recent advances in fully convolutional networks (FCNs) for volumetric image segmentation have demonstrated the importance of leveraging spatial information. In this paper, we propose a novel framework called SequentialSegNet, which efficiently combines features within a single CT image (intra-slice) and among multiple adjacent images (inter-slice) for a multi-organ segmentation. Experimental results show that our approach can effectively improve the segmentation performance on both large-size and small-size abdominal organs including liver, spleen and gallbladder. Yao Zhang 0010, Yang Zhang 0002, Zhongchao Shi, Zhensheng Li, Zhiqiang He 0002 |
ICPR | 5 |
| 2018 | A Diagnostic Report Generator from CT Volumes on Liver Tumor with Semi-supervised Attention Mechanism
Jiang Tian, Zhongchao Shi, Feiyu Xu 0001 |
MICCAI (2) | 3 |
| 2018 | Dynamic Delay Based Cyclic Gradient Update Method for Distributed Training
Wenhui Hu, Peng Wang 0095, Qigang Wang, Zhengdong Zhou, Hui Xiang, Zhongchao Shi |
PRCV (3) | 7 |
| 2017 | Video Captioning with Listwise SupervisionabstractAutomatically describing video content with natural language is a fundamental challenging that has received increasing attention. However, existing techniques restrict the model learning on the pairs of each video and its own sentences, and thus fail to capture more holistically semantic relationships among all sentences. In this paper, we propose to model relative relationships of different video-sentence pairs and present a novel framework, named Long Short-Term Memory with Listwise Supervision (LSTM-LS), for video captioning. Given each video in training data, we obtain a ranking list of sentences w.r.t. a given sentence associated with the video using nearest-neighbor search. The ranking information is represented by a set of rank triplets that can be used to assess the quality of ranking list. The video captioning problem is then solved by learning LSTM model for sentence generation, through maximizing the ranking quality over all the sentences in the list. The experiments on MSVD dataset show that our proposed LSTM-LS produces better performance than the state of the art in generating natural sentences: 51.1% and 32.6% in terms of BLEU@4 and METEOR, respectively. Superior performances are also reported on the movie description M-VAD dataset. Zhongchao Shi |
AAAI | 3 |
| 2016 | Boosting Video Description Generation by Explicitly Translating from Frame-Level CaptionsabstractAutomatically describing video content with natural language is a fundamental challenge of computer vision. The recent advanced technique that approaches this problem is Recurrent Neural Networks (RNN). The need to train RNN on large-scale complex and diverse videos and their associated language, however, makes the task human-labeling intensive and computationally expensive. Moreover, the results can suffer from robustness problem, especially when there are rich of temporal dynamics in the sequence of video frames. We demonstrate in this paper that the above two limitations can be mitigated by jointly exploring the largely available data from image domain and representing each frame by high-level attributes rather than visual features. The former leverages the learnt models on image captioning benchmark to generate caption for each video frame, while the latter explicitly incorporates the obtained captions which are regarded as the attributes of each frame. Specifically, we propose a novel sequence to sequence architecture to generate descriptions for videos, in a sense that the inputs are the captions of sequential frames and it outputs words sequentially. On a widely used YouTube2Text dataset, our proposal is shown to be powerful with superior performance over several state-of-the-art methods including both architectures that are purely developed on video data and RNN-based models which translate directly from visual features to language. Zhongchao Shi |
ACM Multimedia | 2 |
| 2015 | Click-through-based Deep Visual-Semantic Embedding for Image SearchabstractThe problem of image search is mostly considered from the perspectives of feature-based vector model and image ranker learning. A fundamental issue that underlies the success of these approaches is the similarity learning between query and image. The need of image surrounding texts in feature-based vector model, however, makes the similarity sensitive to the quality of text descriptions. On the other, the image ranker learning can suffer from robustness problem, originating from the fact that human labeled query-image pairs do not always predict user search intention precisely. We demonstrate in this paper that the above two issues can be well mitigated by jointly exploring visual-semantic embedding and the use of click-through data. Specifically, we propose a novel click-through-based deep visual-semantic embedding (C-DVSE) model for learning query and image similarity. The proposed model consists of two components: a deep convolutional neural networks followed by an image embedding layer for learning visual embedding, and a deep neural networks for generating query semantic embedding. The objective of our model is to maximize the correlation between semantic (query) and visual (clicked image) embedding. When the visual-semantic embedding is learnt, query-image similarity can be directly computed by cosine similarity on this embedding space. On a large-scale click-based image dataset with 11.7 million queries and one million images, our model is shown to be powerful for keyword-based image search with superior performance over several state-of-the-art methods. Zhongchao Shi |
ACM Multimedia | 2 |
| 2015 | Learning query and image similarities with listwise supervisionabstractOne of the fundamental problems in image search is to learn the ranking functions, i.e., similarity between the textual query and visual image. A number of research paradigms, ranging from feature-based vector model to image ranker learning, have been applied to measure query-image similarity. However, most of the existing similarity learning methods either depend on surrounding texts for ranking images or learn image rankers to satisfy pairwise or tripletwise supervision. In this paper, we propose to leverage listwise supervision into a principled click-through-based query-image similarity learning framework. In particular, the algorithm utilizes click counts for each image in response to a query to get a ranking list. The ranking information is represented by a set of rank triplets that can be used to assess the quality of the ranking list. The image ranking problem is then solved efficiently by learning two linear projections for query and image space respectively, through maximizing the ranking quality over all the training data. When the two linear projections are learnt, query-image similarity can be directly computed by dot product on this projected subspace. On a large-scale click-through-based image dataset with 11.7 million queries and one million images, our learnt model via listwise supervision is shown to be powerful for keyword-based image search with superior performance over several state-of-the-art methods. Zhongchao Shi |
MMSP | 2 |
| 2015 | Fusion of a panoramic camera and 2D laser scanner data for constrained bundle adjustment in GPS-denied environments
Shunping Ji, Xiaowei Shao, Peng Yang 0005, Zhongchao Shi, Ryosuke Shibasaki |
Image Vis. Comput. | 6 |
| 2005 | Remote sensing monitoring of grassland in key region of Northern China
Bin Xu 0008, Zhongchao Shi, Xiuchun Yang, Guixia Yang |
IGARSS | 4 |
| 2004 | A novel fingerprint matching method based on the Hough transform without quantization of the Hough spaceabstractIn contrast to the conventional usage of the Hough transform for fingerprint alignment, our novel Hough transform-based fingerprint registration algorithm proposed in this paper doesn't need to discretize the Hough space when estimating the parameters of the rigid transformation between two fingerprint impressions. In our algorithm, the parameters (rotation, x translation and y translation) of the transform (rigid transform considered) can be estimated separately. In addition, the orientation fields from the fingerprints, instead of the minutiae, are used to estimate the most optimal rotation parameter. The experimental results obtained on the public domain collections of fingerprint images, FVC2002 db2 set A (800 fingerprints) and db3 set A (800 fingerprints), show that the new fingerprint registration algorithm presented in this paper is faster and more effective as compared to the conventional generalized Hough transform-based fingerprint alignment method. Zhongchao Shi, Xuying Zhao, Yangsheng Wang |
ICIG | 2 |
| 2004 | A new segmentation algorithm for low quality fingerprint imageabstractAiming at the segmentation of low quality fingerprint images, a new framework, which is different from traditional methods that usually use some certain features to segment diversified images, is proposed. There are two contributions in this scheme: firstly, we introduce a quality estimation step before segmentation, which can remove a great many false traces effectively; secondly, a new feature eccentric moment is proposed to locate the blurry boundary. Then we segment the image using the new block feature of clarified image. Experimental results show that the proposed method can segment low quality fingerprint images properly. Zhongchao Shi, Yangsheng Wang |
ICIG | 1 |
| 2004 | Remote sensing monitoring on dynamic status of grassland productivity and animal loading balance in Northern ChinaabstractThe study region includes 312 counties of 11 provinces and autonomous regions in Northern China. The methodology is mainly the combination of remote sensing, geographic information system and spatial database technology with field investigation of the grassland. The main conclusions are as follows. (1) the monitoring indicates that grass production of the major grasslands in Northern China is decreasing in recent decades, at a rate of 10-40%. (2) Compared with the grass production, over-grazing is a common phenomenon in grazing region. Among the 153 counties in the grazing region, over 50% is with over-loading of herd. Similar phenomenon is also observed in semi-grazing region where over 80% of the 159 counties have various degrees of herd over-loading. (3) Over-loading is severe in semi-grazing area than in grazing area. Comparison has been done to the loading balance between grazing and semi-grazing regions by the end of 1990's. In grazing region. there were 12 counties (banners) with extreme over-loading, accounting for 7.79% of total county numbers in this region. In semi-grazing region, there were 82 counties showing sign of overloading, accounting for 51.5% of total county numbers in this region. The proportion of extreme over-loading grazing region to the total regions occupies 4.82% in grazing region. This proportion is 34.85% in semi-grazing. Obviously, the severity of extreme overloading in semi-grazing highly surpasses that in grazing region. Understanding this animal loading difference among various types of region may help to facilitate proper administration of grazing for sustainable development of the grassland. Bin Xu 0008, Xiaoping Xin, Zhongchao Shi, Haiqi Liu, Zhongxin Chen, Guixia Yang, Youqi Chen |
IGARSS | 4 |