Jide Li

dblp:157/0936 · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0002-0754-5842ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Unveiling the complementary synergy of CLIP and diffusion models for weakly supervised semantic segmentation
Hang Yao 0002, Yuanchen Wu, Jide Li, Kequan Yang, Jingxin Han, Xiaoqiang Li 0002
Expert Syst. Appl.3
2026 Disentangling co-occurrence with class-specific banks for Weakly Supervised Semantic Segmentation
Hang Yao 0002, Yuanchen Wu, Kequan Yang, Jide Li, Chao Yin 0001, Xiaoqiang Li 0002
Image Vis. Comput.4
2025 Stepwise Decomposition and Dual-stream Focus: A Novel Approach for Training-free Camouflaged Object Segmentation
abstract
While promptable segmentation (e.g., SAM) has shown promise for various segmentation tasks, it still requires manual visual prompts for each object to be segmented. In contrast, task-generic promptable segmentation aims to reduce the need for such detailed prompts by employing only a task-generic prompt to guide segmentation across all test samples. However, when applied to Camouflaged Object Segmentation (COS), current methods still face two critical issues: 1) semantic ambiguity in getting instance-specific text prompts, which arises from insufficient discriminative cues in holistic captions, leading to foreground-background confusion; 2) semantic discrepancy combined with spatial separation in getting instance-specific visual prompts, which results from global background sampling far from object boundaries with low feature correlation, causing SAM to segment irrelevant regions. To mitigate the issues above, we propose RDVP-MSD, a novel training-free test-time adaptation framework that synergizes Region-constrained Dual-stream Visual Prompting (RDVP) via Multimodal Stepwise Decomposition Chain of Thought (MSD-CoT). MSD-CoT progressively disentangles image captions to eliminate semantic ambiguity, while RDVP injects spatial constraints into visual prompting and independently samples visual prompts for foreground and background points, effectively mitigating semantic discrepancy and spatial separation. Without requiring any training or supervision, RDVP-MSD achieves a state-of-the-art segmentation result on multiple COS benchmarks. The codes will be available at https://github.com/ycyinchao/RDVP-MSD.
Chao Yin 0001, Kequan Yang, Jide Li, Pinpin Zhu, Xiaoqiang Li 0002
ACM Multimedia4
2025 See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual understanding and multimodal reasoning. However, LVLMs frequently exhibit hallucination phenomena, manifesting as the generated textual responses that demonstrate inconsistencies with the provided visual content. Existing hallucination mitigation methods are predominantly text-centric, the challenges of visual-semantic alignment significantly limit their effectiveness, especially when confronted with fine-grained visual understanding scenarios. To this end, this paper presents ViHallu, a Vision-Centric Hallucination mitigation framework that enhances visual-semantic alignment through Visual Variation Image Generation and Visual Instruction Construction. ViHallu introduces visual variation images with controllable visual alterations while maintaining the overall image structure. These images, combined with carefully constructed visual instructions, enable LVLMs to better understand fine-grained visual content through fine-tuning, allowing models to more precisely capture the correspondence between visual content and text, thereby enhancing visual-semantic alignment. Extensive experiments on multiple benchmarks show that ViHallu effectively enhances models' fine-grained visual understanding while significantly reducing hallucination tendencies. Furthermore, we release ViHallu-Instruction, a visual instruction dataset specifically designed for hallucination mitigation and visual-semantic alignment. Code is available at https://github.com/oliviadzy/ViHallu.
Ziyun Dai, Xiaoqiang Li 0002, Yuanchen Wu, Jide Li
ACM Multimedia5
2025 AgentMatting: Boosting context aggregation for image matting with context agent
Jide Li, Kequan Yang, Chao Yin 0001, Xiaoqiang Li 0002
Expert Syst. Appl.1
2025 Mutual learning with discrepancy for weakly supervised object detection
Kequan Yang, Xichen Ye, Yuanchen Wu, Jide Li, Xiaoqiang Li 0002, Pinpin Zhu
Expert Syst. Appl.4
2025 ProxyMatting: Transformer-based image matting via region proxy
Jide Li, Kequan Yang, Yuanchen Wu, Xichen Ye, Hanqi Yang, Xiaoqiang Li 0002
Knowl. Based Syst.1
2025 Pseudo-label enhancement for weakly supervised object detection using self-supervised vision transformer
Kequan Yang, Yuanchen Wu, Jide Li, Chao Yin 0001, Xiaoqiang Li 0002
Knowl. Based Syst.3
2025 Mutual Iterative Refinement Network for Scribble-Supervised Camouflaged Object Detection
abstract
Detecting camouflaged objects is challenging due to their high visual similarity to surrounding environments in texture, color, and shape. Traditional Camouflaged Object Detection (COD) methods heavily rely on pixel-level annotations, which are costly and time-consuming. Scribble-Supervised COD (SSCOD) has emerged as a more efficient alternative by using sparse scribble annotations. However, it faces two critical challenges: sparse annotations, compounded by the extreme similarity between foreground and background, cause entangled feature representations and inaccurate predictions in unlabeled regions, and existing SSCOD methods lack robustness to scale variations, resulting in inconsistent predictions across scales. To alleviate these challenges, we propose the Mutual Iterative Refinement Network (MIR-Net), which introduces a cross-branch mutual refinement mechanism to disentangle and enhance foreground and background features. MIR-Net incorporates two novel modules: Background-driven Foreground Feature Enhancement (BFFE) and Foreground-driven Background Feature Enhancement (FBFE), which dynamically suppress irrelevant cues and amplify relevant features. Additionally, we introduce a Scale-Invariant Consistency (SIC) loss that enforces stable and accurate predictions across scales, improving the model's robustness to scale variations. Comprehensive experiments on CAMO, COD10K, and NC4K datasets demonstrate that MIR-Net achieves state-of-the-art performance among SSCOD methods, surpassing all fully supervised CNN-based models and demonstrating competitive performance with fully supervised Transformer-based approaches. These results highlight MIR-Net's potential to advance COD under weak supervision.
Chao Yin 0001, Kequan Yang, Jide Li, Xiaoqiang Li 0002
IEEE Trans. Image Process.3
2024 DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic Segmentation
abstract
Recently, One-stage Weakly Supervised Semantic Segmentation (WSSS) with image-level labels has gained increasing interest due to simplification over its cumbersome multi-stage counterpart. Limited by the inherent ambiguity of Class Activation Map (CAM), we observe that one-stage pipelines often encounter confirmation bias caused by incorrect CAM pseudo-labels, impairing their final segmentation performance. Although recent works discard many unreliable pseudo-labels to implicitly alleviate this issue, they fail to exploit sufficient supervision for their models. To this end, we propose a dual student framework with trustworthy progressive learning (DuPL). Specifically, we propose a dual student network with a discrepancy loss to yield diverse CAMs for each sub-net. The two sub-nets generate supervision for each other, mitigating the confirmation bias caused by learning their own incorrect pseudo-labels. In this process, we progressively introduce more trustworthy pseudo-labels to be involved in the supervision through dynamic threshold adjustment with an adaptive noise filtering strategy. Moreover, we believe that every pixel, even discarded from supervision due to its unreliability, is important for WSSS. Thus, we develop consistency regularization on these discarded regions, providing supervision of every pixel. Experiment results demonstrate the superiority of the proposed DuPL over the recent state-of-the-art alternatives on PASCAL VOC 2012 and MS COCO datasets. Code is available at https://github.com/Wu0409/DuPL.
Yuanchen Wu, Xichen Ye, Kequan Yang, Jide Li, Xiaoqiang Li 0002
CVPR4
2024 DINO is Also a Semantic Guider: Exploiting Class-aware Affinity for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) using image-level labels is a challenging task, with relying on Class Activation Map (CAM) to derive segmentation supervision. Although many efficient single-stage solutions have been proposed, their performance is hindered by the inherent ambiguity of CAM. This paper introduces a new approach, dubbed ECA, to Exploit the self-supervised Vision Transformer, DINO, inducing the Class-aware semantic Affinity to overcome this limitation. Specifically, we introduce a Semantic Affinity Exploitation module (SAE). It establishes the class-agnostic affinity graph through the self-attention of DINO. Using the highly activated patches on CAMs as 'seeds', we propagate them across the affinity graph and yield the Class-aware Affinity Region Map (CARM) as supplementary semantic guidance. Moreover, the selection of reliable 'seeds' is crucial to the CARM generation. Inspired by the observed CAM inconsistency between the global and local views, we develop a CAM Correspondence Enhancement module (CCE) to encourage dense local-to-global CAM correspondences, advancing high-fidelity CAM for seed selection in SAE. Our experimental results demonstrate that ECA effectively improves the model's object pattern understanding. Remarkably, it outperforms state-of-the-art alternatives on the PASCAL VOC 2012 and MS COCO 2014 datasets, achieving 90.1% upper bound performance compared to its fully supervised counterpart. Code is available at https://github.com/Wu0409/ECA.
Yuanchen Wu, Xiaoqiang Li 0002, Jide Li, Kequan Yang, Pinpin Zhu
ACM Multimedia3
2024 Color subspace exploring for natural image matting
abstract
Abstract Deep neural networks have seen a surge of successful methods in natural image matting. However, the overlap of foreground and background color distributions in an image is still troubling in matting. It is observed that the three color channels contain different contrast information of an image: some color channels may provide clearer contrast information for separating the foreground from the image, while the foreground and background color distributions in other channels may heavily overlap, resulting in blurred foreground‐background boundaries. Motivated by this observation, the Color Subspace Exploring Network (CSEMat) is proposed to extract the foreground object from an image by exploring high‐contrast appearance information in individual color spaces. Specifically, a 4‐branch encoder is constructed, with one branch for the RGB image and three branches for subdividing the color space. Each color channel is individually processed by a sub‐encoder. Additionally, the trimap‐based color information aggregation module (CIA) is introduced to integrate the feature maps from the independent sub‐encoders, facilitating the transfer of optimized features to the decoder. Extensive experiments demonstrate that the proposed CSEMat achieves favorable performance on publicly available matting datasets.
Yating Kong, Jide Li, Liangpeng Hu, Xiaoqiang Li 0002
IET Image Process.2
2024 Camouflaged Object Detection via Complementary Information-Selected Network Based on Visual and Semantic Separation
abstract
Camouflaged object detection (COD) is a promising yet challenging task that aims to segment objects concealed within intricate surroundings, a capability crucial for modern industrial applications. Current COD methods primarily focus on the direct fusion of high-level and low-level information, without considering their differences and inconsistencies. Consequently, accurately segmenting highly camouflaged objects in challenging scenarios presents a considerable problem. To mitigate this concern, we propose a novel framework called visual and semantic separation network (VSSNet), which separately extracts low-level visual and high-level semantic cues and adaptively combines them for accurate predictions. Specifically, it features the information extractor module for capturing dimension-aware visual or semantic information from various perspectives. The complementary information-selected module leverages the complementary nature of visual and semantic information for adaptive selection and fusion. In addition, the region disparity weighting strategy encourages the model to prioritize the boundaries of highly camouflaged and difficult-to-predict objects. Experimental results on benchmark datasets show the VSSNet significantly outperforms State-of-the-Art COD approaches without data augmentations and multiscale training techniques. Furthermore, our method demonstrates satisfactory cross-domain generalization performance in real-world industrial environments.
Chao Yin 0001, Kequan Yang, Jide Li, Xiaoqiang Li 0002, Yifan Wu 0011
IEEE Trans. Ind. Informatics3
2023 Hierarchical Semantic Contrast for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) with image-level annotations has achieved great processes through class activation map (CAM). Since vanilla CAMs are hardly served as guidance to bridge the gap between full and weak supervision, recent studies explore semantic representations to make CAM fit for WSSS and demonstrate encouraging results. However, they generally exploit single-level semantics, which may hamper the model to learn a comprehensive semantic structure. Motivated by the prior that each image has multiple levels of semantics, we propose hierarchical semantic contrast (HSC) to ameliorate the above problem. It conducts semantic contrast from coarse-grained to fine-grained perspective, including ROI level, class level, and pixel level, making the model learn a better object pattern understanding. To further improve CAM quality, building upon HSC, we explore consistency regularization of cross supervision and develop momentum prototype learning to utilize abundant semantics across different images. Extensive studies manifest that our plug-and-play learning paradigm, HSC, can significantly boost CAM quality on both non-saliency-guided and saliency-guided baselines, and establish new state-of-the-art WSSS performance on PASCAL VOC 2012 dataset. Code is available at https://github.com/Wu0409/HSC_WSSS.
Yuanchen Wu, Xiaoqiang Li 0002, Songmin Dai, Jide Li, Tong Liu 0001, Shaorong Xie
IJCAI4
2023 Multiscale features integration based multiple-in-single-out network for object detection
Kequan Yang, Jide Li, Songmin Dai, Xiaoqiang Li 0002
Image Vis. Comput.2
2023 Effective Local-Global Transformer for Natural Image Matting
abstract
Learning-based matting methods have been dominated by convolution neural networks for a long time. These methods mainly propagate the alpha matte according to the similarity between unknown and known regions. However, correlations between pixels in unknown and known regions are limited due to the insufficient receptive fields of common convolution neural networks, which leads to inaccurate estimation for pixels in unknown regions that are far away from known regions. In this paper, we propose an Effective Local-Global Transformer for natural image matting (ELGT-Matting), which can further expand receptive fields to establish a wide range of correlations between unknown and known regions. The kernel module is the effective local-global transformer block, and each block consists of two modules: 1) A Window-Level Global MSA (Multi-head Self-Attention) module, which learns global context features among windows. 2) A Local-Global Window MSA, which combines coarse global context features and corresponding fine local window features to help local window self-attention capture both local and context information. Experiments demonstrate that our ELGT-Matting performs outstandingly against other competitive approaches on Composition-1K, Distinctions-646, and real-world AIM-500 datasets. In particular, we achieve a new SOTA result on Composition-1K with MSE 0.00374.
Liangpeng Hu, Yating Kong, Jide Li, Xiaoqiang Li 0002
IEEE Trans. Circuits Syst. Video Technol.3
2023 Multi-Sourced Knowledge Integration for Robust Self-Supervised Facial Landmark Tracking
abstract
Expensive annotation costs significantly hinder the development of facial landmark tracking owing to the frame-by-frame labeling of dense landmarks. The most promising approach to address this problem is to develop a self-supervised tracker for large-scale unlabeled videos. However, existing self-supervised trackers trained using single-sourced knowledge are unstable under unconstrained environments. Herein, we propose multi-sourced knowledge integration (MSKI), a robust self-supervised tracking method. It integrates knowledge from multiple sources to provide supervisory signals, thereby improving the stability of the self-supervised tracker. Specifically, the proposed MSKI comprises two complementary modules: a temporal knowledge reasoning (TempRes) module and an interactive knowledge distillation (KnowDist) module. The TempRes module enforces the tracker to achieve cycle-consistent tracking, allowing the tracker to learn temporal correspondence based on the cycle-consistency of time. To exploit facial geometry knowledge against various occlusions, our tracker imposes a multi-level shape constraint over the structure of facial landmarks by leveraging adversarial shape learning, thereby enabling the tracking of occluded faces. Moreover, the tracker interacts with an initialization detector to further develop complementary knowledge via KnowDist. The KnowDist module distills the spatial and temporal knowledge provided by the detector and tracker to generate plausible labels automatically. Finally, these generated labels are utilized to fine-tune the detector, such that it provides high-quality initial landmarks for the cycle-consistent tracking of the tracker on unlabeled videos. The experimental results show that the proposed MSKI can stabilize the tracking trajectory and improve the robustness against various occlusions.
Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai, Weiqin Tong
IEEE Trans. Multim.3
2022 Robust age estimation model using group-aware contrastive learning
abstract
Abstract Although great efforts have been devoted to developing lightweight models for age estimation in recent works, the robustness is still unsatisfactory in unconstrained environments. This paper proposes a Group‐aware Contrastive Network (GACN), a robust lightweight model, which extracts discriminative features by leveraging contrastive learning rather than increasing model parameters. Specifically, with a carefully designed contrastive loss function, GACN minimizes intra‐class distances and maximizes inter‐class distances between different age groups in feature space. Thus, faces belonging to the same age group are pulled together, while clusters of faces from different age groups are pushed apart. Unlike existing contrastive learning methods, which are separated from the downstream tasks, GACN integrates contrastive learning into age regression and jointly optimizes them for age representation learning. This allows to achieve robust age estimation using a lightweight network that is 1/662 of the model size of VGGNet. Extensive experiments on IMDB‐WIKI, Morph II, and FG‐NET demonstrate that the proposed method has a significant improvement over the baseline model and performs comparably to existing compact and bulky methods.
Xiaoqiang Li 0002, Yifan Wu 0011, Congcong Zhu, Jide Li
IET Image Process.5
2022 Reasoning structural relation for occlusion-robust facial landmark localization
abstract
In facial landmark localization tasks, various occlusions heavily degrade the localization accuracy due to the partial observability of facial features . This paper proposes a structural relation network (SRN) for occlusion-robust landmark localization. Unlike most existing methods that simply exploit the shape constraint, the proposed SRN aims to capture the structural relations among different facial components. These relations can be considered a more powerful shape constraint against occlusion. To achieve this, a hierarchical structural relation module (HSRM) is designed to hierarchically reason the structural relations that represent both long- and short-distance spatial dependencies . Compared with existing network architectures ,the HSRM can efficiently model the spatial relations by leveraging its geometry-aware network architecture, which reduces the semantic ambiguity caused by occlusion. Moreover, the SRN augments the training data by synthesizing occluded faces. To further extend our SRN for occluded video data, we formulate the occluded face synthesis as a Markov decision process (MDP). Specifically, it plans the movement of the dynamic occlusion based on an accumulated reward associated with the performance degradation of the pre-trained SRN. This procedure augments hard samples for robust facial landmark tracking. Extensive experimental results indicate that the proposed method achieves outstanding performance on occluded and masked faces. Code is available at https://github.com/zhuccly/SRN
Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai, Weiqin Tong
Pattern Recognit.3
2021 Improving Robustness of Facial Landmark Detection by Defending against Adversarial Attacks
abstract
Many recent developments in facial landmark detection have been driven by stacking model parameters or augmenting annotations. However, three subsequent challenges remain, including 1) an increase in computational overhead, 2) the risk of overfitting caused by increasing model parameters, and 3) the burden of labor-intensive annotation by humans. We argue that exploring the weaknesses of the detector so as to remedy them is a promising method of robust facial landmark detection. To achieve this, we propose a sample-adaptive adversarial training (SAAT) approach to interactively optimize an attacker and a detector, which improves facial landmark detection as a defense against sample-adaptive black-box attacks. By leveraging adversarial attacks, the proposed SAAT exploits adversarial perturbations beyond the handcrafted transformations to improve the detector. Specifically, an attacker generates adversarial perturbations to reflect the weakness of the detector. Then, the detector must improve its robustness to adversarial perturbations to defend against adversarial attacks. Moreover, a sample-adaptive weight is designed to balance the risks and benefits of augmenting adversarial examples to train the detector. We also introduce a masked face alignment dataset, Masked-300W, to evaluate our method. Experiments show that our SAAT performed comparably to existing state-of-the-art methods. The dataset and model are publicly available at https://github.com/zhuccly/SAAT.
Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai
ICCV3
2020 Unsupervised Tongue Segmentation Using Reference Labels
Kequan Yang, Jide Li, Xiaoqiang Li 0002
ICONIP (1)2
2020 UCCTGAN: Unsupervised Clothing Color Transformation Generative Adversarial Network
abstract
Clothing color transformation refers to changing the clothes color in an original image to the clothes color in a target image. In this paper, we propose an Unsupervised Clothing Color Transformation Generative Adversarial Network (UCCTGAN) for the task. UCCTGAN adopts the color histogram of a target clothes as color guidance and an improved U-net architecture called AntennaNet is put forward to fuse the extracted color information with the original image. Meanwhile, to accomplish unsupervised learning, the loss function is carefully designed according to color moment, which evaluates the chromatic aberration between the target clothing and the generated clothing. Experimental results show that our network has the ability to generate convincing color transformation results.
Shuming Sun, Xiaoqiang Li 0002, Jide Li
ICPR3
2020 Spatial-Temporal Knowledge Integration: Robust Self-Supervised Facial Landmark Tracking
abstract
Diversity of training data significantly affects tracking robustness of model under unconstrained environments. However, existing labeled datasets for facial landmark tracking tend to be large but not diverse, and manually annotating the massive clips of new diverse videos is extremely expensive. To address these problems, we propose a Spatial-Temporal Knowledge Integration (STKI) approach. Unlike most existing methods which rely heavily on labeled data, STKI exploits supervisions from unlabeled data. Specifically, STKI integrates spatial-temporal knowledge from massive unlabeled videos, which has several orders of magnitude more than existing labeled video data on the diversity, for robust tracking. Our framework includes a self-supervised tracker and an image-based detector for tracking initialization. To avoid the distortion of facial shape, the tracker leverages adversarial learning to introduce facial structure prior and temporal knowledge into cycle-consistency tracking. Meanwhile, we design a graph-based knowledge distillation method, which distills the knowledge from tracking and detection results, to improve the generalization of the detector. The fine-tuned detector can provide tracker on unconstrained videos with high-quality tracking initialization. Extensive experimental results show that the proposed method achieves state-of-the-art performance on comprehensive evaluation datasets.
Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Guangtai Ding, Weiqin Tong
ACM Multimedia3
2019 TDCC: Top-Down Semantic Aggregation for Color Constancy
abstract
Color constancy considers the problem of restoring the original color of an illuminated scene. Benefiting from the development of Convolutional Neural Network (CNN), substantial progress on color constancy has been made. High-level features of CNN structure contain semantic information while low-level features show local details. If both are taken into account, they would help achieve a more accurate illuminant estimation. However, previous works paid little attention to the latter for there lacks frameworks which can combine those two kinds of features together. Inspired by the pyramid model, a top-down network that successively propagates high-level information to low-level layers is proposed. This network, named Top-down Semantic Aggregation for Color Constancy (TDCC), takes full advantage of the multi-scale representations with strong semantics. As a result, objects with intrinsic colors are captured and a better estimation is obtained. Experiments on three benchmark datasets demonstrate that TDCC significantly outperforms state-of-the-art color constancy methods.
Xiaoqiang Li 0002, Yaqin Zhu, Jiayue Han, Jide Li, Weiqin Tong
ICME4
2019 TDCC: top-down semantic aggregation for colour constancy
abstract
Images obtained from an illuminated scene often have their original colour contaminated. Colour constancy is a study considering how to restore them. Substantial progress on colour constancy has been made in recent years due to the development of a convolutional neural network (CNN). In a CNN structure, high‐level features contain semantic information while low‐level features show local details. If both are taken into account, they would help achieve a more accurate illuminant estimation. However, previous works paid little attention to the latter for lack of frameworks, which can combine those two kinds of features together. Inspired by the pyramid model, a top‐down network that successively propagates high‐level information to low‐level layers is proposed. This network, named top‐down semantic aggregation for colour constancy (TDCC), takes full advantage of the multi‐scale representations with strong semantics. As a result, objects with intrinsic colours are captured and a better estimation is obtained. Experiments on three benchmark datasets demonstrate that TDCC significantly outperforms state‐of‐the‐art colour constancy methods.
Xiaoqiang Li 0002, Yaqin Zhu, Jiayue Han, Jide Li, Huicheng Lian, Weiqin Tong
IET Image Process.4
2019 Parallel accelerated matting method based on local learning
Xiaoqiang Li 0002, Jide Li, Pin Wu, Huicheng Lian, Weiqin Tong
Neurocomputing2
2018 Automated Tongue Segmentation in Chinese Medicine Based on Deep Learning
Yushan Xue, Xiaoqiang Li 0002, Pin Wu, Jide Li, Weiqin Tong
ICONIP (7)4
2014 Automatic tongue image segmentation based on histogram projection and matting
abstract
This paper mainly discusses how to use histogram projection and LBDM (Learning Based Digital Matting) to extract a tongue from a medical image, which is one of the most important steps in diagnosis of traditional Chinese Medicine. We firstly present an effective method to locate the tongue body, getting the convinced foreground and background area in form of trimap. Then, use this trimap as the input for LBDM algorithm to implement the final segmentation. Experiment was carried out to evaluate the proposed scheme, using 480 samples of pictures with tongue, the results of which were compared with the corresponding ground truth. Experimental results and analysis demonstrated the feasibility and effectiveness of the proposed algorithm.
Xiaoqiang Li 0002, Jide Li, Dan Wang 0013
BIBM2