EDBT 2026 Demo / reviewers in the wild / expert
Xiaoqiang Li 0002
dblp:63/2893-2
· DBLP profile ↗
60ranked-venue papers
13as first author
43since 2021 · last 2026
0000-0001-7243-2783ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 6 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unveiling the complementary synergy of CLIP and diffusion models for weakly supervised semantic segmentation
Hang Yao 0002, Yuanchen Wu, Jide Li, Kequan Yang, Jingxin Han, Xiaoqiang Li 0002 |
Expert Syst. Appl. | 6 |
| 2026 | Label-wise reliability-aware classifier for robust chest X-ray multi-label classification
Wenkai Ye, Xichen Ye, Hang Yao 0002, Kequan Yang, Xiaoqiang Li 0002 |
Expert Syst. Appl. | 5 |
| 2026 | Disentangling co-occurrence with class-specific banks for Weakly Supervised Semantic Segmentation
Hang Yao 0002, Yuanchen Wu, Kequan Yang, Jide Li, Chao Yin 0001, Xiaoqiang Li 0002 |
Image Vis. Comput. | 7 |
| 2026 | ClickVOS: Click Video Object SegmentationabstractVideo Object Segmentation (VOS) task aims to segment objects in videos. However, previous settings either require time-consuming manual masks of target objects at the first frame during inference or lack the flexibility to specify arbitrary objects of interest. To address these limitations, we propose the setting named Click Video Object Segmentation (ClickVOS) which segments objects of interest across the whole video according to a single click per object in the first frame. And we provide the extended datasets DAVIS-P and YouTubeVOS-P that with point annotations to support this task. ClickVOS is of significant practical applications and research implications due to its only 1-2 seconds interaction time for indicating an object, comparing annotating the mask of an object needs several minutes. However, ClickVOS also presents increased challenges. To address this task, we propose an end-to-end baseline approach named called Attention Before Segmentation (ABS), motivated by the attention process of humans. ABS utilizes the given point in the first frame to perceive the target object through a concise yet effective segmentation attention. Although the initial object mask is possibly inaccurate, in our ABS, as the video goes on, the initially imprecise object mask can self-heal instead of deteriorating due to error accumulation, which is attributed to our designed improvement memory that continuously records stable global object memory and updates detailed dense memory. In addition, we conduct various baseline explorations utilizing off-the-shelf algorithms from related fields, which could provide insights for the further exploration of ClickVOS. The experimental results demonstrate the superiority of the proposed ABS approach. Extended datasets and codes will be available at https://github.com/PinxueGuo/ClickVOS. Pinxue Guo, Lingyi Hong, Xinyu Zhou 0006, Shuyong Gao, Wanyun Li, Zhaoyu Chen 0001, Xiaoqiang Li 0002, Wei Zhang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | MFDP: Multi-View Feature Integration and Enhanced Disease Prompting for Radiology Report GenerationabstractRadiology report generation aims to automatically produce diagnostic reports from medical images, reducing radiologists' workload. Most existing models commonly use an encoder-decoder architecture, where the text decoder generates reports based on encoded image tokens. However, these approaches have two major limitations: 1) they always use a single-view feature or simple static fusion multi-view feature, which fails to capture complementary information from multi-view images, and 2) they lack explicit diagnostic information related to the disease during the text decoding process, resulting in reduced clinical accuracy and relevance of the generated report. To deal with the above limitations, this paper proposes a novel framework employing Multi-view Feature Integration and Enhanced Disease Prompting for Radiology Report Generation, called MFDP. Specifically, MFDP introduces two key innovations:1) the Multi-view Feature Fusion (MFF) module is designed to dynamically integrate multi-view images (e.g., frontal and lateral views) through a multi-view attention mechanism that adaptively captures inter-view dependencies, enriching the decoder's input features to generate more comprehensive reports. 2) the Enhanced Disease Prompting (EDP) module is designed to provide explicit diagnostic information by constructing enhanced disease prompts to guide the text decoding process. Experiments on two benchmark datasets, MIMIC-CXR and IU X-Ray, demonstrate that the proposed MFDP is competitive in both Clinical Efficacy (CE) and Natural Language Generation (NLG) metrics. Notably, MFDP achieves a 10% average improvement in CE Recall compared to SOTA models, enabling more precise localization of critical abnormalities while maintaining diagnostic completeness. Yongxu Zhao, Kequan Yang, Yuanchen Wu, Xiaoqiang Li 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Optimized Gradient Clipping for Noisy Label LearningabstractPrevious research has shown that constraining the gradient of loss function w.r.t. model-predicted probabilities can enhance the model robustness against noisy labels. These methods typically specify a fixed optimal threshold for gradient clipping through validation data to obtain the desired robustness against noise. However, this common practice overlooks the dynamic distribution of gradients from both clean and noisy-labeled samples at different stages of training, significantly limiting the model capability to adapt to the variable nature of gradients throughout the training process. To address this issue, we propose a simple yet effective approach called Optimized Gradient Clipping (OGC), which dynamically adjusts the clipping threshold based on the ratio of noise gradients to clean gradients after clipping, estimated by modeling the distributions of clean and noisy samples. This approach allows us to modify the clipping threshold at each training step, effectively controlling the influence of noise gradients. Additionally, we provide statistical analysis to certify the noise-tolerance ability of OGC. Our extensive experiments across various types of label noise, including symmetric, asymmetric, instance-dependent, and real-world noise, demonstrate the effectiveness of our approach. Xichen Ye, Yifan Wu 0011, Xiaoqiang Li 0002, Yifan Chen 0004, Cheng Jin 0001 |
AAAI | 4 |
| 2025 | Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object PerceptionabstractLarge Vision-Language Models (LVLMs) have achieved impressive results across various cross-modal tasks. However, hallucinations, i.e., the models generating counterfactual responses, remain a challenge. Though recent studies have attempted to alleviate object perception hallucinations, they focus on the models’ response generation, and overlooking the task question itself. This paper discusses the vulnerability of LVLMs in solving counterfactual presupposition questions (CPQs), where the models are prone to accept the presuppositions of counterfactual objects and produce severe hallucinatory responses. To this end, we introduce "Antidote", a unified, synthetic data-driven post-training framework for mitigating both types of hallucination above. It leverages synthetic data to incorporate factual priors into questions to achieve self-correction, and decouple the mitigation process into a preference optimization problem. Furthermore, we construct "CP-Bench", a novel benchmark to evaluate LVLMs’ ability to correctly handle CPQs and produce factual responses. Applied to the LLaVA series, Antidote can simultaneously enhance performance on CP-Bench by over 50%, POPE by 1.8-3.3%, and CHAIR & SHR by 30-50%, all without relying on external supervision from stronger LVLMs or human feedback and introducing noticeable catastrophic forgetting issues. Yuanchen Wu, Lu Zhang 0060, Hang Yao 0002, Junlong Du, Shouhong Ding, Yunsheng Wu, Xiaoqiang Li 0002 |
CVPR | 8 |
| 2025 | ToVE: Efficient Vision-Language Learning via Knowledge Transfer from Vision ExpertsabstractVision-language (VL) learning requires extensive visual perception capabilities, such as fine-grained object recognition and spatial perception. Recent works typically rely on training huge models on massive datasets to develop these capabilities. As a more efficient alternative, this paper proposes a new framework that Transfers the knowledge from a hub of Vision Experts (ToVE) for efficient VL learning, leveraging pre-trained vision expert models to promote visual perception capability. Specifically, building on a frozen CLIP image encoder that provides vision tokens for image-conditioned language generation, ToVE introduces a hub of multiple vision experts and a token-aware gating network that dynamically routes expert knowledge to vision tokens. In the transfer phase, we propose a "residual knowledge transfer" strategy, which not only preserves the generalizability of the vision tokens but also allows selective detachment of low-contributing experts to improve inference efficiency. Further, we explore to merge these expert knowledge to a single CLIP encoder, creating a knowledge-merged CLIP that produces more informative vision tokens without expert inference during deployment. Experiment results across various VL tasks demonstrate that the proposed ToVE achieves competitive performance with two orders of magnitude fewer training data. Yuanchen Wu, Junlong Du, Shouhong Ding, Xiaoqiang Li 0002 |
ICLR | 5 |
| 2025 | One Arrow, Two Hawks: Sharpness-aware Minimization for Federated Learning via Global Model TrajectoryabstractFederated learning (FL) presents a promising strategy for distributed and privacy-preserving learning, yet struggles with performance issues in the presence of heterogeneous data distributions. Recently, a series of works based on sharpness-aware minimization (SAM) have emerged to improve local learning generality, proving to be effective in mitigating data heterogeneity effects. However, most SAM-based methods do not directly consider the global objective and require two backward pass per iteration, resulting in diminished effectiveness. To overcome these two bottlenecks, we leverage the global model trajectory to directly measure sharpness for the global objective, requiring only a single backward pass. We further propose a novel and general algorithm FedGMT to overcome data heterogeneity and the pitfalls of previous SAM-based methods. We analyze the convergence of FedGMT and conduct extensive experiments on visual and text datasets in a variety of scenarios, demonstrating that FedGMT achieves competitive accuracy with state-of-the-art FL methods while minimizing computation and communication overhead. Code is available at https://github.com/harrylee999/FL-SAM. Tong Liu 0001, Yangguang Cui, Xiaoqiang Li 0002 |
ICML | 5 |
| 2025 | Towards Rationale-Answer Alignment of LVLMs via Self-Rationale CalibrationabstractLarge Vision-Language Models (LVLMs) have manifested strong visual question answering capability. However, they still struggle with aligning the rationale and the generated answer, leading to inconsistent reasoning and incorrect responses. To this end, this paper introduces Self-Rationale Calibration (SRC) framework to iteratively calibrate the alignment between rationales and answers. SRC begins by employing a lightweight “rationale fine-tuning” approach, which modifies the model’s response format to require a rationale before deriving answer without explicit prompts. Next, SRC searches a diverse set of candidate responses from the fine-tuned LVLMs for each sample, followed by a proposed pairwise scoring strategy using a tailored scoring model, R-Scorer, to evaluate both rationale quality and factual consistency of candidates. Based on a confidence-weighted preference curation process, SRC decouples the alignment calibration into a preference fine-tuning manner, leading to significant improvements of LVLMs in perception, reasoning, and generalization across multiple benchmarks. Our results emphasize the rationale-oriented alignment in exploring the potential of LVLMs. Yuanchen Wu, Shouhong Ding, Ziyin Zhou, Xiaoqiang Li 0002 |
ICML | 5 |
| 2025 | Stepwise Decomposition and Dual-stream Focus: A Novel Approach for Training-free Camouflaged Object SegmentationabstractWhile promptable segmentation (e.g., SAM) has shown promise for various segmentation tasks, it still requires manual visual prompts for each object to be segmented. In contrast, task-generic promptable segmentation aims to reduce the need for such detailed prompts by employing only a task-generic prompt to guide segmentation across all test samples. However, when applied to Camouflaged Object Segmentation (COS), current methods still face two critical issues: 1) semantic ambiguity in getting instance-specific text prompts, which arises from insufficient discriminative cues in holistic captions, leading to foreground-background confusion; 2) semantic discrepancy combined with spatial separation in getting instance-specific visual prompts, which results from global background sampling far from object boundaries with low feature correlation, causing SAM to segment irrelevant regions. To mitigate the issues above, we propose RDVP-MSD, a novel training-free test-time adaptation framework that synergizes Region-constrained Dual-stream Visual Prompting (RDVP) via Multimodal Stepwise Decomposition Chain of Thought (MSD-CoT). MSD-CoT progressively disentangles image captions to eliminate semantic ambiguity, while RDVP injects spatial constraints into visual prompting and independently samples visual prompts for foreground and background points, effectively mitigating semantic discrepancy and spatial separation. Without requiring any training or supervision, RDVP-MSD achieves a state-of-the-art segmentation result on multiple COS benchmarks. The codes will be available at https://github.com/ycyinchao/RDVP-MSD. Chao Yin 0001, Kequan Yang, Jide Li, Pinpin Zhu, Xiaoqiang Li 0002 |
ACM Multimedia | 6 |
| 2025 | See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMsabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual understanding and multimodal reasoning. However, LVLMs frequently exhibit hallucination phenomena, manifesting as the generated textual responses that demonstrate inconsistencies with the provided visual content. Existing hallucination mitigation methods are predominantly text-centric, the challenges of visual-semantic alignment significantly limit their effectiveness, especially when confronted with fine-grained visual understanding scenarios. To this end, this paper presents ViHallu, a Vision-Centric Hallucination mitigation framework that enhances visual-semantic alignment through Visual Variation Image Generation and Visual Instruction Construction. ViHallu introduces visual variation images with controllable visual alterations while maintaining the overall image structure. These images, combined with carefully constructed visual instructions, enable LVLMs to better understand fine-grained visual content through fine-tuning, allowing models to more precisely capture the correspondence between visual content and text, thereby enhancing visual-semantic alignment. Extensive experiments on multiple benchmarks show that ViHallu effectively enhances models' fine-grained visual understanding while significantly reducing hallucination tendencies. Furthermore, we release ViHallu-Instruction, a visual instruction dataset specifically designed for hallucination mitigation and visual-semantic alignment. Code is available at https://github.com/oliviadzy/ViHallu. Ziyun Dai, Xiaoqiang Li 0002, Yuanchen Wu, Jide Li |
ACM Multimedia | 2 |
| 2025 | Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object SegmentationabstractUnsupervised Video Object Segmentation (UVOS) aims to predict pixel-level masks for the most salient objects in videos without any prior annotations. While memory mechanisms have been proven critical in various video segmentation paradigms, their application in UVOS yield only marginal performance gains despite sophisticated design. Our analysis reveals a simple but fundamental flaw in existing methods: over-reliance on memorizing high-level semantic features. UVOS inherently suffers from the deficiency of lacking fine-grained information due to the absence of pixel-level prior knowledge. Consequently, memory design relying solely on high-level features, which predominantly capture abstract semantic cues, is insufficient to generate precise predictions. To resolve this fundamental issue, we propose a novel hierarchical memory architecture to incorporate both shallow- and high-level features for memory, which leverages the complementary benefits of pixel and semantic information. Furthermore, to balance the simultaneous utilization of the pixel and semantic memory features, we propose a heterogeneous interaction mechanism to perform pixel-semantic mutual interactions, which explicitly considers their inherent feature discrepancies. Through the design of Pixel-guided Local Alignment Module (PLAM) and Semantic-guided Global Integration Module (SGIM), we achieve delicate integration of the fine-grained details in shallow-level memory and the semantic representations in high-level memory. Our Hierarchical Memory with Heterogeneous Interaction Network (HMHI-Net) consistently achieves state-of-the-art performance across all UVOS and video saliency detection benchmarks. Moreover, HMHI-Net consistently exhibits high performance across different backbones, further demonstrating its superiority and robustness. Project page: https://github.com/ZhengxyFlow/HMHI-Net . Songcheng He, Wanyun Li, Xiaoqiang Li 0002, Wei Zhang 0016 |
ACM Multimedia | 4 |
| 2025 | AgentMatting: Boosting context aggregation for image matting with context agent
Jide Li, Kequan Yang, Chao Yin 0001, Xiaoqiang Li 0002 |
Expert Syst. Appl. | 4 |
| 2025 | Mutual learning with discrepancy for weakly supervised object detection
Kequan Yang, Xichen Ye, Yuanchen Wu, Jide Li, Xiaoqiang Li 0002, Pinpin Zhu |
Expert Syst. Appl. | 5 |
| 2025 | PrioMatch: Semi-supervised learning guided by prior knowledge
Jiaquan Wang, Yuqing Zou, Yuanchen Wu, Xiaoqiang Li 0002 |
Neurocomputing | 4 |
| 2025 | Dual region mutual enhancement network for camouflaged object detection
Chao Yin 0001, Xiaoqiang Li 0002 |
Image Vis. Comput. | 2 |
| 2025 | ProxyMatting: Transformer-based image matting via region proxy
Jide Li, Kequan Yang, Yuanchen Wu, Xichen Ye, Hanqi Yang, Xiaoqiang Li 0002 |
Knowl. Based Syst. | 6 |
| 2025 | Pseudo-label enhancement for weakly supervised object detection using self-supervised vision transformer
Kequan Yang, Yuanchen Wu, Jide Li, Chao Yin 0001, Xiaoqiang Li 0002 |
Knowl. Based Syst. | 5 |
| 2025 | Self-supervised video object segmentation via pseudo label rectification
Pinxue Guo, Wei Zhang 0016, Xiaoqiang Li 0002, Jianping Fan 0007 |
Pattern Recognit. | 3 |
| 2025 | Mutual Iterative Refinement Network for Scribble-Supervised Camouflaged Object DetectionabstractDetecting camouflaged objects is challenging due to their high visual similarity to surrounding environments in texture, color, and shape. Traditional Camouflaged Object Detection (COD) methods heavily rely on pixel-level annotations, which are costly and time-consuming. Scribble-Supervised COD (SSCOD) has emerged as a more efficient alternative by using sparse scribble annotations. However, it faces two critical challenges: sparse annotations, compounded by the extreme similarity between foreground and background, cause entangled feature representations and inaccurate predictions in unlabeled regions, and existing SSCOD methods lack robustness to scale variations, resulting in inconsistent predictions across scales. To alleviate these challenges, we propose the Mutual Iterative Refinement Network (MIR-Net), which introduces a cross-branch mutual refinement mechanism to disentangle and enhance foreground and background features. MIR-Net incorporates two novel modules: Background-driven Foreground Feature Enhancement (BFFE) and Foreground-driven Background Feature Enhancement (FBFE), which dynamically suppress irrelevant cues and amplify relevant features. Additionally, we introduce a Scale-Invariant Consistency (SIC) loss that enforces stable and accurate predictions across scales, improving the model's robustness to scale variations. Comprehensive experiments on CAMO, COD10K, and NC4K datasets demonstrate that MIR-Net achieves state-of-the-art performance among SSCOD methods, surpassing all fully supervised CNN-based models and demonstrating competitive performance with fully supervised Transformer-based approaches. These results highlight MIR-Net's potential to advance COD under weak supervision. Chao Yin 0001, Kequan Yang, Jide Li, Xiaoqiang Li 0002 |
IEEE Trans. Image Process. | 4 |
| 2024 | Generating and Reweighting Dense Contrastive Patterns for Unsupervised Anomaly DetectionabstractRecent unsupervised anomaly detection methods often rely on feature extractors pretrained with auxiliary datasets or on well-crafted anomaly-simulated samples. However, this might limit their adaptability to an increasing set of anomaly detection tasks due to the priors in the selection of auxiliary datasets or the strategy of anomaly simulation. To tackle this challenge, we first introduce a prior-less anomaly generation paradigm and subsequently develop an innovative unsupervised anomaly detection framework named GRAD, grounded in this paradigm. GRAD comprises three essential components: (1) a diffusion model (PatchDiff) to generate contrastive patterns by preserving the local structures while disregarding the global structures present in normal images, (2) a self-supervised reweighting mechanism to handle the challenge of long-tailed and unlabeled contrastive patterns generated by PatchDiff, and (3) a lightweight patch-level detector to efficiently distinguish the normal patterns and reweighted contrastive patterns. The generation results of PatchDiff effectively expose various types of anomaly patterns, e.g. structural and logical anomaly patterns. In addition, extensive experiments on both MVTec AD and MVTec LOCO datasets also support the aforementioned observation and demonstrate that GRAD achieves competitive anomaly detection accuracy and superior inference speed. Songmin Dai, Yifan Wu 0011, Xiaoqiang Li 0002, Xiangyang Xue 0001 |
AAAI | 3 |
| 2024 | DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic SegmentationabstractRecently, One-stage Weakly Supervised Semantic Segmentation (WSSS) with image-level labels has gained increasing interest due to simplification over its cumbersome multi-stage counterpart. Limited by the inherent ambiguity of Class Activation Map (CAM), we observe that one-stage pipelines often encounter confirmation bias caused by incorrect CAM pseudo-labels, impairing their final segmentation performance. Although recent works discard many unreliable pseudo-labels to implicitly alleviate this issue, they fail to exploit sufficient supervision for their models. To this end, we propose a dual student framework with trustworthy progressive learning (DuPL). Specifically, we propose a dual student network with a discrepancy loss to yield diverse CAMs for each sub-net. The two sub-nets generate supervision for each other, mitigating the confirmation bias caused by learning their own incorrect pseudo-labels. In this process, we progressively introduce more trustworthy pseudo-labels to be involved in the supervision through dynamic threshold adjustment with an adaptive noise filtering strategy. Moreover, we believe that every pixel, even discarded from supervision due to its unreliability, is important for WSSS. Thus, we develop consistency regularization on these discarded regions, providing supervision of every pixel. Experiment results demonstrate the superiority of the proposed DuPL over the recent state-of-the-art alternatives on PASCAL VOC 2012 and MS COCO datasets. Code is available at https://github.com/Wu0409/DuPL. Yuanchen Wu, Xichen Ye, Kequan Yang, Jide Li, Xiaoqiang Li 0002 |
CVPR | 5 |
| 2024 | DINO is Also a Semantic Guider: Exploiting Class-aware Affinity for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) using image-level labels is a challenging task, with relying on Class Activation Map (CAM) to derive segmentation supervision. Although many efficient single-stage solutions have been proposed, their performance is hindered by the inherent ambiguity of CAM. This paper introduces a new approach, dubbed ECA, to Exploit the self-supervised Vision Transformer, DINO, inducing the Class-aware semantic Affinity to overcome this limitation. Specifically, we introduce a Semantic Affinity Exploitation module (SAE). It establishes the class-agnostic affinity graph through the self-attention of DINO. Using the highly activated patches on CAMs as 'seeds', we propagate them across the affinity graph and yield the Class-aware Affinity Region Map (CARM) as supplementary semantic guidance. Moreover, the selection of reliable 'seeds' is crucial to the CARM generation. Inspired by the observed CAM inconsistency between the global and local views, we develop a CAM Correspondence Enhancement module (CCE) to encourage dense local-to-global CAM correspondences, advancing high-fidelity CAM for seed selection in SAE. Our experimental results demonstrate that ECA effectively improves the model's object pattern understanding. Remarkably, it outperforms state-of-the-art alternatives on the PASCAL VOC 2012 and MS COCO 2014 datasets, achieving 90.1% upper bound performance compared to its fully supervised counterpart. Code is available at https://github.com/Wu0409/ECA. Yuanchen Wu, Xiaoqiang Li 0002, Jide Li, Kequan Yang, Pinpin Zhu |
ACM Multimedia | 2 |
| 2024 | Color subspace exploring for natural image mattingabstractAbstract Deep neural networks have seen a surge of successful methods in natural image matting. However, the overlap of foreground and background color distributions in an image is still troubling in matting. It is observed that the three color channels contain different contrast information of an image: some color channels may provide clearer contrast information for separating the foreground from the image, while the foreground and background color distributions in other channels may heavily overlap, resulting in blurred foreground‐background boundaries. Motivated by this observation, the Color Subspace Exploring Network (CSEMat) is proposed to extract the foreground object from an image by exploring high‐contrast appearance information in individual color spaces. Specifically, a 4‐branch encoder is constructed, with one branch for the RGB image and three branches for subdividing the color space. Each color channel is individually processed by a sub‐encoder. Additionally, the trimap‐based color information aggregation module (CIA) is introduced to integrate the feature maps from the independent sub‐encoders, facilitating the transfer of optimized features to the decoder. Extensive experiments demonstrate that the proposed CSEMat achieves favorable performance on publicly available matting datasets. Yating Kong, Jide Li, Liangpeng Hu, Xiaoqiang Li 0002 |
IET Image Process. | 4 |
| 2024 | Camouflaged Object Detection via Complementary Information-Selected Network Based on Visual and Semantic SeparationabstractCamouflaged object detection (COD) is a promising yet challenging task that aims to segment objects concealed within intricate surroundings, a capability crucial for modern industrial applications. Current COD methods primarily focus on the direct fusion of high-level and low-level information, without considering their differences and inconsistencies. Consequently, accurately segmenting highly camouflaged objects in challenging scenarios presents a considerable problem. To mitigate this concern, we propose a novel framework called visual and semantic separation network (VSSNet), which separately extracts low-level visual and high-level semantic cues and adaptively combines them for accurate predictions. Specifically, it features the information extractor module for capturing dimension-aware visual or semantic information from various perspectives. The complementary information-selected module leverages the complementary nature of visual and semantic information for adaptive selection and fusion. In addition, the region disparity weighting strategy encourages the model to prioritize the boundaries of highly camouflaged and difficult-to-predict objects. Experimental results on benchmark datasets show the VSSNet significantly outperforms State-of-the-Art COD approaches without data augmentations and multiscale training techniques. Furthermore, our method demonstrates satisfactory cross-domain generalization performance in real-world industrial environments. Chao Yin 0001, Kequan Yang, Jide Li, Xiaoqiang Li 0002, Yifan Wu 0011 |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | OEST: Outlier Exposure by Simple Transformations for Out-of-Distribution DetectionabstractAlthough the previous works for out-of-distribution(OOD) detection have achieved great improvements, they are still highly dependent on the specific selection of the outliers from external datasets or that transformed by certain data augmentations, and hence cannot be applied in a wide range of domains. To solve this problem, in this paper, we propose a simple, yet effective method called Outlier Exposure by Simple Transformations (OEST), which aims at exposing the outliers by the composition of several simple transformations of data augmentations via energy score. In addition, our training scheme can make full use of nearly all considered data augmentations in previous works, even though some of them are generally regarded as useless. And we also find that, for simple data augmentation, our training scheme is less time-consuming and better in performance than relative works. Furthermore, our experiments validate that our method outperforms the state-of-the-art methods. Yifan Wu 0011, Songmin Dai, Dengye Pan, Xiaoqiang Li 0002 |
ICIP | 4 |
| 2023 | A Spatio-temporal Adaptive Personalized Meta-recommender for Next LocationabstractAs a popular location-based service, next point-of-interest (POI) recommendation predicts users’ next movements based on recent visits. Existing works mostly focus on individual sequential preference mining, ignoring the crucial crowd transition patterns under the effects of corresponding spatio-temporal (ST) contexts. Intuitively, users also follow crowd transition rules. In order to better exploit the extensive ST-specific collaborative signals among users, we propose a novel next POI recommender, including both ST-adaptive crowd-enhanced preference and user-specific multi-semantic interest modeling. We firstly introduce user-agnostic region-and timeslot-level graphs to capture ST-specific crowd transition patterns. As they imply fine-grained general preferences and function as prior meta knowledge, we employ a meta gated recurrent network to make an adaptive prediction for specific ST-context in a meta-learned unified way. Moreover, in modeling user-specific interest, we extract multi-semantic correlations from graph-augmented personal trajectories to obtain high-quality preference representation. Extensive experiments on real-world datasets show the superiority of our proposed model against seven state-of-the-art methods. Tong Liu 0001, Yanmin Zhu 0006, Xiaoqiang Li 0002 |
ICTAI | 5 |
| 2023 | Hierarchical Semantic Contrast for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level annotations has achieved great processes through class activation map (CAM). Since vanilla CAMs are hardly served as guidance to bridge the gap between full and weak supervision, recent studies explore semantic representations to make CAM fit for WSSS and demonstrate encouraging results. However, they generally exploit single-level semantics, which may hamper the model to learn a comprehensive semantic structure. Motivated by the prior that each image has multiple levels of semantics, we propose hierarchical semantic contrast (HSC) to ameliorate the above problem. It conducts semantic contrast from coarse-grained to fine-grained perspective, including ROI level, class level, and pixel level, making the model learn a better object pattern understanding. To further improve CAM quality, building upon HSC, we explore consistency regularization of cross supervision and develop momentum prototype learning to utilize abundant semantics across different images. Extensive studies manifest that our plug-and-play learning paradigm, HSC, can significantly boost CAM quality on both non-saliency-guided and saliency-guided baselines, and establish new state-of-the-art WSSS performance on PASCAL VOC 2012 dataset. Code is available at https://github.com/Wu0409/HSC_WSSS. Yuanchen Wu, Xiaoqiang Li 0002, Songmin Dai, Jide Li, Tong Liu 0001, Shaorong Xie |
IJCAI | 2 |
| 2023 | Active Negative Loss Functions for Learning with Noisy LabelsabstractRobust loss functions are essential for training deep neural networks in the presence of noisy labels. Some robust loss functions use Mean Absolute Error (MAE) as its necessary component. For example, the recently proposed Active Passive Loss (APL) uses MAE as its passive loss function. However, MAE treats every sample equally, slows down the convergence and can make training difficult. In this work, we propose a new class of theoretically robust passive loss functions different from MAE, namely *Normalized Negative Loss Functions* (NNLFs), which focus more on memorized clean samples. By replacing the MAE in APL with our proposed NNLFs, we improve APL and propose a new framework called *Active Negative Loss* (ANL). Experimental results on benchmark and real-world datasets demonstrate that the new set of loss functions created by our ANL framework can outperform state-of-the-art methods. The code is available at
https://github.com/Virusdoll/Active-Negative-Loss. Xichen Ye, Xiaoqiang Li 0002, Songmin Dai, Tong Liu 0001, Weiqin Tong |
NeurIPS | 2 |
| 2023 | Semi-supervised GAN with similarity constraint for mode diversity
Xiaoqiang Li 0002, Yinxiang Luan, Liangbo Chen |
Appl. Intell. | 1 |
| 2023 | Semi-supervised medical imaging segmentation with soft pseudo-label fusion
Xiaoqiang Li 0002, Yuanchen Wu, Songmin Dai |
Appl. Intell. | 1 |
| 2023 | Multiscale features integration based multiple-in-single-out network for object detection
Kequan Yang, Jide Li, Songmin Dai, Xiaoqiang Li 0002 |
Image Vis. Comput. | 4 |
| 2023 | Effective Local-Global Transformer for Natural Image MattingabstractLearning-based matting methods have been dominated by convolution neural networks for a long time. These methods mainly propagate the alpha matte according to the similarity between unknown and known regions. However, correlations between pixels in unknown and known regions are limited due to the insufficient receptive fields of common convolution neural networks, which leads to inaccurate estimation for pixels in unknown regions that are far away from known regions. In this paper, we propose an Effective Local-Global Transformer for natural image matting (ELGT-Matting), which can further expand receptive fields to establish a wide range of correlations between unknown and known regions. The kernel module is the effective local-global transformer block, and each block consists of two modules: 1) A Window-Level Global MSA (Multi-head Self-Attention) module, which learns global context features among windows. 2) A Local-Global Window MSA, which combines coarse global context features and corresponding fine local window features to help local window self-attention capture both local and context information. Experiments demonstrate that our ELGT-Matting performs outstandingly against other competitive approaches on Composition-1K, Distinctions-646, and real-world AIM-500 datasets. In particular, we achieve a new SOTA result on Composition-1K with MSE 0.00374. Liangpeng Hu, Yating Kong, Jide Li, Xiaoqiang Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Deep Reinforcement Learning Based Approach for Online Service Placement and Computation Resource Allocation in Edge ComputingabstractDue to the urgent emergence of computation-intensive intelligent applications on end devices, edge computing has been put forward as an extension of cloud computing, to satisfy the low-latency requirements of these applications. To process heterogenous computation tasks on an edge node, the corresponding services should be placed in advance, including installing softwares and caching databases/libraries. Considering the limited storage space and computation resources on the edge node, services should be elaborately selected and deployed on the edge node and its computation resources should be carefully allocated to placed services, according to the arrivals of computation workloads. The joint service placement and computation resource allocation problem is particularly complicated, in terms of considering the stochastic arrivals of tasks, the additional latency incurred by service migration, and the waiting time of unprocessed tasks. Benefiting from deep reinforcement learning, we propose a novel approach based on parameterized deep Q networks to make the joint service placement and computation resource allocation decisions, with the objective of minimizing the total latency of tasks in a long term. Extensive simulations are conducted to evaluate the convergence and performance achieved by our proposed approach. Tong Liu 0001, Shenggang Ni, Xiaoqiang Li 0002, Yanmin Zhu 0006, Linghe Kong, Yuanyuan Yang 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Multi-Sourced Knowledge Integration for Robust Self-Supervised Facial Landmark TrackingabstractExpensive annotation costs significantly hinder the development of facial landmark tracking owing to the frame-by-frame labeling of dense landmarks. The most promising approach to address this problem is to develop a self-supervised tracker for large-scale unlabeled videos. However, existing self-supervised trackers trained using single-sourced knowledge are unstable under unconstrained environments. Herein, we propose multi-sourced knowledge integration (MSKI), a robust self-supervised tracking method. It integrates knowledge from multiple sources to provide supervisory signals, thereby improving the stability of the self-supervised tracker. Specifically, the proposed MSKI comprises two complementary modules: a temporal knowledge reasoning (TempRes) module and an interactive knowledge distillation (KnowDist) module. The TempRes module enforces the tracker to achieve cycle-consistent tracking, allowing the tracker to learn temporal correspondence based on the cycle-consistency of time. To exploit facial geometry knowledge against various occlusions, our tracker imposes a multi-level shape constraint over the structure of facial landmarks by leveraging adversarial shape learning, thereby enabling the tracking of occluded faces. Moreover, the tracker interacts with an initialization detector to further develop complementary knowledge via KnowDist. The KnowDist module distills the spatial and temporal knowledge provided by the detector and tracker to generate plausible labels automatically. Finally, these generated labels are utilized to fine-tune the detector, such that it provides high-quality initial landmarks for the cycle-consistent tracking of the tracker on unlabeled videos. The experimental results show that the proposed MSKI can stabilize the tracking trajectory and improve the robustness against various occlusions. Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai, Weiqin Tong |
IEEE Trans. Multim. | 2 |
| 2022 | Occlusion-robust Face Alignment using A Viewpoint-invariant Hierarchical Network ArchitectureabstractThe occlusion problem heavily degrades the localization performance of face alignment. Most current solutions for this problem focus on annotating new occlusion data, introducing boundary estimation, and stacking deeper models to improve the robustness of neural networks. However, the performance degradation of models remains under extreme occlusion (i.e. average occlusion of over 50%) because of missing a large amount of facial context information. We argue that exploring neural networks to model the facial hierarchies is a more promising method for dealing with extreme occlusion. Surprisingly, in recent studies, little effort has been devoted to representing the facial hierarchies using neural networks. This paper proposes a new network architecture called GlomFace to model the facial hierarchies against various occlusions, which draws inspiration from the viewpoint-invariant hierarchy of facial structure. Specifically, GlomFace is functionally divided into two modules: the part-whole hierarchical module and the whole-part hierarchical module. The former captures the part-whole hierarchical dependencies of facial parts to suppress multi-scale occlusion information, whereas the latter injects structural reasoning into neural networks by building the whole-part hierarchical relations among facial parts. As a result, GlomFace has a clear topological interpretation due to its correspondence to the facial hierarchies. Extensive experimental results indicate that the proposed GlomFace performs comparably to existing state-of-the-art methods, especially in cases of extreme occlusion. Models are available at https://github.com/zhuccly/GlomFace-Face-Alignment. Congcong Zhu, Xintong Wan, Shaorong Xie, Xiaoqiang Li 0002, Yinzheng Gu |
CVPR | 4 |
| 2022 | PCCN-RE: Point cloud colourisation network based on relevance embeddingabstractAbstract The colour of the point cloud provides abundant information for perception tasks such as autonomous driving and virtual reality. However, few prior work studied the automatic colourisation of colourless point clouds. In this paper, the authors propose a novel method named as Point cloud Colorization Network based on Relevance Embedding (PCCN‐RE) relying on three structures: a relevance embedding structure that efficiently captures local information through the calculation of a covariance matrix within nearby points; a weighted pooling structure designed to facilitate the fusion of features; an enhanced spatial transform network structure that keeps the invariance of input point clouds. On the ShapeNetCore dataset, our PCCN‐RE generates more authentic colour than state‐of‐the‐art methods for colourless point clouds and achieves the highest results by obtaining a Peak Signal to Noise Ratio of 9.40 and a Structural Similarity Index of 0.62. Xiaoqiang Li 0002, Jitao Liu |
IET Comput. Vis. | 2 |
| 2022 | Robust age estimation model using group-aware contrastive learningabstractAbstract Although great efforts have been devoted to developing lightweight models for age estimation in recent works, the robustness is still unsatisfactory in unconstrained environments. This paper proposes a Group‐aware Contrastive Network (GACN), a robust lightweight model, which extracts discriminative features by leveraging contrastive learning rather than increasing model parameters. Specifically, with a carefully designed contrastive loss function, GACN minimizes intra‐class distances and maximizes inter‐class distances between different age groups in feature space. Thus, faces belonging to the same age group are pulled together, while clusters of faces from different age groups are pushed apart. Unlike existing contrastive learning methods, which are separated from the downstream tasks, GACN integrates contrastive learning into age regression and jointly optimizes them for age representation learning. This allows to achieve robust age estimation using a lightweight network that is 1/662 of the model size of VGGNet. Extensive experiments on IMDB‐WIKI, Morph II, and FG‐NET demonstrate that the proposed method has a significant improvement over the baseline model and performs comparably to existing compact and bulky methods. Xiaoqiang Li 0002, Yifan Wu 0011, Congcong Zhu, Jide Li |
IET Image Process. | 1 |
| 2022 | Reasoning structural relation for occlusion-robust facial landmark localizationabstractIn facial landmark localization tasks, various occlusions heavily degrade the localization accuracy due to the partial observability of facial features . This paper proposes a structural relation network (SRN) for occlusion-robust landmark localization. Unlike most existing methods that simply exploit the shape constraint, the proposed SRN aims to capture the structural relations among different facial components. These relations can be considered a more powerful shape constraint against occlusion. To achieve this, a hierarchical structural relation module (HSRM) is designed to hierarchically reason the structural relations that represent both long- and short-distance spatial dependencies . Compared with existing network architectures ,the HSRM can efficiently model the spatial relations by leveraging its geometry-aware network architecture, which reduces the semantic ambiguity caused by occlusion. Moreover, the SRN augments the training data by synthesizing occluded faces. To further extend our SRN for occluded video data, we formulate the occluded face synthesis as a Markov decision process (MDP). Specifically, it plans the movement of the dynamic occlusion based on an accumulated reward associated with the performance degradation of the pre-trained SRN. This procedure augments hard samples for robust facial landmark tracking. Extensive experimental results indicate that the proposed method achieves outstanding performance on occluded and masked faces. Code is available at https://github.com/zhuccly/SRN Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai, Weiqin Tong |
Pattern Recognit. | 2 |
| 2022 | Adaptive Online Mutual Learning Bi-Decoders for Video Object SegmentationabstractOne of the major challenges facing video object segmentation (VOS) is the gap between the training and test datasets due to unseen category in test set, as well as object appearance change over time in the video sequence. To overcome such challenges, an adaptive online framework for VOS is developed with bi-decoders mutual learning. We learn object representation per pixel with bi-level attention features in addition to CNN features, and then feed them into mutual learning bi-decoders whose outputs are further fused to obtain the final segmentation result. We design an adaptive online learning mechanism via a deviation correcting trigger such that bi-decoders online mutual learning will be activated when the previous frame is segmented well meanwhile the current frame is segmented relatively worse. Knowledge distillation from the well segmented previous frames, along with mutual learning between bi-decoders, improves generalization ability and robustness of VOS model. Thus, the proposed model adapts to the challenging scenarios including unseen categories, object deformation, and appearance variation during inference. We extensively evaluate our model on widely-used VOS benchmarks including DAVIS-2016, DAVIS-2017, YouTubeVOS-2018, YouTubeVOS-2019, and UVO. Experimental results demonstrate the superiority of the proposed model over state-of-the-art methods. Pinxue Guo, Wei Zhang 0016, Xiaoqiang Li 0002 |
IEEE Trans. Image Process. | 3 |
| 2021 | Improving Robustness of Facial Landmark Detection by Defending against Adversarial AttacksabstractMany recent developments in facial landmark detection have been driven by stacking model parameters or augmenting annotations. However, three subsequent challenges remain, including 1) an increase in computational overhead, 2) the risk of overfitting caused by increasing model parameters, and 3) the burden of labor-intensive annotation by humans. We argue that exploring the weaknesses of the detector so as to remedy them is a promising method of robust facial landmark detection. To achieve this, we propose a sample-adaptive adversarial training (SAAT) approach to interactively optimize an attacker and a detector, which improves facial landmark detection as a defense against sample-adaptive black-box attacks. By leveraging adversarial attacks, the proposed SAAT exploits adversarial perturbations beyond the handcrafted transformations to improve the detector. Specifically, an attacker generates adversarial perturbations to reflect the weakness of the detector. Then, the detector must improve its robustness to adversarial perturbations to defend against adversarial attacks. Moreover, a sample-adaptive weight is designed to balance the risks and benefits of augmenting adversarial examples to train the detector. We also introduce a masked face alignment dataset, Masked-300W, to evaluate our method. Experiments show that our SAAT performed comparably to existing state-of-the-art methods. The dataset and model are publicly available at https://github.com/zhuccly/SAAT. Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Songmin Dai |
ICCV | 2 |
| 2021 | Point cloud super-resolution based on geometric constraintsabstractAbstract Among all digital representations we have for real physical objects, three‐dimensional ( 3D) is arguably the most expressive encoding. But due to the limitations of 3D scanning equipment, point cloud often becomes sparse or partially missing. A point cloud super‐resolution (PCSR) method based on geometric constraints is proposed to solve the sparse problem of point clouds: it allows dense point clouds to be generated by sparse point clouds. The method is based on the conditional generative adversarial network including redesigned generator and discriminator for point cloud data specially. Moreover, the method can maintain the shape of the dense point cloud by adding geometric constraints. The contributions of our work are as follows: (1) a PCSR method based on geometric constraints is proposed; (2) add a module for obtaining point cloud neighbourhood information in the generator, called K‐nn operation module; and (3) feature aggregation is performed using the weighted pooling to process the neighbourhood information obtained by the K‐nn operation module. Extensive experimental results demonstrate the effectiveness of the proposed method. Xiaoqiang Li 0002, Jitao Liu, Songmin Dai |
IET Comput. Vis. | 1 |
| 2020 | Unsupervised Tongue Segmentation Using Reference Labels
Kequan Yang, Jide Li, Xiaoqiang Li 0002 |
ICONIP (1) | 3 |
| 2020 | UCCTGAN: Unsupervised Clothing Color Transformation Generative Adversarial NetworkabstractClothing color transformation refers to changing the clothes color in an original image to the clothes color in a target image. In this paper, we propose an Unsupervised Clothing Color Transformation Generative Adversarial Network (UCCTGAN) for the task. UCCTGAN adopts the color histogram of a target clothes as color guidance and an improved U-net architecture called AntennaNet is put forward to fuse the extracted color information with the original image. Meanwhile, to accomplish unsupervised learning, the loss function is carefully designed according to color moment, which evaluates the chromatic aberration between the target clothing and the generated clothing. Experimental results show that our network has the ability to generate convincing color transformation results. Shuming Sun, Xiaoqiang Li 0002, Jide Li |
ICPR | 2 |
| 2020 | Spatial-Temporal Knowledge Integration: Robust Self-Supervised Facial Landmark TrackingabstractDiversity of training data significantly affects tracking robustness of model under unconstrained environments. However, existing labeled datasets for facial landmark tracking tend to be large but not diverse, and manually annotating the massive clips of new diverse videos is extremely expensive. To address these problems, we propose a Spatial-Temporal Knowledge Integration (STKI) approach. Unlike most existing methods which rely heavily on labeled data, STKI exploits supervisions from unlabeled data. Specifically, STKI integrates spatial-temporal knowledge from massive unlabeled videos, which has several orders of magnitude more than existing labeled video data on the diversity, for robust tracking. Our framework includes a self-supervised tracker and an image-based detector for tracking initialization. To avoid the distortion of facial shape, the tracker leverages adversarial learning to introduce facial structure prior and temporal knowledge into cycle-consistency tracking. Meanwhile, we design a graph-based knowledge distillation method, which distills the knowledge from tracking and detection results, to improve the generalization of the detector. The fine-tuned detector can provide tracker on unconstrained videos with high-quality tracking initialization. Extensive experimental results show that the proposed method achieves state-of-the-art performance on comprehensive evaluation datasets. Congcong Zhu, Xiaoqiang Li 0002, Jide Li, Guangtai Ding, Weiqin Tong |
ACM Multimedia | 2 |
| 2020 | Scale specified single shot multibox detectorabstractDetecting objects at vastly different scales is a fundamental challenge in computer vision. To solve this, some approaches (e.g. TridentNet) investigate the effect of receptive fields, whereas other approaches (e.g. SNIP, SNIPER) are based on the image pyramid strategy. In this study, a novel single‐shot based detector, called scale specified single‐shot multibox detector (4SD) is proposed. It aims to predict objects of a specific scale range separately by using feature maps of different sizes. First, a parallel multi‐branch architecture with feature maps of different sizes is generated by scale specific inference module. Then, the authors propose a scale specific training scheme to specialise each branch by sampling object instances of proper scales for training. Results are shown on both PASCAL VOC and COCO detection. The proposed method can achieve a mean average precision of 83.1% on PASCAL VOC 2007, and 36.9% on MS‐COCO at a speed of 28 frames per second, which is superior to most single‐stage detectors. Xiaoqiang Li 0002, Chuanwei Liu, Songmin Dai, Huichen Lian, Guangtai Ding |
IET Comput. Vis. | 1 |
| 2020 | Dual attention convolutional network for action recognitionabstractAction recognition has been an active research area for many years. Extracting discriminative spatial and temporal features of different actions plays a key role in accomplishing this task. Current popular methods of action recognition are mainly based on two‐stream Convolutional Networks (ConvNets) or 3D ConvNets. However, the computational cost of two‐stream ConvNets is high for the requirement of optical flow while 3D ConvNets takes too much memory because they have a large amount of parameters. To alleviate such problems, the authors propose a Dual Attention ConvNet (DANet) based on dual attention mechanism which consists of spatial attention and temporal attention. The former concentrates on main motion objects in a video frame by using ConvNet structure and the latter captures related information of multiple video frames by adopting self‐attention. Their network is entirely based on 2D ConvNet and takes in only RGB frames. Experimental results on UCF‐101 and HMDB‐51 benchmarks demonstrate that DANet gets comparable results among leading methods, which proves the effectiveness of the dual attention mechanism. Xiaoqiang Li 0002, Miao Xie, Guangtai Ding, Weiqin Tong |
IET Image Process. | 1 |
| 2020 | Robust landmark-free head pose estimation by learning to crop and background augmentationabstractIt is well known that the performance of head pose estimation is greatly affected by the bounding box margin of the face and its background. Traditionally, researchers will manually choose a suitable bounding box margin to strike a balance between ensuring sufficient information and minimising background noise. However, head pose estimation is still worse when the background is complex in reality or when the box margin changes slightly. To make estimation results more robust, the authors propose two methods to improve it: (i) a convolutional cropping module that can learn to crop the input image to an attentional area for head pose regression. (ii) Background augmentation that can make the network more robust to the background noise. Rather than using the face landmarking to calculate head pose angles, they use another convolutional neural network to regress the head pose angles, which is independent of the landmark detection results. They evaluate the method on BIWI and AFLW2000 dataset and experimental results show that their approach outperforms many other methods. Besides, they evaluate the method on Pointing′04 dataset using head pose accuracy. Furthermore, the approach is more robust and has a lower variance in realistic scenarios. Aoru Xue, Songmin Dai, Xiaoqiang Li 0002 |
IET Image Process. | 4 |
| 2019 | TDCC: Top-Down Semantic Aggregation for Color ConstancyabstractColor constancy considers the problem of restoring the original color of an illuminated scene. Benefiting from the development of Convolutional Neural Network (CNN), substantial progress on color constancy has been made. High-level features of CNN structure contain semantic information while low-level features show local details. If both are taken into account, they would help achieve a more accurate illuminant estimation. However, previous works paid little attention to the latter for there lacks frameworks which can combine those two kinds of features together. Inspired by the pyramid model, a top-down network that successively propagates high-level information to low-level layers is proposed. This network, named Top-down Semantic Aggregation for Color Constancy (TDCC), takes full advantage of the multi-scale representations with strong semantics. As a result, objects with intrinsic colors are captured and a better estimation is obtained. Experiments on three benchmark datasets demonstrate that TDCC significantly outperforms state-of-the-art color constancy methods. Xiaoqiang Li 0002, Yaqin Zhu, Jiayue Han, Jide Li, Weiqin Tong |
ICME | 1 |
| 2019 | Tongue Coating Classification Based on Multiple-Instance Learning and Deep Features
Xiaoqiang Li 0002, Yonghui Tang |
ICONIP (4) | 1 |
| 2019 | Bigdata logs analysis based on seq2seq networks for cognitive Internet of Things
Pin Wu, Zhihui Lu 0002, Zhidan Lei, Xiaoqiang Li 0002, Meikang Qiu, Patrick C. K. Hung |
Future Gener. Comput. Syst. | 5 |
| 2019 | TDCC: top-down semantic aggregation for colour constancyabstractImages obtained from an illuminated scene often have their original colour contaminated. Colour constancy is a study considering how to restore them. Substantial progress on colour constancy has been made in recent years due to the development of a convolutional neural network (CNN). In a CNN structure, high‐level features contain semantic information while low‐level features show local details. If both are taken into account, they would help achieve a more accurate illuminant estimation. However, previous works paid little attention to the latter for lack of frameworks, which can combine those two kinds of features together. Inspired by the pyramid model, a top‐down network that successively propagates high‐level information to low‐level layers is proposed. This network, named top‐down semantic aggregation for colour constancy (TDCC), takes full advantage of the multi‐scale representations with strong semantics. As a result, objects with intrinsic colours are captured and a better estimation is obtained. Experiments on three benchmark datasets demonstrate that TDCC significantly outperforms state‐of‐the‐art colour constancy methods. Xiaoqiang Li 0002, Yaqin Zhu, Jiayue Han, Jide Li, Huicheng Lian, Weiqin Tong |
IET Image Process. | 1 |
| 2019 | Parallel accelerated matting method based on local learning
Xiaoqiang Li 0002, Jide Li, Pin Wu, Huicheng Lian, Weiqin Tong |
Neurocomputing | 1 |
| 2018 | Automated Tongue Segmentation in Chinese Medicine Based on Deep Learning
Yushan Xue, Xiaoqiang Li 0002, Pin Wu, Jide Li, Weiqin Tong |
ICONIP (7) | 2 |
| 2015 | A new unsupervised model of action recognitionabstractThe hand-craft feature description such as blob detection (SIFT), the gradient of edges (HOG) are widely used in the action recognition, while they are not suitable to all kinds of videos and increase the manual intervention. And the deep learning based on unsupervised method have a lot of parameter tuning and iterations. Our paper proposes a new model of action recognition combine the hand-craft spatiotemporal interest points with unsupervised descriptors. This model uses STIP as the spatiotemporal interest points extractor and improved K-means as the unsupervised learning method to build the descriptor. This K-Means based unsupervised descriptor provides higher accuracy than hand-craft descriptors and lower training time than multi-layer unsupervised learning methods. Moreover, we update the BoF model in recognition framework, which constructs local vocabularies to each category. Experimental results indicate that this proposed framework works well. Dan Wang 0013, Qing Shao, Xiaoqiang Li 0002 |
ICIP | 3 |
| 2014 | Automatic tongue image segmentation based on histogram projection and mattingabstractThis paper mainly discusses how to use histogram projection and LBDM (Learning Based Digital Matting) to extract a tongue from a medical image, which is one of the most important steps in diagnosis of traditional Chinese Medicine. We firstly present an effective method to locate the tongue body, getting the convinced foreground and background area in form of trimap. Then, use this trimap as the input for LBDM algorithm to implement the final segmentation. Experiment was carried out to evaluate the proposed scheme, using 480 samples of pictures with tongue, the results of which were compared with the corresponding ground truth. Experimental results and analysis demonstrated the feasibility and effectiveness of the proposed algorithm. Xiaoqiang Li 0002, Jide Li, Dan Wang 0013 |
BIBM | 1 |
| 2009 | A Robust Mesh Watermarking Scheme Based on PCAabstractThis paper proposed a novel robust oblivious watermarking scheme in the spatial domain suitable for 3-D mesh object, which combines the principal component analysis (PCA) method with the construction of cone bins. Firstly, PCA is used to calculate the eigenvectors of covariance matrix of vertex coordinates. Then the object is rotated and translated so that its center of mass and the three eigenvectors coincide with the origin and the three axes of the Cartesian coordinate system. Subsequently, many cone bins are constructed, each bin center according to two angle parameters of spherical coordinate produced by pseudorandom number generators and the vertices are classified into the appropriate cone bins. Each cone bin is divided into a number of sub-bins in order to embed the watermark bit, and the size of sub-bin can make tradeoff between invisible and robustness. Experiment results show the remarkable ability of the proposed mechanism to resist against various attacks such as adding noise, clipping, similarity transform and vertex re-ordering. Xiaoqiang Li 0002, Wei Li 0012 |
ICIG | 2 |
| 2004 | Improved Robust Watermarking in DCT Domain for Color ImagesabstractIn the field of color images watermarking, many methods are accomplished by marking the image luminance, or by processing each color channel separately. This paper proposes a new DCT domain watermarking expressly devised for RGB color images based on the diversity technique in communication system. The watermark is hidden within the data in the same sequence by modifying a subset of block DCT coefficients of each color channel. Detection is based on a combination method by taking into account the information conveyed by three color channels. Even if a particular channel is severely faded, we may still be able to recover a reliable estimated of transmitted watermark through other propagation channel. Experimental results, as well as theoretical analysis, are presented to demonstrate the validity of the new approach with respect to algorithm operating on image luminance only. Xiaoqiang Li 0002, Xiangyang Xue 0001 |
AINA (1) | 1 |
| 2003 | An Optimized Multi-bits Blind Watermarking Scheme
Xiaoqiang Li 0002, Xiangyang Xue 0001, Wei Li 0012 |
ICICS | 1 |