Tianyu Li 0003

dblp:92/9835-3 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0001-5069-1493ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Seeing Beyond Illusion: Generalized and Efficient Mirror Detection
abstract
Reflective imaging enables the mirror imagings and physical entities to possess identical attributes, e.g., color and shape. Current mirror detection (MD) methods primarily rely on designing functional components to establish the correlation and disparities between the imagings and entities, thereby identifying the mirror regions. However, the exploration of extended scenes with dynamic content changes is rarely investigated. Therefore, we propose the MirrorSAM designed for MD based on the Segment Anything Model (SAM). Specifically, due to the varying reflections produced by mirrors in different positions and the complex visual space that interferes with localization, we design the hierarchical mixture of direction experts (HMDE) in the low-rank space to reduce biases towards entities in SAM and dynamically adjust experts based on the input scene. We observe differences in depth between mirrors and adjacent areas, and propose the depth token calibration (DTC), which introduces a learnable depth token to generate the depth map and serve as an error correction factor. We further formulate the selective pixel-prototype contrastive (SPPC) loss, selecting partially confusable samples to promote the decoupling of mirror and non-mirror representations. Extensive experiments conducted on four mirror benchmarks and two settings demonstrate that our approach surpasses state-of-the-art methods with few trainable parameters and FLOPs. We further extend to four transparent surface benchmarks to validate generalization.
Mingfeng Zha, Guoqing Wang 0001, Tianyu Li 0003, Wei Dong 0010, Peng Wang 0023, Yang Yang 0002
AAAI3
2026 RIR-Agent: An interactive framework for effective and adaptive restoration of remote sensing imagery
Junyu Liu, Tianyu Li 0003, Lanyue Liang, Gang Fu 0003, Guoqing Wang 0001, Quan Rui, Xiongxin Tang, Shuyuan Zhu, Yang Yang 0002
Expert Syst. Appl.2
2026 Think Twice Before Determining: Toward Scene-Aware Visual Reasoning for Mirror Detection
abstract
Mirror detection (MD) aims to overcome interference caused by reflections and locate mirror regions. Existing methods focus on designing components to explicitly establish the associations between physical entities and corresponding imagings, or utilizing rotation to construct symmetric consistency. We observe that: a) incomplete and incorrect correspondence between entities and imagings; b) other physical materials (e.g., glass) exhibit characteristics partially similar to mirrors, causing confusion when they co-occur; c) complex interfering factors (e.g., occlusion) and reflection mechanisms may expand vector space several times over. To address these issues in a unified manner, we formulate the scene-aware visual reasoning network (SVRNet) based on visual prompts. Specifically, we construct the prototype-guided prompt chain reasoning (PPCR) that generates a mixed chain of thought reasoning based on maximal difference heterogeneous prototypes to construct comprehensive spatial location and semantic perception. Noise may accumulate gradually through the chain, and crucial clues may also disappear. Therefore, we design the prompt evolution (PE) to filter out noise and enhance the coupling between prompts. We further develop the mixture of prompt injection expert (MPIE) to dynamically select the optimal injection strategy in the low-rank space based on specific scene. Due to reflection interference and random parameter space introducing potential ambiguity, we formulate the three-way evidence-aware (TEA) loss to quantify the uncertainty, thereby providing reliable predictions. To leverage historical knowledge and further disentangle representations, we propose the frequency prototype contrastive (FPC) loss for learning more generalizable features across images. Finally, we relabel 25,828 images and formulate the first point-supervised MD framework. Extensive experiments conducted on four mirror benchmarks under three settings demonstrate that our method surpasses state-of-the-art approaches. Promising results are also achieved on six related benchmarks, showing its generality.
Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Jiayi Ma 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.4
2026 Hierarchical Consistency Learning for Test-Time Adaptation in Camouflage Perception
abstract
Camouflaged object detection (COD) aims to localize targets that exhibit minimal perceptual differences from backgrounds through physical attributes. Existing methods, constrained by the static train-then-freeze paradigm, suffer from domain rigidity and annotation dependency, limiting their adaptability to scene variations and unseen camouflage patterns. To overcome these, we propose the hierarchical consistency learning (HCL) framework, which integrates test-time adaptation for dynamic representation recalibration. Specifically, we design the hierarchical representation reconstruction (HRR) to alleviate feature entanglement by synergizing spatial reconstruction with dual-stream frequency-domain decomposition, enhancing robustness against appearance homogenization. The pixel and spectrum inference provide structural and contextual priors. We further introduce task affinity guidance (TAG) to propagate knowledge across branches via channel-wise affinity, aligning local discriminative cues and mitigating semantic drift. To ensure semantic invariance, we formulate the prototype consistency calibration (PCC), which aggregates region features into compact prototypes and establishes prototype-feature similarity. This imposes implicit and hierarchical constraints that bridge task and representation gaps. Extensive experiments across four camouflaged and four underwater object benchmarks, under three degradation settings, demonstrate that our method consistently outperforms state-of-the-art approaches, highlighting its robustness and generalization under distribution shifts.
Mingfeng Zha, Tianyu Li 0003, Guoqing Wang 0001, Yunqiang Pei, Chaofan Qiao, Jiening Zhang, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Image Process.2
2025 Implicit Counterfactual Learning for Audio-Visual Segmentation
abstract
Audio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and imbalances. To overcome this, we propose the implicit counterfactual framework (ICF) to achieve unbiased cross-modal understanding. Due to the lack of semantics, heterogeneous representations may lead to erroneous matches, especially in complex scenes with ambiguous visual content or interference from multiple audio sources. We introduce the multi-granularity implicit text (MIT) involving video-, segment- and frame-level as the bridge to establish the modality-shared space, reducing modality gaps and providing prior guidance. Visual content carries more information and typically dominates, thereby marginalizing audio features in the decision-making. To mitigate knowledge preference, we propose the semantic counterfactual (SC) to learn orthogonal representations in the latent space, generating diverse counterfactual samples, thus avoiding biases introduced by complex functional designs and explicit modifications of text structures or attributes. We further formulate the collaborative distribution-aware contrastive learning (CDCL), incorporating factual-counterfactual and inter-modality contrasts to align representations, promoting cohesion and decoupling. Extensive experiments on three public datasets validate that the proposed method achieves state-of-the-art performance.
Mingfeng Zha, Tianyu Li 0003, Guoyin Wang 0001, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen
ICCV2
2025 Unlocking spatial textures: Gradient-guided pansharpening for enhancing multispectral imagery
Lanyue Liang, Tianyu Li 0003, Guoqing Wang 0001, Lin Mei 0001, Xiongxin Tang, Chaofan Qiao, Dongyu Xie
Neurocomputing2
2025 Toward Generalized and Realistic Unpaired Image Dehazing via Region-Aware Physical Constraints
abstract
Supervised dehazing models, trained on synthetic hazy-clean image pairs, often face a notable decline in performance when applied to real-world scenes. Consequently, CycleGAN-based unpaired dehazing methods are proposed to improve the model’s generalization. One successful approach among these methods involves decomposing the physical properties of the atmospheric scattering model (ASM). However, estimating physical properties individually from input images is difficult without supervised labels, which ignores the semantic consistency between different physical regions. We claim semantic region information can offer additional geometric spatial constraints for estimating physical properties, as natural images can be divided into regions with similar scene depths. Motivated by this, we propose a novel generalized and realistic unpaired image dehazing framework via region-aware physical constraints (RPC-Dehaze). Our approach utilizes fine-grained semantic region maps from the Segment Anything Model (SAM) in a specially designed region prompt enhancement module. This enables the dehazing and hazing cyclic networks to learn region-aware physical constraints, leading to accurate estimation of haze imaging physical properties. In contrast to existing unpaired methods that treat dehazing and hazing networks equally, we incorporate Retinex theory into the hazing network, allowing it to learn diverse illumination effects in different regions. We adaptively refine the Retinex-based illumination component, resulting in more realistic hazy images. To further facilitate unsupervised learning in our framework, we propose a physical consensual contrastive regularization to ensure compact representation constraints in the latent feature space. Extensive experiments on synthetic and real image datasets show our method surpasses state-of-the-art unpaired dehazing methods in both effectiveness and generalization capability.
Kaihao Lin, Guoqing Wang 0001, Tianyu Li 0003, Yuhui Wu 0001, Chongyi Li, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.3
2025 DMM: Disparity-Guided Multispectral Mamba for Oriented Object Detection in Remote Sensing
abstract
Multispectral oriented object detection faces challenges due to both inter-modal and intra-modal discrepancies. Recent studies often rely on transformer-based models to address these issues and achieve cross-modal fusion detection. However, the quadratic computational complexity of transformers limits their performance in remote sensing imagery. Inspired by the efficiency and lower complexity of Mamba in long sequence tasks, we propose Disparity-guided Multispectral Mamba (DMM), a multispectral oriented object detection framework comprised of a Disparity-guided Cross-modal Fusion Mamba (DCFM) module, a Multi-scale Target-aware Attention (MTA) module, and a Target-Prior Aware (TPA) auxiliary task. The DCFM module leverages disparity information between modalities to adaptively merge features from RGB and IR images, mitigating inter-modal conflicts. The MTA module aims to enhance feature representation by focusing on relevant target regions within the RGB modality, addressing intra-modal variations. The TPA auxiliary task utilizes single-modal labels to guide the optimization of the MTA module, ensuring it focuses on targets and their local context. Extensive experiments on the DroneVehicle and VEDAI datasets demonstrate the effectiveness of our method, which outperforms state-of-the-art methods while maintaining computational efficiency. Code will be available at https://github.com/Another-0/DMM.
Minghang Zhou, Tianyu Li 0003, Chaofan Qiao, Dongyu Xie, Guoqing Wang 0001, Ningjuan Ruan, Lin Mei 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Geosci. Remote. Sens.2
2025 Heterogeneous Experts and Hierarchical Perception for Underwater Salient Object Detection
abstract
Existing underwater salient object detection (USOD) methods design fusion strategies to integrate multimodal information, but lack exploration of modal characteristics. To address this, we separately leverage the RGB and depth branches to learn disentangled representations, formulating the heterogeneous experts and hierarchical perception network (HEHP). Specifically, to reduce modal discrepancies, we propose the hierarchical prototype guided interaction (HPI), which achieves fine-grained alignment guided by the semantic prototypes, and then refines with complementary modalities. We further design the mixture of frequency experts (MoFE), where experts focus on modeling high- and low-frequency respectively, collaborating to explicitly obtain hierarchical representations. To efficiently integrate diverse spatial and frequency information, we formulate the four-way fusion experts (FFE), which dynamically selects optimal experts for fusion while being sensitive to scale and orientation. Since depth maps with poor quality inevitably introduce noises, we design the uncertainty injection (UI) to explore high uncertainty regions by establishing pixel-level probability distributions. We further formulate the holistic prototype contrastive (HPC) loss based on semantics and patches to learn compact and general representations across modalities and images. Finally, we employ varying supervision based on branch distinctions to implicitly construct difference modeling. Extensive experiments on two USOD datasets and four relevant underwater scene benchmarks validate the effect of the proposed method, surpassing state-of-the-art binary detection models. Impressive results on seven natural scene benchmarks further demonstrate the scalability.
Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Chongyi Li, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Image Process.4
2024 Weakly-Supervised Mirror Detection via Scribble Annotations
abstract
Mirror detection is of great significance for avoiding false recognition of reflected objects in computer vision tasks. Existing mirror detection frameworks usually follow a supervised setting, which relies heavily on high quality labels and suffers from poor generalization. To resolve this, we instead propose the first weakly-supervised mirror detection framework and also provide the first scribble-based mirror dataset. Specifically, we relabel 10,158 images, most of which have a labeled pixel ratio of less than 0.01 and take only about 8 seconds to label. Considering that the mirror regions usually show great scale variation, and also irregular and occluded, thus leading to issues of incomplete or over detection, we propose a local-global feature enhancement (LGFE) module to fully capture the context and details. Moreover, it is difficult to obtain basic mirror structure using scribble annotation, and the distinction between foreground (mirror) and background (non-mirror) features is not emphasized caused by mirror reflections. Therefore, we propose a foreground-aware mask attention (FAMA), integrating mirror edges and semantic features to complete mirror regions and suppressing the influence of backgrounds. Finally, to improve the robustness of the network, we propose a prototype contrast loss (PCL) to learn more general foreground features across images. Extensive experiments show that our network outperforms relevant state-of-the-art weakly supervised methods, and even some fully supervised methods. The dataset and codes are available at https://github.com/winter-flow/WSMD.
Mingfeng Zha, Yunqiang Pei, Guoqing Wang 0001, Tianyu Li 0003, Yang Yang 0002, Wenbin Qian, Heng Tao Shen
AAAI4
2024 Region-Aware Distribution Contrast: A Novel Approach to Multi-task Partially Supervised Learning
Tianyu Li 0003, Guoqing Wang 0001, Peng Wang 0023, Yang Yang 0002, Jie Zou 0001
ECCV (51)2
2024 Cascaded Adversarial Attack: Simultaneously Fooling Rain Removal and Semantic Segmentation Networks
abstract
When applying high-level visual algorithms to rainy scenes, it is customary to preprocess the rainy images using low-level rain removal networks, followed by visual networks to achieve the desired objectives. Such a setting has never been explored by adversarial attack methods, which are only limited to attacking one kind of them. Considering the deficiency of multi-functional attacking strategies and the significance for open-world perception scenarios, we are the first to propose a Cascaded Adversarial Attack (CAA) setting, where the adversarial example can simultaneously attack different-level tasks, such as rain removal and semantic segmentation in an integrated system. Specifically, our attack on the rain removal network aims to preserve rain streaks in the output image, while for the semantic segmentation network, we employ powerful existing adversarial attack methods to induce misclassification of the image content. Importantly, CAA innovatively utilizes binary masks to effectively concentrate the aforementioned two significantly disparate perturbation distributions on the input image, enabling attacks on both networks. Additionally, we propose two variants of CAA, which minimize the differences between the two generated perturbations by introducing a carefully designed perturbation interaction mechanism, resulting in enhanced attack performance. Extensive experiments validate the effectiveness of our methods, demonstrating their superior ability to significantly degrade the performance of the downstream task compared to methods that solely attack a single network.
Zhiwen Wang 0004, Yuhui Wu 0001, Zheng Wang 0044, Jiwei Wei, Tianyu Li 0003, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen
ACM Multimedia5
2024 JoReS-Diff: Joint Retinex and Semantic Priors in Diffusion Model for Low-light Image Enhancement
abstract
Low-light image enhancement (LLIE) has achieved promising performance by employing conditional diffusion models. Despite the success of some conditional methods, previous methods may neglect the importance of a sufficient formulation of task-specific condition strategy, resulting in suboptimal visual outcomes. In this study, we propose JoReS-Diff, a novel approach that incorporates Retinex- and semantic-based priors as the additional pre-processing condition to regulate the generating capabilities of the diffusion model. We first leverage pre-trained decomposition network to generate the Retinex prior, which is updated with better quality by an adjustment network and integrated into a refinement network to implement Retinex-based conditional generation at both feature- and image-levels. Moreover, the semantic prior is extracted from the input image with an off-the-shelf semantic segmentation model and incorporated through semantic attention layers. By treating Retinex- and semantic-based priors as the condition, JoReS-Diff presents a unique perspective for establishing an diffusion model for LLIE and similar image enhancement tasks. Extensive experiments validate the rationality and superiority of our approach.
Yuhui Wu 0001, Guoqing Wang 0001, Zhiwen Wang 0004, Yang Yang 0002, Tianyu Li 0003, Malu Zhang, Chongyi Li, Heng Tao Shen
ACM Multimedia5
2024 Generalizing ISP Model by Unsupervised Raw-to-raw Mapping
abstract
ISP (Image Signal Processor) serves as a pipeline converting unprocessed raw images to sRGB images, positioned before nearly all visual tasks. Due to the varying spectral sensitivities of cameras, raw images captured by different cameras exist in different color spaces, making it challenging to deploy ISP across cameras with consistent performance. To address this challenge, it is intuitively to incorporate a raw-to-raw mapping (mapping raw images across camera color spaces) module into the ISP. However, the lack of paired data (i.e., images of the same scene captured by different cameras) makes it difficult to train a raw-to-raw model using supervised learning methods. In this paper, we aim to achieve ISP generalization by proposing the first unsupervised raw-to-raw model. To be specific, we propose a CSTPP (Color Space Transformation Parameters Predictor) module to predict the space transformation parameters in a patch-wise manner, which can accurately perform color space transformation and flexibly manage complex lighting conditions. Additionally, we design a CycleGAN-style training framework to realize unsupervised learning, overcoming the deficiency of paired data. Our proposed unsupervised model achieved performance comparable to that of the state-of-the-art semi-supervised method in raw-to-raw task. Furthermore, to assess its ability to generalize the ISP model across different cameras, we for the first formulated cross-camera ISP task and demonstrated the performance of our method through extensive experiments. The codes are released at https://github.com/ydxxxx/Unsupervised-Raw-to-raw-Mapping.
Dongyu Xie, Chaofan Qiao, Lanyue Liang, Zhiwen Wang 0004, Tianyu Li 0003, Qiao Liu 0003, Chongyi Li, Guoqing Wang 0001, Yang Yang 0002
ACM Multimedia5
2024 Physics-Constrained Comprehensive Optical Neural Networks
abstract
With the advantages of low latency, low power consumption, and high parallelism, optical neural networks (ONN) offer a promising solution for time-sensitive and resource-limited artificial intelligence applications. However, the performance of the ONN model is often diminished by the gap between the ideal simulated system and the actual physical system. To bridge the gap, this work conducts extensive experiments to investigate systematic errors in the optical physical system within the context of image classification tasks. Through our investigation, two quantifiable errors—light source instability and exposure time mismatches—significantly impact the prediction performance of ONN. To address these systematic errors, a physics-constrained ONN learning framework is constructed, including a well designed loss function to mitigate the effect of light fluctuations, a CCD adjustment strategy to alleviate the effects of exposure time mismatches and a ’physics-prior based’ error compensation network to manage other systematic errors, ensuring consistent light intensity across experimental results and simulations. In our experiments, the proposed method achieved a test classification accuracy of 96.5% on the MNIST dataset, a substantial improvement over the 61.6% achieved with the original ONN. For the more challenging QuickDraw16 and Fashion MNIST datasets, experimental accuracy improved from 63.0% to 85.7% and from 56.2% to 77.5%, respectively. Moreover, the comparison results further demonstrate the effectiveness of the proposed physics-constrained ONN learning framework over state-of-the-art ONN approaches. This lays the groundwork for more robust and precise optical computing applications.
Yanbing Liu 0006, Jianwei Qin, Xi Yue, Guoqing Wang 0001, Tianyu Li 0003, Fangwei Ye
NeurIPS7
2024 Towards constructing a DOE-based practical optical neural system for ship recognition in remote sensing images
Yanbing Liu 0006, Shaochong Liu, Tianyu Li 0003, Guoqing Wang 0001
Signal Process.4
2024 Dual Domain Perception and Progressive Refinement for Mirror Detection
abstract
Mirror detection aims to discover mirror regions in images to avoid misidentifying reflected objects. Existing methods mainly mine clues from spatial domain. We observe that the frequencies inside and outside the mirror region are distinctive. Besides, the low-frequency representing the feature semantics can help to locate the mirror region, and the high-frequency representing the details can refine it. Motivated by this, we introduce frequency guidance and propose the dual domain perception progressive refinement network (DPRNet) to mine dual-domain information. Specifically, we first decouple the images into high-frequency and low-frequency components by Laplace pyramid and vision Transformer, respectively, and design the frequency interaction alignment (FIA) module to integrate frequency features to initially localize the mirror region. To handle scale variations, we propose the multi-order feature perception (MOFP) module to adaptively aggregate adjacent features with progressive and gating mechanisms. We further propose the separation-based difference fusion (SDF) module to establish associations between entities and imagings and discover the correct boundary to mine the complete mirror region. Extensive experiments show that DPRNet outperforms the state-of-the-art method by an average of 3% with only about one-fifth of the parameters and FLOPs on four datasets. Our DPRNet also achieves promising performance on remote sensing and camouflage scenarios, validating its generalization. The code is available athttps://github.com/winter-flow/DPRNet.
Mingfeng Zha, Feiyang Fu, Yunqiang Pei, Guoqing Wang 0001, Tianyu Li 0003, Xiongxin Tang, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.5
2024 Density-Aware Cloud Removal of Remote Sensing Imagery Using a Global-Local Fusion Transformer
abstract
Cloud cover poses a significant challenge in remote sensing image processing, affecting the extraction and analysis of terrestrial features. Despite advancements in multitemporal cloud removal methods, single-image declouding remains crucial for emergency response and disaster management, where rapid acquisition of cloud-free imagery is essential. Traditional approaches often rely on synthetic aperture radar (SAR) or cloud masks as guidance for cloud removal, introducing additional complexities and dependencies on extensive data. To address these limitations, we propose a density-aware cloud removal using a global-local fusion Transformer (DCR-GLFT), which leverages density information as guidance and does not rely on extensive data. Specifically, our method employs density labels to guide the cloud removal process through two primary stages: cloud density estimation and density-guided cloud removal. A cloud density classifier is proposed in the first stage, trained with roughly estimated ground truth, to generate density labels for guiding subsequent removal processes. The second stage integrates cloud density information with cloud-ground image features using a Transformer-based network, enabling precise and nuanced cloud removal while preserving underlying surface details through the integration of both global and local features. The proposed method achieved the state-of-the-art results (peak signal-to-noise ratio (PSNR) of 28.93 and structural similarity index measure (SSIM) of 0.84) on the renowned cloud-removal dataset SEN12MS-CR, even without utilizing SAR data for guidance. This accomplishment highlights its significant advancement in the single-image cloud removal task. Our code will be made available athttps://github.com/ruiquan1214/DCR-GLFT.git.
Quan Rui, Shiyuan He, Tianyu Li 0003, Guoqing Wang 0001, Ningjuan Ruan, Lin Mei 0001, Yang Yang 0002, Heng Tao Shen
IEEE Trans. Geosci. Remote. Sens.3
2020 An Unsupervised Retinal Vessel Extraction and Segmentation Method Based On a Tube Marked Point Process Model
abstract
Retinal vessel extraction and segmentation is essential for supporting diagnosis of eye-related diseases. In recent years, deep learning has been applied to vessel segmentation and achieved excellent performance. However, these supervised methods require accurate hand-labeled training data, which may not be available. In this paper, we propose an unsupervised segmentation method based on our previous connected tube marked point process (MPP) model. The vessel network is extracted by the connected-tube MPP model first. Then a new tube-based segmentation method is applied to the extracted tubes. We test this method on STARE and DRIVE databases and the results show that not only do we extract the retina vessel network accurately, but we also achieve high G-means score for vessel segmentation, without using labeled training data.
Tianyu Li 0003, Mary L. Comer, Josiane Zerubia
ICASSP1
2019 Feature Extraction and Tracking of CNN Segmentations for Improved Road Detection from Satellite Imagery
abstract
Road detection in high-resolution satellite images is an important and popular research topic in the field of image processing. In this paper, we propose a novel road extraction and tracking method based on road segmentation results from a convolutional network, providing improved road detection. The proposed method incorporates our previously proposed connected-tube marked point process (MPP) model and a post-tracking algorithm. We present experimental results on the Massachusetts roads dataset to show the performance of our method on road detection in remotely-sensed images.
Tianyu Li 0003, Mary L. Comer, Josiane Zerubia
ICIP1
2018 A Connected-Tube MPP Model for Object Detection with Application to Materials and Remotely-Sensed Images
abstract
In this paper, we propose a connected-tube model based on a Marked Point Process (MPP) for strip feature extraction in images. This model incorporates a connection prior that favors certain connections between tubes based on their mutual positional relationship. Moreover, this model can easily be combined with other geometric models to form a mixed MPP model for more complex detection tasks. The proposed tube model is applied to fiber detection in microscopy images by combining connected-tube and ellipse models. The ellipse model is used for detecting short fibers, while the longer fibers are detected by tube model. We also test the model on road and building detection in remotely sensed images.
Tianyu Li 0003, Mary L. Comer, Josiane Zerubia
ICIP1
2012 The Extended Co-learning Framework for Robust Object Tracking
abstract
Recently, object tracking has been widely studied as a binary classification problem. Semi-supervised learning is particularly suitable for improving classification accuracy when large quantities of unlabeled samples are generated (just like tracking procedure). The purpose of this paper is to fulfill robust and stable tracking by using collaborative learning, which belongs to the scope of semi-supervised learning, among three classifiers. Different from [1], random fern classifier is incorporated to deal with 2bitBP feature newly added and certain constraints are specially implemented in our framework. Besides, the way for selecting positive samples is also altered by us in order to achieve more stable tracking. Algorithm proposed in this paper is validated by tracking pedestrian and cup under occlusion. Experiments and comparison show that our algorithm can avoid drifting problem to some degree and make tracking result more robust and adaptive.
Chen Gong 0002, Yang Liu 0007, Tianyu Li 0003, Jie Yang 0002, Xiangjian He
ICME3