EDBT 2026 Demo / reviewers in the wild / expert
Jiandong Tian
dblp:32/7871
· DBLP profile ↗
50ranked-venue papers
7as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 4 first-author · 23 since 2021Artificial intelligence and machine learning · 26 · 5 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SGC: A self-guided cascade multitask model for low-light object detection
Jiakun Jin, Junchao Zhang 0001, Yidong Luo, Jiandong Tian |
Knowl. Based Syst. | 4 |
| 2026 | MGAF: LiDAR-Camera 3D Object Detection With Multiple Guidance and Adaptive FusionabstractRecent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However, most of them overlook the complementary interaction and guidance between LiDAR and camera. In this work, we propose a novel multi-modality 3D objection detection method, with multi-guided global interaction and LiDAR-guided adaptive fusion, named MGAF. Specifically, we introduce sparse depth guidance (SDG) and LiDAR occupancy guidance (LOG) to generate 3D features with sufficient depth and spatial information. The designed semantic segmentation network captures category and orientation prior information for raw point clouds. In the following, an Adaptive Fusion Dual Transformer (AFDT) is developed to adaptively enhance the interaction of different modal BEV features from both global and bidirectional perspectives. Meanwhile, additional downsampling with sparse height compression and multi-scale dual-path transformer (MSDPT) are designed in order to enlarge the receptive fields of different modal features. Finally, a temporal fusion module is introduced to aggregate features from previous frames. Notably, the proposed AFDT is general, which also shows superior performance on other models. Our framework has undergone extensive experimentation on the large-scale nuScenes dataset, Waymo Open Dataset, and long-range Argoverse2 dataset, consistently demonstrating state-of-the-art performance. Baojie Fan, Caixia Xia, Huijie Fan, Fengyu Xu 0001, Jiandong Tian |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Lightweight Temporal-Frequency Perception Sparse State Space Models for Unified Image RestorationabstractUnified image restoration has become a fundamental issue in image processing. State space models have demonstrated significant potential in image restoration. However, their multi-directional scanning mechanism may introduce computational and feature redundancy, failing to satisfy lightweight deployment requirements. Furthermore, state space models have limitations in perceiving local detail features. To address this, we propose a lightweight channel-adaptive temporal-frequency sparse state space model for unified image restoration. This model enhances the local detail perception capability of the state space model using frequency domain features and simplifies the complexity of the network by sparse mechanisms. Specifically, we designed a U-shaped image restoration deep network based on the channel-adaptive temporal-frequency sparse state space module. This module consists of a temporal-domain dynamic sparse visual state space module and a frequency-domain sparse wavelet detail enhancement module in parallel, and uses a channel shuffling operation to realize temporal-frequency feature fusion. The dynamic sparse state space module uses a top-k mechanism to sparsify features across different scan paths for computational efficiency. The frequency-domain sparse wavelet detail enhancement module utilizes wavelet transformation and convolution operations to extract and enhance details in different directions, and then uses a top-k mechanism to perform sparse processing. Moreover, we introduce a degradation semantic perception module at the end of the encoder to guide the restoration network to adaptively learn the semantics of different degradation types, thereby realizing unified image restoration in complex outdoor environments. Extensive experimental results demonstrate that our method significantly outperforms 31 baseline methods in five complex weather and illumination degradation image restoration tasks while maintaining the lowest parameters and FLOPs. Pengyue Li, Yinke Dou, Jiandong Tian, Yandong Tang |
IEEE Trans. Image Process. | 5 |
| 2026 | Unsupervised Domain Adaptive Object Detection via Semantic Consistency and Compactness LearningabstractUnsupervised domain adaptive object detection methods enhance model robustness in the target domain without requiring target-domain annotations. Despite notable progress, existing methods face two major challenges: 1) insufficient and inefficient learning of holistic feature consistency due to cumbersome pixel-level style matching and semantic discrepancy elimination between domains as well as the overlooking of their collaborative effect; and 2) unreliable learning of category feature compactness caused by poor-quality target-domain samples, inaccurate pseudo-labels and noisy cross-domain contrast paradigms. To address these challenges, we propose a novel Semantic Consistency and Compactness Learning (SCCL) network. For consistency learning, we introduce a Visual Adaptation-guided Semantic Alignment (VSA) module that achieves style matching through simple feature adaptation and incorporates a novel adversarial-free self-supervised method for feature disentanglement. The collaboration between these two aspects enables sufficient and efficient consistency learning. For reliable compactness learning, we develop a plug-and-play Instance Center-Contrastive (ICC) head that, for the first time, comprehensively addresses all three potential causes of unreliable learning through three integrated innovations, concerning sample pseudo-label quality enhancement, reliable sample storage and updating, and a robust sample contrast paradigm. Besides, the mutual reinforcement effect of VSA and ICC simultaneously enhances feature transferability and discriminability. Extensive experiments across four UDA object detection benchmarks with two baselines show that SCCL achieves superior adaptability and robustness. Code will be available at https://github.com/TooZE23/SCCL. Yiming Su, Chunhui Hao, Xiyao Liu 0002, Jiandong Tian |
IEEE Trans. Image Process. | 6 |
| 2025 | GLAM: Global-Local Variation Awareness in Mamba-based World ModelabstractMimicking the real interaction trajectory in the inference of the world model has been shown to improve the sample efficiency of model-based reinforcement learning (MBRL) algorithms. Many methods directly use known state sequences for reasoning. However, this approach fails to enhance the quality of reasoning by capturing the subtle variation between states. Much like how humans infer trends in event development from this variation, in this work, we introduce Global-Local variation Awareness Mamba-based world model (GLAM) that improves reasoning quality by perceiving and predicting variation between states. GLAM comprises two Mamba-based parallel reasoning modules, GMamba and LMamba, which focus on perceiving variation from global and local perspectives, respectively, during the reasoning process. GMamba focuses on identifying patterns of variation between states in the input sequence and leverages these patterns to enhance the prediction of future state variation. LMamba emphasizes reasoning about unknown information, such as rewards, termination signals, and visual representations, by perceiving variation in adjacent states. By integrating the strengths of the two modules, GLAM accounts for higher-value variation in environmental changes, providing the agent with more efficient imagination-based training. We demonstrate that our method outperforms existing methods in normalized human scores on the Atari 100k benchmark. Wenqi Liang, Chunhui Hao, Gan Sun, Jiandong Tian |
AAAI | 5 |
| 2025 | RIOcc: Efficient Cross-Modal Fusion Transformer with Collaborative Feature Refinement for 3D Semantic Occupancy Prediction
Baojie Fan, Yuyu Jiang, Jiandong Tian, Huijie Fan |
ICCV | 5 |
| 2025 | D2 ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-Shot Action Recognition
Wenjie Pei, Qizhong Tan, Guangming Lu 0002, Jiandong Tian, Jun Yu 0002 |
ICCV | 4 |
| 2025 | Multimodal prompt state space models for unified adverse weather removal
Pengyue Li, Jiandong Tian, Yandong Tang |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Interaction-Aware Transformer Network for Human-Object Interaction Detectionabstracthuman-object interaction (HOI) detection tackles the problem of joint localization and classification of HOIs. Recent HOI detection methods are mainly based on transformer networks, where the explicit priors at the object level (e.g., scene layout, object appearance, or category) are usually fed into the transformer to improve the object query ability. Though these methods have achieved remarkable results, they did not pay enough attention to the implicit action-level information, which is the fundamental element of HOI. In this work, we propose an interaction-aware transformer network (IATN) to obtain the interaction-aware query, by jointly utilizing implicit action-level priors and explicit object-level priors. Specifically, we design an action-aware module (AAM) to aggregate implicit action priors from the scene level and instance level, respectively. Then, we design an action-oriented graph (AOG), where human feature and object feature are graph nodes and action semantics represent graph edges, to aggregate priors jointly from action level and object level. Afterwards, the interaction-aware query is acquired and finally adopted to obtain the HOI predictions. Besides, we leverage knowledge distillation to enhance the action-level priors by transferring the final HOI predictions to the intermediate features. Extensive experiments on HICO-DET and V-COCO datasets verify the effectiveness of our proposed interaction-aware model. Weibo Jiang, Weihong Ren, Jiandong Tian, Hanwei Ma, Bowen Chen 0004, Honghai Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2025 | LFUID: Light Field-Based Underwater Image Formation, Restoration, and Real-World DatasetabstractUnderwater imaging in turbid environments presents significant challenges for industrial applications due to optical distortions that severely degrade image quality. This article addresses the fundamental information deficit in traditional single-image restoration methods by introducing the first comprehensive Light Field Underwater Image Dataset, comprising 1356 images captured across diverse environments with 1820× 720× 14× 14× 3 resolution. We develop a novel physics-based image formation model that extends scattering principles to the 4-D light field domain, incorporating water attenuation coefficients and scattering properties while accounting for light field camera characteristics. Our model features three key components: an ambient underwater optical constant, a phase function for angular scattering, and a pixel position modulation function for spatial variations in backscatter. Based on this model, we propose a restoration algorithm that leverages complementary information across subaperture views. Experimental results demonstrate that our approach significantly outperforms state-of-the-art underwater image enhancement techniques across multiple metrics and diverse underwater scenes, establishing a promising new direction for underwater imaging applications. Shijun Zhou, Yiming Su, Doneyue Wang, Weihong Ren, Jiandong Tian |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | A Novel Dehazing Approach: Recovery of Color and Polarization Information Using Polarized CharacteristicsabstractPolarization provides valuable physical information, making it beneficial for various computer vision tasks. However, haze reduces both the color and polarization information of a scene. While existing single-image dehazing methods can restore color information, they are poor at recovering polarization information. Furthermore, current polarization-based dehazing approaches neglect the physical mechanisms of polarization degradation, resulting in inaccurate reconstruction of polarization information. In this paper, we propose a novel polarization dehazing algorithm, along with a polarization degradation model, to accurately recover both polarization and color information. First, we combine two key characteristics (the polarization achromatism prior and polarization attenuation prior) with the polarization degradation model to precisely reconstruct the scene's polarization. Then, we utilize the reconstructed polarization information to recover the color information of the scene. Finally, a multi-scale fusion optimization framework is introduced to further enhance the image quality. Our method shows excellent performance on both real-world indoor and outdoor polarized images, outperforming existing dehazing algorithms in both objective evaluation metrics and subjective visual assessment. Zhenshuo Yang, Chunhui Hao, Yiming Su, Yukuan Zhang, Junchao Zhang 0001, Jiandong Tian |
IEEE Trans. Multim. | 7 |
| 2025 | Soft Robotic Fish Actuated by Bionic Muscle With Embedded Sensing for Self-Adaptive Multiple Modes SwimmingabstractFish can adaptively adjust their body kinematics and swimming modes by sensing to realize optimal propulsion. However, most soft robotic fish have an unchangeable swimming mode through simple structure design, making them difficult to adapt to dynamic and complex fluid environments. Here, inspired by the multiple muscle synergy and lateral line sensing function of fish, we developed a soft robotic fish with multiple actuating units and embedded sensing elements. By collaboratively controlling the amplitude and phase of excitation from the multiple flexible actuating units, the soft robotic fish can successfully realize various swimming modes very similar to those of natural fish. Additionally, the embedded flexible sensing elements enable the robotic fish to sense the swimming state and the surrounding fluid environment in real time. The multiple actuation and embedded sensing allow the soft robotic fish to adaptively switch to an optimal swimming mode in a certain fluid environment. The multimode swimming and perception capabilities proposed in this work not only make soft robotic fish more intelligent and adaptable to complex fluid environments, but also contribute to the future implementation of autonomous control capabilities for robotic fish. Ruiqian Wang, Wenjun Tan, Yiwei Zhang 0012, Lianchao Yang, Wenyuan Chen, Jiandong Tian, Lianqing Liu |
IEEE Trans. Robotics | 8 |
| 2024 | Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection plays a vital role in scene understanding, which aims to predict the HOI triplet in the form of . Existing methods mainly extract multi-modal features (e.g., appearance, object semantics, human pose) and then fuse them together to directly predict HOI triplets. However, most of these methods focus on seeking for self-triplet aggregation, but ignore the potential cross-triplet dependencies, resulting in ambiguity of action prediction. In this work, we propose to explore Self- and Cross-Triplet Correlations (SCTC) for HOI detection. Specifically, we regard each triplet proposal as a graph where Human, Object represent nodes and Action indicates edge, to aggregate self-triplet correlation. Also, we try to explore cross-triplet dependencies by jointly considering instance-level, semantic-level, and layout-level relations. Besides, we leverage the CLIP model to assist our SCTC obtain interaction-aware feature by knowledge distillation, which provides useful action clues for HOI detection. Extensive experiments on HICO-DET and V-COCO datasets verify the effectiveness of our proposed SCTC. Weibo Jiang, Weihong Ren, Jiandong Tian, Liangqiong Qu, Zhiyong Wang 0009, Honghai Liu 0001 |
AAAI | 3 |
| 2024 | SA²VP: Spatially Aligned-and-Adapted Visual PromptabstractAs a prominent parameter-efficient fine-tuning technique in NLP, prompt tuning is being explored its potential in computer vision. Typical methods for visual prompt tuning follow the sequential modeling paradigm stemming from NLP, which represents an input image as a flattened sequence of token embeddings and then learns a set of unordered parameterized tokens prefixed to the sequence representation as the visual prompts for task adaptation of large vision models. While such sequential modeling paradigm of visual prompt has shown great promise, there are two potential limitations. First, the learned visual prompts cannot model the underlying spatial relations in the input image, which is crucial for image encoding. Second, since all prompt tokens play the same role of prompting for all image tokens without distinction, it lacks the fine-grained prompting capability, i.e., individual prompting for different image tokens. In this work, we propose the Spatially Aligned-and-Adapted Visual Prompt model (SA^2VP), which learns a two-dimensional prompt token map with equal (or scaled) size to the image token map, thereby being able to spatially align with the image map. Each prompt token is designated to prompt knowledge only for the spatially corresponding image tokens. As a result, our model can conduct individual prompting for different image tokens in a fine-grained manner. Moreover, benefiting from the capability of preserving the spatial structure by the learned prompt token map, our SA^2VP is able to model the spatial relations in the input image, leading to more effective prompting. Extensive experiments on three challenging benchmarks for image classification demonstrate the superiority of our model over other state-of-the-art methods for visual prompt tuning. Code is available at https://github.com/tommy-xq/SA2VP. Wenjie Pei, Tongqi Xia, Fanglin Chen 0001, Jiandong Tian, Guangming Lu 0002 |
AAAI | 5 |
| 2024 | Robust 3D Tracking with Quality-Aware Shape Completionabstract3D single object tracking remains a challenging problem due to the sparsity and incompleteness of the point clouds. Existing algorithms attempt to address the challenges in two strategies. The first strategy is to learn dense geometric features based on the captured sparse point cloud. Nevertheless, it is quite a formidable task since the learned dense geometric features are with high uncertainty for depicting the shape of the target object. The other strategy is to aggregate the sparse geometric features of multiple templates to enrich the shape information, which is a routine solution in 2D tracking. However, aggregating the coarse shape representations can hardly yield a precise shape representation. Different from 2D pixels, 3D points of different frames can be directly fused by coordinate transform, i.e., shape completion. Considering that, we propose to construct a synthetic target representation composed of dense and complete point clouds depicting the target shape precisely by shape completion for robust 3D tracking. Specifically, we design a voxelized 3D tracking framework with shape completion, in which we propose a quality-aware shape completion mechanism to alleviate the adverse effect of noisy historical predictions. It enables us to effectively construct and leverage the synthetic target representation. Besides, we also develop a voxelized relation modeling module and box refinement module to improve tracking performance. Favorable performance against state-of-the-art algorithms on three benchmarks demonstrates the effectiveness and generalization ability of our method. Zikun Zhou, Guangming Lu 0002, Jiandong Tian, Wenjie Pei |
AAAI | 4 |
| 2024 | GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object DetectionabstractRecent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However, most of them overlook the complementary interaction and guidance be-tween LiDAR and camera. In this work, we propose a novel multi-modality 3D objection detection method, named GA-Fusion, with LiDAR-guided global interaction and adaptive fusion. Specifically, we introduce sparse depth guidance (SDG) and LiDAR occupancy guidance (LOG) to generate 3D features with sufficient depth information. In the following, LiDAR-guided adaptive fusion transformer (LGAFT) is developed to adaptively enhance the interaction of different modal BEV features from a global perspective. Meanwhile, additional downsampling with sparse height compression and multi-scale dual-path transformer (MSDPT) are de-signed to enlarge the receptive fields of different modal features. Finally, a temporal fusion module is introduced to ag-gregate features from previous frames. GAFusion achieves state-of-the-art 3D object detection results with 73.6% mAP and 74.9% NDS on the nuScenes test set. Baojie Fan, Jiandong Tian, Huijie Fan |
CVPR | 3 |
| 2024 | Unbiased Faster R-CNN for Single-source Domain Generalized Object DetectionabstractSingle-source domain generalization (SDG) for object detection is a challenging yet essential task as the distribution bias of the unseen domain degrades the algorithm per-formance significantly. However, existing methods attempt to extract domain-invariant features, neglecting that the bi-ased data leads the network to learn biased features that are non-causal and poorly generalizable. To this end, we pro-pose an Unbiased Faster R-CNN (UFR) for generalizable feature learning. Specifically, we formulate SDG in object detection from a causal perspective and construct a Struc-tural Causal Model (SCM) to analyze the data bias andfeature bias in the task, which are caused by scene confounders and object attribute confounders. Based on the SCM, we de-sign a Global-Local Transformation module for data aug-mentation, which effectively simulates domain diversity and mitigates the data bias. Additionally, we introduce a Causal Attention Learning module that incorporates a designed at-tention invariance loss to learn image-level features that are robust to scene confounders. Moreover, we develop a Causal Prototype Learning module with an explicit instance constraint and an implicit prototype constraint, which fur-ther alleviates the negative impact of object attribute con-founders. Experimental results on five scenes demonstrate the prominent generalization ability of our method, with an improvement of 3.9% mAP on the Night-Clear scene. Shijun Zhou, Xiyao Liu 0002, Chunhui Hao, Baojie Fan, Jiandong Tian |
CVPR | 6 |
| 2024 | QO-Net: Query Optimization Underwater Object Detection NetworkabstractUnderwater object detection has attracted increasing interest for its wide application in various underwater tasks. However, due to underwater image quality degradation and the lack of large-scale underwater object datasets, many underwater detectors suffer from low detection performance. To address the issues, we not only propose a novel underwater transformer detector with multi-scale feature enhancement and query optimization, named QO-Net, but also construct a new underwater object detection dataset, called UODD. Specifically, a Conv-Trans Layer is developed as the unit of QO-Net, which effectively learns multi-scale image feature representation through CNN and simultaneously captures the dependencies among different positions in the sequence data through Transformer, enabling QO-Net to process underwater image sequence information over longer distances. An effective combination can enhance the representation of multi-scale features. Then, QO-Net develops a positional query enhancement strategy to optimize the spatial prior of positional queries, thereby speeding up the convergence of the network training. In addition, UODD also contains more than 20,000 underwater images for training and validation, with a variety of rich underwater categories. Extensive experiments on UODD, Brackish, and TrashCan datasets demonstrate that QO-Net presents favorable detection performance against state-of-the-art methods in terms of robustness and accuracy. Jiandong Tian, Hongyang Sun 0003, Baojie Fan, Hongxin Xu |
IROS | 1 |
| 2024 | A Novel Image Formation Model for DescatteringabstractIn the field of image descattering, the image formation models employed for restoration approaches are often simplified. In these models, scattering distribution is uniform in homogeneous media when transmission is fixed. Through specifically designed experiments, we discover that scattering exhibits non-uniform characteristics even in homogeneous media. Neglecting non-uniform scattering in these models limits their accuracy in representing scattering distribution, resulting in existing image descattering approaches inadequate. To tackle these issues, this paper proposes a novel image formation model for image descattering, considering more physical parameters, such as zenith angle, azimuth angle, scattering phase function, and camera focal length. Our model describes the light transfer process in scattering media more accurately. For image descattering, we introduce corresponding algorithms for parameter estimation in our model and simultaneous restoration from degraded images. Experimental evaluations demonstrate the effectiveness of our proposed model in various tasks, including physical parameter estimation, pure-scattering removal, image dehazing, and underwater image restoration. In terms of calculating parameters, our results are close to the real values; in terms of underwater image restoration, our work outperforms the state-of-art methods; in terms of image dehazing, our work promotes the performance of existing methods by replacing previous models with our model. Jiandong Tian, Shijun Zhou, Baojie Fan, Hui Zhang 0023 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | HCPVF: Hierarchical Cascaded Point-Voxel Fusion for 3D Object DetectionabstractWith the astonishing development of 3D sensors, point cloud based 3D object detection is attracting increasing attention from both industry and academia, and widely applied in various fields, such as robotics and autonomous driving. However, how to balance the 3D object detecting accuracy and speed is still a challenging problem. In this paper, we study this issue and propose a novel and effective 3D point cloudy object detection network based on hierarchical cascaded point-voxel fusion, called HCPVF. Firstly, a novel bird’s-eye-view(BEV) attention mechanism with linear complexity is developed to improve point cloud feature backbone network, which can be implemented easily to mine the point-to-point similarity in BEV’s view, by two cascaded linear layers and two normalization layers. This operation captures long-range dependencies and reduces the uneven sampling of sparse BEV features, making the extracted point cloudy features more discriminative. Secondly, the proposed HCPVF module is equipped with dual-level hierarchical cascaded detection head, including voxel level and the following point level. The voxel level is composed of coarse Region of interest(RoI) pooling and fine RoI pooling, which are cooperated to aggregate voxel features from different grid divisions and predict relatively coarse detection boxes. In the following, the point level is based on Key Points Transformer. It firstly encodes the spatial context information between the original point and the voxel level box. And then, a novel dual-weighted decoder is developed to enhance the context interaction by weighting the channel and spatial dimensions to obtain more accurate detection results. This design utilizes the voxel based method with high computational efficiency and the point based method with more complete spatial information, fusing low-level voxel features and high-level point features through hierarchical cascaded strategy. Extensive experiments demonstate that the proposed HCPVF achieves state-of-the-art 3D detection performance while maintaining computational efficiency on both the Waymo Open Dataset and the highly-competitive KITTI benchmark. Baojie Fan, Jiandong Tian |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Learning Self- and Cross-Triplet Context Clues for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection aims to infer interactions between humans and objects, and it is very important for scene analysis and understanding. The existing methods usually focus on exploring instance-level (e.g., object appearance) or interaction-level (e.g., action semantic) features to conduct interaction prediction. However, most of these methods only consider the self-triplet feature aggregation, which may lead to learning ambiguity without exploring the cross-triplet context exchange. In this paper, from both visual and textual perspectives, we propose a novel method to jointly explore self-and cross-triplet interaction context clues for HOI detection. First, we employ a graph neural network to perform self-triplet aggregation, where human and object features represent graph nodes and visual interaction feature and textual prior knowledge are acted as two different edges. Furthermore, we also attempt to explore cross-triplet context exchange by incorporating symbiotic and layout relationships among different HOI triplets. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones and achieves the impressive performance of 40.32 mAP on HICO-DET and 69.1 mAP on V-COCO datasets, respectively. Weihong Ren, Jinguo Luo, Weibo Jiang, Liangqiong Qu, Zhi Han, Jiandong Tian, Honghai Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Learning a Non-Locally Regularized Convolutional Sparse Representation for Joint Chromatic and Polarimetric DemosaickingabstractDivision of focal plane color polarization camera becomes the mainstream in polarimetric imaging for it directly captures color polarization mosaic image by one snapshot, so image demosaicking is an essential task. Current color polarization demosaicking (CPDM) methods are prone to unsatisfied results since it's difficult to recover missed 15 or 14 pixels out of 16 pixels in color polarization mosaic images. To address this problem, a non-locally regularized convolutional sparse regularization model, which is advantaged in denoising and edge maintaining, is proposed to recall more information for CPDM task, and the CPDM task is transformed into an energy function to be solved by ADMM optimization. Finally, the optimal model generates informative and clear results. The experimental results, including reconstructed synthetic and real-world scenes, demonstrate that our proposed method outperforms the current state-of-the-art methods in terms of quantitative measurements and visual quality. The source code is available at https://github.com/roydon-luo/NLCSR-CPDM. Yidong Luo, Junchao Zhang 0001, Jianbo Shao, Jiandong Tian, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | A Shadow Imaging Bilinear Model and Three-Branch Residual Network for Shadow RemovalabstractThe current shadow removal pipeline relies on the detected shadow masks, which have limitations for penumbras and tiny shadows, and results in an excessively long pipeline. To address these issues, we propose a shadow imaging bilinear model and design a novel three-branch residual (TBR) network for shadow removal. Our bilinear model reveals the single-image shadow removal process and can explain why simply increasing the brightness of shadow areas cannot remove shadows without artifacts. We considerably shorten the shadow removal pipeline by modeling illumination compensation and developing a single-stage shadow removal network without additional detection and refinement networks. Specifically, our network consists of three task branches, i.e., shadow image reconstruction, shadow matte estimation, and shadow removal. To merge these three branches and enhance the shadow removal branch, we design a model-based TBR module. Multiple TBR modules are cascaded to generate an intensive information flow and facilitate feature integration among the three branches. Thus, our network ensures the fidelity of nonshadow areas and restores the light intensity of shadow areas through three-branch collaboration. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods. The model and code are available at https://github.com/nachifur/TBRNet. Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Jiandong Tian, Yandong Tang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Progressive feature-aware recurrent net for low-light image enhancement
Pengyue Li, Xiai Chen, Jiandong Tian, Yandong Tang |
Signal Process. Image Commun. | 3 |
| 2023 | Two-Way Complementary Tracking GuidanceabstractRecently, most impressive Siamese network-based trackers are equipped with two independent branches: tracked object classification and bounding box regression. However, there is no tracking information exchange between them during the tracking optimization process. This may lead to the task-mismatch and accuracy inconsistency between both classification and regression branches during inference. To tackle the problems, we propose a novel Mutual Guidance (MG) strategy for visual object tracking, which constructs the bidirectional and complementary tracking information interaction to maintain the tracked object is well-classified to also be well-localized, between classification and regression branches. Specifically, the classification branch can guide the regression one to pay more attention to the sample with high classified scores, by re-weighting the regression loss with the classification confidence. Similarity, the regression branch also guides the classifier optimization process to focus on samples with larger IoU values. And then, the proposed Mutual Guidance is completed by a series of regularization designs on classification score and regression IoU, which dynamically re-assign the adaptive weights to the losses for each sample during the joint tracking optimization. The developed MG is generic and easy to be plugged into various tracking frameworks such as anchor-based, anchor-free based and transformer based, and boost their performance to some extent with negligible additional cost. In addition, we also develop an adaptive localization(L) branch selection scheme to further assist trackers, which determines proper localization branch for different trackers according to the difference in the way of discriminating positive and negative samples. Extensive experiments verify the effectiveness of MGL and its superiority against the state-of-the-art tracking modules on OTB100, GOT-10K, LaSOT, TrackingNet, UAV123, VOT2018 and VOT2019. Baojie Fan, Guoping Jiang, Jiandong Tian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Few-Shot Object Detection by Knowledge Distillation Using Bag-of-Visual-Words Representations
Wenjie Pei, Dianwen Mei, Fanglin Chen 0001, Jiandong Tian, Guangming Lu 0002 |
ECCV (10) | 5 |
| 2022 | Multi-faceted Distillation of Base-Novel Commonality for Few-Shot Object Detection
Wenjie Pei, Dianwen Mei, Fanglin Chen 0001, Jiandong Tian, Guangming Lu 0002 |
ECCV (9) | 5 |
| 2022 | Dual Aligned Siamese Dense Regression TrackerabstractAnchor or anchor-free based Siamese trackers have achieved the astonishing advancement. However, their parallel regression and classification branches lack the tracked target information link and interaction, and the corresponding independent optimization maybe lead to task-misalignment, such as the reliable classification prediction with imprecisely localization and vice versa. To address this problem, we develop a general Siamese dense regression tracker (SDRT) with both task and feature alignments. It consists of two cooperative and mutual-guidance core branches: dense local regression with RepPoint representation, the global and local multi-classifier fusion with aligned features. They complement and boost each other to constrain the results with well-localized followed to also be well-classified. Specifically, a dense local regression with RepPoint representation, directly estimates and averages multiple dense local bounding box offsets for accurate localization. And then, the refined bounding boxes can be used to learn the global and local affine alignment features for reliable multi-classifier fusion. The classified scores in turn guide the assigned positive bounding boxes for the regression task. The mutual guidance operations can bridge the connection between classification and regression substantially, since the assigned labels of one task depend on the prediction quality of the other task. The proposed tracking module is general, and it can boost both the anchor or anchor-free based Siamese trackers to some extent. The extensive tracking comparisons on six tracking benchmarks verify its favorable and competitive performance over states-of-the-arts tracking modules. Baojie Fan, Hui Zhang 0023, Yang Cong, Yandong Tang, Huijie Fan, Jiandong Tian |
IEEE Trans. Image Process. | 6 |
| 2022 | Discriminative Siamese Complementary Tracker With Flexible UpdateabstractThe offline generative Siamese trackers are equipped with the pre-defined anchors and the fixed target template. They overlook the target-background discriminative information, and lack the flexible target-specific update strategy. To overcome above drawbacks, we propose an adaptive and discriminative Siamese complementary tracking network with flexible update scheme. It consists of three collaborate subnetworks: anchor-free Siamese attention classification and regression subnetwork, online discriminative learning with multi-attention and multi-peak suppression, classifier guided template update subnetwork. All of them are interdependent and complementary to enhance each other for accurate target location. Specifically, an anchor-free multi-attention Siamese tracking subnetwork directly classifies the corresponding image patches with reliability assessment, and cascaded regresses the bounding boxes to progressively refine the predicting accuracy. Its evaluation is flexible and general with both proposal and anchor free in per-pixel prediction manner. Then, we integrate an online discriminative classifier optimizing module as a complementary subnetwork. It introduces spatial-temporal attention mechanism to fully explore multi-view multi-scale target-specific features, and evaluates multi-peak suppression to obtain a single centered peak response map. Its classified results can be fused with Siamese classification branch for accurate target location. Finally, the template update subnetwork is guided by the online discriminative classification scores. Extensive experiments on recent tracking datasets verify its top-ranked tracking accuracy and robustness against some state-of-the-art trackers. Baojie Fan, Jiandong Tian, Yan Peng 0001, Yandong Tang |
IEEE Trans. Multim. | 2 |
| 2021 | Dynamic and reliable subtask tracker with general schatten p-norm regularization
Baojie Fan, Yang Cong, Jiandong Tian, Yandong Tang |
Pattern Recognit. | 3 |
| 2021 | Structured and Consistent Multi-Layer Multi-Kernel Subtask Correction Filter TrackerabstractSome multi-task correlation filter trackers achieve the top-ranked performance in terms of accuracy and robustness. However, they directly fuse multiple types of features into a single kernel space. This operation fails to fully explore the discriminative strength and diversity of different features, and also ignores the structured correspondence of different tasks. To solve these issues, we propose a structured multi-kernel subtask correlation filter tracker with temporal-spatial consistency, which enjoys the merits of both layered multi-kernel subtask learning and structured correlation filter. Specifically, we firstly assign one kernel space to each channel feature. Multi-channel features correspond to multi-kernel spaces to boost their powerful discriminability. And then, we divide the target into multi-layer patches with different sizes, and regard the correlation filter trace of each patch with one channel feature as a subtask. In the following, we incorporate globally and locally structured correlation filters into a unified multi-kernel subtask particle tracking framework. The global and local subtasks complement and enhance each other with similar motion model. The proposed tracker not only exploits the cooperation and complementarity of layered multi-kernel subtask correlation filters, but also mines the underlying geometric structure of global subtasks, and the inner spatial locality correspondences of local subtasks inside the target. This operation is achieved by dual group sparsity regularized terms with mixed-norm lp,q, which decomposes the multi-kernel subtask filter matrix into two collaborative components. They correspond to the adaptive filter feature selection and outlier subtask detection, respectively. Besides, the developed tracking model maintains the temporal coherence and spatial consistency of multi-layer subtask filters via the smooth regularizer. Finally, the tracking formulation is optimized by the accelerated proximal gradient approach (APG). Encouraging analyses on six benchmark datasets, verify the favorable effectiveness and robustness of our method against state-of-the-art trackers. Baojie Fan, Yang Cong, Yandong Tang, Jiandong Tian, Chenliang Xu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Deep Retinex Network for Single Image DehazingabstractIn this paper, we propose a retinex-based decomposition model for a hazy image and a novel end-to-end image dehazing network. In the model, the illumination of the hazy image is decomposed into natural illumination for the haze-free image and residual illumination caused by haze. Based on this model, we design a deep retinex dehazing network (RDN) to jointly estimate the residual illumination map and the haze-free image. Our RDN consists of a multiscale residual dense network for estimating the residual illumination map and a U-Net with channel and spatial attention mechanisms for image dehazing. The multiscale residual dense network can simultaneously capture global contextual information from small-scale receptive fields and local detailed information from large-scale receptive fields to precisely estimate the residual illumination map caused by haze. In the dehazing U-Net, we apply the channel and spatial attention mechanisms in the skip connection of the U-Net to achieve a trade-off between overdehazing and underdehazing by automatically adjusting the channel-wise and pixel-wise attention weights. Compared with scattering model-based networks, fully data-driven networks, and prior-based dehazing methods, our RDN can avoid the errors associated with the simplified scattering model and provide better generalization ability with no dependence on prior information. Extensive experiments show the superiority of the RDN to various state-of-the-art methods. Pengyue Li, Jiandong Tian, Yandong Tang, Guolin Wang, Chengdong Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Tracking-by-Counting: Using Network Flows on Crowd Density Maps for Tracking Multiple TargetsabstractState-of-the-art multi-object tracking (MOT) methods follow the tracking-by-detection paradigm, where object trajectories are obtained by associating per-frame outputs of object detectors. In crowded scenes, however, detectors often fail to obtain accurate detections due to heavy occlusions and high crowd density. In this paper, we propose a new MOT paradigm, tracking-by-counting, tailored for crowded scenes. Using crowd density maps, we jointly model detection, counting, and tracking of multiple targets as a network flow program, which simultaneously finds the global optimal detections and trajectories of multiple targets over the whole video. This is in contrast to prior MOT methods that either ignore the crowd density and thus are prone to errors in crowded scenes, or rely on a suboptimal two-step process using heuristic density-aware point-tracks for matching targets. Our approach yields promising results on public benchmarks of various domains including people tracking, cell tracking, and fish tracking. Weihong Ren, Xinchao Wang, Jiandong Tian, Yandong Tang, Antoni B. Chan |
IEEE Trans. Image Process. | 3 |
| 2020 | Dual Refinement Underwater Object Detection Network
Baojie Fan, Yang Cong, Jiandong Tian |
ECCV (20) | 4 |
| 2020 | Dually Connected Deraining Net Using Pixel-Wise AttentionabstractRecent single image deraining methods either use a recurrent mechanism to gradually learn the mapping between clear images and rainy images, or focus on designing various loss functions to supervise the learning process. In this letter, we propose a dually connected deraining net using pixel-wise attention, for single image rain removal. Specifically, the deraining net adopts an encoder-decoder net as a backbone, which can effectively learn a residual rain-streaks map by jointly using skip sum connection and skip concatenation connection. The dual connections enable the deraining net to promote information flow between layers, and thus can allow it to discriminate and localize the rain streaks. To preserve image details, the decoded features are weighted by the learnable pixel-wise attention for adaptively recalibrating their responses. Experimental results on synthetic datasets demonstrate that the proposed model outperforms the recent state-of-the-art deraining methods. Weihong Ren, Jiandong Tian, Qiang Wang 0015, Yandong Tang |
IEEE Signal Process. Lett. | 2 |
| 2020 | Reliable Multi-Kernel Subtask Graph Correlation TrackerabstractMany astonishing correlation filter trackers pay limited concentration on the tracking reliability and locating accuracy. To solve the issues, we propose a reliable and accurate cross correlation particle filter tracker via graph regularized multi-kernel multi-subtask learning. Specifically, multiple non-linear kernels are assigned to multi-channel features with reliable feature selection. Each kernel space corresponds to one type of reliable and discriminative features. Then, we define the trace of each target subregion with one feature as a single view, and their multi-view cooperations and interdependencies are exploited to jointly learn multi-kernel subtask cross correlation particle filters, and make them complement and boost each other. The learned filters consist of two complementary parts: weighted combination of base kernels and reliable integration of base filters. The former is associated to feature reliability with importance map, and the weighted information reflects different tracking contribution to accurate location. The second part is to find the reliable target subtasks via the response map, to exclude the distractive subtasks or backgrounds. Besides, the proposed tracker constructs the Laplacian graph regularization via cross similarity of different subtasks, which not only exploits the intrinsic structure among subtasks, and preserves their spatial layout structure, but also maintains the temporal-spatial consistency of subtasks. Comprehensive experiments on five datasets demonstrate its remarkable and competitive performance against state-of-the-art methods. Baojie Fan, Yang Cong, Jiandong Tian, Yandong Tang |
IEEE Trans. Image Process. | 3 |
| 2019 | Stacked dense networks for single-image snow removal
Pengyue Li, Mengshen Yun, Jiandong Tian, Yandong Tang, Guolin Wang, Chengdong Wu 0001 |
Neurocomputing | 3 |
| 2018 | Evaluation of shadow featuresabstractShadow features such as colour ratio, texture, and chromaticity have proved to be quite effective in shadow detection. Many shadow detection methods have been proposed on the basis of different features. However, previous works for shadow detection mainly focus on designing an effective classifier for existing shadow features, but pay less attention on the analysis of shadow features themselves. The majority of studies simply report the final shadow detection results rather than make an evaluation on each feature. Readers often do not know which features are more effective or whether these shadow features are complementary. The following problems are still unsolved: the robustness of each feature, which feature plays the most important role in a detection method, and what is the best performance that current features can reach. The purpose of this study is to answer these questions, and the authors hope that this study can offer guidance for future shadow detection algorithms via the evaluation of frequently used shadow features. Several useful and interesting conclusions are obtained after conducting extensive comparison experiments on a large dataset. Liangqiong Qu, Jiandong Tian, Huijie Fan, Yandong Tang |
IET Comput. Vis. | 2 |
| 2018 | Snowflake Removal for Videos via Global and Local Low-Rank DecompositionabstractFalling snow not only blocks human vision, but also significantly degrades the effectiveness of computer vision systems in outdoor environment. In this paper, we aim to remove snowflakes in videos by using the global and local low-rank property of snowflake-removed scenes. The stationary background and the mixture of moving foreground as well as falling snowflake are extracted via the global low-rank matrix decomposition. Some snowflake features, such as its color and size, are used to separate out the snowflakes from other moving objects. Then, the mean absolute difference based patch matching is applied to align every same moving object over frames to grab its low-rank structure. As such, the falling snowflake in front of moving objects can be removed via the local low-rank decomposition. Finally, the snowflake removed videos are generated by pasting moving foreground to stationary backgrounds. Experiments show that our method can remove snowflakes effectively and outperforms the comparison methods. Jiandong Tian, Zhi Han, Weihong Ren, Xiai Chen, Yandong Tang |
IEEE Trans. Multim. | 1 |
| 2017 | DeshadowNet: A Multi-context Embedding Deep Network for Shadow RemovalabstractShadow removal is a challenging task as it requires the detection/annotation of shadows as well as semantic understanding of the scene. In this paper, we propose an automatic and end-to-end deep neural network (DeshadowNet) to tackle these problems in a unified manner. DeshadowNet is designed with a multi-context architecture, where the output shadow matte is predicted by embedding information from three different perspectives. The first global network extracts shadow features from a global view. Two levels of features are derived from the global network and transferred to two parallel networks. While one extracts the appearance of the input image, the other one involves semantic understanding for final prediction. These two complementary networks generate multi-context features to obtain the shadow matte with fine local details. To evaluate the performance of the proposed method, we construct the first large scale benchmark with 3088 image pairs. Extensive experiments on two publicly available benchmarks and our large-scale benchmark show that the proposed method performs favorably against several state-of-the-art methods. Liangqiong Qu, Jiandong Tian, Shengfeng He, Yandong Tang, Rynson W. H. Lau |
CVPR | 2 |
| 2017 | Video Desnowing and Deraining Based on Matrix DecompositionabstractThe existing snow/rain removal methods often fail for heavy snow/rain and dynamic scene. One reason for the failure is due to the assumption that all the snowflakes/rain streaks are sparse in snow/rain scenes. The other is that the existing methods often can not differentiate moving objects and snowflakes/rain streaks. In this paper, we propose a model based on matrix decomposition for video desnowing and deraining to solve the problems mentioned above. We divide snowflakes/rain streaks into two categories: sparse ones and dense ones. With background fluctuations and optical flow information, the detection of moving objects and sparse snowflakes/rain streaks is formulated as a multi-label Markov Random Fields (MRFs). As for dense snowflakes/rain streaks, they are considered to obey Gaussian distribution. The snowflakes/rain streaks, including sparse ones and dense ones, in scene backgrounds are removed by low-rank representation of the backgrounds. Meanwhile, a group sparsity term in our model is designed to filter snow/rain pixels within the moving objects. Experimental results show that our proposed model performs better than the state-of-the-art methods for snow and rain removal. Weihong Ren, Jiandong Tian, Zhi Han, Antoni B. Chan, Yandong Tang |
CVPR | 2 |
| 2017 | Depth and Image Restoration from Light Field in a Scattering MediumabstractTraditional imaging methods and computer vision algorithms are often ineffective when images are acquired in scattering media, such as underwater, fog, and biological tissue. Here, we explore the use of light field imaging and algorithms for image restoration and depth estimation that address the image degradation from the medium. Towards this end, we make the following three contributions. First, we present a new single image restoration algorithm which removes backscatter and attenuation from images better than existing methods do, and apply it to each view in the light field. Second, we combine a novel transmission based depth cue with existing correspondence and defocus cues to improve light field depth estimation. In densely scattering media, our transmission depth cue is critical for depth estimation since the images have low signal to noise ratios which significantly degrades the performance of the correspondence and defocus cues. Finally, we propose shearing and refocusing multiple views of the light field to recover a single image of higher quality than what is possible from a single view. We demonstrate the benefits of our method through extensive experimental results in a water tank. Jiandong Tian, Zak Murez, Tong Cui, David J. Kriegman, Ravi Ramamoorthi |
ICCV | 1 |
| 2017 | Single image dehazing by latent region-segmentation based transmission estimation and weighted L 1-norm regularisationabstractImage dehazing is a useful technique which can eliminate the bad effect of haze on images and enhance the performances of image/video processing algorithms in the hazy weather. In this study, a single image dehazing method is proposed. The authors estimate the initial transmission properly based on latent region‐segmentation and refine the estimated initial transmission by an objective function with a novel weighted L 1 ‐norm regularisation term. The half‐quadratic splitting minimisation method is employed to solve this optimisation problem. They also define an evaluation function to estimate the reliable global atmospheric light. With the refined transmission map and atmospheric light they recover the haze‐free image by the haze imaging model. The authors’ method is compared with three state‐of‐the‐art methods and is also validated by two image quality assessment methods. The comparative experimental results and evaluations demonstrate that their method can recover comparable and even better results with clear details, low contrast loss and high contrast in most cases. Tong Cui, Jiandong Tian, Ende Wang, Yandong Tang |
IET Image Process. | 2 |
| 2017 | A New Intrinsic-Lighting Color Space for Daytime Outdoor ImagesabstractExtracting or separating intrinsic information and illumination from natural images is crucial for better solving computer vision tasks. In this paper, we present a new illumination-based color space, the IL (intrinsic information and lighting level) space. Its first two channels represent 2D intrinsic information, and the third channel is for lighting levels. The IL color space has a one-to-one correspondence with the RGB color space. One valuable benefit of the IL color space is that illumination-related processing can be realized by directly operating on the lighting channel. As an example, based on the extracted lighting channel, we propose a new algorithm to estimate the intrinsic lighting level of an image such that the shadow-free color image and relighting series are obtained. In contrast to the existing color spaces for display or printing, the IL color space intuitively shows the information of reflectance and lighting levels for colors separately. Zhi Han, Jiandong Tian, Liangqiong Qu, Yandong Tang |
IEEE Trans. Image Process. | 2 |
| 2017 | RGBD Salient Object Detection via Deep FusionabstractNumerous efforts have been made to design various low-level saliency cues for RGBD saliency detection, such as color and depth contrast features as well as background and color compactness priors. However, how these low-level saliency cues interact with each other and how they can be effectively incorporated to generate a master saliency map remain challenging problems. In this paper, we design a new convolutional neural network (CNN) to automatically learn the interaction mechanism for RGBD salient object detection. In contrast to existing works, in which raw image pixels are fed directly to the CNN, the proposed method takes advantage of the knowledge obtained in traditional saliency detection by adopting various flexible and interpretable saliency feature vectors as inputs. This guides the CNN to learn a combination of existing features to predict saliency more effectively, which presents a less complex problem than operating on the pixels directly. We then integrate a superpixel-based Laplacian propagation framework with the trained CNN to extract a spatially consistent saliency map by exploiting the intrinsic structure of the input image. Extensive quantitative and qualitative experimental evaluations on three data sets demonstrate that the proposed method consistently outperforms the state-of-the-art methods. Liangqiong Qu, Shengfeng He, Jiawei Zhang 0002, Jiandong Tian, Yandong Tang, Qingxiong Yang |
IEEE Trans. Image Process. | 4 |
| 2017 | Specular Reflection Separation With Color-Lines ConstraintabstractAccording to dichromatic reflection model, the previous methods of specular reflection separation in image processing often separate specular reflection from a single image using patch-based priors. Due to lack of global information, these methods often cannot completely separate the specular component of an image and are incline to degrade image textures. In this paper, we derive a global color-lines constraint from dichromatic reflection model to effectively recover specular and diffuse reflection. Our key observation is from that each image pixel lies along a color line in normalized RGB space and the different color lines representing distinct diffuse chromaticities intersect at one point, namely, the illumination chromaticity. For pixels along the same color line, they spread over the entire image and their distances to the illumination chromaticity reflect the amount of specular reflection components. With global (non-local) information from these color lines, our method can effectively separate specular and diffuse reflection components in a pixelwise way for a single image, and it is suitable for real-time applications. Our experimental results on synthetic and real images show that our method performs better than the state-of-the-art methods to separate specular reflection. Weihong Ren, Jiandong Tian, Yandong Tang |
IEEE Trans. Image Process. | 2 |
| 2016 | New spectrum ratio properties and features for shadow detection
Jiandong Tian, Xiaojun Qi 0001, Liangqiong Qu, Yandong Tang |
Pattern Recognit. | 1 |
| 2016 | Blind Deconvolution With Nonlocal Similarity and l0 Sparsity for Noisy ImageabstractThe blind image deconvolution techniques with sparsity prior in gradient domain are sensitive to noise, even a small amount of noise. To address this problem, in this letter, we propose a novel blind deconvolution model that combines low-rank property, nonlocal similarity, and l0sparsity prior. Low-rank property makes the proposed deblurring model robust to image noise. The joint utilization of nonlocal similarity and l0sparsity prior has improved the accuracy of blur kernel estimation and restores the fine image details. A numerical method is also given to solve the proposed problem. Experimental results on synthetic and real data show that our algorithm performs better against with the state-of-the-art methods for both noise and noise-free images. Weihong Ren, Jiandong Tian, Yandong Tang |
IEEE Signal Process. Lett. | 2 |
| 2011 | Linearity of each channel pixel values from a surface in and out of shadows and its applicationsabstractShadows, the common phenomena in most outdoor scenes, are illuminated by diffuse skylight whereas shaded from direct sunlight. Generally shadows take place in sunny weather when the spectral power distributions (SPD) of sunlight, skylight, and daylight show strong regularity: they principally vary with sun angles. In this paper, we first deduce that the pixel values of a surface illuminated by skylight (in shadow region) and by daylight (in non-shadow region) have a linear relationship, and the linearity is independent of surface reflectance and holds in each color channel. We then use six simulated images that contain 1995 surfaces and two real captured images to test the linearity. The results validate the linearity. Based on the deduced linear relationship, we develop three shadow processing applications include intrinsic image deriving, shadow verification, and shadow removal. The results of the applications demonstrate that the linear relationship have practical values. Jiandong Tian, Yandong Tang |
CVPR | 1 |
| 2009 | Tricolor Attenuation Model for Shadow DetectionabstractShadows, the common phenomena in most outdoor scenes, bring many problems in image processing and computer vision. In this paper, we present a novel method focusing on extracting shadows from a single outdoor image. The proposed tricolor attenuation model (TAM) that describe the attenuation relationship between shadow and its nonshadow background is derived based on image formation theory. The parameters of the TAM are fixed by using the spectral power distribution (SPD) of daylight and skylight, which are estimated according to Planck's blackbody irradiance law. Based on the TAM, a multistep shadow detection algorithm is proposed to extract shadows. Compared with previous methods, the algorithm can be applied to process single images gotten in real complex scenes without prior knowledge. The experimental results validate the performance of the model. Jiandong Tian, Yandong Tang |
IEEE Trans. Image Process. | 1 |