VLDB 2026 Research / reviewers in the wild / expert
Robby T. Tan
dblp:t/RobbyTTan
· DBLP profile ↗
97ranked-venue papers
6as first author
53since 2021 · last 2026
0000-0001-7532-6919ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 88 · 6 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 69 · 4 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging Day and Night: Target-Class Hallucination Suppression in Unpaired Image TranslationabstractDay-to-night unpaired image translation is important to downstream tasks but remains challenging due to large appearance shifts and the lack of direct pixel-level supervision. Existing methods often introduce semantic hallucinations, where objects from target classes such as traffic signs and vehicles, as well as man-made light effects, are incorrectly synthesized. These hallucinations significantly degrade downstream performance. We propose a novel framework that detects and suppresses hallucinations of target-class features during unpaired translation. To detect hallucination, we design a dual-head discriminator that additionly performs semantic segmentation to identify hallucinated content in background regions. To suppress these hallucinations, we introduce class-specific prototypes, constructed by aggregating features of annotated target-domain objects, which act as semantic anchors for each class. Built upon a Schrödinger Bridge-based translation model, our framework performs iterative refinement, where detected hallucination features are explicitly pushed away from class prototypes in feature space, thus preserving object semantics across the translation trajectory. Experiments show that our method outperforms existing approaches both qualitatively and quantitatively. On the BDD100K dataset, it improves mAP by 15.5% for day-to-night domain adaptation, with a notable 31.7% gain for classes such as traffic lights that are prone to hallucinations. Shuwei Li, Robby T. Tan |
AAAI | 3 |
| 2026 | Aggregating Diverse Cue Experts for AI-Generated Image DetectionabstractThe rapid emergence of image synthesis models poses challenges to the generalization of AI-generated image detectors. However, existing methods often rely on model-specific features, leading to overfitting and poor generalization. In this paper, we introduce the Multi-Cue Aggregation Network (MCAN), a novel framework that integrates different yet complementary cues as input. MCAN employs a mixture-of-encoders adapter to dynamically process these cues, enabling more adaptive and robust feature representation. Our cues include the input image itself, which represents the overall content, and high-frequency components that emphasize edge details. Additionally, we introduce a Chromatic Inconsistency (CI) cue, which normalizes intensity values and captures noise information introduced during the image acquisition process in real images, making these noise patterns more distinguishable from those in AI-generated content. Unlike prior methods, MCAN employs a multi-cue aggregation strategy, leveraging spatial, frequency, and chromaticity-based cues. These cues are intrinsically more indicative of real images, enhancing cross-model generalization. Extensive experiments on the GenImage, Chameleon, and UniversalFakeDetect benchmark validate the state-of-the-art performance of MCAN. In the GenImage dataset, MCAN outperforms the best state-of-the-art method by up to 7.4\% in average ACC across eight different image generators. Shuwei Li, Mohan Kankanhalli, Robby T. Tan |
AAAI | 4 |
| 2026 | NaturalSloth: Revisiting Denial-of-Service Attacks on Large Language ModelsabstractLLM serving is limited by provider-side resources: longer generations consume more GPU time, increase latency, and reduce throughput in multi-tenant systems.This creates a denial-of-service (DoS) risk, where attackers degrade service by inducing excessive generation.Prior work on LLM DoS primarily relies on adversarial perturbations that delay end-of-sequence termination.We show perturbations are often unnecessary: natural, benignlooking instructions that specify impractical and meaningless tasks can already trigger excessive generation.To study this overlooked vulnerability, we introduce NaturalSloth, an adversarial dataset of natural, instruction-based DoS prompts.Starting from a human-curated seed set spanning diverse attack categories, we design a multi-agent synthesis framework to scale the dataset while preserving malicious intent and increasing semantic diversity.Experiments across a wide range of proprietary and open-source LLMs show that NaturalSloth consistently induces excessive generation, with attack effectiveness further amplified when combined with jailbreak techniques.Our analysis also reveals significant limitations of existing defenses, highlighting the need for dedicated protections against natural DoS attacks. 1 Yiming Chen 0010, Zexin Li 0001, Xianghu Yue, Robby T. Tan, Haizhou Li 0001 |
ACL (1) | 4 |
| 2026 | CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksabstractLarge Language Models (LLMs) are increasingly used not only to generate code, but also to judge it: comparing, ranking, or scoring competing solutions.However, their reliability in this evaluative role remains poorly understood.Inconsistent or flawed judgments can undermine benchmarks and distort training signals.This paper investigates the performance and robustness of LLMs when used as code judges.We introduce CodeJudgeBench, a benchmark explicitly designed to evaluate LLM-as-a-Judge models across three critical coding tasks: code generation, code repair, and unit test generation.We comprehensively benchmark the performance of 26 LLM-as-a-Judge models, encompassing general-purpose, code-tuned, and reasoning models.Our empirical findings reveal that relatively small reasoning models (e.g., Qwen3-8B) can outperform much larger non-reasoning models up to 70B.We further stress-test robustness by applying both general and code-specific perturbations.All models show significant instability and are sensitive to changes such as response ordering, variable naming, and misleading comments.These findings highlight serious concerns about the consistency and robustness of LLM-based judges for coding tasks. Hongchao Jiang, Yiming Chen 0010, Yushi Cao, Hung-yi Lee, Robby T. Tan |
ACL (1) | 5 |
| 2026 | VoiceBench: Benchmarking LLM-Based Voice AssistantsabstractAbstract Recent advancements in large language models (LLMs) like GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering an improved user experience over text-based interactions. However, a suitable benchmark to rigorously evaluate such speech interactions systems is currently lacking. To bridge this gap, we introduce VoiceBench, the first benchmark specifically designed to assess LLM-based voice assistants. VoiceBench comprises 6,783 synthetic and real spoken instructions recorded from diverse speakers across eight distinct tasks. These instructions are meticulously crafted to assess three crucial capability areas: general knowledge, instruction-following, and safety compliance. Furthermore, VoiceBench systematically incorporates realistic variations common in spoken interactions, including differences in speaker characteristics (e.g., accents), heterogeneous environmental conditions (e.g., reverberation), and content complexities such as mispronunciations. Extensive experiments reveal the limitations of current LLM-based voice assistant models and offer valuable insights for future research and development in this field.1 Yiming Chen 0010, Xianghu Yue, Chen Zhang 0020, Xiaoxue Gao, Robby T. Tan, Haizhou Li 0001 |
Trans. Assoc. Comput. Linguistics | 5 |
| 2025 | NightHaze: Nighttime Image Dehazing via Self-Prior LearningabstractMasked autoencoder (MAE) shows that severe augmentation during training produces robust representations for high-level tasks. This paper brings the MAE-like framework to nighttime image enhancement, demonstrating that severe augmentation during training produces strong network priors that are resilient to real-world night haze degradations. We propose a novel nighttime image dehazing method with self-prior learning. Our main novelty lies in the design of severe augmentation, which allows our model to learn robust priors. Unlike MAE that uses masking, we leverage two key challenging factors of nighttime images as augmentation: light effects and noise. During training, we intentionally degrade clear images by blending them with light effects as well as by adding noise, and subsequently restore the clear images. This enables our model to learn clear background priors. By increasing the noise values to approach as high as the pixel intensity values of the glow and light effect blended images, our augmentation becomes severe, resulting in stronger priors. While our self-prior learning is considerably effective in suppressing glow and revealing details of background scenes, in some cases, there are still some undesired artifacts that remain, particularly in the forms of over-suppression. To address these artifacts, we propose a self-refinement module based on the semi-supervised teacher-student framework. Our NightHaze, especially our MAE-like self-prior learning, shows that models trained with severe augmentation effectively improve the visibility of input haze images, approaching the clarity of clear nighttime images. Extensive experiments demonstrate that our NightHaze achieves state-of-the-art performance, outperforming existing nighttime image dehazing methods by a substantial margin of 15.5% for MUSIQ and 23.5% for ClipIQA. Beibei Lin, Yeying Jin, Wending Yan, Wei Ye 0005, Yuan Yuan 0039, Robby T. Tan |
AAAI | 6 |
| 2025 | Semantic Segmentation on Raindrop Degraded Images Using Two-Stage Dual Teacher-Student LearningabstractExisting semantic segmentation methods face challenges when processing input images degraded by raindrops on the lens or windshield. Unlike other adverse conditions such as fog and nighttime, which degrade visual quality, raindrops not only impair visual appearances but also introduce misleading occlusion, leading to significant performance drops in current models. The novelty of our approach lies in our two-stage, dual teacher-student framework. We tackle the complex problem of raindrop degradation by dividing it into two distinct challenges: degraded visual appearance and raindrop occlusion. These challenges are then addressed individually in two stages, utilizing two pairs of teacher-student networks. This division enables the networks to develop specialized expertise in handling each aspect of raindrop degradation, enabling their collaboration to achieve superior performance. In the first stage, one teacher-student pair focuses on learning to extract information from visual degraded areas. Building on this, the second teacher-student pair focuses specially on the raindrop occlusion. As such, unlike the existing methods, our approach employs a collaborative approach to decompose and address raindrop-induced degradations. In the second stage, we introduce a mask-based recovery technique to identify and rectify areas that likely contain misleading information, thus further refining the predictions. Additionally, this stage encourages both pairs to expand knowledge by swapping their specialized expertise. Our method achieves a performance of 60.3 mIoU on Rainy WCity and 72.8 mIoU on ACDC Rainy, representing an improvement of +4.4 mIoU and +2.3 mIoU over the existing state-of-the-art methods, respectively. Xin Yang 0035, Wending Yan, Yuan Yuan 0039, Michael Bi Mi, Robby T. Tan |
AAAI | 5 |
| 2025 | uMedSum: A Unified Framework for Clinical Abstractive SummarizationabstractAishik Nagar, Yutong Liu, Andy T. Liu, Viktor Schlegel, Vijay Prakash Dwivedi, Arun-Kumar Kaliya-Perumal, Guna Pratheep Kalanchiam, Yili Tang, Robby T. Tan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Aishik Nagar, Andy T. Liu, Viktor Schlegel, Vijay Prakash Dwivedi, Arun-Kumar Kaliya-Perumal, Guna Pratheep Kalanchiam, Yili Tang, Robby T. Tan |
ACL (1) | 9 |
| 2025 | Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure PriorsabstractSelf-supervised depth estimation from monocular cameras in diverse outdoor conditions, such as daytime, rain, and nighttime, is challenging due to the difficulty of learning universal representations and the severe lack of labeled real-world adverse data. Previous methods either rely on synthetic inputs and pseudo-depth labels or directly apply daytime strategies to adverse conditions, resulting in suboptimal results. In this paper, we present the first synthetic-to-real robust depth estimation framework, incorporating motion and structure priors to capture real-world knowledge effectively. In the synthetic adaptation, we transfer motion-structure knowledge inside cost volumes for better robust representation, using a frozen daytime model to train a depth estimator in synthetic adverse conditions. In the innovative real adaptation, which targets to fix synthetic-real gaps, models trained earlier identify the weather-insensitive regions with a designed consistency-reweighting strategy to emphasize valid pseudo-labels. We introduce a new regularization by gathering explicit depth distribution to constrain the model facing real-world data. Experiments show that our method outperforms the state-of-the-art across diverse conditions in multi-frame and single-frame evaluations. We achieve improvements of 7.5% and 4.3% in Ab-sRel and RMSE on average for nuScenes and Robotcar datasets (daytime, nighttime, rain). In zero-shot evaluation of DrivingStereo (rain, fog), our method generalizes better than previous ones. The code is at Syn2Real-Depth. Weilong Yan, Shuwei Shao, Robby T. Tan |
CVPR | 5 |
| 2025 | Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic SegmentationabstractVision Foundation Models (VFMs) and Vision-Language Models (VLMs) have gained traction in Domain Generalized Semantic Segmentation (DGSS) due to their strong generalization capabilities1. However, existing DGSS methods often rely exclusively on either VFMs or VLMs, overlooking their complementary strengths. VFMs (e.g., DINOv2) excel at capturing fine-grained features, while VLMs (e.g., CLIP) provide robust text alignment but struggle with coarse granularity. Despite their complementary strengths, effectively integrating VFMs and VLMs with attention mechanisms is challenging, as the increased patch tokens complicate long-sequence modeling. To address this, we propose MFuser, a novel Mamba-based fusion framework that efficiently combines the strengths of VFMs and VLMs while maintaining linear scalability in sequence length. MFuser consists of two key components: MVFuser, which acts as a co-adapter to jointly fine-tune the two models by capturing both sequential and spatial dynamics; and MTEnhancer, a hybrid attention-Mamba module that refines text embeddings by incorporating image priors. Our approach achieves precise feature locality and strong text alignment without incurring significant computational overhead. Extensive experiments demonstrate that MFuser significantly outperforms state-of-the-art DGSS methods, achieving 68.20 mIoU on synthetic-to-real and 71.87 mIoU on real-to-real benchmarks. The code is available at https://github.com/devinxzhang/MFuser. Robby T. Tan |
CVPR | 2 |
| 2025 | HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature AdaptationabstractZero-shot human-object interaction (HOI) detection remains a challenging task, particularly in generalizing to unseen actions. Existing methods address this challenge by tapping Vision-Language Models (VLMs) to access knowledge beyond the training data. However, they either struggle to distinguish actions involving the same object or demonstrate limited generalization to unseen classes. In this paper, we introduce HOLa (Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation), a novel approach that both enhances generalization to unseen classes and improves action distinction. In training, HOLa decomposes VLM text features for given HOI classes via low-rank factorization, producing class-shared basis features and adaptable weights. These features and weights form a compact HOI representation that preserves shared information across classes, enhancing generalization to unseen classes. Subsequently, we refine action distinction by adapting weights for each HOI class and introducing human-object tokens to enrich visual interaction representations. To further distinguish unseen actions, we guide the weight adaptation with LLM-derived action regularization. Experimental results show that our method sets a new state-of-the-art across zero-shot HOI settings on HICO-DET, achieving an unseen-class mAP of 27.91 in the unseen-verb setting. Our code is available at https://github.com/ChelsieLei/HOLa. Qinqian Lei, Bo Wang 0019, Robby T. Tan |
ICCV | 3 |
| 2025 | 3DOT: Texture Transfer for 3DGS Objects from a Single Reference ImageabstractImage-based 3D texture transfer from a single 2D reference image enables practical customization of 3D object appearances with minimal manual effort.
Adapted 2D editing and text-driven 3D editing approaches can serve this purpose. However, 2D editing typically involves frame-by-frame manipulation, often resulting in inconsistencies across views, while text-driven 3D editing struggles to preserve texture characteristics from reference images.
To tackle these challenges, we introduce \textbf{3DOT}, a \textbf{3D} Gaussian Splatting \textbf{O}bject \textbf{T}exture Transfer method based on a single reference image, integrating: 1) progressive generation, 2) view-consistency gradient guidance, and 3) prompt-tuned gradient guidance.
To ensure view consistency, progressive generation starts by transferring texture from the reference image and gradually propagates it to adjacent views.
View-consistency gradient guidance further reinforces coherence by conditioning the generation model on feature differences between consistent and inconsistent outputs.
To preserve texture characteristics, prompt-tuning-based gradient guidance learns a token that describes differences between original and reference textures, guiding the transfer for faithful texture preservation across views.
Overall, 3DOT combines these strategies to achieve effective texture transfer while maintaining structural coherence across viewpoints.
Extensive qualitative and quantitative evaluations confirm that our three components enable convincing and effective 2D-to-3D texture transfer. Our project page is available here: https://massyzs.github.io/3DOT_web/. Xiao Cao, Beibei Lin, Bo Wang 0019, Robby T. Tan |
NeurIPS | 5 |
| 2025 | GeoComplete: Geometry-Aware Diffusion for Reference-Driven Image CompletionabstractReference-driven image completion, which restores missing regions in a target view using additional images, is particularly challenging when the target view differs significantly from the references. Existing generative methods rely solely on diffusion priors and, without geometric cues such as camera pose or depth, often produce misaligned or implausible content. We propose GeoComplete, a novel framework that incorporates explicit 3D structural guidance to enforce geometric consistency in the completed regions, setting it apart from prior image-only approaches. GeoComplete introduces two key ideas: conditioning the diffusion process on projected point clouds to infuse geometric information, and applying target-aware masking to guide the model toward relevant reference cues. The framework features a dual-branch diffusion architecture. One branch synthesizes the missing regions from the masked target, while the other extracts geometric features from the projected point cloud. Joint self-attention across branches ensures coherent and accurate completion. To address regions visible in references but absent in the target, we project the target view into each reference to detect occluded areas, which are then masked during training. This target-aware masking directs the model to focus on useful cues, enhancing performance in difficult scenarios. By integrating a geometry-aware dual-branch diffusion architecture with a target-aware masking strategy, GeoComplete offers a unified and robust solution for geometry-conditioned image completion. Experiments show that GeoComplete achieves a 17.1% PSNR improvement over state-of-the-art methods, significantly boosting geometric accuracy while maintaining high visual quality. Beibei Lin, Robby T. Tan |
NeurIPS | 3 |
| 2024 | DeS3: Adaptive Attention-Driven Self and Soft Shadow Removal Using ViT SimilarityabstractRemoving soft and self shadows that lack clear boundaries from a single image is still challenging. Self shadows are shadows that are cast on the object itself. Most existing methods rely on binary shadow masks, without considering the ambiguous boundaries of soft and self shadows. In this paper, we present DeS3, a method that removes hard, soft and self shadows based on adaptive attention and ViT similarity. Our novel ViT similarity loss utilizes features extracted from a pre-trained Vision Transformer. This loss helps guide the reverse sampling towards recovering scene structures. Our adaptive attention is able to differentiate shadow regions from the underlying objects, as well as shadow regions from the object casting the shadow. This capability enables DeS3 to better recover the structures of objects even when they are partially occluded by shadows. Different from existing methods that rely on constraints during the training phase, we incorporate the ViT similarity during the sampling stage. Our method outperforms state-of-the-art methods on the SRD, AISTD, LRSS, USR and UIUC datasets, removing hard, soft, and self shadows robustly. Specifically, our method outperforms the SOTA method by 16% of the RMSE of the whole image on the LRSS dataset. Yeying Jin, Wei Ye 0005, Wenhan Yang, Yuan Yuan 0039, Robby T. Tan |
AAAI | 5 |
| 2024 | Few-Shot Learning from Augmented Label-Uncertain Queries in Bongard-HOIabstractDetecting human-object interactions (HOI) in a few-shot setting remains a challenge. Existing meta-learning methods struggle to extract representative features for classification due to the limited data, while existing few-shot HOI models rely on HOI text labels for classification. Moreover, some query images may display visual similarity to those outside their class, such as similar backgrounds between different HOI classes. This makes learning more challenging, especially with limited samples. Bongard-HOI epitomizes this HOI few-shot problem, making it the benchmark we focus on in this paper. In our proposed method, we introduce novel label-uncertain query augmentation techniques to enhance the diversity of the query inputs, aiming to distinguish the positive HOI class from the negative ones. As these augmented inputs may or may not have the same class label as the original inputs, their class label is unknown. Those belonging to a different class become hard samples due to their visual similarity to the original ones. Additionally, we introduce a novel pseudo-label generation technique that enables a mean teacher model to learn from the augmented label-uncertain inputs. We propose to augment the negative support set for the student model to enrich the semantic information, fostering diversity that challenges and enhances the student’s learning. Experimental results demonstrate that our method sets a new state-of-the-art (SOTA) performance by achieving 68.74% accuracy on the Bongard-HOI benchmark, a significant improvement over the existing SOTA of 66.59%. In our evaluation on HICO-FS, a more general few-shot recognition dataset, our method achieves 73.27% accuracy, outperforming the previous SOTA of 71.20% in the 5- way 5-shot task. Qinqian Lei, Bo Wang 0019, Robby T. Tan |
AAAI | 3 |
| 2024 | NightRain: Nighttime Video Deraining via Adaptive-Rain-Removal and Adaptive-CorrectionabstractExisting deep-learning-based methods for nighttime video deraining rely on synthetic data due to the absence of real-world paired data. However, the intricacies of the real world, particularly with the presence of light effects and low-light regions affected by noise, create significant domain gaps, hampering synthetic-trained models in removing rain streaks properly and leading to over-saturation and color shifts. Motivated by this, we introduce NightRain, a novel nighttime video deraining method with adaptive-rain-removal and adaptive-correction. Our adaptive-rain-removal uses unlabeled rain videos to enable our model to derain real-world rain videos, particularly in regions affected by complex light effects. The idea is to allow our model to obtain rain-free regions based on the confidence scores. Once rain-free regions and the corresponding regions from our input are obtained, we can have region-based paired real data. These paired data are used to train our model using a teacher-student framework, allowing the model to iteratively learn from less challenging regions to more challenging regions. Our adaptive-correction aims to rectify errors in our model's predictions, such as over-saturation and color shifts. The idea is to learn from clear night input training videos based on the differences or distance between those input videos and their corresponding predictions. Our model learns from these differences, compelling our model to correct the errors. From extensive experiments, our method demonstrates state-of-the-art performance. It achieves a PSNR of 26.73dB, surpassing existing nighttime video deraining methods by a substantial margin of 13.7%. Beibei Lin, Yeying Jin, Wending Yan, Wei Ye 0005, Yuan Yuan 0039, Shunli Zhang 0005, Robby T. Tan |
AAAI | 7 |
| 2024 | Restoring Speaking Lips from Occlusion for Audio-Visual Speech RecognitionabstractPrior studies on audio-visual speech recognition typically assume the visibility of speaking lips, ignoring the fact that visual occlusion occurs in real-world videos, thus adversely affecting recognition performance. To address this issue, we propose a framework that restores occluded lips in a video by utilizing both the video itself and the corresponding noisy audio. Specifically, the framework aims to achieve these three tasks: detecting occluded frames, masking occluded areas, and reconstruction of masked regions. We tackle the first two issues by utilizing the Class Activation Map (CAM) obtained from occluded frame detection to facilitate the masking of occluded areas. Additionally, we introduce a novel synthesis-matching strategy for the reconstruction to ensure the compatibility of audio features with different levels of occlusion. Our framework is evaluated in terms of Word Error Rate (WER) on the original videos, the videos corrupted by concealed lips, and the videos restored using the framework with several existing state-of-the-art audio-visual speech recognition methods. Experimental results substantiate that our framework significantly mitigates performance degradation resulting from lip occlusion. Under -5dB noise conditions, AV-Hubert's WER increases from 10.62% to 13.87% due to lip occlusion, but rebounds to 11.87% in conjunction with the proposed framework. Furthermore, the framework also demonstrates its capacity to produce natural synthesized images in qualitative assessments. Zexu Pan, Malu Zhang, Robby T. Tan, Haizhou Li 0001 |
AAAI | 4 |
| 2024 | Semantic Segmentation in Multiple Adverse Weather Conditions with Domain Knowledge RetentionabstractSemantic segmentation's performance is often compromised when applied to unlabeled adverse weather conditions. Unsupervised domain adaptation is a potential approach to enhancing the model's adaptability and robustness to adverse weather. However, existing methods encounter difficulties when sequentially adapting the model to multiple unlabeled adverse weather conditions. They struggle to acquire new knowledge while also retaining previously learned knowledge. To address these problems, we propose a semantic segmentation method for multiple adverse weather conditions that incorporates adaptive knowledge acquisition, pseudo-label blending, and weather composition replay. Our adaptive knowledge acquisition enables the model to avoid learning from extreme images that could potentially cause the model to forget. In our approach of blending pseudo-labels, we not only utilize the current model but also integrate the previously learned model into the ongoing learning process. This collaboration between the current teacher and the previous model enhances the robustness of the pseudo-labels for the current target. Our weather composition replay mechanism allows the model to continuously refine its previously learned weather information while simultaneously learning from the new target domain. Our method consistently outperforms the state-of-the-art methods, and obtains the best performance with averaged mIoU (%) of 65.7 and the lowest forgetting (%) of 3.6 against 60.1 and 11.3, on the ACDC datsets for a four-target continual multi-target domain adaptation. Xin Yang 0035, Wending Yan, Yuan Yuan 0039, Michael Bi Mi, Robby T. Tan |
AAAI | 5 |
| 2024 | HEAP: Unsupervised Object Discovery and Localization with Contrastive GroupingabstractUnsupervised object discovery and localization aims to detect or segment objects in an image without any supervision. Recent efforts have demonstrated a notable potential to identify salient foreground objects by utilizing self-supervised transformer features. However, their scopes only build upon patch-level features within an image, neglecting region/image-level and cross-image relationships at a broader scale. Moreover, these methods cannot differentiate various semantics from multiple instances. To address these problems, we introduce Hierarchical mErging framework via contrAstive grouPing (HEAP). Specifically, a novel lightweight head with cross-attention mechanism is designed to adaptively group intra-image patches into semantically coherent regions based on correlation among self-supervised features. Further, to ensure the distinguishability among various regions, we introduce a region-level contrastive clustering loss to pull closer similar regions across images. Also, an image-level contrastive loss is present to push foreground and background representations apart, with which foreground objects and background are accordingly discovered. HEAP facilitates efficient hierarchical image decomposition, which contributes to more accurate object discovery while also enabling differentiation among objects of various classes. Extensive experimental results on semantic segmentation retrieval, unsupervised object discovery, and saliency detection tasks demonstrate that HEAP achieves state-of-the-art performance. Jinheng Xie, Yuan Yuan 0039, Michael Bi Mi, Robby T. Tan |
AAAI | 5 |
| 2024 | CAT: Exploiting Inter-Class Dynamics for Domain Adaptive Object DetectionabstractDomain adaptive object detection aims to adapt detection models to domains where annotated data is unavailable. Existing methods have been proposed to address the domain gap using the semi-supervised student-teacher framework. However, a fundamental issue arises from the class imbalance in the labelled training set, which can result in inaccurate pseudo-labels. The relationship between classes, especially where one class is a majority and the other minority, has a large impact on class bias. We propose Class-Aware Teacher (CAT) to address the class bias issue in the domain adaptation setting. In our work, we ap-proximate the class relationships with our Inter-Class Relation module (ICRm) and exploit it to reduce the bias within the model. In this way, we are able to apply augmentations to highly related classes, both inter- and intra-domain, to boost the performance of minority classes while having minimal impact on majority classes. We further reduce the bias by implementing a class-relation weight to our classification loss. Experiments conducted on various datasets and ablation studies show that our method is able to address the class bias in the domain adaptation setting. On the Cityscapes$\rightarrow$Foggy Cityscapes dataset, we attained a 52.5 mAp, a substantial improvement over the 51.2 mAP achieved by the state-of-the-art method.11www.github.com/mecarill/cat Mikhail Kennerley, Jian-Gang Wang 0001, Bharadwaj Veeravalli, Robby T. Tan |
CVPR | 4 |
| 2024 | NightCC: Nighttime Color Constancy via Adaptive Channel MaskingabstractNighttime conditions pose a significant challenge to color constancy due to the diversity of lighting conditions and the presence of substantial low-light noise. Existing color constancy methods struggle with nighttime scenes, frequently leading to imprecise light color estimations. To tackle nighttime color constancy, we propose a novel unsupervised domain adaptation approach that utilizes labeled daytime data to facilitate learning on unlabeled nighttime images. To specifically address the unique lighting conditions of nighttime and ensure the robustness of pseudo labels, we propose adaptive channel masking and light uncertainty. By selectively masking channels that are less sen-sitive to lighting conditions, adaptive channel masking directs the model to progressively focus on features less affected by variations in light colors and noise. Addition-ally, our model leverages light uncertainty to provide a pixel-wise uncertainty estimation regarding light color prediction, which helps avoid learning from incorrect labels. Our model demonstrates a significant improvement in accuracy, achieving 21.5% lower Mean Angular Error (MAE) compared to the state-of-the-art method on our nighttime dataset. Shuwei Li, Robby T. Tan |
CVPR | 2 |
| 2024 | Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions
Yihao Ai, Bo Wang 0019, Yu Cheng 0009, Xinchao Wang, Robby T. Tan |
ECCV (47) | 6 |
| 2024 | Dual-Rain: Video Rain Removal Using Assertive and Gentle Teachers
Beibei Lin, Yeying Jin, Wending Yan, Wei Ye 0005, Yuan Yuan 0039, Robby T. Tan |
ECCV (68) | 7 |
| 2024 | Remote Sensing Domain Adaptive Alignment via Student-Teacher LearningabstractImage Alignment between Synthetic Aperture Radar (SAR) and Electro-Optical (EO) imagery is a task that has comprehensive remote sensing capabilities. Traditional deep learning-based SAR-optical image matching models heavily rely on supervised learning with expensive annotated datasets, leading to reduced accuracy and overfitting when encountering insufficient data during training. To tackle this issue, this paper proposes a student-teacher framework for Domain Adaptation (DA) approach, transferring deep learning models from well-annotated source domains like normal outside optical images to non-annotated SAR-EO target domains. Additionally, in contrast to previous methods which usually use CNN or ordinary Transformer structure to extract features from image pairs, we use self and cross attention mechanisms in Transformer to obtain feature descriptors that are conditioned on both multimodal images. The larger global receptive field and better feature extraction capability provided by this Transformer shows its ability to accommodate large disparities of multimodal data and manage fewer textures such as rural areas with forest or desert in satellite images, where previous backbones usually struggle to produce repeatable and correct interest points. Qiuhang Liu, Weilong Yan, Bo Wang 0019, Bharadwaj Veeravalli, Robby T. Tan |
IGARSS | 5 |
| 2024 | EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI DetectionabstractDetecting Human-Object Interactions (HOI) in zero-shot settings, where models must handle unseen classes, poses significant challenges. Existing methods that rely on aligning visual encoders with large Vision-Language Models (VLMs) to tap into the extensive knowledge of VLMs, require large, computationally expensive models and encounter training difficulties. Adapting VLMs with prompt learning offers an alternative to direct alignment. However, fine-tuning on task-specific datasets often leads to overfitting to seen classes and suboptimal performance on unseen classes, due to the absence of unseen class labels. To address these challenges, we introduce a novel prompt learning-based framework for Efficient Zero-Shot HOI detection (EZ-HOI). First, we introduce Large Language Model (LLM) and VLM guidance for learnable prompts, integrating detailed HOI descriptions and visual semantics to adapt VLMs to HOI tasks. However, because training datasets contain seen-class labels alone, fine-tuning VLMs on such datasets tends to optimize learnable prompts for seen classes instead of unseen ones. Therefore, we design prompt learning for unseen classes using information from related seen classes, with LLMs utilized to highlight the differences between unseen and related seen classes. Quantitative evaluations on benchmark datasets demonstrate that our EZ-HOI achieves state-of-the-art performance across various zero-shot settings with only 10.35\% to 33.95\% of the trainable parameters compared to existing methods. Code is available at https://github.com/ChelsieLei/EZ-HOI. Qinqian Lei, Bo Wang 0019, Robby T. Tan |
NeurIPS | 3 |
| 2024 | End-to-End Video Semantic Segmentation in Adverse Weather using Fusion Blocks and Temporal-Spatial Teacher-Student LearningabstractAdverse weather conditions can significantly degrade the video frames, causing existing video semantic segmentation methods to produce erroneous predictions. In this work, we target adverse weather conditions and introduce an end-to-end domain adaptation strategy that leverages a fusion block, temporal-spatial teacher-student learning, and a temporal weather degradation augmentation approach. The fusion block integrates temporal information from adjacent frames at the feature level, trained end-to-end, eliminating the need for pretrained optical flow, distinguishing our method from existing approaches. Our teacher-student approach involves two teachers: one focuses on exploring temporal information from adjacent frames, and the other harnesses spatial information from the current frame. Finally, we apply temporal weather degradation augmentation to consecutive frames to more accurately represent adverse weather degradations. Our method achieves a performance of 25.4 and 33.0 mIoU on the adaptation from VIPER and Synthia to MVSS, respectively, representing an improvement of 4.3 and 5.8 mIoU over the existing state-of-the-art method. Xin Yang 0035, Wending Yan, Michael Bi Mi, Yuan Yuan 0039, Robby T. Tan |
NeurIPS | 5 |
| 2024 | Learning to Remove Rain in Video With Self-SupervisionabstractIn heavy rain video, rain streak and rain accumulation are the most common causes of degradation. They occlude background information and can significantly impair the visibility. Most existing methods rely heavily on the synthetic training data, and thus raise the domain gap problem that prevents the trained models from performing adequately in real testing cases. Unlike these methods, we introduce a self-learning method to remove both rain streaks and rain accumulation without using any ground-truth clean images in training our model, which consequently can alleviate the domain gap issue. The main idea is based on the assumptions that (1) adjacent clean frames can be aligned or warped from one frame to another frame, (2) rain streaks are distributed randomly in the temporal domain, (3) the rain streak/accumulation related variables/priors can be inferred reliably from the information within the images/sequences. Based on these assumptions, we construct an augmented Self-Learned Deraining Network (SLDNet+) to remove both rain streaks and rain accumulation by utilizing temporal correlation, consistency, and rain-related priors. For the temporal correlation, our SLDNet+ takes rain degraded adjacent frames as its input, aligns them, and learns to predict the clean version of the current frame. For the temporal consistency, a new loss is designed to build a robust mapping between the predicted clean frame and non-rain regions from the adjacent rain frames. For the rain-streak-related prior, the rain streak removal network is optimized jointly with motion estimation and rain region detection; while for the rain-accumulation-related prior, a novel non-local video rain accumulation removal method is developed to estimate the accumulation-lines from the whole input video and to offer better color constancy and temporal smoothness. Extensive experiments show the effectiveness of our approach, which provides superior results compared with the existing state of the art methods both quantitatively and qualitatively. The source code will be made publicly available at: https://github.com/flyywh/CVPR-2020-Self-Rain-Removal-Journal. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Alex Chichung Kot, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Estimating Reflectance Layer from a Single Image: Integrating Reflectance Guidance and Shadow/Specular Aware LearningabstractEstimating the reflectance layer from a single image is a challenging task. It becomes more challenging when the input image contains shadows or specular highlights, which often render an inaccurate estimate of the reflectance layer. Therefore, we propose a two-stage learning method, including reflectance guidance and a Shadow/Specular-Aware (S-Aware) network to tackle the problem. In the first stage, an initial reflectance layer free from shadows and specularities is obtained with the constraint of novel losses that are guided by prior-based shadow-free and specular-free images. To further enforce the reflectance layer to be independent of shadows and specularities in the second-stage refinement, we introduce an S-Aware network that distinguishes the reflectance image from the input image. Our network employs a classifier to categorize shadow/shadow-free, specular/specular-free classes, enabling the activation features to function as attention maps that focus on shadow/specular regions. Our quantitative and qualitative evaluations show that our method outperforms the state-of-the-art methods in the reflectance layer estimation that is free from shadows and specularities. Yeying Jin, Ruoteng Li, Wenhan Yang, Robby T. Tan |
AAAI | 4 |
| 2023 | Dynamic Transformers Provide a False Sense of EfficiencyabstractDespite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference.Multi-exit is a mainstream approach to address this issue by making a tradeoff between efficiency and accuracy, where the saving of computation comes from an early exit.However, whether such saving from earlyexiting is robust remains unknown.Motivated by this, we first show that directly adapting existing adversarial attack approaches targeting model accuracy cannot significantly reduce inference efficiency.To this end, we propose a simple yet effective attacking framework, SAME, a novel slowdown attack framework on multi-exit models, which is specially tailored to reduce the efficiency of the multi-exit models.By leveraging the multi-exit models' design characteristics, we utilize all internal predictions to guide the adversarial sample generation instead of merely considering the final prediction.Experiments on the GLUE benchmark show that SAME can effectively diminish the efficiency gain of various multi-exit models by 80% on average, convincingly validating its effectiveness and generalization ability. 1 Yiming Chen 0010, Zexin Li 0001, Wei Yang 0013, Cong Liu 0005, Robby T. Tan, Haizhou Li 0001 |
ACL (1) | 6 |
| 2023 | 2PCNet: Two-Phase Consistency Training for Day-to-Night Unsupervised Domain Adaptive Object DetectionabstractObject detection at night is a challenging problem due to the absence of night image annotations. Despite several domain adaptation methods, achieving high-precision results remains an issue. False-positive error propagation is still observed in methods using the well-established student-teacher framework, particularly for small-scale and low-light objects. This paper proposes a two-phase consistency unsupervised domain adaptation network, 2PCNet, to address these issues. The network employs high-confidence bounding-box predictions from the teacher in the first phase and appends them to the student's region proposals for the teacher to re-evaluate in the second phase, resulting in a combination of high and low confidence pseudolabels. The night images and pseudo-labels are scaled-down before being used as input to the student, providing stronger small-scale pseudo-labels. To address errors that arise from low-light regions and other night-related attributes in images, we propose a night-specific augmentation pipeline called NightAug. This pipeline involves applying random augmentations, such as glare, blur, and noise, to daytime images. Experiments on publicly available datasets demonstrate that our method achieves superior results to state-of-the-art methods by 20%, and to supervised models trained directly on the target data.11www.github.com/mecarill/2pcnet Mikhail Kennerley, Jian-Gang Wang 0001, Bharadwaj Veeravalli, Robby T. Tan |
CVPR | 4 |
| 2023 | DSFNet: Dual Space Fusion Network for Occlusion-Robust 3D Dense Face AlignmentabstractSensitivity to severe occlusion and large view angles limits the usage scenarios of the existing monocular 3D dense face alignment methods. The state-of-the-art 3DMM-based method, directly regresses the model's coefficients, underutilizing the low-level 2D spatial and semantic information, which can actually offer cues for face shape and orientation. In this work, we demonstrate how modeling 3D facial geometry in image and model space jointly can solve the occlusion and view angle problems. Instead of predicting the whole face directly, we regress image space features in the visible facial region by dense prediction first. Subsequently, we predict our model's coefficients based on the regressed feature of the visible regions, leveraging the prior knowledge of whole face geometry from the morphable models to complete the invisible regions. We further propose a fusion network that combines the advantages of both the image and model space predictions to achieve high robustness and accuracy in unconstrained scenarios. Thanks to the proposed fusion module, our method is robust not only to occlusion and large pitch and roll view angles, which is the bene- fit of our image space approach, but also to noise and large yaw angles, which is the benefit of our model space method. Comprehensive evaluations demonstrate the superior performance of our method compared with the state-of-the-art methods. On the 3D dense face alignment task, we achieve 3.80% NME on the AFLW2000-3D dataset, which outperforms the state-of-the-art method by 5.5%. Code is available at https://github.com/1hyfst/DSFNet. Heyuan Li, Bo Wang 0019, Yu Cheng 0009, Mohan Kankanhalli, Robby T. Tan |
CVPR | 5 |
| 2023 | Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertabstractTalking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite much progress, they hardly focus on the content of lip movements i.e., the visual intelligibility of the spoken words, which is an important aspect of generation quality. To address the problem, we propose using a lipreading expert to improve the intelligibility of the generated lip regions by penalizing the incorrect generation results. Moreover, to compensate for data scarcity, we train the lip-reading expert in an audio-visual self-supervised manner. With a lip-reading expert, we propose a novel contrastive learning to enhance lip-speech synchronization, and a transformer to encode audio synchronically with video, while considering global temporal dependency of audio. For evaluation, we propose a new strategy with two different lip-reading experts to measure intelligibility of the generated videos. Rigorous experiments show that our proposal is superior to other State-of-the-art (SOTA) methods, such as Wav2Lip, in reading intelligibility i.e., over 38% Word Error Rate (WER) on LRS2 dataset and 27.8% accuracy on LRW dataset. We also achieve the SOTA performance in lip-speech synchronization and comparable performances in visual quality. Xinyuan Qian 0001, Malu Zhang, Robby T. Tan, Haizhou Li 0001 |
CVPR | 4 |
| 2023 | EqMotion: Equivariant Multi-Agent Motion Prediction with Invariant Interaction ReasoningabstractLearning to predict agent motions with relationship reasoning is important for many applications. In motion prediction tasks, maintaining motion equivariance under Euclidean geometric transformations and invariance of agent interaction is a critical and fundamental principle. However, such equivariance and invariance properties are overlooked by most existing methods. To fill this gap, we propose Eq-Motion, an efficient equivariant motion prediction model with invariant interaction reasoning. To achieve motion equivariance, we propose an equivariant geometric feature learning module to learn a Euclidean transformable feature through dedicated designs of equivariant operations. To reason agent's interactions, we propose an invariant interaction reasoning module to achieve a more stable interaction modeling. To further promote more comprehensive motion features, we propose an invariant pattern feature learning module to learn an invariant pattern feature, which cooperates with the equivariant geometric feature to enhance network expressiveness. We conduct experiments for the proposed model on four distinct scenarios: particle dynamics, molecule dynamics, human skeleton motion prediction and pedestrian trajectory prediction. Experimental results show that our method is not only generally applicable, but also achieves state-of-the-art prediction performances on all the four tasks, improving by 24.0/30.1/8.6/9.2%. Code is available at https://github.com/MediaBrain-SJTU/EqMotion. Chenxin Xu, Robby T. Tan, Yuhong Tan, Siheng Chen, Yu Guang Wang 0001, Xinchao Wang, Yanfeng Wang 0001 |
CVPR | 2 |
| 2023 | Auxiliary Tasks Benefit 3D Skeleton-based Human Motion PredictionabstractExploring spatial-temporal dependencies from observed motions is one of the core challenges of human motion prediction. Previous methods mainly focus on dedicated network structures to model the spatial and temporal dependencies. This paper considers a new direction by introducing a model learning framework with auxiliary tasks. In our auxiliary tasks, partial body joints’ coordinates are corrupted by either masking or adding noise and the goal is to recover corrupted coordinates depending on the rest coordinates. To work with auxiliary tasks, we propose a novel auxiliary-adapted transformer, which can handle incomplete, corrupted motion data and achieve coordinate recovery via capturing spatial-temporal dependencies. Through auxiliary tasks, the auxiliary-adapted transformer is promoted to capture more comprehensive spatial-temporal dependencies among body joints’ coordinates, leading to better feature learning. Extensive experimental results have shown that our method outperforms state-of-the-art methods by remarkable margins of 7.2%, 3.7%, and 9.4% in terms of 3D mean per joint position error (MPJPE) on the Human3.6M, CMU Mocap, and 3DPW datasets, respectively. We also demonstrate that our method is more robust under data missing cases and noisy data cases. Code is available at https://github.com/MediaBrain-SJTU/AuxFormer. Chenxin Xu, Robby T. Tan, Yuhong Tan, Siheng Chen, Xinchao Wang, Yanfeng Wang 0001 |
ICCV | 2 |
| 2023 | Deep Homography Mixture for Single Image Rolling Shutter CorrectionabstractWe present a deep homography mixture motion model for single image rolling shutter correction. Rolling shutter (RS) effects are often caused by row-wise exposure delay in the widely adopted CMOS sensor. Previous methods often require more than one frame for the correction, leading to data quality requirements. Few approaches address the more challenging task of single image RS correction, which often adopt designs like trajectory estimation or long rectangular kernels, to learn the camera motion parameters of an RS image, to restore the global shutter (GS) image. In this work, we adopt a more straightforward method to learn deep homography mixture motion between an RS image and its corresponding GS image, without large solution space or strict restrictions on image features. We show that dividing an image into blocks with a Gaussian weight of block scanlines fits well for the RS setting. Moreover, instead of directly learning the motion mapping, we learn coefficients that assemble several motion bases to produce the correction motion, where these bases are learned from the consecutive frames of natural videos beforehand. Experiments show that our method outperforms existing single RS methods statistically and visually, in both synthesized and real RS images. Our code and dataset are available at https://github.com/DavidYan2001/Deep_HM. Weilong Yan, Robby T. Tan, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 2 |
| 2023 | Enhancing Visibility in Nighttime Haze Images Using Guided APSF and Gradient Adaptive ConvolutionabstractVisibility in hazy nighttime scenes is frequently reduced by multiple factors, including low light, intense glow, light scattering, and the presence of multicolored light sources. Existing nighttime dehazing methods often struggle with handling glow or low-light conditions, resulting in either excessively dark visuals or unsuppressed glow outputs. In this paper, we enhance the visibility from a single nighttime haze image by suppressing glow and enhancing low-light regions. To handle glow effects, our framework learns from the rendered glow pairs. Specifically, a light source aware network is proposed to detect light sources of night images, followed by the APSF (Angular Point Spread Function)-guided glow rendering. Our framework is then trained on the rendered images, resulting in glow suppression. Moreover, we utilize gradient-adaptive convolution, to capture edges and textures in hazy scenes. By leveraging extracted edges and textures, we enhance the contrast of the scene without losing important structural details. To boost low-light intensity, our network learns an attention map, then adjusted by gamma correction. This attention has high values on low-light regions and low values on haze and glow regions. Extensive evaluation on real nighttime haze images, demonstrates the effectiveness of our method. Our experiments demonstrate that our method achieves a PSNR of 30.38dB, outperforming state-of-the-art methods by 13% on GTA5 nighttime haze dataset. Our data and code is available at: https://github.com/jinyeying/nighttime_dehaze. Yeying Jin, Beibei Lin, Wending Yan, Yuan Yuan 0039, Wei Ye 0005, Robby T. Tan |
ACM Multimedia | 6 |
| 2023 | Dual Networks Based 3D Multi-Person Pose Estimation From Monocular VideoabstractMonocular 3D human pose estimation has made progress in recent years. Most of the methods focus on single persons, which estimate the poses in the person-centric coordinates, i.e., the coordinates based on the center of the target person. Hence, these methods are inapplicable for multi-person 3D pose estimation, where the absolute coordinates (e.g., the camera coordinates) are required. Moreover, multi-person pose estimation is more challenging than single pose estimation, due to inter-person occlusion and close human interactions. Existing top-down multi-person methods rely on human detection (i.e., top-down approach), and thus suffer from the detection errors and cannot produce reliable pose estimation in multi-person scenes. Meanwhile, existing bottom-up methods that do not use human detection are not affected by detection errors, but since they process all persons in a scene at once, they are prone to errors, particularly for persons in small scales. To address all these challenges, we propose the integration of top-down and bottom-up approaches to exploit their strengths. Our top-down network estimates human joints from all persons instead of one in an image patch, making it robust to possible erroneous bounding boxes. Our bottom-up network incorporates human-detection based normalized heatmaps, allowing the network to be more robust in handling scale variations. Finally, the estimated 3D poses from the top-down and bottom-up networks are fed into our integration network for final 3D poses. To address the common gaps between training and testing data, we do optimization during the test time, by refining the estimated 3D human poses using high-order temporal constraint, re-projection loss, and bone length regularizations. We also introduce a two-person pose discriminator that enforces natural two-person interactions. Finally, we apply a semi-supervised method to overcome the 3D ground-truth data scarcity. Our evaluations demonstrate the effectiveness of the proposed method and its individual components. Our code and pretrained models are available publicly: https://github.com/3dpose/3D-Multi-Person-Pose. Yu Cheng 0009, Bo Wang 0019, Robby T. Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Bottom-up 2D pose estimation via dual anatomical centers for small-scale persons
Yu Cheng 0009, Yihao Ai, Bo Wang 0019, Xinchao Wang, Robby T. Tan |
Pattern Recognit. | 5 |
| 2022 | Structure Representation Network and Uncertainty Feedback Learning for Dense Non-uniform Fog Removal
Yeying Jin, Wending Yan, Wenhan Yang, Robby T. Tan |
ACCV (3) | 4 |
| 2022 | Object Detection in Foggy Scenes by Embedding Depth and Reconstruction into Domain Adaptation
Xin Yang 0035, Michael Bi Mi, Yuan Yuan 0039, Robby T. Tan |
ACCV (6) | 5 |
| 2022 | Unsupervised Night Image Enhancement: When Layer Decomposition Meets Light-Effects Suppression
Yeying Jin, Wenhan Yang, Robby T. Tan |
ECCV (37) | 3 |
| 2022 | Recurrent Multi-Frame Deraining: Combining Physics Guidance and Adversarial LearningabstractExisting video rain removal methods mainly focus on rain streak removal and are solely trained based on the synthetic data, which neglect more complex degradation factors, e.g., rain accumulation, and the prior knowledge in real rain data. Thus, in this paper, we build a more comprehensive rain model with several degradation factors and construct a novel two-stage video rain removal method that combines the power of synthetic videos and real data. Specifically, a novel two-stage progressive network is proposed: recovery guided by a physics model, and further restoration by adversarial learning. The first stage performs an inverse recovery process guided by our proposed rain model. An initially estimated background frame is obtained based on the input rain frame. The second stage employs adversarial learning to refine the result, i.e., recovering the overall color and illumination distributions of the frame, the background details that are failed to be recovered in the first stage, and removing the artifacts generated in the first stage. Furthermore, we also introduce a more comprehensive rain model that includes degradation factors, e.g., occlusion and rain accumulation, which appear in real scenes yet ignored by existing methods. This model, which generates more realistic rain images, will train and evaluate our models better. Extensive evaluations on synthetic and real videos show the effectiveness of our method in comparisons to the state-of-the-art methods. Our datasets, results and code are available at: https://github.com/flyywh/Recurrent-Multi-Frame-Deraining. Wenhan Yang, Robby T. Tan, Jiashi Feng, Shiqi Wang 0001, Bin Cheng 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Human object interaction detection using two-direction spatial enhancement and exclusive object prior
Robby T. Tan |
Pattern Recognit. | 2 |
| 2022 | Feature-Aligned Video Raindrop Removal With Temporal ConstraintsabstractExisting adherent raindrop removal methods focus on the detection of the raindrop locations, and then use inpainting techniques or generative networks to recover the background behind raindrops. Yet, as adherent raindrops are diverse in sizes and appearances, the detection is challenging for both single image and video. Moreover, unlike rain streaks, adherent raindrops tend to cover the same area in several frames. Addressing these problems, our method employs a two-stage video-based raindrop removal method. The first stage is the single image module, which generates initial clean results. The second stage is the multiple frame module, which further refines the initial results using temporal constraints, namely, by utilizing multiple input frames in our process and applying temporal consistency between adjacent output frames. Our single image module employs a raindrop removal network to generate initial raindrop removal results, and create a mask representing the differences between the input and initial output. Once the masks and initial results for consecutive frames are obtained, our multiple-frame module aligns the frames in both the image and feature levels and then obtains the clean background. Our method initially employs optical flow to align the frames, and then utilizes deformable convolution layers further to achieve feature-level frame alignment. To remove small raindrops and recover correct backgrounds, a target frame is predicted from adjacent frames. A series of unsupervised losses are proposed so that our second stage, which is the video raindrop removal module, can self-learn from video data without ground truths. Experimental results on real videos demonstrate the state-of-art performance of our method both quantitatively and qualitatively. Wending Yan, Wenhan Yang, Robby T. Tan |
IEEE Trans. Image Process. | 4 |
| 2021 | Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular VideosabstractDespite the recent progress, 3D multi-person pose estimation from monocular videos is still challenging due to the commonly encountered problem of missing information caused by occlusion, partially out-of-frame target persons, and inaccurate person detection. To tackle this problem, we propose a novel framework integrating graph convolutional networks (GCNs) and temporal convolutional networks (TCNs) to robustly estimate camera-centric multi-person 3D poses that does not require camera parameters. In particular, we introduce a human-joint GCN, which unlike the existing GCN, is based on a directed graph that employs the 2D pose estimator's confidence scores to improve the pose estimation results. We also introduce a human-bone GCN, which models the bone connections and provides more information beyond human joints. The two GCNs work together to estimate the spatial frame-wise 3D poses and can make use of both visible joint and bone information in the target frame to estimate the occluded or missing human-part information. To further refine the 3D pose estimation, we use our temporal convolutional networks (TCNs) to enforce the temporal and human-dynamics constraints. We use a joint-TCN to estimate person-centric 3D poses across frames, and propose a velocity-TCN to estimate the speed of 3D joints to ensure the consistency of the 3D pose estimation in consecutive frames. Finally, to estimate the 3D human poses for multiple persons, we propose a root-TCN that estimates camera-centric 3D poses without requiring camera parameters. Quantitative and qualitative evaluations demonstrate the effectiveness of the proposed method. Yu Cheng 0009, Bo Wang 0019, Bo Yang 0070, Robby T. Tan |
AAAI | 4 |
| 2021 | Monocular 3D Multi-Person Pose Estimation by Integrating Top-Down and Bottom-Up NetworksabstractIn monocular video 3D multi-person pose estimation, inter-person occlusion and close interactions can cause human detection to be erroneous and human-joints grouping to be unreliable. Existing top-down methods rely on human detection and thus suffer from these problems. Existing bottom-up methods do not use human detection, but they process all persons at once at the same scale, causing them to be sensitive to multiple-persons scale variations. To address these challenges, we propose the integration of top-down and bottom-up approaches to exploit their strengths. Our top-down network estimates human joints from all persons instead of one in an image patch, making it robust to possible erroneous bounding boxes. Our bottom-up network incorporates human-detection based normalized heatmaps, allowing the network to be more robust in handling scale variations. Finally, the estimated 3D poses from the top-down and bottom-up networks are fed into our integration network for final 3D poses. Besides the integration of top-down and bottom-up networks, unlike existing pose discriminators that are designed solely for a single person, and consequently cannot assess natural inter-person interactions, we propose a two-person pose discriminator that enforces natural two-person interactions. Lastly, we also apply a semi-supervised method to overcome the 3D ground-truth data scarcity. Quantitative and qualitative evaluations show the effectiveness of the proposed method. Our code is available publicly.1 Yu Cheng 0009, Bo Wang 0019, Bo Yang 0070, Robby T. Tan |
CVPR | 4 |
| 2021 | Nighttime Visibility Enhancement by Increasing the Dynamic Range and Suppression of Light EffectsabstractMost existing nighttime visibility enhancement methods focus on low light. Night images, however, do not only suffer from low light, but also from man-made light effects such as glow, glare, floodlight, etc. Hence, when the existing nighttime visibility enhancement methods are applied to these images, they intensify the effects, degrading the visibility even further. High dynamic range (HDR) imaging methods can address the low light and over-exposed regions, however they cannot remove the light effects, and thus cannot enhance the visibility in the affected regions. In this paper, given a single nighttime image as input, our goal is to enhance its visibility by increasing the dynamic range of the intensity, and thus can boost the intensity of the low light regions, and at the same time, suppress the light effects (glow, glare) simultaneously. First, we use a network to estimate the camera response function (CRF) from the input image to linearise the image. Second, we decompose the linearised image into low-frequency (LF) and high-frequency (HF) feature maps that are processed separately through two networks for light effects suppression and noise removal respectively. Third, we use a network to increase the dynamic range of the processed LF feature maps, which are then combined with the processed HF feature maps to generate the final output that has increased dynamic range and suppressed light effects. Our experiments show the effectiveness of our method in comparison with the state-of-the-art nighttime visibility enhancement methods. Aashish Sharma, Robby T. Tan |
CVPR | 2 |
| 2021 | Self-Aligned Video Deraining With Transmission-Depth ConsistencyabstractIn this paper, we address the problem of rain streaks and rain accumulation removal in video, by developing a self-alignment network with transmission-depth consistency. Existing video based deraining methods focus only on rain streak removal, and commonly use optical flow to align the rain video frames. However, besides rain streaks, rain accummulation can considerably degrade visibility; and, optical flow estimation in a rain video is still erroneous, making the deraining performance tend to be inaccurate. Our method employs deformable convolution layers in our encoder to achieve feature-level frame alignment, and hence avoids using optical flow. For rain streaks, our method predicts the current frame from its adjacent frames, such that rain streaks that appear randomly in the temporal domain can be removed. For rain accumulation, our method employs a transmission-depth consistency loss to resolve the ambiguity between the depth and water-droplet density. Our network estimates the depth from consecutive rain-accumulation-removal outputs, and calculates the transmission map using a commonly used physics model. To ensure photometric-temporal and depth-temporal consistencies, our method estimates the camera poses, so that it can warp one frame to its adjacent frames. Experimental results show that our method is effective in removing both rain streaks and rain accumulation, outperforming those of state-of-the-art methods quantitatively and qualitatively. Wending Yan, Robby T. Tan, Wenhan Yang, Dengxin Dai |
CVPR | 2 |
| 2021 | DC-ShadowNet: Single-Image Hard and Soft Shadow Removal Using Unsupervised Domain-Classifier Guided NetworkabstractShadow removal from a single image is generally still an open problem. Most existing learning-based methods use supervised learning and require a large number of paired images (shadow and corresponding non-shadow images) for training. A recent unsupervised method, Mask-ShadowGAN [13], addresses this limitation. However, it requires a binary mask to represent shadow regions, making it inapplicable to soft shadows. To address the problem, in this paper, we propose an unsupervised domain-classifier guided shadow removal network, DC-ShadowNet. Specifically, we propose to integrate a shadow/shadow-free domain classifier into a generator and its discriminator, enabling them to focus on shadow regions. To train our network, we introduce novel losses based on physics-based shadow-free chromaticity, shadow-robust perceptual features, and boundary smoothness. Moreover, we show that our unsupervised network can be used for test-time training that further improves the results. Our experiments show that all these novel components allow our method to handle soft shadows, and also to perform better on hard shadows both quantitatively and qualitatively than the existing state-of-the-art shadow removal methods. Yeying Jin, Aashish Sharma, Robby T. Tan |
ICCV | 3 |
| 2021 | Continuous-time Radar-inertial Odometry for Automotive RadarsabstractWe present an approach for radar-inertial odometry which uses a continuous-time framework to fuse measurements from multiple automotive radars and an inertial measurement unit (IMU). Adverse weather conditions do not have a significant impact on the operating performance of radar sensors unlike that of camera and LiDAR sensors. Radar’s robustness in such conditions and the increasing prevalence of radars on passenger vehicles motivate us to look at the use of radar for ego-motion estimation. A continuous-time trajectory representation is applied not only as a framework to enable heterogeneous and asynchronous multi-sensor fusion, but also, to facilitate efficient optimization by being able to compute poses and their derivatives in closed-form and at any given time along the trajectory. We compare our continuous-time estimates to those from a discrete-time radar-inertial odometry approach and show that our continuous-time method outperforms the discrete-time method. To the best of our knowledge, this is the first time a continuous-time framework has been applied to radar-inertial odometry. Yin Zhi Ng, Benjamin Choi, Robby T. Tan, Lionel Heng |
IROS | 3 |
| 2021 | Guest Editorial: Special Issue on "Computer Vision for All Seasons: Adverse Weather and Lighting Conditions"
Dengxin Dai, Robby T. Tan, Vishal M. Patel, Jiri Matas, Bernt Schiele, Luc Van Gool |
Int. J. Comput. Vis. | 2 |
| 2021 | Single Image Deraining: From Model-Based to Data-Driven and BeyondabstractThe goal of single-image deraining is to restore the rain-free background scenes of an image degraded by rain streaks and rain accumulation. The early single-image deraining methods employ a cost function, where various priors are developed to represent the properties of rain and background layers. Since 2017, single-image deraining methods step into a deep-learning era, and exploit various types of networks, i.e., convolutional neural networks, recurrent neural networks, generative adversarial networks, etc., demonstrating impressive performance. Given the current rapid development, in this paper, we provide a comprehensive survey of deraining methods over the last decade. We summarize the rain appearance models, and discuss two categories of deraining approaches: model-based and data-driven approaches. For the former, we organize the literature based on their basic models and priors. For the latter, we discuss the developed ideas related to architectures, constraints, loss functions, and training datasets. We present milestones of single-image deraining methods, review a broad selection of previous works in different categories, and provide insights on the historical development route from the model-based to data-driven methods. We also summarize performance comparisons quantitatively and qualitatively. Beyond discussing the technicality of deraining methods, we also discuss the future possible directions. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Yuming Fang 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Certainty driven consistency loss on multi-teacher networks for semi-supervised learning
Robby T. Tan |
Pattern Recognit. | 2 |
| 2020 | Nighttime Stereo Depth Estimation using Joint Translation-Stereo Learning: Light Effects and Uninformative RegionsabstractNighttime stereo depth estimation is still challenging, as assumptions associated with daytime lighting conditions do not hold any longer. Nighttime is not only about lowlight and dense noise, but also about glow/glare, flares, non-uniform distribution of light, etc. One of the possible solutions is to train a network on night stereo images in a fully supervised manner. However, to obtain proper disparity ground-truths that are dense, independent from glare/glow, and have sufficiently far depth ranges is extremely intractable. To address the problem, we introduce a network joining day/night translation and stereo. In training the network, our method does not require ground-truth disparities of the night images, or paired day/night images. We utilize a translation network that can render realistic night stereo images from day stereo images. We then train a stereo network on the rendered night stereo images using the available disparity supervision from the corresponding day stereo images, and simultaneously also train the day/night translation network. We handle the fake depth problem, which occurs due to the unsupervised/unpaired translation, for light effects (e.g., glow/glare) and uninformative regions (e.g., low-light and saturated regions), by adding structure-preservation and weighted-smoothness constraints. Our experiments show that our method outperforms the baseline methods on night images. Aashish Sharma, Loong Fah Cheong, Lionel Heng, Robby T. Tan |
3DV | 4 |
| 2020 | 3D Human Pose Estimation Using Spatio-Temporal Networks with Explicit Occlusion TrainingabstractEstimating 3D poses from a monocular video is still a challenging task, despite the significant progress that has been made in the recent years. Generally, the performance of existing methods drops when the target person is too small/large, or the motion is too fast/slow relative to the scale and speed of the training data. Moreover, to our knowledge, many of these methods are not designed or trained under severe occlusion explicitly, making their performance on handling occlusion compromised. Addressing these problems, we introduce a spatio-temporal network for robust 3D human pose estimation. As humans in videos may appear in different scales and have various motion speeds, we apply multi-scale spatial features for 2D joints or keypoints prediction in each individual frame, and multi-stride temporal convolutional networks (TCNs) to estimate 3D joints or keypoints. Furthermore, we design a spatio-temporal discriminator based on body structures as well as limb motions to assess whether the predicted pose forms a valid pose and a valid movement. During training, we explicitly mask out some keypoints to simulate various occlusion cases, from minor to severe occlusion, so that our network can learn better and becomes robust to various degrees of occlusion. As there are limited 3D ground truth data, we further utilize 2D video data to inject a semi-supervised learning capability to our network. Experiments on public data sets validate the effectiveness of our method, and our ablation studies show the strengths of our network's individual submodules. Yu Cheng 0009, Bo Yang 0070, Bo Wang 0019, Robby T. Tan |
AAAI | 4 |
| 2020 | Single-Image Camera Response Function Using Prediction Consistency and Gradual Refinement
Aashish Sharma, Robby T. Tan, Loong Fah Cheong |
ACCV (6) | 2 |
| 2020 | All in One Bad Weather Removal Using Architectural SearchabstractMany methods have set state-of-the-art performance on restoring images degraded by bad weather such as rain, haze, fog, and snow, however they are designed specifically to handle one type of degradation. In this paper, we propose a method that can handle multiple bad weather degradations: rain, fog, snow and adherent raindrops using a single network. To achieve this, we first design a generator with multiple task-specific encoders, each of which is associated with a particular bad weather degradation type. We utilize a neural architecture search to optimally process the image features extracted from all encoders. Subsequently, to convert degraded image features to clean background features, we introduce a series of tensor-based operations encapsulating the underlying physics principles behind the formation of rain, fog, snow and adherent raindrops. These operations serve as the basic building blocks for our architectural search. Finally, our discriminator simultaneously assesses the correctness and classifies the degradation type of the restored image. We design a novel adversarial learning scheme that only backpropagates the loss of a degradation type to the respective task-specific encoder. Despite being designed to handle different types of bad weather, extensive experiments demonstrate that our method performs competitively to the individual and dedicated state-of-the-art image restoration methods. Ruoteng Li, Robby T. Tan, Loong Fah Cheong |
CVPR | 2 |
| 2020 | Optical Flow in Dense Foggy Scenes Using Semi-Supervised LearningabstractIn dense foggy scenes, existing optical flow methods are erroneous. This is due to the degradation caused by dense fog particles that break the optical flow basic assumptions such as brightness and gradient constancy. To address the problem, we introduce a semi-supervised deep learning technique that employs real fog images without optical flow ground-truths in the training process. Our network integrates the domain transformation and optical flow networks in one framework. Initially, given a pair of synthetic fog images, its corresponding clean images and optical flow ground-truths, in one training batch we train our network in a supervised manner. Subsequently, given a pair of real fog images and a pair of clean images that are not corresponding to each other (unpaired), in the next training batch, we train our network in an unsupervised manner. We then alternate the training of synthetic and real data iteratively. We use real data without ground-truths, since to have ground-truths in such conditions is intractable, and also to avoid the overfitting problem of synthetic data training, where the knowledge learned on synthetic data cannot be generalized to real data testing. Together with the network architecture design, we propose a new training strategy that combines supervised synthetic-data training and unsupervised real-data training. Experimental results show that our method is effective and outperforms the state-of-the-art methods in estimating optical flow in dense foggy scenes. Wending Yan, Aashish Sharma, Robby T. Tan |
CVPR | 3 |
| 2020 | Self-Learning Video Rain Streak Removal: When Cyclic Consistency Meets Temporal CorrespondenceabstractIn this paper, we address the problem of rain streaks removal in video by developing a self-learned rain streak removal method, which does not require any clean groundtruth images in the training process. The method is inspired by fact that the adjacent frames are highly correlated and can be regarded as different versions of identical scene, and rain streaks are randomly distributed along the temporal dimension. With this in mind, we construct a two-stage Self-Learned Deraining Network (SLDNet) to remove rain streaks based on both temporal correlation and consistency. In the first stage, SLDNet utilizes the temporal correlations and learns to predict the clean version of the current frame based on its adjacent rain video frames. In the second stage, SLDNet enforces the temporal consistency among different frames. It takes both the current rain frame and adjacent rain video frames to recover structural details. The first stage is responsible for reconstructing main structures, and the second stage is responsible for extracting structural details. We build our network architecture with two sub-tasks, i.e. motion estimation, and rain region detection, and optimize them jointly. Our extensive experiments demonstrate the effectiveness of our method, offering better results both quantitatively and qualitatively. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Jiaying Liu 0001 |
CVPR | 2 |
| 2020 | Object Tracking Using Spatio-Temporal Networks for Future Prediction Location
Yuan Liu 0015, Ruoteng Li, Yu Cheng 0009, Robby T. Tan, Xiubao Sui |
ECCV (22) | 4 |
| 2020 | Nighttime Defogging Using High-Low Frequency Decomposition and Grayscale-Color Networks
Wending Yan, Robby T. Tan, Dengxin Dai |
ECCV (12) | 2 |
| 2020 | Joint Rain Detection and Removal from a Single Image with Contextualized Deep NetworksabstractRain streaks, particularly in heavy rain, not only degrade visibility but also make many computer vision algorithms fail to function properly. In this paper, we address this visibility problem by focusing on single-image rain removal, even in the presence of dense rain streaks and rain-streak accumulation, which is visually similar to mist or fog. To achieve this, we introduce a new rain model and a deep learning architecture. Our rain model incorporates a binary rain map indicating rain-streak regions, and accommodates various shapes, directions, and sizes of overlapping rain streaks, as well as rain accumulation, to model heavy rain. Based on this model, we construct a multi-task deep network, which jointly learns three targets: the binary rain-streak map, rain streak layers, and clean background, which is our ultimate output. To generate features that can be invariant to rain steaks, we introduce a contextual dilated network, which is able to exploit regional contextual information. To handle various shapes and directions of overlapping rain streaks, our strategy is to utilize a recurrent process that progressively removes rain streaks. Our binary map provides a constraint and thus additional information to train our network. Extensive evaluation on real images, particularly in heavy rain, shows the effectiveness of our model and architecture. Wenhan Yang, Robby T. Tan, Jiashi Feng, Zongming Guo, Shuicheng Yan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Heavy Rain Image Restoration: Integrating Physics Model and Conditional Adversarial LearningabstractMost deraining works focus on rain streaks removal but they cannot deal adequately with heavy rain images. In heavy rain, streaks are strongly visible, dense rain accumulation or rain veiling effect significantly washes out the image, further scenes are relatively more blurry, etc. In this paper, we propose a novel method to address these problems. We put forth a 2-stage network: a physics-based backbone followed by a depth-guided GAN refinement. The first stage estimates the rain streaks, the transmission, and the atmospheric light governed by the underlying physics. To tease out these components more reliably, a guided filtering framework is used to decompose the image into its low- and high-frequency components. This filtering is guided by a rain-free residue image - its content is used to set the passbands for the two channels in a spatially-variant manner so that the background details do not get mixed up with the rain-streaks. For the second stage, the refinement stage, we put forth a depth-guided GAN to recover the background details failed to be retrieved by the first stage, as well as correcting artefacts introduced by that stage. We have evaluated our method against state of the art methods. Extensive experiments show that our method outperforms them on real rain image data, recovering visually clean images with good details. Ruoteng Li, Loong Fah Cheong, Robby T. Tan |
CVPR | 3 |
| 2019 | Occlusion-Aware Networks for 3D Human Pose Estimation in VideoabstractOcclusion is a key problem in 3D human pose estimation from a monocular video. To address this problem, we introduce an occlusion-aware deep-learning framework. By employing estimated 2D confidence heatmaps of keypoints and an optical-flow consistency constraint, we filter out the unreliable estimations of occluded keypoints. When occlusion occurs, we have incomplete 2D keypoints and feed them to our 2D and 3D temporal convolutional networks (2D and 3D TCNs) that enforce temporal smoothness to produce a complete 3D pose. By using incomplete 2D keypoints, instead of complete but incorrect ones, our networks are less affected by the error-prone estimations of occluded keypoints. Training the occlusion-aware 3D TCN requires pairs of a 3D pose and a 2D pose with occlusion labels. As no such a dataset is available, we introduce a ``Cylinder Man Model'' to approximate the occupation of body parts in 3D space. By projecting the model onto a 2D plane in different viewing angles, we obtain and label the occluded keypoints, providing us plenty of training data. In addition, we use this model to create a pose regularization constraint, preferring the 2D estimations of unreliable keypoints to be occluded. Our method outperforms state-of-the-art methods on Human 3.6M and HumanEva-I datasets. Yu Cheng 0009, Bo Yang 0070, Bo Wang 0019, Wending Yan, Robby T. Tan |
ICCV | 5 |
| 2019 | RainFlow: Optical Flow Under Rain Streaks and Rain Veiling EffectabstractOptical flow in heavy rainy scenes is challenging due to the presence of both rain steaks and rain veiling effect, which break the existing optical flow constraints. Concerning this, we propose a deep-learning based optical flow method designed to handle heavy rain. We introduce a feature multiplier in our network that transforms the features of an image affected by the rain veiling effect into features that are less affected by it, which we call veiling-invariant features. We establish a new mapping operation in the feature space to produce streak-invariant features. The operation is based on a feature pyramid structure of the input images, and the basic idea is to preserve the chromatic features of the background scenes while canceling the rain-streak patterns. Both the veiling-invariant and streak-invariant features are computed and optimized automatically based on the the accuracy of our optical flow estimation. Our network is end-to-end, and handles both rain streaks and the veiling effect in an integrated framework. Extensive experiments show the effectiveness of our method, which outperforms the state of the art method and other baseline methods. We also show that our network can robustly maintain good performance on clean (no rain) images even though it is trained under rain image data. Ruoteng Li, Robby T. Tan, Loong Fah Cheong, Angelica I. Avilés-Rivero, Qingnan Fan, Carola-Bibiane Schönlieb |
ICCV | 2 |
| 2019 | GraphX $$^\mathbf{\small NET } -$$ -Chest X-Ray Classification Under Extreme Minimal Supervision
Angelica I. Avilés-Rivero, Nicolas Papadakis, Ruoteng Li, Philip Sellars, Qingnan Fan, Robby T. Tan, Carola-Bibiane Schönlieb |
MICCAI (6) | 6 |
| 2018 | Loss Guided Activation for Action Recognition in Still Images
Robby T. Tan, Shaodi You |
ACCV (5) | 2 |
| 2018 | Attentive Generative Adversarial Network for Raindrop Removal From a Single ImageabstractRaindrops adhered to a glass window or camera lens can severely hamper the visibility of a background scene and degrade an image considerably. In this paper, we address the problem by visually removing raindrops, and thus transforming a raindrop degraded image into a clean one. The problem is intractable, since first the regions occluded by raindrops are not given. Second, the information about the background scene of the occluded regions is completely lost for most part. To resolve the problem, we apply an attentive generative network using adversarial training. Our main idea is to inject visual attention into both the generative and discriminative networks. During the training, our visual attention learns about raindrop regions and their surroundings. Hence, by injecting this information, the generative network will pay more attention to the raindrop regions and the surrounding structures, and the discriminative network will be able to assess the local consistency of the restored regions. This injection of visual attention to both generative and discriminative networks is the main contribution of this paper. Our experiments show the effectiveness of our approach, which outperforms the state of the art methods quantitatively and qualitatively. Rui Qian 0003, Robby T. Tan, Wenhan Yang, Jiajun Su, Jiaying Liu 0001 |
CVPR | 2 |
| 2018 | Robust Optical Flow in Rainy Scenes
Ruoteng Li, Robby T. Tan, Loong Fah Cheong |
ECCV (15) | 2 |
| 2018 | Video Foreground Cosegmentation Based on Common FateabstractExtracting an object of interest from a single video still faces significant difficulties when the object has variegated appearance, manifests articulated motion, or experiences occlusions by other objects. In this paper, we present a video cosegmentation method to address the aforementioned challenges. Departing from the objectness attributes and motion coherence used by traditional foreground-background separation and video segmentation methods, we place central importance in the role of “common fate.” Specifically, the different parts of the object should persist together in all the videos despite the possible presence of incoherent (e.g., articulated) motions. To accomplish this idea, we first extract seed superpixels by a motion-based foreground segmentation method. We next formulate a set of initial to-link constraints between these superpixels based on whether they exhibit the characteristics of common fate. An iterative manifold ranking algorithm is then proposed to trim away the incorrect and accidental linkage relationships. Having discovered the parts that should cohere together, we next perform clustering to extract the entire object and to handle the case in which there might be multiple objects present. This clustering is performed at two levels: the superpixel level and the object level. This two-level clustering algorithm also performs automatic model selection to estimate the number of object classes extracted. Finally, a multiclass labeling Markov random field is used to obtain a refined segmentation result. To evaluate the performance of our framework, we introduce a new data set in which the videos have complex form and motion that are liable to ambiguity in interpretation. Our experimental results on this data set show that our method successfully addresses the challenges in the extraction of complex objects and outperforms the state-of-the-art video segmentation and cosegmentation methods in our data set. Jiaming Guo, Loong Fah Cheong, Robby T. Tan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Deep Joint Rain Detection and Removal from a Single ImageabstractIn this paper, we address a rain removal problem from a single image, even in the presence of heavy rain and rain streak accumulation. Our core ideas lie in our new rain image model and new deep learning architecture. We add a binary map that provides rain streak locations to an existing model, which comprises a rain streak layer and a background layer. We create a model consisting of a component representing rain streak accumulation (where individual streaks cannot be seen, and thus visually similar to mist or fog), and another component representing various shapes and directions of overlapping rain streaks, which usually happen in heavy rain. Based on the model, we develop a multi-task deep learning architecture that learns the binary rain streak map, the appearance of rain streaks, and the clean background, which is our ultimate output. The additional binary map is critically beneficial, since its loss function can provide additional strong information to the network. To handle rain streak accumulation (again, a phenomenon visually similar to mist or fog) and various shapes and directions of overlapping rain streaks, we propose a recurrent rain detection and removal network that removes rain streaks and clears up the rain accumulation iteratively and progressively. In each recurrence of our method, a new contextualized dilated network is developed to exploit regional contextual information and to produce better representations for rain detection. The evaluation on real images, particularly on heavy rain, shows the effectiveness of our models and architecture. Wenhan Yang, Robby T. Tan, Jiashi Feng, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
CVPR | 2 |
| 2017 | Haze visibility enhancement: A Survey and quantitative benchmarking
Yu Li 0003, Shaodi You, Michael S. Brown, Robby T. Tan |
Comput. Vis. Image Underst. | 4 |
| 2017 | Water detection through spatio-temporal invariant descriptorsabstractIn this work, we aim to segment and detect water in videos. Water detection is beneficial for appllications such as video search, outdoor surveillance, and systems such as unmanned ground vehicles and unmanned aerial vehicles . The specific problem, however, is less discussed compared to general texture recognition. Here, we analyze several motion properties of water. First, we describe a video pre-processing step, to increase invariance against water reflections and water colours. Second, we investigate the temporal and spatial properties of water and derive corresponding local descriptors . The descriptors are used to locally classify the presence of water and a binary water detection mask is generated through spatio-temporal Markov Random Field regularization of the local classifications. Third, we introduce the Video Water Database, containing several hours of water and non-water videos, to validate our algorithm. Experimental evaluation on the Video Water Database and the DynTex database indicates the effectiveness of the proposed algorithm, outperforming multiple algorithms for dynamic texture recognition and material recognition. Pascal Mettes, Robby T. Tan, Remco C. Veltkamp |
Comput. Vis. Image Underst. | 2 |
| 2017 | Single Image Rain Streak Decomposition Using Layer PriorsabstractRain streaks impair visibility of an image and introduce undesirable interference that can severely affect the performance of computer vision and image analysis systems. Rain streak removal algorithms try to recover a rain streak free background scene. In this paper, we address the problem of rain streak removal from a single image by formulating it as a layer decomposition problem, with a rain streak layer superimposed on a background layer containing the true scene content. Existing decomposition methods that address this problem employ either sparse dictionary learning methods or impose a low rank structure on the appearance of the rain streaks. While these methods can improve the overall visibility, their performance can often be unsatisfactory, for they tend to either over-smooth the background images or generate -images that still contain noticeable rain streaks. To address the problems, we propose a method that imposes priors for both the background and rain streak layers. These priors are based on Gaussian mixture models learned on small patches that can accommodate a variety of background appearances as well as the appearance of the rain streaks. Moreover, we introduce a structure residue recovery step to further separate the background residues and improve the decomposition quality. Quantitative evaluation shows our method outperforms existing methods by a large margin. We overview our method and demonstrate its effectiveness over prior work on a number of examples. Yu Li 0003, Robby T. Tan, Xiaojie Guo 0001, Jiangbo Lu, Michael S. Brown |
IEEE Trans. Image Process. | 2 |
| 2016 | Rain Streak Removal Using Layer PriorsabstractThis paper addresses the problem of rain streak removal from a single image. Rain streaks impair visibility of an image and introduce undesirable interference that can severely affect the performance of computer vision algorithms. Rain streak removal can be formulated as a layer decomposition problem, with a rain streak layer superimposed on a background layer containing the true scene content. Existing decomposition methods that address this problem employ either dictionary learning methods or impose a low rank structure on the appearance of the rain streaks. While these methods can improve the overall visibility, they tend to leave too many rain streaks in the background image or over-smooth the background image. In this paper, we propose an effective method that uses simple patch-based priors for both the background and rain layers. These priors are based on Gaussian mixture models and can accommodate multiple orientations and scales of the rain streaks. This simple approach removes rain streaks better than the existing methods qualitatively and quantitatively. We overview our method and demonstrate its effectiveness over prior work on a number of examples. Yu Li 0003, Robby T. Tan, Xiaojie Guo 0001, Jiangbo Lu, Michael S. Brown |
CVPR | 2 |
| 2016 | Robust Optical Flow Estimation of Double-Layer Images under Transparency or ReflectionabstractThis paper deals with a challenging, frequently encountered, yet not properly investigated problem in two-frame optical flow estimation. That is, the input frames are compounds of two imaging layers - one desired background layer of the scene, and one distracting, possibly moving layer due to transparency or reflection. In this situation, the conventional brightness constancy constraint - the cornerstone of most existing optical flow methods - will no longer be valid. In this paper, we propose a robust solution to this problem. The proposed method performs both optical flow estimation, and image layer separation. It exploits a generalized double-layer brightness consistency constraint connecting these two tasks, and utilizes the priors for both of them. Experiments on both synthetic data and real images have confirmed the efficacy of the proposed method. To the best of our knowledge, this is the first attempt towards handling generic optical flow fields of two-frame images containing transparency or reflection. Jiaolong Yang, Hongdong Li, Yuchao Dai, Robby T. Tan |
CVPR | 4 |
| 2016 | Adherent Raindrop Modeling, Detectionand Removal in VideoabstractRaindrops adhered to a windscreen or window glass can significantly degrade the visibility of a scene. Modeling, detecting and removing raindrops will, therefore, benefit many computer vision applications, particularly outdoor surveillance systems and intelligent vehicle systems. In this paper, a method that automatically detects and removes adherent raindrops is introduced. The core idea is to exploit the local spatio-temporal derivatives of raindrops. To accomplish the idea, we first model adherent raindrops using law of physics, and detect raindrops based on these models in combination with motion and intensity temporal derivatives of the input video. Having detected the raindrops, we remove them and restore the images based on an analysis that some areas of raindrops completely occludes the scene, and some other areas occlude only partially. For partially occluding areas, we restore them by retrieving as much as possible information of the scene, namely, by solving a blending function on the detected partially occluding areas using the temporal intensity derivative. For completely occluding areas, we recover them by using a video completion technique. Experimental results using various real videos show the effectiveness of our method. Shaodi You, Robby T. Tan, Rei Kawakami, Yasuhiro Mukaigawa, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Understanding image concepts using ISTOP model
Mohsen Sardari Zarchi, Robby T. Tan, Coert Van Gemeren, S. Amirhassan Monadjemi, Remco C. Veltkamp |
Pattern Recognit. | 2 |
| 2015 | Simultaneous video defogging and stereo reconstructionabstractWe present a method to jointly estimate scene depth and recover the clear latent image from a foggy video sequence. In our formulation, the depth cues from stereo matching and fog information reinforce each other, and produce superior results than conventional stereo or defogging algorithms. We first improve the photo-consistency term to explicitly model the appearance change due to the scattering effects. The prior matting Laplacian constraint on fog transmission imposes a detail-preserving smoothness constraint on the scene depth. We further enforce the ordering consistency between scene depth and fog transmission at neighboring points. These novel constraints are formulated together in an MRF framework, which is optimized iteratively by introducing auxiliary variables. The experiment results on real videos demonstrate the strength of our method. Zhuwen Li, Robby T. Tan, Danping Zou, Steven Zhiying Zhou, Loong Fah Cheong |
CVPR | 3 |
| 2015 | Nighttime Haze Removal with Glow and Multiple Light ColorsabstractThis paper focuses on dehazing nighttime images. Most existing dehazing methods use models that are formulated to describe haze in daytime. Daytime models assume a single uniform light color attributed to a light source not directly visible in the scene. Nighttime scenes, however, commonly include visible lights sources with varying colors. These light sources also often introduce noticeable amounts of glow that is not present in daytime haze. To address these effects, we introduce a new nighttime haze model that accounts for the varying light sources and their glow. Our model is a linear combination of three terms: the direct transmission, airlight and glow. The glow term represents light from the light sources that is scattered around before reaching the camera. Based on the model, we propose a framework that first reduces the effect of the glow in the image, resulting in a nighttime image that consists of direct transmission and airlight only. We then compute a spatially varying atmospheric light map that encodes light colors locally. This atmospheric map is used to predict the transmission, which we use to obtain our nighttime scene reflection image. We demonstrate the effectiveness of our nighttime haze model and correction method on a number of examples and compare our results with existing daytime and nighttime dehazing methods' results. Yu Li 0003, Robby T. Tan, Michael S. Brown |
ICCV | 2 |
| 2014 | Consistent Foreground Co-segmentation
Jiaming Guo, Loong Fah Cheong, Robby T. Tan, Steven Zhiying Zhou |
ACCV (4) | 3 |
| 2014 | Raindrop Detection and Removal from Long Range Trajectories
Shaodi You, Robby T. Tan, Rei Kawakami, Yasuhiro Mukaigawa, Katsushi Ikeuchi |
ACCV (2) | 2 |
| 2014 | A Contrast Enhancement Framework with JPEG Artifacts Suppression
Yu Li 0003, Fangfang Guo, Robby T. Tan, Michael S. Brown |
ECCV (2) | 3 |
| 2013 | Adherent Raindrop Detection and Removal in VideoabstractRaindrops adhered to a windscreen or window glass can significantly degrade the visibility of a scene. Detecting and removing raindrops will, therefore, benefit many computer vision applications, particularly outdoor surveillance systems and intelligent vehicle systems. In this paper, a method that automatically detects and removes adherent raindrops is introduced. The core idea is to exploit the local spatio-temporal derivatives of raindrops. First, it detects raindrops based on the motion and the intensity temporal derivatives of the input video. Second, relying on an analysis that some areas of a raindrop completely occludes the scene, yet the remaining areas occludes only partially, the method removes the two types of areas separately. For partially occluding areas, it restores them by retrieving as much as possible information of the scene, namely, by solving a blending function on the detected partially occluding areas using the temporal intensity change. For completely occluding areas, it recovers them by using a video completion technique. Experimental results using various real videos show the effectiveness of the proposed method. Shaodi You, Robby T. Tan, Rei Kawakami, Katsushi Ikeuchi |
CVPR | 2 |
| 2013 | Camera Spectral Sensitivity and White Balance Estimation from Sky ImagesabstractPhotometric camera calibration is often required in physics-based computer vision. There have been a number of studies to estimate camera response functions (gamma function), and vignetting effect from images. However less attention has been paid to camera spectral sensitivities and white balance settings. This is unfortunate, since those two properties significantly affect image colors. Motivated by this, a method to estimate camera spectral sensitivities and white balance setting jointly from images with sky regions is introduced. The basic idea is to use the sky regions to infer the sky spectra. Given sky images as the input and assuming the sun direction with respect to the camera viewing direction can be extracted, the proposed method estimates the turbidity of the sky by fitting the image intensities to a sky model. Subsequently, it calculates the sky spectra from the estimated turbidity. Having the sky $$RGB$$ values and their corresponding spectra, the method estimates the camera spectral sensitivities together with the white balance setting. Precomputed basis functions of camera spectral sensitivities are used in the method for robust estimation. The whole method is novel and practical since, unlike existing methods, it uses sky images without additional hardware, assuming the geolocation of the captured sky is known. Experimental results using various real images show the effectiveness of the method. Rei Kawakami, Hongxun Zhao, Robby T. Tan, Katsushi Ikeuchi |
Int. J. Comput. Vis. | 3 |
| 2011 | Multi-person tracking based on vertical reference lines and dynamic visibility analysisabstractMultiple people tracking from multiple cameras can suffer from various problems, particularly from inter-person occlusions. This paper attempts to solve the problems by analyzing the view visibility and ranking the reliability of the cues from 2D views. It combines the visibility with the smoothness constraints into a probability framework, which offers a more flexible and robust estimation. Moreover, it introduces 3D reference lines to estimate the 2D position of every individual in the input images. These lines can estimate more accurate and robust 2D positions. The experimental results and quantitative evaluations on the standard data set show the effectiveness of the method. Xinghan Luo, Robby T. Tan, Remco C. Veltkamp |
ICIP | 2 |
| 2010 | Estimating optical properties of layered surfaces using the spider modelabstractMany object surfaces are composed of layers of different physical substances, known as layered surfaces. These surfaces, such as patinas, water colors, and wall paintings, have more complex optical properties than diffuse surfaces. Although the characteristics of layered surfaces, like layer opacity, mixture of colors, and color gradations, are significant, they are usually ignored in the analysis of many methods in computer vision, causing inaccurate or even erroneous results. Therefore, the main goals of this paper are twofold: to solve problems of layered surfaces by focusing mainly on surfaces with two layers (i.e., top and bottom layers), and to introduce a decomposition method based on a novel representation of a nonlinear correlation in the color space that we call the “spider” model. When we plot a mixture of colors of one bottom layer and n different top layers into the RGB color space, then we will have n different curves intersecting at one point, resembling the shape of a spider. Hence, given a single input image containing one bottom layer and at least one top layer, we can fit their color distributions by using the spider model and then decompose those layered surfaces. The last step is equivalent to extracting the approximated optical properties of the two layers: the top layer's opacity, and the top and bottom layers' reflections. Experiments with real images, which include the photographs of ancient wall paintings, show the effectiveness of our method. Tetsuro Morimoto, Robby T. Tan, Rei Kawakami, Katsushi Ikeuchi |
CVPR | 2 |
| 2010 | Human Pose Estimation for Multiple Persons Based on Volume ReconstructionabstractMost of the development of pose recognition focused on a single person. However, many applications of computer vision essentially require the estimation of multiple people. Hence, in this paper, we address the problems of estimating poses of multiple persons using volumes estimated from multiple cameras. One of the main issues that causes the multiple person from multiple cameras to be problematic is the present of `ghost' volumes. This problem arises when the projections of two different silhouettes of two different persons onto the 3D world overlap in a place where in fact there is no person in it. To solve this problem, we first introduce a novel principal axis-based framework to estimate the 3D ground plane positions of multiple people, and then use the position cues to label the multi-person volumes (voxels), while considering the voxel connectivity. Having labeled the voxels, we fit the volume of each person with a body model, and determine the pose of the person based on the model. The results on real videos demonstrate the accuracy and efficiency of our approach. Xinghan Luo, Berend Berendsen, Robby T. Tan, Remco C. Veltkamp |
ICPR | 3 |
| 2008 | Visibility in bad weather from a single imageabstractBad weather, such as fog and haze, can significantly degrade the visibility of a scene. Optically, this is due to the substantial presence of particles in the atmosphere that absorb and scatter light. In computer vision, the absorption and scattering processes are commonly modeled by a linear combination of the direct attenuation and the airlight. Based on this model, a few methods have been proposed, and most of them require multiple input images of a scene, which have either different degrees of polarization or different atmospheric conditions. This requirement is the main drawback of these methods, since in many situations, it is difficult to be fulfilled. To resolve the problem, we introduce an automated method that only requires a single input image. This method is based on two basic observations: first, images with enhanced visibility (or clear-day images) have more contrast than images plagued by bad weather; second, airlight whose variation mainly depends on the distance of objects to the viewer, tends to be smooth. Relying on these two observations, we develop a cost function in the framework of Markov random fields, which can be efficiently optimized by various techniques, such as graph-cuts or belief propagation. The method does not require the geometrical information of the input image, and is applicable for both color and gray images. Robby T. Tan |
CVPR | 1 |
| 2007 | On Automatic Absorption Detection for Imaging Spectroscopy: A Comparative StudyabstractIn this paper, we aim at presenting a survey on automatic absorption recovery methods for imaging spectroscopy. We commence by viewing the algorithms in the literature from a technical perspective and presenting an overview of the derivative analysis, fingerprint, and maximum modulus wavelet transform techniques. In addition to these methods, we also present a novel absorption recovery approach based upon unimodal regression and continuum removal. With this technical review of the methods under study, we perform a complexity analysis and examine the implementation issues pertaining to each of the alternatives. We show how detected absorption bands can be used for purposes of material identification. We conclude this paper by providing a performance study and providing identification results on hyperspectral imagery. To this end, we make use of a number of distance measures to evaluate the quality of the recovered absorptions, as compared to continuum-removed spectra. Zhouyu Fu, Antonio Robles-Kelly, Terry Caelli, Robby T. Tan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2005 | Reflection Components Decomposition of Textured Surfaces Using Linear Basis FunctionsabstractMost existing methods of reflection components decomposition using a single color image require color segmentation. Few methods that employ local operations are able to avoid the requirement; however, they usually suffer from color discontinuity problems. In this paper, we introduce a decomposition method using a single color image that does not require (global) color segmentation or (local) color discontinuity detection. The method principally utilizes the coefficients of the reflectance basis functions of input image and its specular-free image. Combining those coefficients enables us to find the diffuse coefficients of the specular pixels for every surface color. As a result, the decomposition becomes a well-posed problem and able to be solved in closed-form equations. Our experimental results on real complex textured images show the effectiveness of our proposed method. Robby T. Tan, Katsushi Ikeuchi |
CVPR (1) | 1 |
| 2005 | Consistent Surface Color for Texturing Large Objects in Outdoor ScenesabstractColor appearance of an object is significantly influenced by the color of the illumination. When the illumination color changes, the color appearance of the object change accordingly, causing its appearance to be inconsistent. To arrive at color constancy, we have developed a physics-based method of estimating and removing the illumination color. In this paper, we focus on the use of this method to deal with outdoor scenes, since very few physics-based methods have successfully handled outdoor color constancy. Our method is principally based on shadowed and non-shadowed regions. Previously researchers have discovered that shadowed regions are illuminated by sky light, while non-shadowed regions are illuminated by a combination of sky light and sunlight. Based on this difference of illumination, we estimate the illumination colors (both the sunlight and the sky light) and then remove them. To reliably estimate the illumination colors in outdoor scenes, we include the analysis of noise, since the presence of noise is inevitable in natural images. As a result, compared to existing methods, the proposed method is more effective and robust in handling outdoor scenes. In addition, the proposed method requires only a single input image, making it useful for many applications of computer vision Rei Kawakami, Katsushi Ikeuchi, Robby T. Tan |
ICCV | 3 |
| 2005 | Separating Reflection Components of Textured Surfaces Using a Single ImageabstractIn inhomogeneous objects, highlights are linear combinations of diffuse and specular reflection components. A number of methods have been proposed to separate or decompose these two components. To our knowledge, all methods that use a single input image require explicit color segmentation to deal with multicolored surfaces. Unfortunately, for complex textured images, current color segmentation algorithms are still problematic to segment correctly. Consequently, a method without explicit color segmentation becomes indispensable and this paper presents such a method. The method is based solely on colors, particularly chromaticity, without requiring any geometrical information. One of the basic ideas is to iteratively compare the intensity logarithmic differentiation of an input image and its specular-free image. A specular-free image is an image that has exactly the same geometrical profile as the diffuse component of the input image and that can be generated by shifting each pixel's intensity and maximum chromaticity nonlinearly. Unlike existing methods using a single image, all processes in the proposed method are done locally, involving a maximum of only two neighboring pixels. This local operation is useful for handling textured objects with complex multicolored scenes. Evaluations by comparison with the results of polarizing filters demonstrate the effectiveness of the proposed method. Robby T. Tan, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Separating Reflection Components Based on Chromaticity and Noise AnalysisabstractMany algorithms in computer vision assume diffuse only reflections and deem specular reflections to be outliers. However, in the real world, the presence of specular reflections is inevitable since there are many dielectric inhomogeneous objects which have both diffuse and specular reflections. To resolve this problem, we present a method to separate the two reflection components. The method is principally based on the distribution of specular and diffuse points in a two-dimensional maximum chromaticity-intensity space. We found that, by utilizing the space and known illumination color, the problem of reflection component separation can be simplified into the problem of identifying diffuse maximum chromaticity. To be able to identify the diffuse maximum chromaticity correctly, an analysis of the noise is required since most real images suffer from it. Unlike existing methods, the proposed method can separate the reflection components robustly for any kind of surface roughness and light direction. Robby T. Tan, Ko Nishino, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Illumination Chromaticity Estimation using Inverse-Intensity Chromaticity SpaceabstractExisting color constancy methods cannot handle both uniform colored surfaces and highly textured surfaces in a single integrated framework. Statistics-based methods require many surface colors, and become error prone when there are only few surface colors. In contrast, dichromatic-based methods can successfully handle uniformly colored surfaces, but cannot be applied to highly textured surfaces since they require precise color segmentation. In this paper, we present a single integrated method to estimate illumination chromaticity from single/multi-colored surfaces. Unlike the existing dichromatic-based methods, the proposed method requires only rough highlight regions, without segmenting the colors inside them. We show that, by analyzing highlights, a direct correlation between illumination chromaticity and image chromaticity can be obtained. This correlation is clearly described in "inverse-intensity chromaticity space", a new two-dimensional space we introduce. In addition, by utilizing the Hough transform and histogram analysis in this space, illumination chromaticity can be estimated robustly, even for a highly textured surface. Experimental results on real images show the effectiveness of the method. Robby T. Tan, Ko Nishino, Katsushi Ikeuchi |
CVPR (1) | 1 |
| 2003 | Polarization-based Inverse Rendering from a Single ViewabstractThis paper presents a method to estimate geometrical, photometrical, and environmental information of a single-viewed object in one integrated framework under fixed viewing position and fixed illumination direction. These three types of information are important to render a photorealistic image of a real object. Photometrical information represents the texture and the surface roughness of an object, while geometrical and environmental information represent the 3D shape of an object and the illumination distribution, respectively. The proposed method estimates the 3D shape by computing the surface normal from polarization data, calculates the texture of the object from the diffuse only reflection component, determines the illumination directions from the position of the brightest intensity in the specular reflection component, and finally computes the surface roughness of the object by using the estimated illumination distribution. Daisuke Miyazaki, Robby T. Tan, Kenji Hara, Katsushi Ikeuchi |
ICCV | 2 |
| 2003 | Separating Reflection Components of Textured Surfaces using a Single ImageabstractThe presence of highlights, which in dielectric inhomogeneous objects are linear combination of specular and diffuse reflection components, is inevitable. A number of methods have been developed to separate these reflection components. To our knowledge, all methods that use a single input image require explicit color segmentation to deal with multicolored surfaces. Unfortunately, for complex textured images, current color segmentation algorithms are still problematic to segment correctly. Consequently, a method without explicit color segmentation becomes indispensable, and this paper presents such a method. The method is based solely on colors, particularly chromaticity, without requiring any geometrical parameter information. One of the basic ideas is to compare the intensity logarithmic differentiation of specular-free images and input images iteratively. The specular-free image is a pseudo-code of diffuse components that can be generated by shifting a pixel's intensity and chromaticity nonlinearly while retaining its hue. All processes in the method are done locally, involving a maximum of only two pixels. The experimental results on natural images show that the proposed method is accurate and robust under known scene illumination chromaticity. Unlike the existing methods that use a single image, our method is effective for textured objects with complex multicolored scenes. Robby T. Tan, Katsushi Ikeuchi |
ICCV | 1 |