EDBT 2026 Demo / reviewers in the wild / expert
Yuewei Lin
dblp:41/1100
· DBLP profile ↗
39ranked-venue papers
6as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 11 since 2021Computer networks · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FCC: Fully Connected Correlation for One-Shot SegmentationabstractOne-shot segmentation (OSS) aims to segment the target object in a query image using only one set of support image and mask. Therefore, having strong prior information for the target object using the support set is essential to guide the initial training of OSS, which leads to the success of one-shot segmentation in challenging cases, such as when the target object shows considerable variation in appearance, texture, or scale across the support and query images. To enrich this prior knowledge, we introduce FCC (Fully Connected Correlation) which integrates pixel-level correlations between support and query features, capturing associations that reveal target-specific patterns and correspondences in both same-layers and cross-layers. FCC captures previously inaccessible target information, effectively addressing the limitations of support mask. Our approach consistently demonstrates state-of-the-art performance in the PASCAL, COCO, and domain shift tests, while also notably accelerating model convergence. We conducted an ablation study and cross-layer correlation analysis to validate FCC’s core methodology. These findings reveal the effectiveness of FCC in enhancing prior information and overall model performance for OSS1. Seonghyeon Moon, Haein Kong, Muhammad Haris Khan, Mubbasir Kapadia, Yuewei Lin |
WACV | 5 |
| 2025 | PRVQL: Progressive Knowledge-Guided Refinement for Robust Egocentric Visual Query LocalizationabstractEgocentric visual query localization (EgoVQL) focuses on localizing the target of interest in space and time from first-person videos, given a visual query. Despite recent progressive, existing methods often struggle to handle severe object appearance changes and cluttering background in the video due to lacking sufficient target cues, leading to degradation. Addressing this, we introduce PRVQL, a novel Progressive knowledge-guided Refinement framework for EgoVQL. The core is to continuously exploit target-relevant knowledge directly from videos and utilize it as guidance to refine both query and video features for improving target localization. Our PRVQL contains multiple processing stages. The target knowledge from one stage, comprising appearance and spatial knowledge extracted via two specially designed knowledge learning modules, are utilized as guidance to refine the query and videos features for the next stage, which are used to generate more accurate knowledge for further feature refinement. With such a progressive process, target knowledge in PRVQL can be gradually improved, which, in turn, leads to better refined query and video features for localization in the final stage. Compared to previous methods, our PRVQL, besides the given object cues, enjoys additional crucial target information from a video as guidance to refine features, and hence enhances EgoVQL in complicated scenes. In our experiments on challenging Ego4D, PRVQL achieves state-of-the-art result and largely surpasses other methods, showing its efficacy. Our code, model and results will be released at https://github.com/fb-reps/PRVQL. Bing Fan, Yunhe Feng, Yapeng Tian, James Liang, Yuewei Lin, Yan Huang 0002, Heng Fan 0001 |
ICCV | 5 |
| 2025 | Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video GroundingabstractTransformer has attracted increasing interest in spatio-temporal video grounding, or STVG, owing to its end-to-end pipeline and promising result. Existing Transformer-based STVG approaches often leverage a set of object queries, which are initialized simply using zeros and then gradually learn target position information via iterative interactions with multimodal features, for spatial and temporal localization. Despite simplicity, these zero object queries, due to lacking target-specific cues, are hard to learn discriminative target information from interactions with multimodal features in complicated scenarios (e.g., with distractors or occlusion), resulting in degradation. Addressing this, we introduce a novel $\textbf{T}$arget-$\textbf{A}$ware Transformer for $\textbf{STVG}$ ($\textbf{TA-STVG}$), which seeks to adaptively generate object queries via exploring target-specific cues from the given video-text pair, for improving STVG. The key lies in two simple yet effective modules, comprising text-guided temporal sampling (TTS) and attribute-aware spatial activation (ASA), working in a cascade. The former focuses on selecting target-relevant temporal cues from a video utilizing holistic text information, while the latter aims at further exploiting the fine-grained visual attribute information of the object from previous target-aware temporal cues, which is applied for object query initialization. Compared to existing methods leveraging zero-initialized queries, object queries in our TA-STVG, directly generated from a given video-text pair, naturally carry target-specific cues, making them adaptive and better interact with multimodal features for learning more discriminative information to improve STVG. In our experiments on three benchmarks, including HCSTVG-v1/-v2 and VidSTG, TA-STVG achieves state-of-the-art performance and significantly outperforms the baseline, validating its efficacy. Moreover, TTS and ASA are designed for general purpose. When applied to existing methods such as TubeDETR and STCAT, we show substantial performance gains, verifying its generality. Code is released at https://github.com/HengLan/TA-STVG. Xin Gu 0003, Yaojie Shen, Chenxi Luo, Tiejian Luo, Yan Huang 0002, Yuewei Lin, Heng Fan 0001, Libo Zhang 0001 |
ICLR | 6 |
| 2025 | Efficient and Accurate Low-Resolution Transformer TrackingabstractHigh-performance Transformer trackers have exhibited excellent results, yet they often bear a heavy computational load. Observing that a smaller input can immediately and conveniently reduce computations without changing the model, an easy solution is to adopt a low-resolution input for efficient Transformer tracking. Albeit faster, this hurts tracking accuracy much due to the information loss in low resolution tracking. In this paper, we aim to mitigate such information loss to boost performance of low-resolution Transformer tracking via dual knowledge distillation from a frozen high-resolution (but not a larger) Transformer tracker. The core lies in two simple yet effective distillation modules, including query-key-value knowledge distillation (QKV-KD) and discrimination knowledge distillation (Disc-KD), across resolutions. The former, from the global view, allows the low-resolution tracker to inherit features and interactions from the high-resolution tracker, while the later, from the target-aware view, enhances the target-background distinguishing capacity via imitating discriminative regions from its high-resolution counterpart. With dual knowledge distillation, our Low-Resolution Transformer Tracker, dubbed LoReTrack, enjoys not only high efficiency owing to reduced computation but also enhanced accuracy by distilling knowledge from the high-resolution tracker. In extensive experiments, LoReTrack with a 2562resolution consistently improves baseline with the same resolution, and shows competitive or better results compared to the 3842high-resolution Transformer tracker, while running 52% faster and saving 56% MACs. Moreover, LoReTrack is resolution-scalable. With a 1282resolution, it runs 25 fps on a CPU with SUC scores of 64.9%/46.4% on LaSOT/LaSOText, surpassing other CPU real-time trackers. Code is released at https://github.com/ShaohuaDong2021/LoReTrack. Shaohua Dong, Yunhe Feng, James Liang, Qing Yang 0003, Yuewei Lin, Heng Fan 0001 |
IROS | 5 |
| 2025 | VAPO: Visibility-Aware Keypoint Localization for Efficient 6DoF Object Pose EstimationabstractLocalizing predefined 3D keypoints in a 2D image is an effective way to establish 3D-2D correspondences for instance-level 6DoF object pose estimation. However, unreliable localization results of invisible keypoints degrade the quality of correspondences. In this paper, we address this issue by localizing the important keypoints in terms of visibility. Since keypoint visibility information is currently missing in the dataset collection process, we propose an efficient way to generate binary visibility labels from available object-level annotations, for keypoints of both asymmetric objects and symmetric objects. We further derive real-valued visibility-aware importance from binary labels based on the PageRank algorithm. Taking advantage of the flexibility of our visibility-aware importance, we construct VAPO (Visibility-Aware POse estimator) by integrating the visibility-aware importance with a state-of-the-art pose estimation algorithm, along with additional positional encoding. VAPO can work in both CAD-based and CAD-free settings. Extensive experiments are conducted on popular pose estimation benchmarks including Linemod, Linemod-Occlusion, and YCB-V, demonstrating that VAPO clearly achieves state-of-the-art performances. Project page: https://github.com/RuyiLian/VAPO. Ruyi Lian, Yuewei Lin, Longin Jan Latecki, Haibin Ling |
IROS | 2 |
| 2025 | DD-RobustBench: An Adversarial Robustness Benchmark for Dataset DistillationabstractDataset distillation techniques have revolutionized the way of utilizing large datasets by compressing them into smaller, yet highly effective subsets that preserve the original datasets' accuracy. However, while these methods have proven effective in reducing data size and training times, the robustness of these distilled datasets against adversarial attacks remains underexplored. This vulnerability poses significant risks, particularly in security-sensitive applications. To address this critical gap, we introduce DD-RobustBench, a novel and comprehensive benchmark specifically designed to evaluate the adversarial robustness of distilled datasets. Our benchmark is the most extensive of its kind and integrates a variety of dataset distillation techniques, including recent advancements such as TESLA, DREAM, SRe2L, and D4M, which have shown promise in enhancing model performance. DD-RobustBench also rigorously tests these datasets against a diverse array of adversarial attack methods to ensure broad applicability. Our evaluations cover a wide spectrum of datasets, including but not limited to, the widely used ImageNet-1K. This allows us to assess the robustness of distilled datasets in scenarios mirroring real-world applications. Furthermore, our detailed quantitative analysis investigates how different components involved in the distillation process, such as data augmentation, downsampling, and clustering, affect dataset robustness. Our findings provide critical insights into which techniques enhance or weaken the resilience of distilled datasets against adversarial threats, offering valuable guidelines for developing more robust distillation methods in the future. Through DD-RobustBench, we aim not only to benchmark but also to push the boundaries of dataset distillation research by highlighting areas for improvement and suggesting pathways for future innovations in creating datasets that are not only compact and efficient but also secure and resilient to adversarial challenges. The implementation details and essential instructions are available on DD-RobustBench. Yifan Wu 0037, Jiawei Du 0002, Ping Liu 0004, Yuewei Lin, Wei Xu 0038, Wenqing Cheng |
IEEE Trans. Image Process. | 4 |
| 2024 | AesFA: An Aesthetic Feature-Aware Arbitrary Neural Style TransferabstractNeural style transfer (NST) has evolved significantly in recent years. Yet, despite its rapid progress and advancement, existing NST methods either struggle to transfer aesthetic information from a style effectively or suffer from high computational costs and inefficiencies in feature disentanglement due to using pre-trained models. This work proposes a lightweight but effective model, AesFA---Aesthetic Feature-Aware NST. The primary idea is to decompose the image via its frequencies to better disentangle aesthetic styles from the reference image while training the entire model in an end-to-end manner to exclude pre-trained models at inference completely. To improve the network's ability to extract more distinct representations and further enhance the stylization quality, this work introduces a new aesthetic feature: contrastive loss. Extensive experiments and ablations show the approach not only outperforms recent NST methods in terms of stylization quality, but it also achieves faster inference. Codes are available at https://github.com/Sooyyoungg/AesFA. Joonwoo Kwon, Yuewei Lin, Shinjae Yoo, Jiook Cha |
AAAI | 3 |
| 2024 | EVD4UAV: An Altitude-Sensitive Benchmark to Evade Vehicle Detection in UAVabstractVehicle detection in Unmanned Aerial Vehicle (UAV) captured images has wide applications in aerial photography and remote sensing. There are many public benchmark datasets proposed for the vehicle detection and tracking in UAV images. Recent studies show that adding an adversarial patch on objects can fool the well-trained deep neural networks based object detectors, posing security concerns to the downstream tasks. However, the current public UAV datasets might ignore the diverse altitudes, vehicle attributes, fine-grained instance-level annotation in mostly side view with blurred vehicle roof, so none of them is good to study the adversarial patch based vehicle detection attack problem. In this paper, we propose a new dataset named EVD4UAV as an altitude-sensitive benchmark to evade vehicle detection in UAV with 6,284 images and 90,886 fine-grained annotated vehicles. The EVD4UAV dataset has diverse altitudes (50m, 70m, 90m), vehicle attributes (color, type), fine-grained annotation (horizontal and rotated bounding boxes, instance-level mask) in top view with clear vehicle roof. One white-box and two black-box patch based attack methods are implemented to attack three classic deep neural networks based object detectors on EVD4UAV. The experimental results show that these representative attack methods could not achieve the robust altitude-insensitive attack performance. Huiming Sun, Jiacheng Guo, Zibo Meng, Tianyun Zhang, Jianwu Fang, Yuewei Lin, Hongkai Yu |
IV | 6 |
| 2024 | Efficient Temporal Action Segmentation via Boundary-aware Query VotingabstractAlthough the performance of Temporal Action Segmentation (TAS) has been improved in recent years, achieving promising results often comes with a high computational cost due to dense inputs, complex model structures, and resource-intensive post-processing requirements. To improve the efficiency while keeping the high performance, we present a novel perspective centered on per-segment classification. By harnessing the capabilities of Transformers, we tokenize each video segment as an instance token, endowed with intrinsic instance segmentation. To realize efficient action segmentation, we introduce BaFormer, a boundary-aware Transformer network. It employs instance queries for instance segmentation and a global query for class-agnostic boundary prediction, yielding continuous segment proposals. During inference, BaFormer employs a simple yet effective voting strategy to classify boundary-wise segments based on instance segmentation. Remarkably, as a single-stage approach, BaFormer significantly reduces the computational costs, utilizing only 6% of the running time compared to the state-of-the-art method DiffAct, while producing better or comparable accuracy over several popular benchmarks. The code for this project is publicly available at https://github.com/peiyao-w/BaFormer. Yuewei Lin, Erik Blasch, Haibin Ling |
NeurIPS | 2 |
| 2024 | CLIPCEIL: Domain Generalization through CLIP via Channel rEfinement and Image-text aLignmentabstractDomain generalization (DG) is a fundamental yet challenging topic in machine learning. Recently, the remarkable zero-shot capabilities of the large pre-trained vision-language model (e.g., CLIP) have made it popular for various downstream tasks. However, the effectiveness of this capacity often degrades when there are shifts in data distribution during testing compared to the training data. In this paper, we propose a novel method, known as CLIPCEIL, a model that utilizes Channel rEfinement and Image-text aLignment to facilitate the CLIP to the inaccessible $\textit{out-of-distribution}$ test datasets that exhibit domain shifts. Specifically, we refine the feature channels in the visual domain to ensure they contain domain-invariant and class-relevant features by using a lightweight adapter. This is achieved by minimizing the inter-domain variance while maximizing the inter-class variance. In the meantime, we ensure the image-text alignment by aligning text embeddings of the class descriptions and their corresponding image embedding while further removing the domain-specific features. Moreover, our model integrates multi-scale CLIP features by utilizing a self-attention fusion module, technically implemented through one Transformer layer. Extensive experiments on five widely used benchmark datasets demonstrate that CLIPCEIL outperforms the existing state-of-the-art methods. The source code is available at \url{https://github.com/yuxi120407/CLIPCEIL}. Shinjae Yoo, Yuewei Lin |
NeurIPS | 3 |
| 2024 | Defense against Adversarial Cloud Attack on Remote Sensing Salient Object DetectionabstractDetecting the salient objects in a remote sensing image has wide applications. Many existing deep learning methods have been proposed for Salient Object Detection (SOD) in remote sensing images with remarkable results. However, the recent adversarial attack examples, generated by changing a few pixel values on the original image, could result in a collapse for the well-trained deep learning model. Different with existing methods adding perturbation to original images, we propose to jointly tune adversarial exposure and additive perturbation for attack and constrain image close to cloudy image as Adversarial Cloud. Cloud is natural and common in remote sensing images, however, camouflaging cloud based adversarial attack and defense for remote sensing images are not well studied before. Furthermore, we design DefenseNet as a learnable pre-processing to the adversarial cloudy images to preserve the performance of the deep learning based remote sensing SOD model, without tuning the already deployed deep SOD model. By considering both regular and generalized adversarial examples, the proposed DefenseNet can defend the proposed Adversarial Cloud in white-box setting and other attack methods in black-box setting. Experimental results on a synthesized benchmark from the public remote sensing dataset (EORSSD) show the promising defense against adversarial cloud attacks. Huiming Sun, Lan Fu, Qing Guo 0005, Zibo Meng, Tianyun Zhang, Yuewei Lin, Hongkai Yu |
WACV | 7 |
| 2024 | Exploring Robust Features for Improving Adversarial RobustnessabstractWhile deep neural networks (DNNs) have revolutionized many fields, their fragility to carefully designed adversarial attacks impedes the usage of DNNs in safety-critical applications. In this article, we strive to explore the robust features that are not affected by the adversarial perturbations, that is, invariant to the clean image and its adversarial examples (AEs), to improve the model's adversarial robustness. Specifically, we propose a feature disentanglement model to segregate the robust features from nonrobust features and domain-specific features. The extensive experiments on five widely used datasets with different attacks demonstrate that robust features obtained from our model improve the model's adversarial robustness compared to the state-of-the-art approaches. Moreover, the trained domain discriminator is able to identify the domain-specific features from the clean images and AEs almost perfectly. This enables AE detection without incurring additional computational costs. With that, we can also specify different classifiers for clean images and AEs, thereby avoiding any drop in clean image accuracy. Hong Wang 0024, Yuefan Deng, Shinjae Yoo, Yuewei Lin |
IEEE Trans. Cybern. | 4 |
| 2024 | INSURE: An Information Theory iNspired diSentanglement and pURification modEl for Domain GeneralizationabstractDomain Generalization (DG) aims to learn a generalizable model on the unseen target domain by only training on the multiple observed source domains. Although a variety of DG methods have focused on extracting domain-invariant features, the domain-specific class-relevant features have attracted attention and been argued to benefit generalization to the unseen target domain. To take into account the class-relevant domain-specific information, in this paper we propose an Information theory iNspired diSentanglement and pURification modEl (INSURE) to explicitly disentangle the latent features to obtain sufficient and compact (necessary) class-relevant feature for generalization to the unseen domain. Specifically, we first propose an information theory inspired loss function to ensure the disentangled class-relevant features contain sufficient class label information and the other disentangled auxiliary feature has sufficient domain information. We further propose a paired purification loss function to let the auxiliary feature discard all the class-relevant information and thus the class-relevant feature will contain sufficient and compact (necessary) class-relevant information. Moreover, instead of using multiple encoders, we propose to use a learnable binary mask as our disentangler to make the disentanglement more efficient and make the disentangled features complementary to each other. We conduct extensive experiments on five widely used DG benchmark datasets including PACS, VLCS, OfficeHome, TerraIncognita, and DomainNet. The proposed INSURE achieves state-of-the-art performance. We also empirically show that domain-specific class-relevant features are beneficial for domain generalization. The code is available at https://github.com/yuxi120407/INSURE. Huan-Hsin Tseng, Shinjae Yoo, Haibin Ling, Yuewei Lin |
IEEE Trans. Image Process. | 5 |
| 2023 | Transferable Adversarial Attack on 3D Object Tracking in Point Cloud
Xiaoqiong Liu, Yuewei Lin, Qing Yang 0003, Heng Fan 0001 |
MMM (2) | 2 |
| 2023 | Performance Analysis of Coordinated Interference Mitigation Approach for Automotive RadarabstractMillimeter automotive radar has great potential in advanced driver assistance systems (ADASs) to enable safety features, such as adaptive cruise control and collision avoidance. However, with widely deployment of millimeter radars on vehicles, the risk of radar mutual interference becomes a major factor limiting the high performance of radar detection. In this article, we analyze the mutual interference among multiple frequency modulated continuous wave (FMCW) radars. On the one hand, we study the interference in detail by considering co-channel interference (CCI) and adjacent channel interference (ACI) simultaneously. Besides, the CCI is analyzed by employing stochastic geometry model while the ACI is assessed by the deterministic analysis method. On the other hand, we propose a time-frequency division multiple access (TFDMA) scheme to mitigate the interference in a coordinated manner and evaluate it in terms of mitigation delay, the probability of interference, effective detectable density, maximum number of interference-free radar, and control signaling overhead. Finally, we study the power allocation strategy to enable the effectiveness of the coordinated interference mitigation approach based on the interference analysis. Simulation results verify the proposed framework for interference analysis by employing Monte Carlo method, and the performance improvement of the coordinated interference mitigation approach is 3.5 dB. Yi Wang 0011, Qixun Zhang, Zhiqing Wei, Yuewei Lin, Zhiyong Feng 0001 |
IEEE Internet Things J. | 4 |
| 2023 | Coarse-to-Fine Task-Driven Inpainting for Geoscience ImagesabstractThe processing and recognition of geoscience images have wide applications. Most of existing researches focus on understanding the high-quality geoscience images by assuming that all the images are clear. However, in many real-world cases, the geoscience images might contain occlusions during the image acquisition. This problem actually implies the image inpainting problem in computer vision and multimedia. As far as we know, all the existing image inpainting algorithms learn to repair the occluded regions for a better visualization quality, they are excellent for natural images but not good enough for geoscience images, and they never consider the following gescience task when developing inpainting methods. This paper aims to repair the occluded regions for a better geoscience task performance and advanced visualization quality simultaneously, without changing the current deployed deep learning based geoscience models. Because of the complex context of geoscience images, we propose a coarse-to-fine encoder-decoder network with the help of designed coarse-to-fine adversarial context discriminators to reconstruct the occluded image regions. Due to the limited data of geoscience images, we propose a MaskMix based data augmentation method, which augments inpainting masks instead of augmenting original images, to exploit the limited geoscience image data. The experimental results on three public geoscience datasets for remote sensing scene recognition, cross-view geolocation and semantic segmentation tasks respectively show the effectiveness and accuracy of the proposed method. The code is available at:https://github.com/HMS97/Task-driven-Inpainting. Huiming Sun, Jin Ma 0005, Qing Guo 0005, Qin Zou 0001, Shaoyue Song, Yuewei Lin, Hongkai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Design and Performance Analysis of 3-D Markov-Chain-Model-Based Fair Spectrum-Sharing Access for IoT ServicesabstractThe spectrum-sharing access technology using unlicensed spectrum bands is considered as a promising solution to solve the spectrum deficiency problem for various Internet of Things services. But, the efficient and fairness-guaranteed spectrum-sharing access among different systems is challenging by using the listen-before-talk (LBT) technology due to the uncertain channel quality and the channel access collision issues in unlicensed spectrum bands. To solve these problems, a novel three dimension-based Markov chain model is designed to formulate the collision probability of the spectrum-sharing access process using the contention window (CW) back-off algorithm based on the channel quality indicator feedback information. The key reasons for the packet transmission failure are comprehensively analyzed by considering both the channel collision and the channel quality deterioration conditions. A fairness-based spectrum-sharing access algorithm is proposed by assigning the optimal CW for the LBT system to minimize the collision probability and guarantee the fairness among LBT and wireless fidelity coexisted systems. Hardware platform-based evaluation results prove that our proposed algorithms can improve the system throughput and guarantee the fairness among different systems in the unlicensed spectrum band. Qixun Zhang, Zhu Han 0001, Yuewei Lin |
IEEE Internet Things J. | 4 |
| 2022 | Point Adversarial Self-Mining: A Simple Method for Facial Expression RecognitionabstractIn this article, we propose a simple yet effective approach, called point adversarial self mining (PASM), to improve the recognition accuracy in facial expression recognition (FER). Unlike previous works focusing on designing specific architectures or loss functions to solve this problem, PASM boosts the network capability by simulating human learning processes: providing updated learning materials and guidance from more capable teachers. Specifically, to generate new learning materials, PASM leverages a point adversarial attack method and a trained teacher network to locate the most informative position related to the target task, generating harder learning samples to refine the network. The searched position is highly adaptive since it considers both the statistical information of each sample and the teacher network capability. Other than being provided new learning materials, the student network also receives guidance from the teacher network. After the student network finishes training, the student network changes its role and acts as a teacher, generating new learning materials and providing stronger guidance to train a better student network. The adaptive learning materials generation and teacher/student update can be conducted more than one time, improving the network capability iteratively. Extensive experimental results validate the efficacy of our method over the existing state of the arts for FER. Ping Liu 0004, Yuewei Lin, Zibo Meng, Weihong Deng, Joey Tianyi Zhou, Yi Yang 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | Transparent Object Tracking BenchmarkabstractVisual tracking has achieved considerable progress in recent years. However, current research in the field mainly focuses on tracking of opaque objects, while little attention is paid to transparent object tracking. In this paper, we make the first attempt in exploring this problem by proposing a Transparent Object Tracking Benchmark (TOTB). Specifically, TOTB consists of 225 videos (86K frames) from 15 diverse transparent object categories. Each sequence is manually labeled with axis-aligned bounding boxes. To the best of our knowledge, TOTB is the first benchmark dedicated to transparent object tracking. In order to understand how existing trackers perform and to provide comparison for future research on TOTB, we extensively evaluate 25 state-of-the-art tracking algorithms. The evaluation results exhibit that more efforts are needed to improve transparent object tracking. Besides, we observe some nontrivial findings from the evaluation that are discrepant with some common beliefs in opaque object tracking. For example, we find that deeper features are not always good for improvements. Moreover, to encourage future research, we introduce a novel tracker, named TransATOM, which leverages transparency features for tracking and surpasses all 25 evaluated approaches by a large margin. By releasing TOTB, we expect to facilitate future research and application of transparent object tracking in both the academia and industry. The TOTB and evaluation results as well as TransATOM are available at https: //hengfan2010.github.io/projects/TOTB/. Heng Fan 0001, Halady Akhilesha Miththanthaya, Siranjiv Ramana Rajan, Xiaoqiong Liu, Zhilin Zou, Yuewei Lin, Haibin Ling |
ICCV | 7 |
| 2021 | AGKD-BML: Defense Against Adversarial Attack by Attention Guided Knowledge Distillation and Bi-directional Metric LearningabstractWhile deep neural networks have shown impressive performance in many tasks, they are fragile to carefully de-signed adversarial attacks. We propose a novel adversarial training-based model by Attention Guided Knowledge Distillation and Bi-directional Metric Learning (AGKD-BML). The attention knowledge is obtained from a weight-fixed model trained on a clean dataset, referred to as a teacher model, and transferred to a model that is under training on adversarial examples (AEs), referred to as a student model. In this way, the student model is able to focus on the correct region, as well as correcting the intermediate features corrupted by AEs to eventually improve the model accuracy. Moreover, to efficiently regularize the representation in feature space, we propose a bidirectional metric learning. Specifically, given a clean image, it is first attacked to its most confusing class to get the forward AE. A clean image in the most confusing class is then randomly picked and attacked back to the original class to get the backward AE. A triplet loss is then used to shorten the representation distance between original image and its AE, while enlarge that between the forward and backward AEs. We conduct extensive adversarial robustness experiments on two widely used datasets with different attacks. Our proposed AGKD-BML model consistently outperforms the state-of-the-art approaches. The code of AGKD-BML will be available at: https://github.com/hongw579/AGKD-BML. Hong Wang 0024, Yuefan Deng, Shinjae Yoo, Haibin Ling, Yuewei Lin |
ICCV | 5 |
| 2021 | TracKlinic: Diagnosis of Challenge Factors in Visual TrackingabstractGeneric visual object tracking is difficult due to many challenge factors (e.g., occlusion, blur, etc.). Each of these factors may cause serious problems for a tracker, and when they work together can make things even more complicated. Despite a great amount of efforts devoted to understanding the behavior of trackers, reliable and quantifiable ways for studying the per factor tracking behavior remain barely available. Addressing this issue, in this paper we contribute to the community a tracking diagnosis toolkit, TracKlinic, for diagnosis of challenge factors of tracking algorithms.TracKlinic consists of two novel components focusing on the data and analysis aspects, respectively. For the data component, we carefully prepare a set of 2,390 annotated videos, each involving one and only one major challenge factor. When analyzing an algorithm for a specific challenge factor, such one-factor-per-sequence rule greatly inhibits the disturbance from other factors and consequently leads to more faithful analysis. For the analysis component, given the tracking results on all sequences, it investigates the behavior of the tracker under each individual factor and generates the report automatically. With TracKlinic, a thorough study is conducted on ten state-of-the-art trackers on nine challenge factors (including two compound ones). The results suggest that, heavy shape variation and occlusion are the two most challenging factors faced by most trackers. Besides, out-of-view, though does not happen frequently, is often fatal. By sharing TracKlinic1, we expect to make it much easier for diagnosing tracking algorithms, and to thus facilitate developing better ones. Heng Fan 0001, Fan Yang 0035, Peng Chu, Yuewei Lin, Haibin Ling |
WACV | 4 |
| 2020 | Improved Deep Hashing With Soft Pairwise Similarity for Multi-Label Image RetrievalabstractHash coding has been widely used in the approximate nearest neighbor search for large-scale image retrieval. Recently, many deep hashing methods have been proposed and shown largely improved performance over traditional feature-learning methods. Most of these methods examine the pairwise similarity on the semantic-level labels, where the pairwise similarity is generally defined in a hard-assignment way. That is, the pairwise similarity is “1” if they share no less than one class label and “0” if they do not share any. However, such similarity definition cannot reflect the similarity ranking for pairwise images that hold multiple labels. In this paper, an improved deep hashing method is proposed to enhance the ability of multi-label image retrieval. We introduce a pairwise quantified similarity calculated on the normalized semantic labels. Based on this, we divide the pairwise similarity into two situations-“hard similarity” and “soft similarity,” where cross-entropy loss and mean square error loss are adapted respectively for more robust feature learning and hash coding. Experiments on four popular datasets demonstrate that the proposed method outperforms the competing methods and achieves the state-of-the-art performance in multi-label image retrieval. Zheng Zhang 0036, Qin Zou 0001, Yuewei Lin, Long Chen 0005, Song Wang 0002 |
IEEE Trans. Multim. | 3 |
| 2019 | Domain Adaptation for Convolutional Neural Networks-Based Remote Sensing Scene ClassificationabstractRemote sensing (RS) scene classification plays an important role in the field of earth observation. With the rapid development of the RS techniques, a large number of RS scene images are available. As manually labeling large-scale RS scene images is both labor and time consuming, when a new unlabeled data set is obtained, how to use the existing labeled data sets to classify the new unlabeled images is an important research direction. Different RS scene image data sets may be taken from different type of sensors, and the images may vary from imaging modalities, spatial resolutions, and image scales, so the distribution discrepancy exists among different image data sets. As a result, simply applying convolutional neural networks (CNN) trained on source domain cannot accurately classify the images on target domain. Domain adaptation (DA) can be helpful to solve this problem. In this letter, we design a subspace alignment (SA) and CNN-based framework to solve the DA problem in RS scene image classification. A new SA layer is proposed and added into CNN models for DA, which could align the source and target domains in some feature subspace. Fine-tuning the modified CNN model with the added SA layer makes the CNN model adapt to the aligned feature subspace and helps to relieve the domain distribution discrepancy. The experiments conducted on two public data sets show that adding the SA layer into CNN improves the scene classification on the target domain. Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Qiang Zhang 0030, Yuewei Lin, Song Wang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2017 | Learning View-Invariant Features for Person Identification in Temporally Synchronized Videos Taken by Wearable CamerasabstractIn this paper, we study the problem of Cross-View Person Identification (CVPI), which aims at identifying the same person from temporally synchronized videos taken by different wearable cameras. Our basic idea is to utilize the human motion consistency for CVPI, where human motion can be computed by optical flow. However, optical flow is view-variant - the same person's optical flow in different videos can be very different due to view angle change. In this paper, we attempt to utilize 3D human-skeleton sequences to learn a model that can extract view-invariant motion features from optical flows in different views. For this purpose, we use 3D Mocap database to build a synthetic optical flow dataset and train a Triplet Network (TN) consisting of three sub-networks: two for optical flow sequences from different views and one for the underlying 3D Mocap skeleton sequence. Finally, sub-networks for optical flows are used to extract view-invariant features for CVPI. Experimental results show that, using only the motion information, the proposed method can achieve comparable performance with the state-of-the-art methods. Further combination of the proposed method with an appearance-based method achieves new state-of-the-art performance. Xiaochuan Fan, Yuewei Lin, Hao Guo 0002, Hongkai Yu, Dazhou Guo, Song Wang 0002 |
ICCV | 3 |
| 2017 | Visual-Attention-Based Background Modeling for Detecting Infrequently Moving ObjectsabstractMotion is one of the most important cues to separate foreground objects from the background in a video. Using a stationary camera, it is usually assumed that the background is static, while the foreground objects are moving most of the time. However, in practice, the foreground objects may show infrequent motions, such as abandoned objects and sleeping persons. Meanwhile, the background may contain frequent local motions, such as waving trees and/or grass. Such complexities may prevent the existing background subtraction algorithms from correctly identifying the foreground objects. In this paper, we propose a new approach that can detect the foreground objects with frequent and/or infrequent motions. Specifically, we use a visual-attention mechanism to infer a complete background from a subset of frames and then propagate it to the other frames for accurate background subtraction. Furthermore, we develop a feature-matching-based local motion stabilization algorithm to identify frequent local motions in the background for reducing false positives in the detected foreground. The proposed approach is fully unsupervised, without using any supervised learning for object detection and tracking. Extensive experiments on a large number of videos have demonstrated that the proposed approach outperforms the state-of-the-art motion detection and background subtraction methods in comparison. Yuewei Lin, Yu Cao 0003, Youjie Zhou, Song Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Cross-Domain Recognition by Identifying Joint Subspaces of Source Domain and Target DomainabstractThis paper introduces a new method to solve the cross-domain recognition problem. Different from the traditional domain adaption methods which rely on a global domain shift for all classes between the source and target domains, the proposed method is more flexible to capture individual class variations across domains. By adopting a natural and widely used assumption that the data samples from the same class should lay on an intrinsic low-dimensional subspace, even if they come from different domains, the proposed method circumvents the limitation of the global domain shift, and solves the cross-domain recognition by finding the joint subspaces of the source and target domains. Specifically, given labeled samples in the source domain, we construct a subspace for each of the classes. Then we construct subspaces in the target domain, called anchor subspaces, by collecting unlabeled samples that are close to each other and are highly likely to belong to the same class. The corresponding class label is then assigned by minimizing a cost function which reflects the overlap and topological structure consistency between subspaces across the source and target domains, and within the anchor subspaces, respectively. We further combine the anchor subspaces to the corresponding source subspaces to construct the joint subspaces. Subsequently, one-versus-rest support vector machine classifiers are trained using the data samples belonging to the same joint subspaces and applied to unlabeled data in the target domain. We evaluate the proposed method on two widely used datasets: 1) object recognition dataset for computer vision tasks and 2) sentiment classification dataset for natural language processing tasks. Comparison results demonstrate that the proposed method outperforms the comparison methods on both datasets. Yuewei Lin, Jing Chen 0008, Yu Cao 0003, Youjie Zhou, Lingfeng Zhang 0001, Yuan Yan Tang, Song Wang 0002 |
IEEE Trans. Cybern. | 1 |
| 2016 | Groupwise Tracking of Crowded Similar-Appearance Targets from Low-Continuity Image SequencesabstractAutomatic tracking of large-scale crowded targets are of particular importance in many applications, such as crowded people/vehicle tracking in video surveillance, fiber tracking in materials science, and cell tracking in biomedical imaging. This problem becomes very challenging when the targets show similar appearance and the interslice/ inter-frame continuity is low due to sparse sampling, camera motion and target occlusion. The main challenge comes from the step of association which aims at matching the predictions and the observations of the multiple targets. In this paper we propose a new groupwise method to explore the target group information and employ the within-group correlations for association and tracking. In particular, the within-group association is modeled by a nonrigid 2D Thin-Plate transform and a sequence of group shrinking, group growing and group merging operations are then developed to refine the composition of each group. We apply the proposed method to track large-scale fibers from microscopy material images and compare its performance against several other multi-target tracking methods. We also apply the proposed method to track crowded people from videos with poor inter-frame continuity. Hongkai Yu, Youjie Zhou, Jeff P. Simmons, Craig Przybyla, Yuewei Lin, Xiaochuan Fan, Yang Mi, Song Wang 0002 |
CVPR | 5 |
| 2015 | Combining local appearance and holistic view: Dual-Source Deep Neural Networks for human pose estimationabstractWe propose a new learning-based method for estimating 2D human pose from a single image, using Dual-Source Deep Convolutional Neural Networks (DS-CNN). Recently, many methods have been developed to estimate human pose by using pose priors that are estimated from physiologically inspired graphical models or learned from a holistic perspective. In this paper, we propose to integrate both the local (body) part appearance and the holistic view of each local part for more accurate human pose estimation. Specifically, the proposed DS-CNN takes a set of image patches (category-independent object proposals for training and multi-scale sliding windows for testing) as the input and then learns the appearance of each local part by considering their holistic views in the full body. Using DS-CNN, we achieve both joint detection, which determines whether an image patch contains a body joint, and joint localization, which finds the exact location of the joint in the image patch. Finally, we develop an algorithm to combine these joint detection/localization results from all the image patches for estimating the human pose. The experimental results show the effectiveness of the proposed method by comparing to the state-of-the-art human-pose estimation methods based on pose priors that are estimated from physiologically inspired graphical models or learned from a holistic perspective. Xiaochuan Fan, Yuewei Lin, Song Wang 0002 |
CVPR | 3 |
| 2015 | Co-Interest Person Detection from Multiple Wearable Camera VideosabstractWearable cameras, such as Google Glass and Go Pro, enable video data collection over larger areas and from different views. In this paper, we tackle a new problem of locating the co-interest person (CIP), i.e., the one who draws attention from most camera wearers, from temporally synchronized videos taken by multiple wearable cameras. Our basic idea is to exploit the motion patterns of people and use them to correlate the persons across different videos, instead of performing appearance-based matching as in traditional video co-segmentation/localization. This way, we can identify CIP even if a group of people with similar appearance are present in the view. More specifically, we detect a set of persons on each frame as the candidates of the CIP and then build a Conditional Random Field (CRF) model to select the one with consistent motion patterns in different videos and high spacial-temporal consistency in each video. We collect three sets of wearable-camera videos for testing the proposed algorithm. All the involved people have similar appearances in the collected videos and the experiments demonstrate the effectiveness of the proposed algorithm. Yuewei Lin, Kareem Abdelfatah, Youjie Zhou, Xiaochuan Fan, Hongkai Yu, Hui Qian 0001, Song Wang 0002 |
ICCV | 1 |
| 2015 | Cross-domain recognition by identifying compact joint subspacesabstractThis paper introduces a new method to solve the cross-domain recognition problem. Different from the traditional domain adaption methods which rely on a global domain shift for all classes between source and target domain, the proposed method is more flexible to capture individual class variations across domains. We propose to solves the problem by finding the compact joint subspaces of source and target domain. We evaluate the proposed method on two widely used datasets and comparison results demonstrates that the proposed method outperforms the comparison methods. Yuewei Lin, Jing Chen 0008, Yu Cao 0003, Youjie Zhou, Lingfeng Zhang 0001, Song Wang 0002 |
ICIP | 1 |
| 2015 | NNMap: A method to construct a good embedding for nearest neighbor classification
Jing Chen 0008, Yuan Yan Tang, C. L. Philip Chen, Bin Fang 0001, Zhaowei Shang, Yuewei Lin |
Neurocomputing | 6 |
| 2014 | Dual Fuzzy Hypergraph Regularized Multi-label Learning for Protein Subcellular Location PredictionabstractWith the explosion of newly found proteins, it is necessary and urgent to develop automated computational methods for protein sub cellular location prediction. In particular, the problem of predictor construction for multi-location proteins is challenging. Considering the main limitations of the existing methods, we propose a hierarchical multi-label learning model FHML for both single-location proteins and multi-location proteins. In this model, feature space is firstly decomposed onto a set of nonnegative bases under the nonnegative data factorization framework. The nonnegative bases act as latent feature concepts and the corresponding coefficients on these bases are views as the new feature representation on the latent feature concepts. The similar decomposition is later performed in label space, and then the latent label concepts are extracted. Using these latent concepts as hyper edges, we construct dual fuzzy hyper graphs to exploit the intrinsic high-order relations embedded in both feature space and label space. Finally, the sub cellular location annotation information is propagated from the labeled proteins to the unlabeled proteins by performing dual fuzzy hyper graph Laplacian regularization. In this work, our proposed method is evaluated on eukaryotic protein benchmark dataset, and the experimental results have shown its effectiveness. Jing Gien, Yuan Yan Tang, C. L. Philip Chen, Yuewei Lin |
ICPR | 4 |
| 2013 | Recognize Human Activities from Partially Observed VideosabstractRecognizing human activities in partially observed videos is a challenging problem and has many practical applications. When the unobserved subsequence is at the end of the video, the problem is reduced to activity prediction from unfinished activity streaming, which has been studied by many researchers. However, in the general case, an unobserved subsequence may occur at any time by yielding a temporal gap in the video. In this paper, we propose a new method that can recognize human activities from partially observed videos in the general case. Specifically, we formulate the problem into a probabilistic framework: 1) dividing each activity into multiple ordered temporal segments, 2) using spatiotemporal features of the training video samples in each segment as bases and applying sparse coding (SC) to derive the activity likelihood of the test video sample at each segment, and 3) finally combining the likelihood at each segment to achieve a global posterior for the activities. We further extend the proposed method to include more bases that correspond to a mixture of segments with different temporal lengths (MSSC), which can better represent the activities with large intra-class variations. We evaluate the proposed methods (SC and MSSC) on various real videos. We also evaluate the proposed methods on two special cases: 1) activity prediction where the unobserved subsequence is at the end of the video, and 2) human activity recognition on fully observed videos. Experimental results show that the proposed methods outperform existing state-of-the-art comparison methods. Yu Cao 0003, Daniel Paul Barrett, Andrei Barbu, N. Siddharth 0001, Haonan Yu, Aaron Michaux, Yuewei Lin, Sven J. Dickinson, Jeffrey Mark Siskind, Song Wang 0002 |
CVPR | 7 |
| 2013 | Visual saliency detection with center shift
Weibin Yang, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang, Yuewei Lin |
Neurocomputing | 5 |
| 2013 | A Visual-Attention Model Using Earth Mover's Distance-Based Saliency Measurement and Nonlinear Feature CombinationabstractThis paper introduces a new computational visual-attention model for static and dynamic saliency maps. First, we use the Earth Mover's Distance (EMD) to measure the center-surround difference in the receptive field, instead of using the Difference-of-Gaussian filter that is widely used in many previous visual-attention models. Second, we propose to take two steps of biologically inspired nonlinear operations for combining different features: combining subsets of basic features into a set of super features using the Lm-norm and then combining the super features using the Winner-Take-All mechanism. Third, we extend the proposed model to construct dynamic saliency maps from videos by using EMD for computing the center-surround difference in the spatiotemporal receptive field. We evaluate the performance of the proposed model on both static image data and video data. Comparison results show that the proposed model outperforms several existing models under a unified evaluation setting. Yuewei Lin, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang, Song Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | A Computational Model for Saliency Maps by Using Local EntropyabstractThis paper presents a computational framework for saliency maps. It employs the Earth Mover's Distance based on weighted-Histogram (EMD-wH) to measure the center-surround difference, instead of the Difference-of-Gaussian (DoG) filter used by traditional models. In addition, the model employs not only the traditional features such as colors, intensity and orientation but also the local entropy which expresses the local complexity. The major advantage of combining the local entropy map is that it can detect the salient regions which are not complex regions. Also, it uses a general framework to integrate the feature dimensions instead of summing the features directly. This model considers both local and global salient information, in contrast to the existing models that consider only one or the other. Furthermore, the "large scale bias" and "central bias" hypotheses are used in this model to select the fixation locations in the saliency map of different scales. The performance of this model is assessed by comparing their saliency maps and human fixation density. The results from this model are finally compared to those from other bottom-up models for reference. Yuewei Lin, Bin Fang 0001, Yuan Yan Tang |
AAAI | 1 |
| 2008 | A Cell Based Dynamic Spectrum Management Scheme with Interference Mitigation for Cognitive NetworksabstractThe scarcity of spectrum resource in future wireless networks has inspired the demand of dynamic spectrum management (DSM). This paper investigates a cell based dynamic spectrum management (CBDSM) scheme to enhance the spectrum utilization and maximize the profit of operators for cognitive networks. In this scheme, the economic factor of the spectrum is taken into account in order to guarantee the rationality for spectrum trading. Especially, we focus on the interference mitigation method, which is considered as a fundamental issue for applying DSM to wireless systems. As a potential tool for promoting the distributed autonomous radio resource optimization algorithms, game theory is applied in the DSM scheme to investigate a win-win solution for spectrum trading between RATs. The simulation results reveal that the proposed CBDSM scheme improves the spectrum utilization and the profit of operators while effectively mitigating mutual interference between wireless networks. Vanbien Le, Yuewei Lin, Zhiyong Feng 0001, Ping Zhang 0003 |
VTC Spring | 2 |
| 2008 | An Auction Based Joint Radio Resource Management Scheme and Architecture in a Multi-Operator ScenarioabstractThis article proposes an auction mechanism based scheme for joint radio resource management (JRRM) in a reconfigurable system in a multi-operator scenario. Through the periodical auction and transaction for the radio resource between different radio access technologies (RATs) of multiple operators, the spare radio resource can be fully utilized to meet the demand of the RATs being short of radio resource. In order to rationally handle the profit assignment between multiple operators, a specific pricing strategy is put forward. In addition, a novel architecture is also presented to support this JRRM scheme. Simulation results reveal that the proposed scheme not only effectively reduces the total session blocking probability, but also greatly improves the total radio resource utilization ratio and the profits of operators. Xian Zeng, Zhiyong Feng 0001, Vanbien Le, Yuewei Lin |
VTC Spring | 5 |
| 2008 | Autonomic Joint Session Scheduling Strategies for Heterogeneous Wireless NetworksabstractIn order to optimize usage of radio resource for heterogeneous radio access technologies (RATs) and jointly designed from the user perspective, the joint session scheduling (JOSCH) mechanism has been introduced to split traffic over tightly coupled radio network. This paper presents distributed reinforcement learning (RL) as an autonomic approach for the JOSCH. Through the "trial-and-error" interaction with its radio environment, the JOSCH agent learns to split the traffic in a best way and allocate sub-streams in the proper RATs. A backpropagation neural network is adopted to generalize the large input state space of the RL algorithm to reduce memory requirement. Extensive simulations show that the proposed algorithm not only realizes the autonomy of JOSCH through the online learning process, but also improves the service quality at user side and the spectrum utility at operator side base on the suitable strategies. Yuewei Lin, Zhiyong Feng 0001, Huying Cai |
WCNC | 2 |