Yang Wang 0003

dblp:w/YangWang3 · DBLP profile ↗
← Back
126ranked-venue papers
14as first author
53since 2021 · last 2026
0000-0001-9447-1791ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 93 · 8 first-author · 36 since 2021Artificial intelligence and machine learning · 84 · 12 first-author · 35 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Agent-SAMA: State-Aware Mobile Assistant
abstract
Mobile Graphical User Interface (GUI) agents aim to autonomously complete tasks within or across apps based on user instructions. While recent Multimodal Large Language Models (MLLMs) enable these agents to interpret UI screens and perform actions, existing agents remain fundamentally reactive. They reason over the current UI screen but lack a structured representation of the app navigation flow, lim- iting GUI agents’ ability to understand execution context, detect unexpected execution results, and recover from errors. We introduce Agent-SAMA, a state-aware multi-agent framework that models app execution as a Finite State Machine (FSM), treating UI screens as states and user actions as transitions. Agent-SAMA implements four specialized agents that collaboratively construct and use FSMs in real time to guide task planning, execution verification, and recovery. We evaluate Agent-SAMA on two types of benchmarks: cross- app (Mobile-Eval-E, SPA-Bench) and mostly single-app (AndroidWorld). On Mobile-Eval-E, Agent-SAMA achieves an 84.0% success rate and a 71.9% recovery rate. On SPA-Bench, it reaches an 80.0% success rate with a 66.7% recovery rate. Compared to prior methods, Agent-SAMA improves task success by up to 12% and recovery success by 13.8%. On AndroidWorld, Agent-SAMA achieves a 63.7% success rate, outperforming the baselines. Our results demonstrate that structured state modeling enhances robustness and can serve as a lightweight, model-agnostic memory layer for future GUI agents.
Linqiang Guo, Wei Liu 0155, Yi Wen Heng, Tse-Hsun (Peter) Chen, Yang Wang 0003
AAAI5
2026 ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
abstract
Chain-of-thought (CoT) reasoning improves large language model performance on complex tasks, but often produces excessively long and inefficient reasoning traces.Existing methods shorten CoTs using length penalties or global entropy reduction, implicitly assuming that low uncertainty is desirable throughout reasoning.We show instead that reasoning efficiency is governed by the trajectory of uncertainty.CoTs with dominant downward entropy trends are substantially shorter.Motivated by this insight, we propose Entropy Trend Reward (ETR), a trajectory-aware objective that encourages progressive uncertainty reduction while allowing limited local exploration.We integrate ETR into Group Relative Policy Optimization (GRPO) and evaluate it across multiple reasoning models and challenging benchmarks.ETR consistently achieves a superior accuracy-efficiency trade-off, improving DeepSeek-R1-Distill-7B by +9.9% accuracy while reducing CoT length by 67% across four benchmarks.
Xuan Xiong, Huan Liu 0014, Li Gu, Zhixiang Chi, Yuanhao Yu, Yang Wang 0003
ACL (1)7
2026 Test-time Adaptation for 3D Human Pose and Shape Estimation
Zhixiang Chi, Sen Wang 0003, Yang Wang 0003, Xinxin Zuo
FG4
2026 An Empirical Study of Self-supervised Pretraining in X-Ray Security Screening
Niloofar Akbari, Yang Wang 0003, Xinxin Zuo
ICPR (8)2
2026 Controllable Diffusion-Based Data Augmentation for X-Ray Object Detection
Jacob Kingi, Yang Wang 0003, Xinxin Zuo
ICPR (8)2
2026 ProtoAug: Prototype-Guided Uncertainty-Aware Augmentation for Long-Tail Motion Prediction
Ziheng Lu, Yingfeng Cai, Hai Wang 0003, Long Chen 0003, Yang Wang 0003, Xinxin Zuo
IEEE Internet Things J.5
2025 MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt Tuning
abstract
Recent advancements in handwritten text recognition (HTR) have enabled the effective conversion of handwritten text to digital formats. However, achieving robust recognition across diverse writing styles remains challenging. Traditional HTR methods lack writer-specific personalization at test time due to limitations in model architecture and training strategies. Existing attempts to bridge this gap, through gradient-based meta-learning, still require labeled examples and suffer from parameter-inefficient fine-tuning, leading to substantial computational and memory overhead. To overcome these challenges, we propose an efficient framework that formulates personalization as prompt tuning, incorporating an auxiliary image reconstruction task with a self-supervised loss to guide prompt adaptation with unlabeled test-time examples. To ensure self-supervised loss effectively minimizes text recognition error, we leverage meta-learning to learn the optimal initialization of the prompts. As a result, our method allows the model to efficiently capture unique writing styles by updating less than 1% of its parameters and eliminating the need for time-intensive annotation processes. We validate our approach on the RIMES and IAM Handwriting Database benchmarks, where it consistently outperforms previous state-ofthe-art methods while using 20x fewer parameters. We believe this represents a significant advancement in personalized handwritten text recognition, paving the way for more reliable and practical deployment in resource-constrained scenarios.
Wenhao Gu, Li Gu, Chingyee Yee Suen, Yang Wang 0003
CVPR4
2025 Plug-in Feedback Self-Adaptive Attention in CLIP for Training-Free Open-Vocabulary Segmentation
abstract
CLIP exhibits strong visual-textual alignment but struggle with open-vocabulary segmentation due to poor localization. Prior methods enhance spatial coherence by modifying intermediate attention. But, this coherence isn't consistently propagated to the final output due to subsequent operations such as projections. Additionally, intermediate attention lacks direct interaction with text representations, such semantic discrepancy limits the full potential of CLIP. In this work, we propose a training-free, feedback-driven self-adaptive framework that adapts output-based patch-level correspondences back to the intermediate attention. The output predictions, being the culmination of the model's processing, encapsulate the most comprehensive visual and textual semantics about each patch. Our approach enhances semantic consistency between internal representations and final predictions by leveraging the model's outputs as a stronger spatial coherence prior. We design key modules, including attention isolation, confidence-based pruning for sparse adaptation, and adaptation ensemble, to effectively feedback the output coherence cues. Our method functions as a plug-in module, seamlessly integrating into four state-of-the-art approaches with three backbones (ViT-B, ViT-L, ViT-H). We further validate our framework across multiple attention types (Q-K, self-self, and Proxy augmented with MAE, SAM, and DINO). Our approach consistently improves their performance across eight benchmarks.
Zhixiang Chi, Li Gu, Huan Liu 0014, Ziqiang Wang 0003, Yang Wang 0003, Konstantinos N. Plataniotis
ICCV7
2025 Learning to Adapt Frozen CLIP for Few-Shot Test-Time Domain Adaptation
abstract
Few-shot Test-Time Domain Adaptation focuses on adapting a model at test time to a specific domain using only a few unlabeled examples, addressing domain shift. Prior methods leverage CLIP's strong out-of-distribution (OOD) abilities by generating domain-specific prompts to guide its generalized, frozen features. However, since downstream datasets are not explicitly seen by CLIP, solely depending on the feature space knowledge is constrained by CLIP's prior knowledge. Notably, when using a less robust backbone like ViT-B/16, performance significantly drops on challenging real-world benchmarks. Departing from the state-of-the-art of inheriting the intrinsic OOD capability of CLIP, this work introduces learning directly on the input space to complement the dataset-specific knowledge for frozen CLIP. Specifically, an independent side branch is attached in parallel with CLIP and enforced to learn exclusive knowledge via revert attention. To better capture the dataset-specific label semantics for downstream adaptation, we propose to enhance the inter-dispersion among text features via greedy text ensemble and refinement. The text and visual features are then progressively fused in a domain-aware manner by a generated domain prompt to adapt toward a specific domain. Extensive experiments show our method's superiority on 5 large-scale benchmarks (WILDS and DomainNet), notably improving over smaller networks like ViT-B/16 with gains of \textbf{+5.1} in F1 for iWildCam and \textbf{+3.1\%} in WC Acc for FMoW. \href{https://github.com/chi-chi-zx/L2C}{Our Code: L2C}
Zhixiang Chi, Li Gu, Huan Liu 0014, Ziqiang Wang 0003, Yang Wang 0003, Konstantinos N. Plataniotis
ICLR6
2025 Collaborative Cloud-edge Generalized Category Discovery
abstract
Generalized category discovery (GCD) aims to group unlabeled samples from known and unknown classes when only part of the labeled data in the known classes is given. It allows the model to adapt to dynamic environments by discovering novel categories. However, when we applied the GCD approach to the decentralized open world, we still encountered the following challenges: (1) none of labeled data easily obtained in the open world, (2) heterogeneous label spaces across different environments, (3)representation degradation caused by fine-tuning models with limited data in specific environments. To address the above challenges, we introduce a new and practical task, namely Cloud-edge GCD (CE-GCD). Different from semi-supervised GCD, CE-GCD assumes that we only have a base model trained on common public categories, and aims to perform personalized unsupervised novel category discovery in multiple environments with heterogeneous label spaces. Data from different environments or clients cannot be shared, only model parameters can be transferred. To tackle this problem, we propose a novel GCD framework based on energy-guided known class discrimination and multi-level contrastive learning. In each client, we first use the classifier of the base model to distinguish between known and unknown classes, and then perform unsupervised learning on the unknown classes. Each client transfers category information through prototypes to assist learning. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach.
Yingbing Liu, Fei Ma 0001, Xinxin Zuo, Fan Zhang 0007, Yang Wang 0003
ACM Multimedia6
2025 PointMAC: Meta-Learned Adaptation for Robust Test-Time Point Cloud Completion
abstract
Point cloud completion is essential for robust 3D perception in safety-critical applications such as robotics and augmented reality. However, existing models perform static inference and rely heavily on inductive biases learned during training, limiting their ability to adapt to novel structural patterns and sensor-induced distortions at test time. To address this limitation, we propose PointMAC, a meta-learned framework for robust test-time adaptation in point cloud completion. It enables sample-specific refinement without requiring additional supervision. Our method optimizes the completion model under two self-supervised auxiliary objectives that simulate structural and sensor-level incompleteness. A meta-auxiliary learning strategy based on Model-Agnostic Meta-Learning (MAML) ensures that adaptation driven by auxiliary objectives is consistently aligned with the primary completion task. During inference, we adapt the shared encoder on-the-fly by optimizing auxiliary losses, with the decoder kept fixed. To further stabilize adaptation, we introduce Adaptive $\lambda$-Calibration, a meta-learned mechanism for balancing gradients between primary and auxiliary objectives. Extensive experiments on synthetic, simulated, and real-world datasets demonstrate that PointMAC achieves state-of-the-art results by refining each sample individually to produce high-quality completions. To the best of our knowledge, this is the first work to apply meta-auxiliary test-time adaptation to point cloud completion.
Linlian Jiang, Li Gu, Ziqiang Wang 0003, Xinxin Zuo, Yang Wang 0003
NeurIPS6
2025 DocTTT: Test-Time Training for Handwritten Document Recognition Using Meta-Auxiliary Learning
abstract
Despite recent significant advancements in Handwritten Document Recognition (HDR), the efficient and accurate recognition of text against complex backgrounds, diverse handwriting styles, and varying document layouts remains a practical challenge. Moreover, this issue is seldom addressed in academic research, particularly in scenarios with minimal annotated data available. In this paper, we introduce the DocTTT framework to address these challenges. The key innovation of our approach is that it uses test-time training to adapt the model to each specific input during testing. We propose a novel Meta-Auxiliary learning approach that combines Meta-learning and self-supervised Masked Autoencoder (MAE). During testing, we adapt the visual representation parameters using a self-supervised MAE loss. During training, we learn the model parameters using a meta-learning framework, so that the model parameters are learned to adapt to a new input effectively. Experimental results show that our proposed method significantly outperforms existing state-of-the-art approaches on benchmark datasets.
Wenhao Gu, Li Gu, Ziqiang Wang 0003, Ching Y. Suen, Yang Wang 0003
WACV5
2024 Test-Time Personalization with Meta Prompt for Gaze Estimation
abstract
Despite the recent remarkable achievement in gaze estimation, efficient and accurate personalization of gaze estimation without labels is a practical problem but rarely touched on in the literature. To achieve efficient personalization, we take inspiration from the recent advances in Natural Language Processing (NLP) by updating a negligible number of parameters, "prompts", at the test time. Specifically, the prompt is additionally attached without perturbing original network and can contain less than 1% of a ResNet-18's parameters. Our experiments show high efficiency of the prompt tuning approach. The proposed one can be 10 times faster in terms of adaptation speed than the methods compared. However, it is non-trivial to update the prompt for personalized gaze estimation without labels. At the test time, it is essential to ensure that the minimizing of particular unsupervised loss leads to the goals of minimizing gaze estimation error. To address this difficulty, we propose to meta-learn the prompt to ensure that its updates align with the goal. Our experiments show that the meta-learned prompt can be effectively adapted even with a simple symmetry loss. In addition, we experiment on four cross-dataset validations to show the remarkable advantages of the proposed method.
Huan Liu 0014, Julia Qi, Mohammad Hassanpour, Yang Wang 0003, Konstantinos N. Plataniotis, Yuanhao Yu
AAAI5
2024 Test-Time Domain Adaptation by Learning Domain-Aware Batch Normalization
abstract
Test-time domain adaptation aims to adapt the model trained on source domains to unseen target domains using a few unlabeled images. Emerging research has shown that the label and domain information is separately embedded in the weight matrix and batch normalization (BN) layer. Previous works normally update the whole network naively without explicitly decoupling the knowledge between label and domain. As a result, it leads to knowledge interference and defective distribution adaptation. In this work, we propose to reduce such learning interference and elevate the domain knowledge learning by only manipulating the BN layer. However, the normalization step in BN is intrinsically unstable when the statistics are re-estimated from a few samples. We find that ambiguities can be greatly reduced when only updating the two affine parameters in BN while keeping the source domain statistics. To further enhance the domain knowledge extraction from unlabeled data, we construct an auxiliary branch with label-independent self-supervised learning (SSL) to provide supervision. Moreover, we propose a bi-level optimization based on meta-learning to enforce the alignment of two learning objectives of auxiliary and main branches. The goal is to use the auxiliary branch to adapt the domain and benefit main task for subsequent inference. Our method keeps the same computational cost at inference as the auxiliary branch can be thoroughly discarded after adaptation. Extensive experiments show that our method outperforms the prior works on five WILDS real-world domain shift datasets. Our method can also be integrated with methods with label-dependent optimization to further push the performance boundary. Our code is available at https://github.com/ynanwu/MABN.
Zhixiang Chi, Yang Wang 0003, Konstantinos N. Plataniotis, Songhe Feng
AAAI3
2024 Distribution Alignment for Fully Test-Time Adaptation with Dynamic Online Data Streams
Ziqiang Wang 0003, Zhixiang Chi, Li Gu, Zhi Liu 0003, Konstantinos N. Plataniotis, Yang Wang 0003
ECCV (24)7
2024 Adapting to Distribution Shift by Visual Domain Prompt Generation
abstract
In this paper, we aim to adapt a model at test-time using a few unlabeled data to address distribution shifts. To tackle the challenges of extracting domain knowledge from a limited amount of data, it is crucial to utilize correlated information from pre-trained backbones and source domains. Previous studies fail to utilize recent foundation models with strong out-of-distribution generalization. Additionally, domain-centric designs are not flavored in their works. Furthermore, they employ the process of modelling source domains and the process of learning to adapt independently into disjoint training stages. In this work, we propose an approach on top of the pre-computed features of the foundation model. Specifically, we build a knowledge bank to learn the transferable knowledge from source domains. Conditioned on few-shot target data, we introduce a domain prompt generator to condense the knowledge bank into a domain-specific prompt. The domain prompt then directs the visual features towards a particular domain via a guidance module. Moreover, we propose a domain-aware contrastive loss and employ meta-learning to facilitate domain knowledge extraction. Extensive experiments are conducted to validate the domain knowledge extraction. The proposed method outperforms previous work on 5 large-scale benchmarks including WILDS and DomainNet.
Zhixiang Chi, Li Gu, Tao Zhong 0003, Huan Liu 0014, Yuanhao Yu, Konstantinos N. Plataniotis, Yang Wang 0003
ICLR7
2024 ELF-UA: Efficient Label-Free User Adaptation in Gaze Estimation
Yong Wu 0007, Yang Wang 0003, Sanqing Qu, Zhijun Li 0001, Guang Chen 0001
IJCAI2
2024 Visually Guided Audio Source Separation with Meta Consistency Learning
abstract
In this paper, we tackle the problem of visually guided audio source separation in the context of both known and unknown objects (e.g., musical instruments). Recent successful end-to-end deep learning approaches adopt a single network with fixed parameters to generalize across unseen test videos. However, it can be challenging to generalize in cases where the distribution shift between training and test videos is higher as they fail to utilize internal information of unknown test videos. Based on this observation, we introduce a meta-consistency driven test time adaptation scheme that enables the pretrained model to quickly adapt to known and unknown test music videos in order to bring substantial improvements. In particular, we design a self-supervised audio-visual consistency objective as an auxiliary task that learns the synchronization between audio and its corresponding visual embedding. Concretely, we apply a meta-consistency training scheme to further optimize the pretrained model for effective and faster test time adaptation. We obtain substantial performance gains with only a smaller number of gradient updates and without any additional parameters for the task of audio source separation. Extensive experimental results across datasets demonstrate the effectiveness of our proposed method.
Md. Amirul Islam, Seyed Shahabeddin Nabavi, Irina Kezele, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
WACV4
2024 MetaUSACC: Unlabeled scene adaptation for crowd counting via meta-auxiliary learning
Penghui Shao, Anyong Qing, Yang Wang 0003
Expert Syst. Appl.5
2024 Crowd Counting Using Meta-Test-Time Adaptation
abstract
Machine learning algorithms are commonly used for quickly and efficiently counting people from a crowd. Test-time adaptation methods for crowd counting adjust model parameters and employ additional data augmentation to better adapt the model to the specific conditions encountered during testing. The majority of current studies concentrate on unsupervised domain adaptation. These approaches commonly perform hundreds of epochs of training iterations, requiring a sizable number of unannotated data of every new target domain apart from annotated data of the source domain. Unlike these methods, we propose a meta-test-time adaptive crowd counting approach called CrowdTTA, which integrates the concept of test-time adaptation into the meta-learning framework and makes it easier for the counting model to adapt to the unknown test distributions. To facilitate the reliable supervision signal at the pixel level, we introduce uncertainty by inserting the dropout layer into the counting model. The uncertainty is then used to generate valuable pseudo labels, serving as effective supervisory signals for adapting the model. In the context of meta-learning, one image can be regarded as one task for crowd counting. In each iteration, our approach is a dual-level optimization process. In the inner update, we employ a self-supervised consistency loss function to optimize the model so as to simulate the parameters update process that occurs during the test phase. In the outer update, we authentically update the parameters based on the image with ground truth, improving the model's performance and making the pseudo labels more accurate in the next iteration. At test time, the input image is used for adapting the model before testing the image. In comparison to various supervised learning and domain adaptation methods, our results via extensive experiments on diverse datasets showcase the general adaptive capability of our approach across datasets with varying crowd densities and scales.
Ferrante Neri, Li Gu, Ziqiang Wang 0003, Jian Wang 0110, Anyong Qing, Yang Wang 0003
Int. J. Neural Syst.7
2024 TTAGaze: Self-Supervised Test-Time Adaptation for Personalized Gaze Estimation
abstract
In this paper, we address the problem of personalized gaze estimation. Due to the anatomical differences between individuals, current personalized gaze models often rely on fine-tuning or fully-supervised methods with labeled calibration samples, which may not be practical in real-world applications. To tackle this limitation, we propose an approach called Self-Supervised Test-Time Adaptation for Personalized Gaze Estimation (TTAGaze), which enables adaptation with small unlabeled data at test time. Our goal is to develop a gaze estimation model specifically adapted to a target person using only a few unlabeled images. We call this setting as unsupervised few-shot personalized adaptation in gaze estimation, which is more aligned with real-world scenarios compared to existing approaches. Additionally, Our approach leverages self-supervised learning and meta-learning. The model consists of the main task (gaze estimation) and a self-supervised auxiliary task. During training, the two task are trained using a coupled method. At test time, adaptation is achieved by optimizing the self-supervised loss adapted to an unseen person with a few unlabeled data. The model parameters are learned via model-agnostic meta-learning (MAML) to facilitate effective unsupervised few-shot personalized adaptation in gaze estimation. Experimental results demonstrate that the proposed method outperforms alternative approaches on several widely-used benchmark datasets.
Yong Wu 0007, Guang Chen 0001, Linwei Ye, Yuanning Jia, Zhi Liu 0003, Yang Wang 0003
IEEE Trans. Circuits Syst. Video Technol.6
2023 MetaZSCIL: A Meta-Learning Approach for Generalized Zero-Shot Class Incremental Learning
abstract
Generalized zero-shot learning (GZSL) aims to recognize samples whose categories may not have been seen at training. Standard GZSL cannot handle dynamic addition of new seen and unseen classes. In order to address this limitation, some recent attempts have been made to develop continual GZSL methods. However, these methods require end-users to continuously collect and annotate numerous seen class samples, which is unrealistic and hampers the applicability in the real-world. Accordingly, in this paper, we propose a more practical and challenging setting named Generalized Zero-Shot Class Incremental Learning (CI-GZSL). Our setting aims to incrementally learn unseen classes without any training samples, while recognizing all classes previously encountered. We further propose a bi-level meta-learning based method called MetaZSCIL to directly optimize the network to learn how to incrementally learn. Specifically, we sample sequential tasks from seen classes during the offline training to simulate the incremental learning process. For each task, the model is learned using a meta-objective such that it is capable to perform fast adaptation without forgetting. Note that our optimization can be flexibly equipped with most existing generative methods to tackle CI-GZSL. This work introduces a feature generative framework that leverages visual feature distribution alignment to produce replayed samples of previously seen classes to reduce catastrophic forgetting. Extensive experiments conducted on five widely used benchmarks demonstrate the superiority of our proposed method.
Tengfei Liang, Songhe Feng, Yi Jin 0001, Gengyu Lyu, Haojun Fei, Yang Wang 0003
AAAI7
2023 Privacy-Preserving Learning via Data and Knowledge Distillation
abstract
In the current era of data science, deep learning, computer vision and image analysis have become ubiquitous across various sectors, ranging from government agencies and large corporations to small end devices, due to their ability to simplify people’s lives. However, the widespread use of sensitive image data and the high memorization capacity of deep learning present significant privacy risks. Now, a simple Google search can yield numerous images of a person, and the knowledge that a specific patient’s record was utilized for training a specific model associated with a disease may reveal the patient’s ailment, potentially leading to membership privacy leakage and other advanced attacks in the future. Furthermore, these unprotected models may also suffer from poor generalization due to this overfitting to train data. Previous state-of-the-art methods like differential privacy (DP) and regularizer-based defenses compromised functionality, i.e., task accuracy, to preserve privacy. Such an imbalanced trade-off raises concerns about the practicability of such defenses. Other existing knowledge-transfer-based methods either reuse private data or require more public data, which could compromise privacy and may not be viable in certain domains. To address these challenges, where membership privacy is of utmost importance and utility cannot be compromised, we propose a novel collaborative distillation approach that transfers the private model’s knowledge based on a minimal amount of distilled synthetic data, leading to a compact private model in an end-to-end fashion. Empirically, our proposed method guarantees superior performance compared to most advanced models currently in use, increasing utility by almost 8%, 34%, and 6% for CIFAR-10, CIFAR-100, and MNIST, respectively. The utility resembles non-private counterparts almost closely while maintaining a respectable level of membership privacy leakage of 50-53.5%, despite employing a smaller model with 50% fewer parameters.
Fahim Faisal, Carson K. Leung, Noman Mohammed, Yang Wang 0003
DSAA4
2023 Point-TTA: Test-Time Adaptation for Point Cloud Registration Using Multitask Meta-Auxiliary Learning
abstract
We present Point-TTA, a novel test-time adaptation framework for point cloud registration (PCR) that improves the generalization and the performance of registration models. While learning-based approaches have achieved impressive progress, generalization to unknown testing environments remains a major challenge due to the variations in 3D scans. Existing methods typically train a generic model and the same trained model is applied on each instance during testing. This could be sub-optimal since it is difficult for the same model to handle all the variations during testing. In this paper, we propose a test-time adaptation approach for PCR. Our model can adapt to unseen distributions at test-time without requiring any prior knowledge of the test data. Concretely, we design three self-supervised auxiliary tasks that are optimized jointly with the primary PCR task. Given a test instance, we adapt our model using these auxiliary tasks and the updated model is used to perform the inference. During training, our model is trained using a meta-auxiliary learning approach, such that the adapted model via auxiliary tasks improves the accuracy of the primary task. Experimental results demonstrate the effectiveness of our approach in improving generalization of point cloud registration and outperforming other state-of-the-art approaches.
Ahmed Hatem, Yiming Qian, Yang Wang 0003
ICCV3
2023 MetaGCD: Learning to Continually Learn in Generalized Category Discovery
abstract
In this paper, we consider a real-world scenario where a model that is trained on pre-defined classes continually encounters unlabeled data that contains both known and novel classes. The goal is to continually discover novel classes while maintaining the performance in known classes. We name the setting Continual Generalized Category Discovery (C-GCD). Existing methods for novel class discovery cannot directly handle the C-GCD setting due to some unrealistic assumptions, such as the unlabeled data only containing novel classes. Furthermore, they fail to discover novel classes in a continual fashion. In this work, we lift all these assumptions and propose an approach, called MetaGCD, to learn how to incrementally discover with less forgetting. Our proposed method uses a meta-learning framework and leverages the offline labeled data to simulate the testing incremental learning process. A meta-objective is defined to revolve around two conflicting learning objectives to achieve novel class discovery without forgetting. Furthermore, a soft neighborhood-based contrastive network is proposed to discriminate uncorrelated images while attracting correlated images. We build strong baselines and conduct extensive experiments on three widely used benchmarks to demonstrate the superiority of our method.
Zhixiang Chi, Yang Wang 0003, Songhe Feng
ICCV3
2023 Test-Time Adaptation for Point Cloud Upsampling Using Meta-Learning
abstract
Affordable 3D scanners often produce sparse and non-uniform point clouds that negatively impact downstream applications in robotic systems. While existing point cloud upsampling architectures have demonstrated promising results on standard benchmarks, they tend to experience significant performance drops when the test data have different distributions from the training data. To address this issue, this paper proposes a test-time adaption approach to enhance model generality of point cloud upsampling. The proposed approach leverages meta-learning to explicitly learn network parameters for test-time adaption. Our method does not require any prior information about the test data. During meta-training, the model parameters are learned from a collection of instance-level tasks, each of which consists of a sparse-dense pair of point clouds from the training data. During meta-testing, the trained model is fine-tuned with a few gradient updates to produce a unique set of network parameters for each test instance. The updated model is then used for the final prediction. Our framework is generic and can be applied in a plug-and-play manner with existing backbone networks in point cloud upsampling. Extensive experiments demonstrate that our approach improves the performance of state-of-the-art models.
Ahmed Hatem, Yiming Qian, Yang Wang 0003
IROS3
2023 Meta-Auxiliary Learning for Future Depth Prediction in Videos
abstract
We consider a new problem of future depth prediction in videos. Given a sequence of observed frames in a video, the goal is to predict the depth map of a future frame that has not been observed yet. Depth estimation plays a vital role for scene understanding and decision-making in intelligent systems. Predicting future depth maps can be valuable for autonomous vehicles to anticipate the behaviours of their surrounding objects. Our proposed model for this problem has a two-branch architecture. One branch is for the primary task of future depth prediction. The other branch is for an auxiliary task of image reconstruction. The auxiliary branch can act as a regularization. Inspired by some recent work on test-time adaption, we use the auxiliary task during testing to adapt the model to a specific test video. We also propose a novel meta-auxiliary learning that learns the model specifically for the purpose of effective test-time adaptation. Experimental results demonstrate that our proposed approach outperforms other alternative methods.
Huan Liu 0014, Zhixiang Chi, Yuanhao Yu, Yang Wang 0003, Jun Chen 0005, Jin Tang 0005
WACV4
2023 Few-Shot Learning of Compact Models via Task-Specific Meta Distillation
abstract
We consider a new problem of few-shot learning of com-pact models. Meta-learning is a popular approach for few-shot learning. Previous work in meta-learning typically assumes that the model architecture during meta-training is the same as the model architecture used for final deployment. In this paper, we challenge this basic assumption. For final deployment, we often need the model to be small. But small models usually do not have enough capacity to effectively adapt to new tasks. In the mean time, we often have access to the large dataset and extensive computing power during meta-training since meta-training is typically per-formed on a server. In this paper, we propose task-specific meta distillation that simultaneously learns two models in meta-learning: a large teacher model and a small student model. These two models are jointly learned during meta-training. Given a new task during meta-testing, the teacher model is first adapted to this task, then the adapted teacher model is used to guide the adaptation of the student model. The adapted student model is used for final deployment. We demonstrate the effectiveness of our approach in few-shot image classification using model-agnostic meta-learning (MAML). Our proposed method outperforms other alternatives on several benchmark datasets.
Yong Wu 0007, Shekhor Chanda, Mehrdad Hosseinzadeh, Zhi Liu 0003, Yang Wang 0003
WACV5
2023 Semantic-Aware Graph Matching Mechanism for Multi-Label Image Recognition
abstract
Multi-label image recognition aims to predict a set of labels that present in an image. The key to deal with such problem is to mine the associations between image contents and labels, and further obtain the correct assignments between images and their labels. In this paper, we treat each image as a bag of instances, and formulate the task of multi-label image recognition as an instance-label matching selection problem. To model such problem, we propose an innovative Semantic-aware Graph Matching framework for Multi-Label image recognition (ML-SGM), in which Graph Matching mechanism is introduced owing to its good performance of excavating the instance and label relationship. The framework explicitly establishes category correlations and instance-label correspondences by modeling the relation among content-aware (instance) and semantic-aware (label) category representations, to facilitate multi-label image understanding and reduce the dependency of large amounts of training samples for each category. Specifically, we first construct an instance spatial graph and a label semantic graph respectively and then incorporate them into a constructed assignment graph by connecting each instance to all labels. Subsequently, the graph network block is adopted to aggregate and update all nodes and edges state on the assignment graph to form structured representations for each instance and label. Our network finally derives a prediction score for each instance-label correspondence and optimizes such correspondence with a weighted cross-entropy loss. Empirical results conducted on generic multi-label image recognition demonstrate the superiority of our proposed method. Moreover, the proposed method also shows advantages in multi-label recognition with partial labels and multi-label few-shot learning, as well as outperforms current state-of-the-art methods with a clear margin. Our code is available athttps://github.com/yananwu0510/ML-SGM.
Songhe Feng, Yang Wang 0003
IEEE Trans. Circuits Syst. Video Technol.3
2023 Test-Time Adaptation for Optical Flow Estimation Using Motion Vectors
abstract
Due to the prohibitive cost as well as technical challenges in annotating ground-truth optical flow for large-scale realistic video datasets, the existing deep learning models for optical flow estimation mostly rely on synthetic data for training, which in turn may lead to significant performance degradation under test-data distribution shift in real-world environments. In this work, we propose the methodology to tackle this important problem. We design a self-supervised learning task for adjusting the optical flow estimation model at test time. We exploit the fact that most videos are stored in compressed formats, from which compact information on motion, in the form of motion vectors and residuals, can be made readily available. We formulate the self-supervised task as motion vector prediction, and link this task to optical flow estimation. To the best of our knowledge, our Test-Time Adaption guided with Motion Vectors (TTA-MV), is the first work to perform such adaptation for optical flow. The experimental results demonstrate that TTA-MV can improve the generalization capability of various well-known deep learning methods for optical flow estimation, such as FlowNet, PWCNet, and RAFT.
Seyed Mehdi Ayyoubzadeh, Irina Kezele, Yuanhao Yu, Xiaolin Wu 0001, Yang Wang 0003, Jin Tang 0005
IEEE Trans. Image Process.6
2023 Stress Detection Through Wrist-Based Electrodermal Activity Monitoring and Machine Learning
abstract
Stress is an inevitable part of modern life. While stress can negatively impact a person's life and health, positive and under-controlled stress can also enable people to generate creative solutions to problems encountered in their daily lives. Although it is hard to eliminate stress, we can learn to monitor and control its physical and psychological effects. It is essential to provide feasible and immediate solutions for more mental health counselling and support programs to help people relieve stress and improve their mental health. Popular wearable devices, such as smartwatches with several sensing capabilities, including physiological signal monitoring, can alleviate the problem. This work investigates the feasibility of using wrist-based electrodermal activity (EDA) signals collected from wearable devices to predict people's stress status and identify possible factors impacting stress classification accuracy. We use data collected from wrist-worn devices to examine the binary classification discriminating stress from non-stress. For efficient classification, five machine learning-based classifiers were examined. We explore the classification performance on four available EDA databases under different feature selections. According to the results, Support Vector Machine (SVM) outperforms the other machine learning approaches with an accuracy of 92.9 for stress prediction. Additionally, when the subject classification included gender information, the performance analysis showed significant differences between males and females. We further examine a multimodal approach for stress classifications. The results indicate that wearable devices with EDA sensors have a great potential to provide helpful insight for improved mental health monitoring.
Lili Zhu, Petros Spachos, Pai Chet Ng, Yuanhao Yu, Yang Wang 0003, Konstantinos N. Plataniotis, Dimitrios Hatzinakos
IEEE J. Biomed. Health Informatics5
2023 Adaptive Group-Wise Consistency Network for Co-Saliency Detection
abstract
Co-saliency detection focuses on detecting common and salient objects among a group of images. With the application of deep learning in co-saliency detection, more accurate and more effective models are proposed in an end-to-end manner. However, two major drawbacks in these models hinder the further performance improvement of co-saliency detection: 1) the static manner-based inference, and 2) the constant quantity of input images. To address these limitations, we present a novel Adaptive Group-wise Consistency Network (AGCNet) with the ability of content-adaptive adjustment for a given image group with random quantity of images. In AGCNet, we first introduce intra-saliency priors generated from any off-the-shelf salient object detection model. Then, an Adaptive Group-wise Consistency (AGC) module is proposed to capture group consistency for each individual image, and is applied on three-scale features to capture the group consistency from different perspectives. This module is composed of two key components, where the content-adaptive group consistency block breaks the above limitations to adaptively capture the global group consistency with the assistance of intra-saliency priors and the ranking-based fusion block combines the consistency with individual attributes of each image feature to generate discriminative group consistency feature for each image. Following AGC modules, a specially designed Aggregated Decoder aggregates the three-scale group consistency features to adapt to co-salient objects with diverse scales for preliminary detection. Finally, we incorporate two normal decoders to progressively refine the preliminary detection and generate the final co-saliency maps. Extensive experiments on four benchmark datasets demonstrate that our AGCNet achieves competitive performance as compared with 19 state-of-the-art models, and the proposed modules experimentally show substantial practical merits.
Zhen Bai 0001, Zhi Liu 0003, Gongyang Li, Yang Wang 0003
IEEE Trans. Multim.4
2023 Fast Human Pose Estimation in Compressed Videos
abstract
Current approaches for human pose estimation in videos can be categorized into per-frame and warping-based methods. Both approaches have their pros and cons. For example, per-frame methods are generally more accurate, but they are often slow. Warping-based approaches are more efficient, but the performance is usually not good. To bridge the gap, in this paper, we propose a novel fast framework for human pose estimation to meet the real-time inference with controllable accuracy degradation in compressed video domain. Our approach takes advantage of the motion representation (called “motion vector”) that is readily available in a compressed video. Pose joints in a frame are obtained by directly warping the pose joints from the previous frame using the motion vectors. We also propose modules to correct possible errors introduced by the pose warping when needed. Extensive experimental results demonstrate the effectiveness of our proposed framework for accelerating the speed of top-down human pose estimation in videos.
Huan Liu 0014, Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jun Chen 0005, Jin Tang 0005
IEEE Trans. Multim.4
2023 Spatio-Temporal Self-Attention Network for Video Saliency Prediction
abstract
3D convolutional neural networks have achieved promising results for video tasks in computer vision, including video saliency prediction that is explored in this paper. However, 3D convolution encodes visual representation merely on fixed local spacetime according to its kernel size, while human attention is always attracted by relational visual features at different time. To overcome this limitation, we propose a novel Spatio-Temporal Self-Attention 3D Network (STSANet) for video saliency prediction, in which multiple Spatio-Temporal Self-Attention (STSA) modules are employed at different levels of 3D convolutional backbone to directly capture long-range relations between spatio-temporal features of different time steps. Besides, we propose an Attentional Multi-Scale Fusion (AMSF) module to integrate multi-level features with the perception of context in semantic and spatio-temporal subspaces. Extensive experiments demonstrate the contributions of key components of our method, and the results on DHF1K, Hollywood-2, UCF, and DIEM benchmark datasets clearly prove the superiority of the proposed model compared with all state-of-the-art models.
Ziqiang Wang 0003, Zhi Liu 0003, Gongyang Li, Yang Wang 0003, Tianhong Zhang, Lihua Xu, Jijun Wang 0003
IEEE Trans. Multim.4
2022 Self-Supervised Spatiotemporal Representation Learning by Exploiting Video Continuity
abstract
Recent self-supervised video representation learning methods have found significant success by exploring essential properties of videos, e.g. speed, temporal order, etc. This work exploits an essential yet under-explored property of videos, the \textit{video continuity}, to obtain supervision signals for self-supervised representation learning. Specifically, we formulate three novel continuity-related pretext tasks, i.e. continuity justification, discontinuity localization, and missing section approximation, that jointly supervise a shared backbone for video representation learning. This self-supervision approach, termed as Continuity Perception Network (CPNet), solves the three tasks altogether and encourages the backbone network to learn local and long-ranged motion and context representations. It outperforms prior arts on multiple downstream tasks, such as action recognition, video retrieval, and action localization. Additionally, the video continuity can be complementary to other coarse-grained video properties for representation learning, and integrating the proposed pretext task to prior arts can yield much performance gains.
Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen, Peng Dai 0002, Juwei Lu, Yang Wang 0003
AAAI7
2022 Contrastive Learning for Unsupervised Video Highlight Detection
abstract
Video highlight detection can greatly simplify video browsing, potentially paving the way for a wide range of ap-plications. Existing efforts are mostly fully-supervised, requiring humans to manually identify and label the interesting moments (called highlights) in a video. Recent weakly supervised methods forgo the use of highlight annotations, but typically require extensive efforts in collecting external data such as web-crawled videos for model learning. This observation has inspired us to consider unsupervised highlight detection where neither frame-level nor video-level annotations are available in training. We propose a simple contrastive learning framework for unsupervised highlight detection. Our framework encodes a video into a vector representation by learning to pick video clips that help to distinguish it from other videos via a contrastive objective using dropout noise. This inherently allows our framework to identify video clips corresponding to highlight of the video. Extensive empirical evaluations on three highlight detection benchmarks demonstrate the superior performance of our approach.
Taivanbat Badamdorj, Mrigank Rochan, Yang Wang 0003
CVPR3
2022 MetaFSCIL: A Meta-Learning Approach for Few-Shot Class Incremental Learning
abstract
In this paper, we tackle the problem of few-shot class incremental learning (FSCIL). FSCIL aims to incrementally learn new classes with only a few samples in each class. Most existing methods only consider the incremental steps at test time. The learning objective of these methods is often hand-engineered and is not directly tied to the objective (i.e. incrementally learning new classes) during testing. Those methods are sub-optimal due to the misalignment between the training objectives and what the methods are expected to do during evaluation. In this work, we proposed a bi-level optimization based on meta-learning to directly optimize the network to learn how to incrementally learn in the setting of FSCIL. Concretely, we propose to sample sequences of incremental tasks from base classes for training to simulate the evaluation protocol. For each task, the model is learned using a meta-objective such that it is capable to perform fast adaptation without forgetting. Furthermore, we propose a bi-directional guided modulation, which is learned to automatically modulate the activations to reduce catastrophic forgetting. Extensive experimental results demonstrate that the proposed method outperforms the baseline and achieves the state-of-the-art results on CIFARIOO, MiniImageNet, and CUB200 datasets.
Zhixiang Chi, Li Gu, Huan Liu 0014, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
CVPR4
2022 Generating Privacy Preserving Synthetic Medical Data
abstract
Due to the recent development in the deep learning community and the availability of state-of-the-art models, medical practitioners are getting more interested in computer vision and deep learning for diagnosis tasks. Moreover, those medical diagnostic models can also increase the reliability of conventional findings. As radiology images can convey a lot of information for a patient’s diagnosis task, the problem is that such medical data may contain sensitive private information in their content header. De-anonymization (i.e., removal of sensitive header information) does not work well due to the re-identification risk, which may link those images to essential details (e.g., birth date, SSN, institution name, etc.), and such an approach can also reduce utility. In the medical domain, utility is significant because a less accurate diagnosis may lead to the wrong course of treatment and/or loss of life. In this paper, we developed a differentially private approach that can generate high-quality and high dimensional synthetic medical image data with guaranteed differential privacy. It can be used to create sufficient quality data to train a deep model. Moreover, we used W-GAN for bounded gradient guarantee, which eliminates the need for an extensive clipping hyperparameter search. We also added noise selectively to the generator to maintain the privacy-utility trade-off. Due to a noise-free discriminator and such selective noise addition to the generator, high-quality and reliable generated radiology images can be utilized for diagnosis tasks. Moreover, our approach can work in a distributed system where different hospitals can contain their private images in the local server and use a central server to generate synthetic radiology images without storing patient data.
Fahim Faisal, Noman Mohammed, Carson K. Leung, Yang Wang 0003
DSAA4
2022 Few-Shot Class-Incremental Learning via Entropy-Regularized Data-Free Replay
Huan Liu 0014, Li Gu, Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jun Chen 0005, Jin Tang 0005
ECCV (24)4
2022 Hierarchical Deep Learning Model with Inertial and Physiological Sensors Fusion for Wearable-Based Human Activity Recognition
abstract
This paper presents a human activity recognition (HAR) system with wearable devices. While various approaches have been suggested for HAR, most of them focus on either 1) the inertial sensors to capture the physical movement or 2) subject-dependent evaluations that are less practical to real world cases. To this end, our work integrates sensing in-puts from physiological sensors to compensate the limitation of inertial sensors in capturing the human activities with less physical movements. Physiological sensors can capture physiological responses reflecting human behaviors in executing daily activities. To simulate a realistic application, three different evaluation scenarios are considered, namely All-access, Cross-subject and Cross-activity. Lastly, we propose a Hierarchical Deep Learning (HDL) model, which improves the accuracy and stability of HAR, compared to conventional models. Our proposed HDL with fusion of inertial and physiological sensing inputs achieves 97.16%, 92.23%, 90.18% average accuracy in All-access, Cross-subject, Cross-activity scenarios, which confirms the effectiveness of our approach.
Dae Yon Hwang, Pai Chet Ng, Yuanhao Yu, Yang Wang 0003, Petros Spachos, Dimitrios Hatzinakos, Konstantinos N. Plataniotis
ICASSP4
2022 Feasibility Study of Stress Detection with Machine Learning through EDA from Wearable Devices
abstract
The recent pandemic has brought tremendous changes to everyone’s life, causing stress about losing loved ones, losing jobs, and having changes in sleep or eating habits. This study investigates the feasibility of utilizing Electrodermal Activity (EDA) collected from wearable devices to detect people’s stress. EDA can quantify the changes in sympathetic dynamics by measuring sweat produced by our sweat glands. Currently, the adoption of EDA sensors to commercially off-the-shelf smart-watches is still in the infancy stage, and only a few brands have the EDA sensors implemented into their smartwatch. To facilitate our feasibility study, we need the datasets that contain the EDA signals collected from wearable devices. This paper uses two publicly available datasets containing the EDA signals collected from research-grade wearable devices. We cast the stress detection problem as a binary classification problem and trained the classifiers with three popular machine learning methods: K-Nearest Neighbor, Logistic Regression, and Random Forests. According to experimental results, Random Forests achieves an accuracy of 85.7% to classify stress from non-stress status. The results verified that wearable devices with EDA sensors have the potential to predict stress status.
Lili Zhu, Pai Chet Ng, Yuanhao Yu, Yang Wang 0003, Petros Spachos, Dimitrios Hatzinakos, Konstantinos N. Plataniotis
ICC4
2022 Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-Experts
abstract
In this paper, we tackle the problem of domain shift. Most existing methods perform training on multiple source domains using a single model, and the same trained model is used on all unseen target domains. Such solutions are sub-optimal as each target domain exhibits its own specialty, which is not adapted. Furthermore, expecting single-model training to learn extensive knowledge from multiple source domains is counterintuitive. The model is more biased toward learning only domain-invariant features and may result in negative knowledge transfer. In this work, we propose a novel framework for unsupervised test-time adaptation, which is formulated as a knowledge distillation process to address domain shift. Specifically, we incorporate Mixture-of-Experts (MoE) as teachers, where each expert is separately trained on different source domains to maximize their specialty. Given a test-time target domain, a small set of unlabeled data is sampled to query the knowledge from MoE. As the source domains are correlated to the target domains, a transformer-based aggregator then combines the domain knowledge by examining the interconnection among them. The output is treated as a supervision signal to adapt a student prediction network toward the target domain. We further employ meta-learning to enforce the aggregator to distill positive knowledge and the student network to achieve fast adaptation. Extensive experiments demonstrate that the proposed method outperforms the state-of-the-art and validates the effectiveness of each proposed component. Our code is available at https://github.com/n3il666/Meta-DMoE.
Tao Zhong 0003, Zhixiang Chi, Li Gu, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
NeurIPS4
2022 Region-of-interest and channel attention-based joint optimization of image compression and computer vision
Linwei Ye, Jie Liang 0001, Yang Wang 0003, Jingning Han
Neurocomputing4
2022 Referring Segmentation in Images and Videos With Cross-Modal Self-Attention Network
abstract
We consider the problem of referring segmentation in images and videos with natural language. Given an input image (or video) and a referring expression, the goal is to segment the entity referred by the expression in the image or video. In this paper, we propose a cross-modal self-attention (CMSA) module to utilize fine details of individual words and the input image or video, which effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the visual input. We further propose a gated multi-level fusion (GMLF) module to selectively integrate self-attentive cross-modal features corresponding to different levels of visual features. This module controls the feature fusion of information flow of features at different levels with high-level and low-level semantic information related to different attentive words. Besides, we introduce cross-frame self-attention (CFSA) module to effectively integrate temporal information in consecutive frames which extends our method in the case of referring segmentation in videos. Experiments on benchmark datasets of four referring image datasets and two actor and action video segmentation datasets consistently demonstrate that our proposed approach outperforms existing state-of-the-art methods.
Linwei Ye, Mrigank Rochan, Zhi Liu 0003, Xiaoqin Zhang 0002, Yang Wang 0003
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Gaze Estimation via Modulation-Based Adaptive Network With Auxiliary Self-Learning
abstract
Given a face image, most of previous works in gaze estimation infer the gaze via a well-trained model with supervised training. However, the distribution of test data may be very different compared to that of training data since samples might be corrupted in real-world scenarios (e.g., taking a photo in strong light). This will lead to a gap between source domain (i.e., training data) and target domain (i.e., test data). In this paper, we first introduce self-supervised learning into our method for addressing challenging situations in gaze estimation. Moreover, existing appearance-based gaze estimation methods focus on directing towards the development of powerful regressors, which mainly utilize face and eye images simultaneously or face (eye) images only. However, the problem of inter cues between face and eye features has been largely overlooked. To this end, we propose a novel Modulation-based Adaptive Network (MANet) for gaze estimation, which uses high-level knowledge to filter the distractive information and bridges the intrinsic relationship between face and eye features. Further, we combine self-supervised learning and MANet to learn to adapt to challenging cases, such as abnormal lighting conditions and poor-quality images, by minimizing a self-supervised loss and a supervised loss jointly. The experimental results on several datasets demonstrate the effectiveness of our proposed approach with a real-time speed of 900fpson a PC with an NVIDIA Titan RTX GPU.
Yong Wu 0007, Gongyang Li, Zhi Liu 0003, Mengke Huang, Yang Wang 0003
IEEE Trans. Circuits Syst. Video Technol.5
2022 AdaCrowd: Unlabeled Scene Adaptation for Crowd Counting
abstract
We address the problem of image-based crowd counting. In particular, we propose a new problem calledunlabeled scene-adaptive crowd counting. Given a new target scene, we would like to have a crowd counting model specifically adapted to this particular scene based on the target data that capture some information about the new scene. In this paper, we propose to use one or more unlabeled images from the target scene to perform the adaptation. In comparison with the existing problem setups (e.g. fully supervised), our proposed problem setup is closer to the real-world applications of crowd counting systems. We introduce a novelAdaCrowdframework to solve this problem. Our framework consists of a crowd counting network and a guiding network. The guiding network predicts some parameters in the crowd counting network based on the unlabeled images from a particular scene. This allows our model to adapt to different target scenes. The experimental results on several challenging benchmark datasets demonstrate the effectiveness of our proposed approach compared with other alternative methods. Code is available athttps://github.com/maheshkkumar/adacrowd
Mahesh Kumar Krishna Reddy, Mrigank Rochan, Yiwei Lu 0001, Yang Wang 0003
IEEE Trans. Multim.4
2021 Test-Time Fast Adaptation for Dynamic Scene Deblurring via Meta-Auxiliary Learning
abstract
In this paper, we tackle the problem of dynamic scene deblurring. Most existing deep end-to-end learning approaches adopt the same generic model for all unseen test images. These solutions are sub-optimal, as they fail to utilize the internal information within a specific image. On the other hand, a self-supervised approach, SelfDeblur, enables internal training within a test image from scratch, but it does not fully take advantage of large external datasets. In this work, we propose a novel self-supervised meta-auxiliary learning to improve the performance of deblurring by integrating both external and internal learning. Concretely, we build a self-supervised auxiliary reconstruction task that shares a portion of the network with the primary deblurring task. The two tasks are jointly trained on an external dataset. Furthermore, we propose a meta-auxiliary training scheme to further optimize the pretrained model as a base learner, which is applicable for fast adaptation at test time. During training, the performance of both tasks is coupled. Therefore, we are able to exploit the internal information at test time via the auxiliary task to enhance the performance of deblurring. Extensive experimental results across evaluation datasets demonstrate the effectiveness of test-time adaptation of the proposed method.
Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
CVPR2
2021 Image Change Captioning by Learning From an Auxiliary Task
abstract
We tackle the challenging task of image change captioning. The goal is to describe the subtle difference between two very similar images by generating a sentence caption. While the recent methods mainly focus on proposing new model architectures for this problem, we instead focus on an alternative training scheme. Inspired by the success of multi-task learning, we formulate a training scheme that uses an auxiliary task to improve the training of the change captioning network. We argue that the task of composed query image retrieval is a natural choice as the auxiliary task. Given two almost similar images as the input, the primary network generates a caption describing the fine change between those two images. Next, the auxiliary network is provided with the generated caption and one of those two images. It then tries to pick the second image among a set of candidates. This forces the primary network to generate detailed and precise captions via having an extra supervision loss by the auxiliary network. Furthermore, we propose a new scheme for selecting a negative set of candidates for the retrieval task that can effectively improve the performance. We show that the proposed training strategy performs well on the task of change captioning on benchmark datasets.
Mehrdad Hosseinzadeh, Yang Wang 0003
CVPR2
2021 Toward Personalized Emotion Recognition: A Face Recognition Based Attention Method for Facial Emotion Recognition
abstract
This paper aims to address the subject-dependent challenge of the facial emotion recognition (FER) task. To accomplish this, we propose a novel face recognition based attention FER (FRA-FER) framework which propagates subtle face recognition (FR) features through the FER network. Particularly, first a spatial attention map from the feature maps of an FR convolutional neural network (CNN) is created and then it is fused into the FER-CNN. By doing this FR feature propagation, the FER network is personalized as it takes the advantage of the FR features learned from large-scale face recognition datasets. Experiments on the two challenging datasets AffectNet and AFEW demonstrate the superiority of our proposed FRA-FER network to the state-of-the-art work.
Mostafa Shahabinejad, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005
FG2
2021 Joint Visual and Audio Learning for Video Highlight Detection
abstract
In video highlight detection, the goal is to identify the interesting moments within an unedited video. Although the audio component of the video provides important cues for highlight detection, the majority of existing efforts focus almost exclusively on the visual component. In this paper, we argue that both audio and visual components of a video should be modeled jointly to retrieve its best moments. To this end, we propose an audio-visual network for video highlight detection. At the core of our approach lies a bimodal attention mechanism, which captures the interaction between the audio and visual components of a video, and produces fused representations to facilitate highlight detection. Furthermore, we introduce a noise sentinel technique to adaptively discount a noisy visual or audio modality. Empirical evaluations on two benchmark datasets demonstrate the superior performance of our approach over the state-of-the-art methods.
Taivanbat Badamdorj, Mrigank Rochan, Yang Wang 0003
ICCV3
2021 Video Captioning of Future Frames
abstract
Being able to anticipate and describe what may happen in the future is a fundamental ability for humans. Given a short clip of a scene about "a person is sitting behind a piano", humans can describe what will happen afterward, i.e. "the person is playing the piano". In this paper, we consider the task of captioning future events to assess the performance of intelligent models on anticipation and video description generation tasks simultaneously. More specifically, given only the frames relating to an occurring event (activity), the goal is to generate a sentence describing the most likely next event in the video. We tackle the problem by first predicting the next event in the semantic space of convolutional features, then fusing contextual information into those features, and feeding them to a captioning module. Departing from using recurrent units allows us to train the network in parallel. We compare the proposed method with a baseline and an oracle method on the ActivityNet-Captions dataset. Experimental results demonstrate that the proposed method outperforms the baseline and is comparable to the oracle method. We perform additional ablation study to further analyze our approach.
Mehrdad Hosseinzadeh, Yang Wang 0003
WACV2
2021 Circular Complement Network for RGB-D Salient Object Detection
Zhen Bai 0001, Zhi Liu 0003, Gongyang Li, Linwei Ye, Yang Wang 0003
Neurocomputing5
2021 ATCC: Accurate tracking by criss-cross location attention
Yong Wu 0007, Zhi Liu 0003, Xiaofei Zhou 0003, Linwei Ye, Yang Wang 0003
Image Vis. Comput.5
2020 Sentence Guided Temporal Modulation for Dynamic Video Thumbnail Generation
Mrigank Rochan, Mahesh Kumar Krishna Reddy, Yang Wang 0003
BMVC3
2020 Composed Query Image Retrieval Using Locally Bounded Features
abstract
Composed query image retrieval is a new problem where the query consists of an image together with a requested modification expressed via a textual sentence. The goal is then to retrieve the images that are generally similar to the query image, but differ according to the requested modification. Previous methods usually consider the image as a whole. In this paper, we propose a novel method that represents the image using a set of local areas in the image. The relationship between each word in the modification text and each area in the image is then explicitly established, allowing the model to accurately correlate the modification text to parts of the image. We conduct extensive experiments on three benchmark datasets. The results show that our method outperforms other state-of-the-art approaches by a considerable margin.
Mehrdad Hosseinzadeh, Yang Wang 0003
CVPR2
2020 Deep Learning-Based Image Compression with Trellis Coded Quantization
abstract
Recently many works attempt to develop image compression models based on deep learning architectures, where the uniform scalar quantizer (SQ) is commonly applied to the feature maps between the encoder and decoder. In this paper, we propose to incorporate trellis coded quantizer (TCQ) into a deep learning based image compression framework. A soft-to-hard strategy is applied to allow for back propagation during training. We develop a simple image compression model that consists of three subnetworks (encoder, decoder and entropy estimation), and optimize all of the components in an end-to-end manner. We experiment on two high resolution image datasets and both show that our model can achieve superior performance at low bit rates. We also show the comparisons between TCQ and SQ based on our proposed baseline model and demonstrate the advantage of TCQ.
Jie Liang 0001, Yang Wang 0003
DCC4
2020 Cross-Modal Weighting Network for RGB-D Salient Object Detection
Gongyang Li, Zhi Liu 0003, Linwei Ye, Yang Wang 0003, Haibin Ling
ECCV (17)4
2020 Few-Shot Scene-Adaptive Anomaly Detection
Yiwei Lu 0001, Frank Yu, Mahesh Kumar Krishna Reddy, Yang Wang 0003
ECCV (5)4
2020 Adaptive Video Highlight Detection by Learning from User History
Mrigank Rochan, Mahesh Kumar Krishna Reddy, Linwei Ye, Yang Wang 0003
ECCV (21)4
2020 Unsupervised Learning of Camera Pose with Compositional Re-estimation
abstract
We consider the problem of unsupervised camera pose estimation. Given an input video sequence, our goal is to estimate the camera pose (i.e. the camera motion) between consecutive frames. Traditionally, this problem is tackled by placing strict constraints on the transformation vector or by incorporating optical flow through a complex pipeline. We propose an alternative approach that utilizes a compositional re-estimation process for camera pose estimation. Given an input, we first estimate a depth map. Our method then iteratively estimates the camera motion based on the estimated depth map. Our approach significantly improves the predicted camera motion both quantitatively and visually. Furthermore, the re-estimation resolves the problem of out-of-boundaries pixels in a novel and simple way. Another advantage of our approach is that it is adaptable to other camera pose estimation approaches. Experimental analysis on KITTI benchmark dataset demonstrates that our method outperforms existing state-of-the-art approaches in unsupervised camera ego-motion estimation.
Seyed Shahabeddin Nabavi, Mehrdad Hosseinzadeh, Ramin Fahimi, Yang Wang 0003
WACV4
2020 Few-Shot Scene Adaptive Crowd Counting Using Meta-Learning
abstract
We consider the problem of few-shot scene adaptive crowd counting. Given a target camera scene, our goal is to adapt a model to this specific scene with only a few labeled images of that scene. The solution to this problem has potential applications in numerous real-world scenarios, where we ideally like to deploy a crowd counting model specially adapted to a target camera. We accomplish this challenge by taking inspiration from the recently introduced learning-to-learn paradigm in the context of few-shot regime. In training, our method learns the model parameters in a way that facilitates the fast adaptation to the target scene. At test time, given a target scene with a small number of labeled data, our method quickly adapts to that scene with a few gradient updates to the learned parameters. Our extensive experimental results show that the proposed approach outperforms other alternatives in few-shot scene adaptive crowd counting.
Mahesh Kumar Krishna Reddy, Mohammad Asiful Hossain, Mrigank Rochan, Yang Wang 0003
WACV4
2020 Attentional coarse-and-fine generative adversarial networks for image inpainting
Minyu Chen 0001, Zhi Liu 0003, Linwei Ye, Yang Wang 0003
Neurocomputing4
2020 Weakly supervised instance segmentation using multi-stage erasing refinement and saliency-guided proposals ordering
Zhi Liu 0003, Gongyang Li, Linwei Ye, Lei Zhou 0003, Yang Wang 0003
J. Vis. Commun. Image Represent.6
2020 Dual Convolutional LSTM Network for Referring Image Segmentation
abstract
We consider referring image segmentation. It is a problem at the intersection of computer vision and natural language understanding. Given an input image and a referring expression in the form of a natural language sentence, the goal is to segment the object of interest in the image referred by the linguistic query. To this end, we propose a dual convolutional LSTM (ConvLSTM) network to tackle this problem. Our model consists of an encoder network and a decoder network, where ConvLSTM is used in both encoder and decoder networks to capture spatial and sequential information. The encoder network extracts visual and linguistic features for each word in the expression sentence, and adopts an attention mechanism to focus on words that are more informative in the multimodal interaction. The decoder network integrates the features generated by the encoder network at multiple levels as its input and produces the final precise segmentation mask. Experimental results on four challenging datasets demonstrate that the proposed network achieves superior segmentation performance compared with other state-of-the-art methods.
Linwei Ye, Zhi Liu 0003, Yang Wang 0003
IEEE Trans. Multim.3
2019 Future Frame Prediction Using Convolutional VRNN for Anomaly Detection
abstract
Anomaly detection in videos aims at reporting anything that does not conform the normal behaviour or distribution. However, due to the sparsity of abnormal video clips in real life, collecting annotated data for supervised learning is exceptionally cumbersome. Inspired by the practicability of generative models for semi-supervised learning, we propose a novel sequential generative model based on variational autoencoder (VAE) for future frame prediction with convolutional LSTM (ConvLSTM). To the best of our knowledge, this is the first work that considers temporal information in future frame prediction based anomaly detection framework from the model perspective. Our experiments demonstrate that our approach is superior to the state-of-the-art methods on three benchmark datasets.
Yiwei Lu 0001, K. Mahesh Kumar, Seyed Shahabeddin Nabavi, Yang Wang 0003
AVSS4
2019 Video-Based Person Re-Identification using Refined Attention Networks
abstract
We consider the problem of video-based person reidentification. The goal is to identify a person from videos captured under different cameras. In this paper, we propose an efficient attention based model for person re-identifying from videos. Our method generates an attention score for each frame based on frame-level features. The attention scores of all frames in a video are used to produce a weighted feature vector for the input video. This video-level feature vector is refined iteratively for re-identifying persons from videos. Unlike most existing deep learning methods that use global or spatial representation, our approach focuses on attention scores. Extensive experiments on three benchmark datasets demonstrate that our method achieves the state-of-the-art performance.
Tanzila Rahman, Mrigank Rochan, Yang Wang 0003
AVSS3
2019 Non-Local Attentive Temporal Network for Video-Based Person Re-Identification
abstract
Given a video containing a person, the goal of person re-identification is to identify the same person from videos captured under different cameras. A common approach for tackling this problem is to first extract image features for all frames in the video. These frame-level features are then combined (e.g. via temporal pooling) to form a video-level feature vector. The video-level features of two input videos are then compared by calculating the distance between them. More recently, attention-based learning mechanism has been proposed for this problem. In particular, recurrent neural networks have been used to generate the attention scores of frames in a video. However, the limitation of RNN-based approach is that it is difficult for RNNs to capture long-range dependencies in videos. Inspired by the success of non-local neural networks, we propose a novel non-local temporal attention model in this paper. Our model can effectively capture long-range and global dependencies among the frames of the videos. Extensive experiments on three different benchmark datasets (i.e. iLIDS-VID, PRID-2011 and SDU-VID) show that our proposed method outperforms other state-of-the-art approaches.
Shivansh Rao, Tanzila Rahman, Mrigank Rochan, Yang Wang 0003
AVSS5
2019 One-Shot Scene-Specific Crowd Counting
Mohammad Asiful Hossain, K. Mahesh Kumar, Mehrdad Hosseinzadeh, Omit Chanda, Yang Wang 0003
BMVC5
2019 Video Summarization by Learning From Unpaired Data
abstract
We consider the problem of video summarization. Given an input raw video, the goal is to select a small subset of key frames from the input video to create a shorter summary video that best describes the content of the original video. Most of the current state-of-the-art video summarization approaches use supervised learning and require labeled training data. Each training instance consists of a raw input video and its ground truth summary video curated by human annotators. However, it is very expensive and difficult to create such labeled training examples. To address this limitation, we propose a novel formulation to learn video summarization from unpaired data. We present an approach that learns to generate optimal video summaries using a set of raw videos (V) and a set of summary videos (S), where there exists no correspondence between V and S. We argue that this type of data is much easier to collect. Our model aims to learn a mapping function F : V -> S such that the distribution of resultant summary videos from F(V) is similar to the distribution of S with the help of an adversarial objective. In addition, we enforce a diversity constraint on F(V) to ensure that the generated video summaries are visually diverse. Experimental results on two benchmark datasets indicate that our proposed approach significantly outperforms other alternative methods.
Mrigank Rochan, Yang Wang 0003
CVPR2
2019 Cross-Modal Self-Attention Network for Referring Image Segmentation
abstract
We consider the problem of referring image segmentation. Given an input image and a natural language expression, the goal is to segment the object referred by the language expression in the image. Existing works in this area treat the language expression and the input image separately in their representations. They do not sufficiently capture long-range correlations between these two modalities. In this paper, we propose a cross-modal self-attention (CMSA) module that effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the input image. In addition, we propose a gated multi-level fusion module to selectively integrate self-attentive cross-modal features corresponding to different levels in the image. This module controls the information flow of features at different levels. We validate the proposed approach on four evaluation datasets. Our proposed approach consistently outperforms existing state-of-the-art methods.
Linwei Ye, Mrigank Rochan, Zhi Liu 0003, Yang Wang 0003
CVPR4
2019 Compression Artifact Removal with Stacked Multi-Context Channel-Wise Attention Network
abstract
Image compression plays an important role in saving disk storage and transmission bandwidth. Among traditional compression standards, JPEG is one of the commonly used standards in lossy image compression. However, the decompressed JPEG images usually have inevitable artifacts due to the quantization step, especially at low bitrate. Many recent works leverage deep learning networks to remove the JPEG artifacts and have achieved notable progress. In this paper, we propose a stacked multi-context channel-wise attention model. The channel-wise attention adaptively integrates features along the channel dimension given a set of feature maps. We apply multiple context-based channel attentions to enable the network to capture features from different resolutions. The entire architecture is trained progressively from the image space of low quality factor to that of high quality factor. Experiments show that we can achieve the state-of-the-art performance with lower complexity.
Jie Liang 0001, Yang Wang 0003
ICIP3
2019 Convolutional Temporal Attention Model for Video-Based Person Re-Identification
abstract
The goal of video-based person re-identification is to match two input videos, so that the distance of the two videos is small if two videos contain the same person. A common approach for person re-identification is to first extract image features for all frames in the video, then aggregate all the features to form a video-level feature. The video-level features of two videos can then be used to calculate the distance of the two videos. In this paper, we propose a temporal attention approach for aggregating frame-level features into a video-level feature vector for re-identification. Our method is motivated by the fact that not all frames in a video are equally informative. We propose a fully convolutional temporal attention model for generating the attention scores. Fully convolutional network (FCN) has been widely used in semantic segmentation for generating 2D output maps. In this paper, we formulate video based person reidentification as a sequence labeling problem like semantic segmentation. We establish a connection between them and modify FCN to generate attention scores to represent the importance of each frame. Extensive experiments on three different benchmark datasets (i.e. iLIDS-VID, PRID-2011 and SDU-VID) show that our proposed method outperforms other state-of-the-art approaches.
Tanzila Rahman, Mrigank Rochan, Yang Wang 0003
ICME3
2019 Semantic Segmentation in Compressed Videos
abstract
Existing approaches for semantic segmentation in videos usually extract each frame as an RGB image, then apply standard image-based semantic segmentation models on each frame. This is time-consuming. In this paper, we tackle this problem by exploring the nature of video compression techniques. A compressed video contains three types of frames, I-frames, P-frames, and B-frames. I-frames are represented as regular images, P-frames are represented as motion vectors and residual errors, and B-frames are bidirectionally frames that can be regarded as a special case of a P frame. We propose a method that directly operates on I-frames (as RGB images) and P-frames (motion vectors and residual errors) in a video. Our proposed model uses a ConvLSTM model to capture the temporal information in the video required for producing the semantic segmentation on P-frames. Our experimental results show that our method performs much faster than other alternatives while achieveing similar performance in terms of accuracies.
Yiwei Lu 0001, Yang Wang 0003
MMSP3
2019 Crowd Counting Using Scale-Aware Attention Networks
abstract
In this paper, we consider the problem of crowd counting in images. Given an image of a crowded scene, our goal is to estimate the density map of this image, where each pixel value in the density map corresponds to the crowd density at the corresponding location in the image. Given the estimated density map, the final crowd count can be obtained by summing over all values in the density map. One challenge of crowd counting is the scale variation in images. In this work, we propose a novel scale-aware attention network to address this challenge. Using the attention mechanism popular in recent deep learning architectures, our model can automatically focus on certain global and local scales appropriate for the image. By combining these global and local scale attentions, our model outperforms other state-of-the-art methods for crowd counting on several benchmark datasets.
Mohammad Asiful Hossain, Mehrdad Hosseinzadeh, Omit Chanda, Yang Wang 0003
WACV4
2019 Saliency detection via multi-level integration and multi-scale fusion neural networks
Mengke Huang, Zhi Liu 0003, Linwei Ye, Xiaofei Zhou 0003, Yang Wang 0003
Neurocomputing5
2019 Weakly labeled fine-grained classification with hierarchy relationship of fine and coarse labels
Qihan Jiao, Zhi Liu 0003, Linwei Ye, Yang Wang 0003
J. Vis. Commun. Image Represent.4
2018 Future Semantic Segmentation with Convolutional LSTM
Seyed Shahabeddin Nabavi, Mrigank Rochan, Yang Wang 0003
BMVC3
2018 Video Summarization Using Fully Convolutional Sequence Networks
Mrigank Rochan, Linwei Ye, Yang Wang 0003
ECCV (12)3
2018 Visual Relationship Detection Using Joint Visual-Semantic Embedding
abstract
Visual relationship detection can serve as the intermediate building block for higher level tasks such as image captioning, visual question answering, image-text matching. Due to the long tail of relationship distribution in real world images, zero-shot predication of relationships that it has never seen before can alleviate stress of collecting every possible relationship. Following zero-shot learning (ZSL) strategies, we propose a joint visual-semantic embedding model for visual relationship detection. In our model, the visual vector and semantic vector are projected to a shared latent space to learn the similarity between the two branches. In the semantic embedding, sequential features in terms ofare learned to provide the context information and then concatenated with corresponding component vector of the relationship triplet. Experiments show that the proposed model achieves superior performance in zero-shot visual relationship detection and comparable results in non-zero-shot scenario.
Yang Wang 0003
ICPR2
2018 Privacy-Preserving Age Estimation for Content Rating
abstract
Content rating (aka. maturity rating) rates the suitability of kinds of media (e.g., movies and video games) to its audience. It is essential to prevent a specific age group of people such as children from inappropriate information. However, in practice the administration of content rating system is usually suggestion-based declaration by media sources or key-based password which can easily fail if someone ignores the suggestions or somehow knows the keys. In this paper, we propose to estimate user's age in a privacy-preserving manner for automatic content rating. Several privacy-preserving approaches on facial images with different degree of privacy are proposed and evaluated on a deep neural network architecture for age estimation accuracy. We also introduce an attention mechanism which can adaptively learn discriminative features from the processed facial images. Experiments show that the proposed attention-based model performs better than the baseline model and achieves a reasonable performance to that with raw images in testing.
Linwei Ye, Noman Mohammed, Yang Wang 0003, Jie Liang 0001
MMSP4
2018 Learning Semantic Segmentation with Diverse Supervision
abstract
Models based on deep convolutional neural networks (CNN) have significantly improved the performance of semantic segmentation. However, learning these models requires a large amount of training images with pixel-level labels, which are very costly and time-consuming to collect. In this paper, we propose a method for learning CNNbased semantic segmentation models from images with several types of annotations that are available for various computer vision tasks, including image-level labels for classification, box-level labels for object detection and pixel-level labels for semantic segmentation. The proposed method is flexible and can be used together with any existing CNNbased semantic segmentation networks. Experimental evaluation on the challenging PASCAL VOC 2012 and SIFTflow benchmarks demonstrate that the proposed method can effectively make use of diverse training data to improve the performance of the learned models.
Linwei Ye, Zhi Liu 0003, Yang Wang 0003
WACV3
2017 Adapting Object Detectors from Images to Weakly Labeled Videos
Omit Chanda, Eu Wern Teh, Mrigank Rochan, Yang Wang 0003
BMVC5
2017 Salient Object Detection using a Context-Aware Refinement Network
Md. Amirul Islam, Mahmoud Kalash, Mrigank Rochan, Neil D. B. Bruce, Yang Wang 0003
BMVC5
2017 Person Re-Identification by Localizing Discriminative Regions
Tanzila Rahman, Mrigank Rochan, Yang Wang 0003
BMVC3
2017 Gated Feedback Refinement Network for Dense Image Labeling
abstract
Effective integration of local and global contextual information is crucial for dense labeling problems. Most existing methods based on an encoder-decoder architecture simply concatenate features from earlier layers to obtain higher-frequency details in the refinement stages. However, there are limits to the quality of refinement possible if ambiguous information is passed forward. In this paper we propose Gated Feedback Refinement Network (G-FRNet), an end-to-end deep learning framework for dense labeling tasks that addresses this limitation of existing methods. Initially, G-FRNet makes a coarse prediction and then it progressively refines the details by efficiently integrating local and global contextual information during the refinement stages. We introduce gate units that control the information passed forward in order to filter out ambiguity. Experiments on three challenging dense labeling datasets (CamVid, PASCAL VOC 2012, and Horse-Cow Parsing) show the effectiveness of our method. Our proposed approach achieves state-of-the-art results on the CamVid and Horse-Cow Parsing datasets, and produces competitive results on the PASCAL VOC 2012 dataset.
Md. Amirul Islam, Mrigank Rochan, Neil D. B. Bruce, Yang Wang 0003
CVPR4
2017 Depth-aware object instance segmentation
abstract
We consider the problem of object instance segmentation. The goal is to label each pixel in an image according to its object class as well as its object instance. The proposed approach consists of three steps including object instance detection, category-specific instance segmentation and depth-aware ordering. The novelty of the proposed approach is that it uses the depth information to resolve the ambiguity of pixel labels when two object instances are overlapping. Experimental results on the PASCAL VOC 2012 benchmark demonstrate the competitive performance of the proposed approach compared with other state-of-the-art methods.
Linwei Ye, Zhi Liu 0003, Yang Wang 0003
ICIP3
2017 Object localization in weakly labeled data using regularized attention networks
abstract
We consider the problem of weakly supervised object localization. For an object of interest (e.g. “car”), an image is weakly labeled when its label only indicates the presence/absence of this object, but not the exact location of the object in the image. Given a collection of weakly labeled images for an object, our goal is to localize the object of interest in each image. We propose a novel architecture called the regularized attention network for this problem. Our work builds upon the attention network proposed in [1]. We extend the standard attention network by incorporating a regularization term that encourages the attention scores of object proposals to mimic the scoring distribution of a strong fully supervised object detector. Despite of the simplicity of our approach, our proposed architecture achieves the state-of-the-art results on several benchmark datasets.
Eu Wern Teh, Yang Wang 0003
VCIP3
2017 Salient Object Segmentation via Effective Integration of Saliency and Objectness
abstract
This paper proposes an effective salient object segmentation method via the graph-based integration of saliency and objectness. Based on the superpixel segmentation result of the input image, a graph is built to represent superpixels using regular vertex, background seed vertex with the addition of a terminal vertex. The edge weights on the graph are defined by integrating the difference of appearance, saliency, and objectness between superpixels. Then, the object probability of each superpixel is measured by finding the shortest path from the corresponding vertex to the terminal vertex on the graph, and the resultant object probability map can generally better highlight salient objects and suppress background regions compared to both saliency map and objectness map. Finally, the object probability map is used to initialize salient object and background, and effectively incorporated into the framework of graph cut to obtain the final salient object segmentation result. Extensive experimental results on three public benchmark datasets show that the proposed method consistently improves the salient object segmentation performance and outperforms the state-of-the-art salient object segmentation methods. Furthermore, experimental results also demonstrate that the proposed graph-based integration method is more effective than other fusion schemes and robust to saliency maps generated using various saliency models.
Linwei Ye, Zhi Liu 0003, Liquan Shen, Cong Bai, Yang Wang 0003
IEEE Trans. Multim.6
2016 Attention Networks for Weakly Supervised Object Localization
Eu Wern Teh, Mrigank Rochan, Yang Wang 0003
BMVC3
2016 LSTM for Image Annotation with Relative Visual Importance
Geng Yan, Yang Wang 0003, Zicheng Liao
BMVC2
2016 Beyond verbs: Understanding actions in videos with text
abstract
We consider the problem of joint modeling of videos and their corresponding textual descriptions (e.g. sentences or phrases). Our approach consists of three components: the video representation, the textual representation, and a joint model that links videos and text. Our video representation uses the state-of-the-art deep 3D ConvNet to capture the semantic information in the video. Our textual representation uses the recent advancement in learning word and sentence vectors from large text corpus. The joint model is learned to score the correct (video, text) pairs higher than the incorrect ones. We demonstrate our approach in several applications: 1) retrieving sentences given a video; 2) retrieving videos given a sentence; 3) zero-shot action recognition in videos.
Shujon Naha, Yang Wang 0003
ICPR2
2016 Object figure-ground segmentation using zero-shot learning
abstract
We consider the problem of object figure-ground segmentation when the object categories are not available during training (i.e. zero-shot). During training, we learn standard segmentation models for a handful of object categories (called “source objects”) using existing semantic segmentation datasets. During testing, we are given images of objects (called “target objects”) that are unseen during training. Our goal is to segment the target objects from the background. Our method learns to transfer the knowledge from the source objects to the target objects. Our experimental results demonstrate the effectiveness of our approach.
Shujon Naha, Yang Wang 0003
ICPR2
2016 Weakly supervised object localization and segmentation in videos
Mrigank Rochan, Shafin Rahman, Neil D. B. Bruce, Yang Wang 0003
Image Vis. Comput.4
2015 Weakly supervised localization of novel objects using appearance transfer
abstract
We consider the problem of localizing unseen objects in weakly labeled image collections. Given a set of images annotated at the image level, our goal is to localize the object in each image. The novelty of our proposed work is that, in addition to building object appearance model from the weakly labeled data, we also make use of existing detectors of some other object classes (which we call “familiar objects”). We propose a method for transferring the appearance models of the familiar objects to the unseen object. Our experimental results on both image and video datasets demonstrate the effectiveness of our approach.
Mrigank Rochan, Yang Wang 0003
CVPR2
2014 Examining visual saliency prediction in naturalistic scenes
abstract
Given the significant number of potential applications, visual saliency has increasingly become an area of interest in image and vision research. Many different strategies for predicting visual saliency have been proposed, that differ in their composition or rationale, and with a significant focus on improving performance across standard benchmarks. Recent benchmarks considering a large number of algorithms have further provided an understanding of the behavior of different algorithms. Performance evaluation has primarily focused on indoor and outdoor images of urban environments, many of which are composed, and contain salient objects. In this work, we test the performance of a number of the better performing algorithms on data derived from naturalistic scenes. In addition, given the strong connection to human vision, we test a putative model for early visual processing in primates tied to spectral energy and normalization. Results demonstrate significant differences between common datasets, and natural images. Performance analysis of the second-order contrast model also provides additional insight concerning the role of spectral energy in determining saliency. Finally we include analysis that demonstrates statistical properties of images that tend to imply common gaze patterns across observers.
Shafin Rahman, Mrigank Rochan, Yang Wang 0003, Neil D. B. Bruce
ICIP3
2014 Human parsing with a cascade of hierarchical poselet based pruners
abstract
We address the problem of human parsing using part-based models. In particular, we consider part-based models that exploit rich pairwise relationship between parts, e.g. the color symmetry between left/right limbs. This poses a computational challenge since the state space of each part is very large, and algorithmic tricks (e.g. the distance transform) cannot be applied to handle these types of pairwise relationships. We propose to prune the state space of each part using a cascade of pruners. These pruners can filter out 99.6% of the states per part to about 500 states per part, while keeping the ground-truth states in the pruned state most of the time. In the pruned space, we can afford to apply human parsing models with more complex pairwise relationships between parts, such as the color symmetry. We demonstrate our method on a challenging human parsing dataset.
Duan Tran, Yang Wang 0003, David A. Forsyth
ICME2
2014 A machine learning approach for stock price prediction
abstract
Data mining and machine learning approaches can be incorporated into business intelligence (BI) systems to help users for decision support in many real-life applications. Here, in this paper, we propose a machine learning approach for BI applications. Specifically, we apply structural support vector machines (SSVMs) to perform classification on complex inputs such as the nodes of a graph structure. We connect collaborating companies in the information technology sector in a graph structure and use an SSVM to predict positive or negative movement in their stock prices. The complexity of the SSVM cutting plane optimization problem is determined by the complexity of the separation oracle. It is shown that (i) the separation oracle performs a task equivalent to maximum a posteriori (MAP) inference and (ii) a minimum graph cutting algorithm can solve this problem in the stock price case in polynomial time. Experimental results show the practicability of our proposed machine learning approach in predicting stock prices.
Carson K. Leung, Richard Kyle MacKinnon, Yang Wang 0003
IDEAS3
2014 Learning mid-level features from object hierarchy for image classification
abstract
We propose a new approach for constructing mid-level visual features for image classification. We represent an image using the outputs of a collection of binary classifiers. These binary classifiers are trained to differentiate pairs of object classes in an object hierarchy. Our feature representation implicitly captures the hierarchical structure in object classes. We show that our proposed approach outperforms other baseline methods in image classification.
Somayah Albaradei, Yang Wang 0003, Liangliang Cao, Li-Jia Li 0001
WACV2
2013 Non-parametric Filtering for Geometric Detail Extraction and Material Representation
abstract
Geometric detail is a universal phenomenon in real world objects. It is an important component in object modeling, but not accounted for in current intrinsic image works. In this work, we explore using a non-parametric method to separate geometric detail from intrinsic image components. We further decompose an image as albedo * (coarse-scale shading + shading detail). Our decomposition offers quantitative improvement in albedo recovery and material classification. Our method also enables interesting image editing activities, including bump removal, geometric detail smoothing/enhancement and material transfer.
Zicheng Liao, Jason Rock, Yang Wang 0003, David A. Forsyth
CVPR3
2013 Large multi-class image categorization with ensembles of label trees
abstract
We consider sublinear test-time algorithms for image categorization when the number of classes is very large. Our method builds upon the label tree approach proposed in [1], which decomposes the label set into a tree structure and classify a test example by traversing the tree. Even though this method achieves logarithmic run-time, its performance is limited by the fact that any errors made in an internal node of the tree cannot be recovered. In this paper, we propose label forests - ensembles of label trees. Each tree in a label forest will decompose the label set in a slightly different way. The final classification decision is made by aggregating information across all trees in the label forest. The test running time of label forest is still logarithmic in the number of categories. But using an ensemble of label trees achieves much better performance in terms of accuracies. We demonstrate our approach on an image classification task that involves 1000 categories.
Yang Wang 0003, David A. Forsyth
ICME1
2013 Discovering Latent Clusters from Geotagged Beach Images
Yang Wang 0003, Liangliang Cao
MMM (2)1
2013 Optimizing Nondecomposable Loss Functions in Structured Prediction
abstract
We develop an algorithm for structured prediction with nondecomposable performance measures. The algorithm learns parameters of Markov Random Fields (MRFs) and can be applied to multivariate performance measures. Examples include performance measures such as Fβ score (natural language processing), intersection over union (object category segmentation), Precision/Recall at k (search engines), and ROC area (binary classifiers). We attack this optimization problem by approximating the loss function with a piecewise linear function. The loss augmented inference forms a Quadratic Program (QP), which we solve using LP relaxation. We apply this approach to two tasks: object class-specific segmentation and human action retrieval from videos. We show significant improvement over baseline approaches that either use simple loss functions or simple scoring functions on the PASCAL VOC and H3D Segmentation datasets, and a nursing home action recognition dataset.
Mani Ranjbar, Tian Lan 0006, Yang Wang 0003, Stephen N. Robinovitch, Ze-Nian Li, Greg Mori
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Building a dictionary of image fragments
abstract
We show how to build large dictionaries of meaningful image fragments. These fragments could represent objects, objects in a local context, or parts of scenes. Our fragments operate as region-based exemplars, and we show how they can be used for image classification, to localize objects, and to compose new images. While each of these activities has been demonstrated before, each has required manually extracted fragments. Because our method for fragment extraction is automatic it can operate at a large scale. Our method uses recent advances in generic object detection techniques, together with discriminative tests to obtain good, clean fragment sets with extensive diversity. Our fragments are organized by the tags of the source images to build a semantically organized fragment table. A good set of fragment exemplars describes only the object, rather than object+context. Context could help identify an object; but it could also contribute noise, because other objects might appear in the same context. We show a slight improvement in classification performance by two standard exemplar matching methods using our fragment dictionary over such methods using image exemplars. This suggests that knowing the support of an exemplar is valuable. Furthermore, we demonstrate our automatically built fragment dictionary is capable of good localization. Finally, our fragment dictionary supports a keyword based fragment search system, which allows artists to get the fragments they need to make image collages.
Zicheng Liao, Ali Farhadi, Yang Wang 0003, Ian Endres, David A. Forsyth
CVPR3
2012 Image Retrieval with Structured Object Queries Using Latent Ranking SVM
Tian Lan 0006, Weilong Yang, Yang Wang 0003, Greg Mori
ECCV (6)3
2012 Kernel Latent SVM for Visual Recognition
abstract
Latent SVMs (LSVMs) are a class of powerful tools that have been successfully applied to many applications in computer vision. However, a limitation of LSVMs is that they rely on linear models. For many computer vision tasks, linear models are suboptimal and nonlinear models learned with kernels typically perform much better. Therefore it is desirable to develop the kernel version of LSVM. In this paper, we propose kernel latent SVM (KLSVM) -- a new learning framework that combines latent SVMs and kernel methods. We develop an iterative training algorithm to learn the model parameters. We demonstrate the effectiveness of KLSVM using three different applications in visual recognition. Our KLSVM formulation is very general and can be applied to solve a wide range of applications in computer vision and machine learning.
Weilong Yang, Yang Wang 0003, Arash Vahdat, Greg Mori
NIPS2
2012 Discriminative hierarchical part-based models for human parsing and action recognition
Yang Wang 0003, Duan Tran, Zicheng Liao, David A. Forsyth
J. Mach. Learn. Res.1
2012 Discriminative Latent Models for Recognizing Contextual Group Activities
abstract
In this paper, we go beyond recognizing the actions of individuals and focus on group activities. This is motivated from the observation that human actions are rarely performed in isolation; the contextual information of what other people in the scene are doing provides a useful cue for understanding high-level activities. We propose a novel framework for recognizing group activities which jointly captures the group activity, the individual person actions, and the interactions among them. Two types of contextual information, group-person interaction and person-person interaction, are explored in a latent variable framework. In particular, we propose three different approaches to model the person-person interaction. One approach is to explore the structures of person-person interaction. Differently from most of the previous latent structured models, which assume a predefined structure for the hidden layer, e.g., a tree structure, we treat the structure of the hidden layer as a latent variable and implicitly infer it during learning and inference. The second approach explores person-person interaction in the feature level. We introduce a new feature representation called the action context (AC) descriptor. The AC descriptor encodes information about not only the action of an individual person in the video, but also the behavior of other people nearby. The third approach combines the above two. Our experimental results demonstrate the benefit of using contextual information for disambiguating group activities.
Tian Lan 0006, Yang Wang 0003, Weilong Yang, Stephen N. Robinovitch, Greg Mori
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Max-margin Latent Dirichlet Allocation for Image Classification and Annotation
abstract
We present the max-margin latent Dirichlet allocation, a max-margin variant of supervised topic models, for image classification and annotation. Our model for image classification (called MMLDA c) integrates discriminative classification with generative topic models. Our model for image annotation (called MMLDA a) extends MMLDA c to the case of multi-label problems, where each image can be associated with more than one annotation terms. We derive efficient learning algorithms for both models and demonstrate experimentally the advantages of our proposed models over other baseline methods. 1
Yang Wang 0003, Greg Mori
BMVC1
2011 Latent Boosting for Action Recognition
abstract
In this thesis, we present work towards addressing a grand challenge of computer vision, human action recognition and detection. In particular, we focus on the problem of recognizing and detecting the actions of a person from a video sequence. To recognize human actions in a video, a typical approach involves first detecting and tracking people, followed by classification. However, accurate tracking is challenging, and the state-of-art tracking methods are not reliable. Since accurate tracking is not a direct end-goal of action recognition, we consider tracking as a latent variable and train a model focused on action recognition. We propose a novel learning algorithm for training models with latent variables in a boosting framework. Moreover, we show that the algorithm can be used to train an action recognition model in which the tracking trajectory of a person is a latent variable. This new model outperforms baselines on a variety of datasets.
Zhi Feng Huang, Weilong Yang, Yang Wang 0003, Greg Mori
BMVC3
2011 Learning hierarchical poselets for human parsing
abstract
We consider the problem of human parsing with part-based models. Most previous work in part-based models only considers rigid parts (e.g. torso, head, half limbs) guided by human anatomy. We argue that this representation of parts is not necessarily appropriate for human parsing. In this paper, we introduce hierarchical poselets-a new representation for human parsing. Hierarchical poselets can be rigid parts, but they can also be parts that cover large portions of human bodies (e.g. torso + left arm). In the extreme case, they can be the whole bodies. We develop a structured model to organize poselets in a hierarchical way and learn the model parameters in a max-margin framework. We demonstrate the superior performance of our proposed approach on two datasets with aggressive pose variations.
Yang Wang 0003, Duan Tran, Zicheng Liao
CVPR1
2011 Discriminative figure-centric models for joint action localization and recognition
abstract
In this paper we develop an algorithm for action recognition and localization in videos. The algorithm uses a figure-centric visual word representation. Different from previous approaches it does not require reliable human detection and tracking as input. Instead, the person location is treated as a latent variable that is inferred simultaneously with action recognition. A spatial model for an action is learned in a discriminative fashion under a figure-centric representation. Temporal smoothness over video sequences is also enforced. We present results on the UCF-Sports dataset, verifying the effectiveness of our model in situations where detection and tracking of individuals is challenging.
Tian Lan 0006, Yang Wang 0003, Greg Mori
ICCV2
2011 Hidden Part Models for Human Action Recognition: Probabilistic versus Max Margin
abstract
We present a discriminative part-based approach for human action recognition from video sequences using motion features. Our model is based on the recently proposed hidden conditional random field (HCRF) for object recognition. Similarly to HCRF for object recognition, we model a human action by a flexible constellation of parts conditioned on image observations. Differently from object recognition, our model combines both large-scale global features and local patch features to distinguish various actions. Our experimental results show that our model is comparable to other state-of-the-art approaches in action recognition. In particular, our experimental results demonstrate that combining large-scale global features and local patch features performs significantly better than directly applying HCRF on local patches alone. We also propose an alternative for learning the parameters of an HCRF model in a max-margin framework. We call this method the max-margin hidden conditional random field (MMHCRF). We demonstrate that MMHCRF outperforms HCRF in human action recognition. In addition, MMHCRF can handle a much broader range of complex hidden structures arising in various problems in computer vision.
Yang Wang 0003, Greg Mori
IEEE Trans. Pattern Anal. Mach. Intell.1
2010 Recognizing human actions from still images with latent poses
abstract
We consider the problem of recognizing human actions from still images. We propose a novel approach that treats the pose of the person in the image as latent variables that will help with recognition. Different from other work that learns separate systems for pose estimation and action recognition, then combines them in an ad-hoc fashion, our system is trained in an integrated fashion that jointly considers poses and actions. Our learning objective is designed to directly exploit the pose information for action recognition. Our experimental results demonstrate that by inferring the latent poses, we can improve the final action recognition results.
Weilong Yang, Yang Wang 0003, Greg Mori
CVPR2
2010 Optimizing Complex Loss Functions in Structured Prediction
Mani Ranjbar, Greg Mori, Yang Wang 0003
ECCV (2)3
2010 A Discriminative Latent Model of Object Classes and Attributes
Yang Wang 0003, Greg Mori
ECCV (5)1
2010 Beyond Actions: Discriminative Models for Contextual Group Activities
abstract
We propose a discriminative model for recognizing group activities. Our model jointly captures the group activity, the individual person actions, and the interactions among them. Two new types of contextual information, group-person interaction and person-person interaction, are explored in a latent variable framework. Different from most of the previous latent structured models which assume a predefined structure for the hidden layer, e.g. a tree structure, we treat the structure of the hidden layer as a latent variable and implicitly infer it during learning and inference. Our experimental results demonstrate that by inferring this contextual information together with adaptive structures, the proposed model can significantly improve activity recognition performance.
Tian Lan 0006, Yang Wang 0003, Weilong Yang, Greg Mori
NIPS2
2010 A Discriminative Latent Model of Image Region and Object Tag Correspondence
abstract
We propose a discriminative latent model for annotating images with unaligned object-level textual annotations. Instead of using the bag-of-words image representation currently popular in the computer vision community, our model explicitly captures more intricate relationships underlying visual and textual information. In particular, we model the mapping that translates image regions to annotations. This mapping allows us to relate image regions to their corresponding annotation terms. We also model the overall scene label as latent information. This allows us to cluster test images. Our training data consist of images and their associated annotations. But we do not have access to the ground-truth region-to-annotation mapping or the overall scene label. We develop a novel variant of the latent SVM framework to model them as latent variables. Our experimental results demonstrate the effectiveness of the proposed model compared with other baseline methods.
Yang Wang 0003, Greg Mori
NIPS1
2009 Efficient Human Action Detection Using a Transferable Distance Function
Weilong Yang, Yang Wang 0003, Greg Mori
ACCV (2)2
2009 Max-margin hidden conditional random fields for human action recognition
abstract
We present a new method for classification with structured latent variables. Our model is formulated using the max-margin formalism in the discriminative learning literature. We propose an efficient learning algorithm based on the cutting plane method and decomposed dual optimization. We apply our model to the problem of recognizing human actions from video sequences, where we model a human action as a global root template and a constellation of several “parts”. We show that our model outperforms another similar method that uses hidden conditional random fields, and is comparable to other state-of-the-art approaches. More importantly, our proposed work is quite general and can potentially be applied in a wide variety of vision problems that involve various complex, interdependent latent structures.
Yang Wang 0003, Greg Mori
CVPR1
2009 A Rate Distortion Approach for Semi-Supervised Conditional Random Fields
abstract
We propose a novel information theoretic approach for semi-supervised learning of conditional random fields. Our approach defines a training objective that combines the conditional likelihood on labeled data and the mutual information on unlabeled data. Different from previous minimum conditional entropy semi-supervised discriminative learning methods, our approach can be naturally cast into the rate distortion theory framework in information theory. We analyze the tractability of the framework for structured prediction and present a convergent variational training algorithm to defy the combinatorial explosion of terms in the sum over label configurations. Our experimental results show that the rate distortion approach outperforms standard $l_2$ regularization and minimum conditional entropy regularization on both multi-class classification and sequence labeling problems.
Yang Wang 0003, Gholamreza Haffari, Greg Mori
NIPS1
2009 Human Action Recognition by Semilatent Topic Models
abstract
We propose two new models for human action recognition from video sequences using topic models. Video sequences are represented by a novel "bag-of-words" representation, where each frame corresponds to a "word." Our models differ from previous latent topic models for visual recognition in two major aspects: first of all, the latent topics in our models directly correspond to class labels; second, some of the latent variables in previous topic models become observed in our case. Our models have several advantages over other latent topic models used in visual recognition. First of all, the training is much easier due to the decoupling of the model parameters. Second, it alleviates the issue of how to choose the appropriate number of latent topics. Third, it achieves much better performance by utilizing the information provided by the class labels in the training set. We present action classification results on five different data sets. Our results are either comparable to, or significantly better than previously published results on these data sets.
Yang Wang 0003, Greg Mori
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Multiple Tree Models for Occlusion and Spatial Constraints in Human Pose Estimation
Yang Wang 0003, Greg Mori
ECCV (3)1
2008 Boosting with incomplete information
abstract
In real-world machine learning problems, it is very common that part of the input feature vector is incomplete: either not available, missing, or corrupted. In this paper, we present a boosting approach that integrates features with incomplete information and those with complete information to form a strong classifier. By introducing hidden variables to model missing information, we form loss functions that combine fully labeled data with partially labeled data to effectively learn normalized and unnormalized models. The primal problems of the proposed optimization problems with these loss functions are provided to show their close relationship and the motivations behind them. We use auxiliary functions to bound the change of the loss functions and derive explicit parameter update rules for the learning algorithms. We demonstrate encouraging results on two real-world problems --- visual object recognition in computer vision and named entity recognition in natural language processing --- to show the effectiveness of the proposed boosting approach.
Gholamreza Haffari, Yang Wang 0003, Greg Mori, Feng Jiao
ICML2
2008 Learning a discriminative hidden part model for human action recognition
abstract
We present a discriminative part-based approach for human action recognition from video sequences using motion features. Our model is based on the recently proposed hidden conditional random field~(hCRF) for object recognition. Similar to hCRF for object recognition, we model a human action by a flexible constellation of parts conditioned on image observations. Different from object recognition, our model combines both large-scale global features and local patch features to distinguish various actions. Our experimental results show that our model is comparable to other state-of-the-art approaches in action recognition. In particular, our experimental results demonstrate that combining large-scale global features and local patch features performs significantly better than directly applying hCRF on local patches alone.
Yang Wang 0003, Greg Mori
NIPS1
2006 Unsupervised Discovery of Action Classes
abstract
In this paper we consider the problem of describing the action being performed by human figures in still images. We will attack this problem using an unsupervised learning approach, attempting to discover the set of action classes present in a large collection of training images. These action classes will then be used to label test images. Our approach uses the coarse shape of the human figures to match pairs of images. The distance between a pair of images is computed using a linear programming relaxation technique. This is a computationally expensive process, and we employ a fast pruning method to enable its use on a large collection of images. Spectral clustering is then performed using the resulting distances. We present clustering and image labeling results on a variety of datasets.
Yang Wang 0003, Hao Jiang 0007, Mark S. Drew, Ze-Nian Li, Greg Mori
CVPR (2)1
2005 Fast Krylov Methods for N-Body Learning
abstract
This paper addresses the issue of numerical computation in machine learning domains based on similarity metrics, such as kernel methods, spectral techniques and Gaussian processes. It presents a general solution strategy based on Krylov subspace iteration and fast N-body learning methods. The experiments show significant gains in computation and storage on datasets arising in image segmentation, object detection and dimensionality reduction. The paper also presents theoretical bounds on the stability of these methods.
Nando de Freitas, Yang Wang 0003, Maryam Mahdaviani, Dustin Lang
NIPS2