VLDB 2026 Research / reviewers in the wild / expert
Tieniu Tan
dblp:t/TieniuTan · also T. N. Tan
· DBLP profile ↗
473ranked-venue papers
32as first author
68since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 310 · 26 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 296 · 21 first-author · 29 since 2021Security and privacy · 23 · 1 since 2021Human-computer interaction and ubiquitous computing · 19Databases, data management, data science and information retrieval · 10Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSystems, architecture and hardware · 3Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks (Abstract Reprint)abstractDeep learning-based face recognition models are vulnerable to adversarial attacks. In contrast to general noises, the presence of imperceptible adversarial noises can lead to catastrophic errors in deep face recognition models. The primary difference between adversarial noise and general noise lies in its specificity. Adversarial attack methods give rise to noises tailored to the characteristics of the individual image and recognition model at hand. Diverse samples and recognition models can engender specific adversarial noise patterns, which pose significant challenges for adversarial defense. Addressing this challenge in the realm of face recognition presents a more formidable endeavor due to the inherent nature of face recognition as an open set task. In order to tackle this challenge, it is imperative to employ customized processing for each individual input sample. Drawing inspiration from the biological immune system, which can identify and respond to various threats, this paper aims to create an artificial immune system to provide adversarial defense for face recognition. The proposed defense model incorporates the principles of antibody cloning, mutation, selection, and memory mechanisms to generate a distinct antibody for each input sample, wherein the term antibody refers to a specialized noise removal manner. Furthermore, we introduce a self-supervised adversarial training mechanism that serves as a simulated rehearsal of immune system invasions. Extensive experimental results demonstrate the efficacy of the proposed method, surpassing state-of-the-art adversarial defense methods. Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
AAAI | 7 |
| 2026 | Delphi: A Neuro-Symbolic Framework for Individualized, Safe and Interpretable Treatment RecommendationabstractClinical reinforcement learning (RL) holds promise for treatment recommendation but remains hindered by black-box decision processes, limited safety guarantees, and lack of individualized reasoning. We introduce Delphi Engine, the first fully trainable neuro-symbolic causal RL framework for dynamic treatment planning, designed to answer three core clinical questions in real time: Why this action? Why is it safe? Why for this patient? Specifically, Delphi integrates: (1) causality-aware state modeling using discretized physiological variables and subtype-specific causal graphs; (2) adaptive symbolic rule constraints, combining clinical guidelines and behavior-derived rules into soft differentiable logic; and (3) interpretable decision fusion, where actions are selected based on joint neural-symbolic Q-values and explained via structured LLM-based justifications. We evaluate Delphi on the MIMIC-III sepsis cohort using both standard off-policy evaluations (WIS↑1.47, DR↑1.29, RMSE↓0.207) and the first blinded physician evaluation of an explainable RL system in healthcare. Delphi consistently outperforms historical physicians' treatments in safety (+10.4%), understandability (+8.9%), and adoption rate (+5.75%) across six clinical axes. These results highlight Delphi’s potential as a safe, interpretable, and patient-specific AI assistant for critical care medicine. Muchan Tao, Yuqi Fang, Caifeng Shan, Tieniu Tan |
AAAI | 5 |
| 2026 | What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test TimeabstractTest-Time Reinforcement Learning (TTRL) enables Large Language Models (LLMs) to enhance reasoning capabilities on unlabeled test streams by deriving pseudo-rewards from majority voting consensus.However, existing TTRL methods rely exclusively on positive pseudo-labeling strategies.Such reliance becomes vulnerable under challenging scenarios where answer distributions are highly dispersed, resulting in weak consensus that inadvertently reinforces incorrect trajectories as supervision signals.In this paper, we propose SCRL (Selective-Complementary Reinforcement Learning), a robust test-time reinforcement learning framework that effectively mitigates label noise amplification.SCRL develops Selective Positive Pseudo-Labeling, which enforces strict consensus criteria to filter unreliable majorities.Complementarily, SCRL introduces Entropy-Gated Negative Pseudo-Labeling, the first negative supervision mechanism in TTRL, to reliably prune incorrect trajectories based on generation uncertainty.Extensive experiments on multiple reasoning benchmarks demonstrate that SCRL achieves substantial improvements over baselines, while maintaining robust generalization and training stability under constrained rollout budgets.Our code is available at https://github.com/Jasper Jian Liang 0001, Yanbo Wang 0004, Shuo Lu, Ran He 0001, Tieniu Tan |
ACL (1) | 6 |
| 2026 | CAS-AIR-3D: A Large-scale Low-quality Multi-modal Face Database
Qi Li 0005, Xiaoxiao Dong, Weining Wang 0001, Zhenan Sun, Tieniu Tan, Caifeng Shan |
Int. J. Comput. Vis. | 5 |
| 2026 | Dark Miner: Towards combating residuals in concept erasure for text-to-image diffusion models
Zheling Meng, Bo Peng 0002, Xiaochuan Jin, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Neurocomputing | 7 |
| 2026 | Layer-wise contrastive network for unsupervised graph representation learningabstractUnsupervised graph representation learning has emerged as a cornerstone for extracting meaningful insights from complex relational data. While contrastive learning paradigms have achieved notable success, they traditionally rely on stochastic data augmentations to generate multiple input views—a process that often incurs significant computational overhead and sensitivity to augmentation quality. To transcend these limitations, we propose the Layer-wise Contrastive Network (LCN), a novel and efficient paradigm that redefines the construction of contrastive views. Unlike conventional methods that rely on extrinsic data perturbations, LCN exploits the intrinsic architectural hierarchy of Graph Convolutional Networks. By treating distinct neural layers as different views of the same graph instance, we introduce a contrastive objective that enforces consistency between shallow and deep representations. This mechanism not only eliminates the need for expensive augmentation operations but also distills and preserves fundamental node characteristics from the original graph throughout deeper layers. Furthermore, LCN serves as a flexible plug-and-play framework, exemplified by its extension into Wide LCN, which integrates with traditional augmentation-based methods. Extensive evaluations across both transductive and inductive benchmarks demonstrate that our method achieves superior representational robustness and computational efficiency, offering a scalable and principled perspective for future graph-based contrastive learning. Our code is available at https://github.com/XiangluZhu/LCN.git . • We introduce a novel contrastive loss to learn graph representations by contrasting shallow and deep features. • Our method is flexible and can be readily combined with existing graph contrastive learning techniques that utilize data augmentation. • We demonstrate outstanding performance with our method across four benchmark datasets for node classification. Xianglu Zhu, Zhang Zhang 0001, Zilei Wang, Liang Wang 0001, Tieniu Tan |
Neurocomputing | 5 |
| 2025 | Exploring Vacant Classes in Label-Skewed Federated LearningabstractLabel skews, characterized by disparities in local label distribution across clients, pose a significant challenge in federated learning. As minority classes suffer from worse accuracy due to overfitting on local imbalanced data, prior methods often incorporate class-balanced learning techniques during local training. Although these methods improve the mean accuracy across all classes, we observe that vacant classes—referring to categories absent from a client's data distribution—remain poorly recognized. Besides, there is still a gap in the accuracy of local models on minority classes compared to the global model. This paper introduces FedVLS, a novel approach to label-skewed federated learning that integrates both vacant-class distillation and logit suppression simultaneously. Specifically, vacant-class distillation leverages knowledge distillation during local training on each client to retain essential information related to vacant classes from the global model. Moreover, logit suppression directly penalizes network logits for non-label classes, effectively addressing misclassifications in minority classes that may be biased toward majority classes. Extensive experiments validate the efficacy of FedVLS, demonstrating superior performance compared to previous state-of-the-art (SOTA) methods across diverse datasets with varying degrees of label skews. Kuangpu Guo, Yuhe Ding, Jian Liang 0001, Zilei Wang, Ran He 0001, Tieniu Tan |
AAAI | 6 |
| 2025 | Protecting Model Adaptation from Trojans in the Unlabeled DataabstractModel adaptation tackles the distribution shift problem with a pre-trained model instead of raw data, which has become a popular paradigm due to its great privacy protection. Existing methods always assume adapting to a clean target domain, overlooking the security risks of unlabeled samples. This paper for the first time explores the potential trojan attacks on model adaptation launched by well-designed poisoning target data. Concretely, we provide two trigger patterns with two poisoning strategies for different prior knowledge owned by attackers. These attacks achieve a high success rate while maintaining the normal performance on clean samples in the test stage. To defend against such backdoor injection, we propose a plug-and-play method named DiffAdapt, which can be seamlessly integrated with existing adaptation algorithms. Experiments across commonly used benchmarks and adaptation methods demonstrate the effectiveness of DiffAdapt. We hope this work will shed light on the safety of transfer learning with unlabeled data. Lijun Sheng, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan |
AAAI | 5 |
| 2025 | Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability TheoryabstractRecently, scaling test-time compute on Large Language Models (LLM) has garnered wide attention. However, there has been limited investigation of how various reasoning prompting strategies perform as scaling. In this paper, we focus on a standard and realistic scaling setting: majority voting. We systematically conduct experiments on 6 LLMs $\times$ 8 prompting strategies $\times$ 6 benchmarks. Experiment results consistently show that as the sampling time and computational overhead increase, complicated prompting strategies with superior initial performance gradually fall behind simple Chain-of-Thought. We analyze this phenomenon and provide theoretical proofs. Additionally, we propose a probabilistic method to efficiently predict scaling performance and identify the best prompting strategy under large sampling times, eliminating the need for resource-intensive inference processes in practical applications. Furthermore, we introduce two ways derived from our theoretical analysis to significantly improve the scaling performance. We hope that our research can promote to re-examine the role of complicated prompting, unleash the potential of simple prompting strategies, and provide new insights for enhancing test-time scaling performance. Code is available at https://github.com/MraDonkey/rethinking_prompting. Yexiang Liu, Zekun Li 0007, Zhi Fang, Nan Xu 0014, Ran He 0001, Tieniu Tan |
ACL (1) | 6 |
| 2025 | SHARP: Steering Hallucination in LVLMs via Representation EngineeringabstractJunfei Wu, Yue Ding, Guofan Liu, Tianze Xia, Ziyue Huang, Dianbo Sui, Qiang Liu, Shu Wu, Liang Wang, Tieniu Tan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Junfei Wu, Yue Ding 0009, Guofan Liu, Tianze Xia, Dianbo Sui, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
EMNLP | 10 |
| 2025 | REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge EditingabstractLarge language model editing methods frequently suffer from overfitting, wherein factual updates can propagate beyond their intended scope, overemphasizing the edited target even when it's contextually inappropriate.To address this challenge, we introduce REACT (Representation Extraction And Controllable Tuning), a unified two-phase framework designed for precise and controllable knowledge editing.In the initial phase, we utilize tailored stimuli to extract latent factual representations and apply Principal Component Analysis with a simple learnbale linear transformation to compute a directional "belief shift" vector for each instance.In the second phase, we apply controllable perturbations to hidden states using the obtained vector with a magnitude scalar, gated by a pre-trained classifier that permits edits only when contextually necessary.Relevant experiments on EVOKE benchmarks demonstrate that REACT significantly reduces overfitting across nearly all evaluation metrics, and experiments on COUNTERFACT and MQuAKE shows that our method preserves balanced basic editing performance (reliability, locality, and generality) under diverse editing scenarios. Haitian Zhong, Yuhuan Liu, Guofan Liu, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
EMNLP | 9 |
| 2025 | Enhancing End-to-End Autonomous Driving with Latent World ModelabstractIn autonomous driving, end-to-end planners directly utilize raw sensor data, enabling them to extract richer scene features and reduce information loss compared to traditional planners. This raises a crucial research question: how can we develop better scene feature representations to fully leverage sensor data in end-to-end driving? Self-supervised learning methods show great success in learning rich feature representations in NLP and computer vision. Inspired by this, we propose a novel self-supervised learning approach using the LAtent World model (LAW) for end-to-end driving. LAW predicts future latent scene features based on current features and ego trajectories. This self-supervised task can be seamlessly integrated into perception-free and perception-based frameworks, improving scene feature learning while optimizing trajectory prediction. LAW achieves state-of-the-art performance across multiple benchmarks, including real-world open-loop benchmark nuScenes, NAVSIM, and simulator-based closed-loop benchmark CARLA. The code will be released. Yingyan Li, Lue Fan, Jiawei He 0002, Yuqi Wang 0001, Yuntao Chen, Zhaoxiang Zhang 0001, Tieniu Tan |
ICLR | 7 |
| 2025 | Breaking Mental Set to Improve Reasoning through Diverse Multi-Agent DebateabstractLarge Language Models (LLMs) have seen significant progress but continue to struggle with persistent reasoning mistakes.
Previous methods of *self-reflection* have been proven limited due to the models’ inherent fixed thinking patterns.
While Multi-Agent Debate (MAD) attempts to mitigate this by incorporating multiple agents, it often employs the same reasoning methods, even though assigning different personas to models. This leads to a "fixed mental set", where models rely on homogeneous thought processes without exploring alternative perspectives.
In this paper, we introduce Diverse Multi-Agent Debate (DMAD), a method that encourages agents to think with distinct reasoning approaches. By leveraging diverse problem-solving strategies, each agent can gain insights from different perspectives, refining its responses through discussion and collectively arriving at the optimal solution. DMAD effectively breaks the limitations of fixed mental sets. We evaluate DMAD against various prompting techniques, including *self-reflection* and traditional MAD, across multiple benchmarks using both LLMs and Multimodal LLMs. Our experiments show that DMAD consistently outperforms other methods, delivering better results than MAD in fewer rounds. Code is available at https://github.com/MraDonkey/DMAD. Yexiang Liu, Jie Cao 0002, Zekun Li 0001, Ran He 0001, Tieniu Tan |
ICLR | 5 |
| 2025 | LoRA-Pro: Are Low-Rank Adapters Properly Optimized?abstractLow-rank adaptation, also known as LoRA, has emerged as a prominent method for parameter-efficient fine-tuning of foundation models.
Despite its computational efficiency, LoRA still yields inferior performance compared to full fine-tuning.
In this paper, we first uncover a fundamental connection between the optimization processes of LoRA and full fine-tuning: using LoRA for optimization is mathematically equivalent to full fine-tuning using a low-rank gradient for parameter updates.
And this low-rank gradient can be expressed in terms of the gradients of the two low-rank matrices in LoRA.
Leveraging this insight, we introduce LoRA-Pro, a method that enhances LoRA's performance by strategically adjusting the gradients of these low-rank matrices.
This adjustment allows the low-rank gradient to more accurately approximate the full fine-tuning gradient, thereby narrowing the performance gap between LoRA and full fine-tuning.
Furthermore, we theoretically derive the optimal solutions for adjusting the gradients of the low-rank matrices, applying them during fine-tuning in LoRA-Pro.
We conduct extensive experiments across natural language understanding, dialogue generation, mathematical reasoning, code generation, and image classification tasks, demonstrating that LoRA-Pro substantially improves LoRA's performance, effectively narrowing the gap with full fine-tuning.
Our code is publicly available at https://github.com/mrflogs/LoRA-Pro. Zhengbo Wang, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan |
ICLR | 5 |
| 2025 | TEST-V: TEst-time Support-set Tuning for Zero-shot Video ClassificationabstractRecently, adapting Vision Language Models (VLMs) to zero-shot visual classification by tuning class embedding with a few prompts (Test-time Prompt Tuning, TPT) or replacing class names with generated visual samples (support-set) has shown promising results. However, TPT cannot avoid the semantic gap between modalities while the support-set cannot be tuned. To this end, we draw on each other's strengths and propose a novel framework, namely TEst-time Support-set Tuning for zero-shot Video Classification (TEST-V). It first dilates the support-set with multiple prompts (Multi-prompting Support-set Dilation, MSD) and then erodes the support-set via learnable weights to mine key cues dynamically (Temporal-aware Support-set Erosion, TSE). Specifically, i) MSD expands the support samples for each class based on multiple prompts inquired from LLMs to enrich the diversity of the support-set. ii) TSE tunes the support-set with factorized learnable weights according to the temporal prediction consistency in a self-supervised manner to dig pivotal supporting cues for each class. TEST-V achieves state-of-the-art results across four benchmarks and shows good interpretability. Rui Yan 0010, Hongyu Qu, Xiaoyu Du 0002, Jinhui Tang 0001, Tieniu Tan |
IJCAI | 7 |
| 2025 | BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language ModelsabstractRecently, leveraging pre-trained vision-language models (VLMs) for building vision-language-action (VLA) models has emerged as a promising approach to effective robot manipulation learning. However, only few methods incorporate 3D signals into VLMs for action prediction, and they do not fully leverage the spatial structure inherent in 3D data, leading to low data efficiency. In this paper, we introduce a new paradigm for constructing 3D VLAs. Specifically, we first pre-train the VLM backbone to take 2D images as input and produce 2D heatmaps as output. Using this pre-trained VLM as the backbone, we then fine-tune the entire VLA model while maintaining alignment between inputs and outputs by: (1) projecting raw point cloud inputs into multi-view images, and (2) predicting heatmaps before generating the final action. Extensive experiments show that the resulting model, BridgeVLA, can learn 3D manipulation both efficiently and effectively. BridgeVLA outperforms state-of-the-art baselines across three simulation benchmarks. In RLBench, it improves the average success rate from 81.4\% to 88.2\%. In COLOSSEUM, it demonstrates significantly better performance in challenging generalization settings, boosting the average success rate from 56.7\% to 64.0\%. In GemBench, it surpasses all the comparing baseline methods in terms of average success rate. In real-robot experiments, BridgeVLA outperforms a state-of-the-art baseline method by 32\% on average. It generalizes robustly in multiple out-of-distribution settings, including visual disturbances and unseen instructions. Remarkably, it is able to achieve a success rate of 95.4\% on 10+ tasks with only 3 trajectories per task, while other VLA methods such as $\pi_{0}$ fail completely. Project Website: https://bridgevla.github.io/. Peiyan Li 0001, Xiangnan Wu, Yan Huang 0008, Liang Wang 0001, Tao Kong, Tieniu Tan |
NeurIPS | 9 |
| 2025 | The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language ModelsabstractTest-time adaptation (TTA) methods have gained significant attention for enhancing the performance of vision-language models (VLMs) such as CLIP during inference, without requiring additional labeled data. However, current TTA researches generally suffer from major limitations such as duplication of baseline results, limited evaluation metrics, inconsistent experimental settings, and insufficient analysis. These problems hinder fair comparisons between TTA methods and make it difficult to assess their practical strengths and weaknesses. To address these challenges, we introduce TTA-VLM, a comprehensive benchmark for evaluating TTA methods on VLMs. Our benchmark implements 8 episodic TTA and 7 online TTA methods within a unified and reproducible framework, and evaluates them across 15 widely used datasets. Unlike prior studies focused solely on CLIP, we extend the evaluation to SigLIP—a model trained with a Sigmoid loss—and include training-time tuning methods such as CoOp, MaPLe, and TeCoA to assess generality. Beyond classification accuracy, TTA-VLM incorporates various evaluation metrics, including robustness, calibration, out-of-distribution detection, and stability, enabling a more holistic assessment of TTA methods. Through extensive experiments, we find that 1) existing TTA methods produce limited gains compared to the previous pioneering work; 2) current TTA methods exhibit poor collaboration with training-time fine-tuning methods; 3) accuracy gains frequently come at the cost of reduced model trustworthiness. We release TTA-VLM to provide fair comparison and comprehensive evaluation of TTA methods for VLMs, and we hope it encourages the community to develop more reliable and generalizable TTA strategies. The code is available in https://github.com/TomSheng21/tta-vlm. Lijun Sheng, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan |
NeurIPS | 5 |
| 2025 | Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual DrawingabstractAs textual reasoning with large language models (LLMs) has advanced significant, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods primarily approach multimodal reasoning in a straightforward, text-centric manner, where both reasoning and answer derivation are conducted purely through text, with the only difference being the presence of multimodal input. As a result, these methods often encounter fundamental limitations in spatial reasoning tasks that demand precise geometric understanding and continuous spatial tracking\textemdash capabilities that humans achieve through mental visualization and manipulation. To address the limitations, we propose drawing to reason in space, a novel paradigm that enables LVLMs to reason through elementary drawing operations in the visual space. By equipping models with basic drawing operations including annotating bounding boxes and drawing auxiliary lines, we empower them to express and analyze spatial relationships through direct visual manipulation, meanwhile avoiding the performance ceiling imposed by specialized perception tools in previous tool-integrated reasoning approaches. To cultivate this capability, we develop a three-stage training framework: cold-start training with synthetic data to establish basic drawing abilities, reflective rejection sampling to enhance self-reflection behaviors, and reinforcement learning to directly optimize for target rewards. Extensive experiments demonstrate that our model, named \textsc{Spark}, consistently outperforms existing methods across diverse spatial reasoning benchmarks involving maze navigation, static spatial reasoning, video-based reasoning and multi-view-based reasoning tasks, with an average improvement of 11.5\%. Ablation studies reveal the critical role of each training stage, with reflective rejection sampling particularly enhancing the model's self-correction capabilities and reasoning potential. Junfei Wu, Jian Guan 0002, Kaituo Feng, Qiang Liu 0006, Liang Wang 0001, Wei Wu 0014, Tieniu Tan |
NeurIPS | 8 |
| 2025 | Concept Corrector: Erase Concepts on the Fly for Text-to-Image Diffusion Models
Zheling Meng, Bo Peng 0002, Xiaochuan Jin, Yueming Lyu, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
PRCV (5) | 7 |
| 2025 | A Comprehensive Survey on Test-Time Adaptation Under Distribution Shifts
Jian Liang 0001, Ran He 0001, Tieniu Tan |
Int. J. Comput. Vis. | 3 |
| 2025 | Correction: Open-Vocabulary Text-Driven Human Image Generation
Kaiduo Zhang, Muyi Sun, Jianxin Sun 0003, Kunbo Zhang, Zhenan Sun, Tieniu Tan |
Int. J. Comput. Vis. | 6 |
| 2024 | AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head SynthesisabstractAudio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies. However, when it comes to few-shot talking head generation, a practical scenario where only few seconds of talking video is available for one identity, two limitations emerge: 1) they either have no base model, which serves as a facial prior for fast convergence, or ignore the importance of audio when building the prior; 2) most of them overlook the degree of correlation between different face regions and audio, e.g., mouth is audio related, while ear is audio independent. In this paper, we present Audio Enhanced Neural Radiance Field (AE-NeRF) to tackle the above issues, which can generate realistic portraits of a new speaker with few-shot dataset. Specifically, we introduce an Audio Aware Aggregation module into the feature fusion stage of the reference scheme, where the weight is determined by the similarity of audio between reference and target image. Then, an Audio-Aligned Face Generation strategy is proposed to model the audio related and audio independent regions respectively, with a dual-NeRF framework. Extensive experiments have shown AE-NeRF surpasses the state-of-the-art on image fidelity, audio-lip synchronization, and generalization ability, even in limited training set or training iterations. Wei Wang 0025, Bo Peng 0002, Yingya Zhang, Jing Dong 0003, Tieniu Tan |
AAAI | 7 |
| 2024 | A Hard-to-Beat Baseline for Training-free CLIP-based AdaptationabstractContrastive Language-Image Pretraining (CLIP) has gained popularity for its remarkable zero-shot capacity.
Recent research has focused on developing efficient fine-tuning methods, such as prompt learning and adapter, to enhance CLIP's performance in downstream tasks.
However, these methods still require additional training time and computational resources, which is undesirable for devices with limited resources.
In this paper, we revisit a classical algorithm, Gaussian Discriminant Analysis (GDA), and apply it to the downstream classification of CLIP.
Typically, GDA assumes that features of each class follow Gaussian distributions with identical covariance.
By leveraging Bayes' formula, the classifier can be expressed in terms of the class means and covariance, which can be estimated from the data without the need for training.
To integrate knowledge from both visual and textual modalities, we ensemble it with the original zero-shot classifier within CLIP.
Extensive results on 17 datasets validate that our method surpasses or achieves comparable results with state-of-the-art methods on few-shot classification, imbalanced learning, and out-of-distribution generalization.
In addition, we extend our method to base-to-new generalization and unsupervised learning, once again demonstrating its superiority over competing approaches.
Our code is publicly available at https://github.com/mrflogs/ICLR24. Zhengbo Wang, Jian Liang 0001, Lijun Sheng, Ran He 0001, Zilei Wang, Tieniu Tan |
ICLR | 6 |
| 2024 | Realistic Unsupervised CLIP Fine-tuning with Universal Entropy OptimizationabstractThe emergence of vision-language models, such as CLIP, has spurred a significant research effort towards their application for downstream supervised learning tasks. Although some previous studies have explored the unsupervised fine-tuning of CLIP, they often rely on prior knowledge in the form of class names associated with ground truth labels. This paper explores a realistic unsupervised fine-tuning scenario, considering the presence of out-of-distribution samples from unknown classes within the unlabeled data. In particular, we focus on simultaneously enhancing out-of-distribution detection and the recognition of instances associated with known classes. To tackle this problem, we present a simple, efficient, and effective approach called Universal Entropy Optimization (UEO). UEO leverages sample-level confidence to approximately minimize the conditional entropy of confident instances and maximize the marginal entropy of less confident instances. Apart from optimizing the textual prompt, UEO incorporates optimization of channel-wise affine transformations within the visual branch of CLIP. Extensive experiments across 15 domains and 4 different types of prior knowledge validate the effectiveness of UEO compared to baseline methods. The code is at https://github.com/tim-learn/UEO. Jian Liang 0001, Lijun Sheng, Zhengbo Wang, Ran He 0001, Tieniu Tan |
ICML | 5 |
| 2024 | Connecting the Dots: Collaborative Fine-tuning for Black-Box Vision-Language ModelsabstractWith the emergence of pretrained vision-language models (VLMs), considerable efforts have been devoted to fine-tuning them for downstream tasks. Despite the progress made in designing efficient fine-tuning methods, such methods require access to the model’s parameters, which can be challenging as model owners often opt to provide their models as a black box to safeguard model ownership. This paper proposes a Collaborative Fine-Tuning (CraFT) approach for fine-tuning black-box VLMs to downstream tasks, where one only has access to the input prompts and the output predictions of the model. CraFT comprises two modules, a prompt generation module for learning text prompts and a prediction refinement module for enhancing output predictions in residual style. Additionally, we introduce an auxiliary prediction-consistent loss to promote consistent optimization across these modules. These modules are optimized by a novel collaborative training algorithm. Extensive experiments on few-shot classification over 15 datasets demonstrate the superiority of CraFT. The results show that CraFT achieves a decent gain of about 12% with 16-shot datasets and only 8,000 queries. Moreover, CraFT trains faster and uses only about 1/80 of the memory footprint for deployment, while sacrificing only 1.62% compared to the white-box method. Our code is publicly available at https://github.com/mrflogs/CraFT. Zhengbo Wang, Jian Liang 0001, Ran He 0001, Zilei Wang, Tieniu Tan |
ICML | 5 |
| 2024 | DTS-TPT: Dual Temporal-Sync Test-time Prompt Tuning for Zero-shot Activity Recognition
Rui Yan 0010, Hongyu Qu, Xiangbo Shu, Jinhui Tang 0001, Tieniu Tan |
IJCAI | 6 |
| 2024 | VLKEB: A Large Vision-Language Model Knowledge Editing BenchmarkabstractRecently, knowledge editing on large language models (LLMs) has received considerable attention. Compared to this, editing Large Vision-Language Models (LVLMs) faces extra challenges from diverse data modalities and complicated model components, and data for LVLMs editing are limited. The existing LVLM editing benchmark, which comprises three metrics (Reliability, Locality, and Generality), falls short in the quality of synthesized evaluation images and cannot assess whether models apply edited knowledge in relevant content. Therefore, we employ more reliable data collection methods to construct a new Large $\textbf{V}$ision-$\textbf{L}$anguage Model $\textbf{K}$nowledge $\textbf{E}$diting $\textbf{B}$enchmark, $\textbf{VLKEB}$, and extend the Portability metric for more comprehensive evaluation. Leveraging a multi-modal knowledge graph, our image data are bound with knowledge entities. This can be further used to extract entity-related knowledge, which constitutes the base of editing data. We conduct experiments of different editing methods on five LVLMs, and thoroughly analyze how do they impact the models. The results reveal strengths and deficiencies of these methods and hopefully provide insights for future research. The codes and dataset are available at: https://github.com/VLKEB/VLKEB. Haitian Zhong, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
NeurIPS | 7 |
| 2024 | ST-SBV: Spatial-Temporal Self-Blended Videos for Deepfake Detection
Weinan Guan, Wei Wang 0025, Bo Peng 0002, Jing Dong 0003, Tieniu Tan |
PRCV (5) | 5 |
| 2024 | Artifact feature purification for cross-domain detection of AI-generated images
Zheling Meng, Bo Peng 0002, Jing Dong 0003, Tieniu Tan, Haonan Cheng |
Comput. Vis. Image Underst. | 4 |
| 2024 | Artificial Immune System of Secure Face Recognition Against Adversarial Attacks
Yunlong Wang 0003, Yuhao Zhu 0003, Yongzhen Huang, Zhenan Sun, Qi Li 0005, Tieniu Tan |
Int. J. Comput. Vis. | 7 |
| 2024 | Open-Vocabulary Text-Driven Human Image Generation
Kaiduo Zhang, Muyi Sun, Jianxin Sun 0003, Kunbo Zhang, Zhenan Sun, Tieniu Tan |
Int. J. Comput. Vis. | 6 |
| 2023 | Semantic 3D-Aware Portrait Synthesis and Manipulation Based on Compositional Neural Radiance FieldabstractRecently 3D-aware GAN methods with neural radiance field have developed rapidly. However, current methods model the whole image as an overall neural radiance field, which limits the partial semantic editability of synthetic results. Since NeRF renders an image pixel by pixel, it is possible to split NeRF in the spatial dimension. We propose a Compositional Neural Radiance Field (CNeRF) for semantic 3D-aware portrait synthesis and manipulation. CNeRF divides the image by semantic regions and learns an independent neural radiance field for each region, and finally fuses them and renders the complete image. Thus we can manipulate the synthesized semantic regions independently, while fixing the other parts unchanged. Furthermore, CNeRF is also designed to decouple shape and texture within each semantic region. Compared to state-of-the-art 3D-aware GAN methods, our approach enables fine-grained semantic region manipulation, while maintaining high-quality 3D-consistent synthesis. The ablation studies show the effectiveness of the structure and loss function used by our method. In addition real image inversion and cartoon portrait 3D editing experiments demonstrate the application potential of our method. Tianxiang Ma, Bingchuan Li, Jing Dong 0003, Tieniu Tan |
AAAI | 5 |
| 2023 | CFFT-GAN: Cross-Domain Feature Fusion Transformer for Exemplar-Based Image TranslationabstractExemplar-based image translation refers to the task of generating images with the desired style, while conditioning on certain input image. Most of the current methods learn the correspondence between two input domains and lack the mining of information within the domain. In this paper, we propose a more general learning approach by considering two domain features as a whole and learning both inter-domain correspondence and intra-domain potential information interactions. Specifically, we propose a Cross-domain Feature Fusion Transformer (CFFT) to learn inter- and intra-domain feature fusion. Based on CFFT, the proposed CFFT-GAN works well on exemplar-based image translation. Moreover, CFFT-GAN is able to decouple and fuse features from multiple domains by cascading CFFT modules. We conduct rich quantitative and qualitative experiments on several image translation tasks, and the results demonstrate the superiority of our approach compared to state-of-the-art methods. Ablation studies show the importance of our proposed CFFT. Application experimental results reflect the potential of our method. Tianxiang Ma, Bingchuan Li, Wei Liu 0035, Miao Hua, Jing Dong 0003, Tieniu Tan |
AAAI | 6 |
| 2023 | Notice of Removal: Exploiting Semantic Attributes for Transductive Zero-Shot LearningabstractRemoved. Zhengbo Wang, Jian Liang 0001, Zilei Wang, Tieniu Tan |
ICASSP | 4 |
| 2023 | Free Lunch for Domain Adversarial Training: Environment Label Smoothing
Yifan Zhang 0004, Xue Wang 0010, Jian Liang 0001, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001, Tieniu Tan |
ICLR | 7 |
| 2023 | AdaNPC: Exploring Non-Parametric Classifier for Test-Time AdaptationabstractMany recent machine learning tasks focus to develop models that can generalize to unseen distributions. Domain generalization (DG) has become one of the key topics in various fields. Several literatures show that DG can be arbitrarily hard without exploiting target domain information. To address this issue, test-time adaptive (TTA) methods are proposed. Existing TTA methods require offline target data or extra sophisticated optimization procedures during the inference stage. In this work, we adopt Non-Parametric Classifier to perform the test-time Adaptation (AdaNPC). In particular, we construct a memory that contains the feature and label pairs from training domains. During inference, given a test instance, AdaNPC first recalls $k$ closed samples from the memory to vote for the prediction, and then the test feature and predicted label are added to the memory. In this way, the sample distribution in the memory can be gradually changed from the training distribution towards the test distribution with very little extra computation cost. We theoretically justify the rationality behind the proposed method. Besides, we test our model on extensive numerical experiments. AdaNPC significantly outperforms competitive baselines on various DG benchmarks. In particular, when the adaptation target is a series of domains, the adaptation accuracy of AdaNPC is $50$% higher than advanced TTA methods. Yifan Zhang 0004, Xue Wang 0010, Kexin Jin, Kun Yuan 0001, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001, Tieniu Tan |
ICML | 8 |
| 2023 | OneNet: Enhancing Time Series Forecasting Models under Concept Drift by Online EnsemblingabstractOnline updating of time series forecasting models aims to address the concept drifting problem by efficiently updating forecasting models based on streaming data. Many algorithms are designed for online time series forecasting, with some exploiting cross-variable dependency while others assume independence among variables. Given every data assumption has its own pros and cons in online time series modeling, we propose **On**line **e**nsembling **Net**work (**OneNet**). It dynamically updates and combines two models, with one focusing on modeling the dependency across the time dimension and the other on cross-variate dependency. Our method incorporates a reinforcement learning-based approach into the traditional online convex programming framework, allowing for the linear combination of the two models with dynamically adjusted weights. OneNet addresses the main shortcoming of classical online learning methods that tend to be slow in adapting to the concept drift. Empirical results show that OneNet reduces online forecasting error by more than $\mathbf{50}\\%$ compared to the State-Of-The-Art (SOTA) method. Yifan Zhang 0004, Qingsong Wen, Xue Wang 0010, Liang Sun 0001, Zhang Zhang 0001, Liang Wang 0001, Rong Jin 0001, Tieniu Tan |
NeurIPS | 9 |
| 2023 | End-to-End Alternating Optimization for Real-World Blind Super Resolution
Zhengxiong Luo 0001, Yan Huang 0008, Liang Wang 0001, Tieniu Tan |
Int. J. Comput. Vis. | 5 |
| 2023 | Dual-focus transfer network for zero-shot learning
Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001, Tieniu Tan |
Neurocomputing | 5 |
| 2023 | Few-shot learning with unsupervised part discovery and part-aligned similarity
Zhang Zhang 0001, Wei Wang 0025, Liang Wang 0001, Zilei Wang, Tieniu Tan |
Pattern Recognit. | 6 |
| 2023 | Temporal sparse adversarial attack on sequence-based gait recognition
Ziwen He, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Pattern Recognit. | 4 |
| 2023 | Learning Domain Invariant Representations for Generalizable Person Re-IdentificationabstractGeneralizable person Re-Identification (ReID) aims to learn ready-to-use cross-domain representations for direct cross-data evaluation, which has attracted growing attention in the recent computer vision (CV) community. In this work, we construct a structural causal model (SCM) among identity labels, identity-specific factors (clothing/shoes color etc.), and domain-specific factors (background, viewpoints etc.). According to the causal analysis, we propose a novel Domain Invariant Representation Learning for generalizable person Re-Identification (DIR-ReID) framework. Specifically, we propose to disentangle the identity-specific and domain-specific factors into two independent feature spaces, based on which an effective backdoor adjustment approximate implementation is proposed for serving as a causal intervention towards the SCM. Extensive experiments have been conducted, showing that DIR-ReID outperforms state-of-the-art (SOTA) methods on large-scale domain generalization (DG) ReID benchmarks. Yifan Zhang 0004, Zhang Zhang 0001, Da Li 0003, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Image Process. | 6 |
| 2022 | Generalizable Person Re-identification via Self-Supervised Batch Norm Test-Time AdaptionabstractIn this paper, we investigate the generalization problem of person re-identification (re-id), whose major challenge is the distribution shift on an unseen domain. As an important tool of regularizing the distribution, batch normalization (BN) has been widely used in existing methods. However, they neglect that BN is severely biased to the training domain and inevitably suffers the performance drop if directly generalized without being updated. To tackle this issue, we propose Batch Norm Test-time Adaption (BNTA), a novel re-id framework that applies the self-supervised strategy to update BN parameters adaptively. Specifically, BNTA quickly explores the domain-aware information within unlabeled target data before inference, and accordingly modulates the feature distribution normalized by BN to adapt to the target domain. This is accomplished by two designed self-supervised auxiliary tasks, namely part positioning and part nearest neighbor matching, which help the model mine the domain-aware information with respect to the structure and identity of body parts, respectively. To demonstrate the effectiveness of our method, we conduct extensive experiments on three re-id datasets and confirm the superior performance to the state-of-the-art methods. Chenyang Si, Yan Huang 0008, Liang Wang 0001, Tieniu Tan |
AAAI | 5 |
| 2022 | 3D Shape Temporal Aggregation for Video-Based Clothing-Change Person Re-identification
Yan Huang 0008, Shaogang Gong, Liang Wang 0001, Tieniu Tan |
ACCV (5) | 5 |
| 2022 | Cross-Domain Cross-Set Few-Shot Learning via Learning Compact and Aligned Representations
Zhang Zhang 0001, Wei Wang 0115, Liang Wang 0001, Zilei Wang, Tieniu Tan |
ECCV (20) | 6 |
| 2022 | Pointly-Supervised Panoptic Segmentation
Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan |
ECCV (30) | 3 |
| 2022 | Disentangled Federated Learning for Tackling Attributes Skew via Invariant Aggregation and Diversity TransferringabstractAttributes skew hinders the current federated learning (FL) frameworks from consistent optimization directions among the clients, which inevitably leads to performance reduction and unstable convergence. The core problems lie in that: 1) Domain-specific attributes, which are non-causal and only locally valid, are indeliberately mixed into global aggregation. 2) The one-stage optimizations of entangled attributes cannot simultaneously satisfy two conflicting objectives, i.e., generalization and personalization. To cope with these, we proposed disentangled federated learning (DFL) to disentangle the domain-specific and cross-invariant attributes into two complementary branches, which are trained by the proposed alternating local-global optimization independently. Importantly, convergence analysis proves that the FL system can be stably converged even if incomplete client models participate in the global aggregation, which greatly expands the application scope of FL. Extensive experiments verify that DFL facilitates FL with higher performance, better interpretability, and faster convergence rate, compared with SOTA FL methods on both manually synthesized and realistic attributes skew datasets. Zhengquan Luo, Yunlong Wang 0003, Zilei Wang, Zhenan Sun, Tieniu Tan |
ICML | 5 |
| 2022 | GraphDIVE: Graph Classification by Mixture of Diverse ExpertsabstractGraph classification is a challenging research task in many applications across a broad range of domains. Recently, Graph Neural Network (GNN) models have achieved superior performance on various real-world graph datasets. Despite their successes, most of current GNN models largely suffer from the ubiquitous class imbalance problem, which typically results in prediction bias towards majority classes. Although many imbalanced learning methods have been proposed, they mainly focus on regular Euclidean data and cannot well utilize topological structure of graph (non-Euclidean) data. To boost the performance of GNNs and investigate the relationship between topological structure and class imbalance, we propose GraphDIVE, which learns multi-view graph representations and combine multi-view experts (i.e., classifiers). Specifically, multi-view graph representations correspond to the intrinsic diverse graph topological structure characteristics. Extensive experiments on molecular benchmark datasets demonstrate the effectiveness of the proposed approach. Fenyu Hu, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
IJCAI | 6 |
| 2022 | Defeating DeepFakes via Adversarial Visual ReconstructionabstractExisting DeepFake detection methods focus on passive detection, i.e., they detect fake face images by exploiting the artifacts produced during DeepFake manipulation. These detection-based methods have their limitation that they only work for ex-post forensics but cannot erase the negative influences of DeepFakes. In this work, we propose a proactive framework for combating DeepFake before the data manipulations. The key idea is to find a well defined substitute latent representation to reconstruct target facial data, leading the reconstructed face to disable the DeepFake generation. To this end, we invert face images into latent codes with a well trained auto-encoder, and search the adversarial face embeddings in their neighbor with the gradient descent method. Extensive experiments on three typical DeepFake manipulation methods, facial attribute editing, face expression manipulation, and face swapping, have demonstrated the effectiveness of our method in different settings. Ziwen He, Wei Wang 0025, Weinan Guan, Jing Dong 0003, Tieniu Tan |
ACM Multimedia | 5 |
| 2022 | Focal and efficient IOU loss for accurate bounding box regression
Yifan Zhang 0004, Weiqiang Ren, Zhang Zhang 0001, Liang Wang 0001, Tieniu Tan |
Neurocomputing | 6 |
| 2022 | Identifying the key frames: An attention-aware sampling method for action recognition
Wenkai Dong, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan |
Pattern Recognit. | 4 |
| 2022 | MonoPoly: A practical monocular 3D object detector
He Guan, Chunfeng Song, Zhaoxiang Zhang 0001, Tieniu Tan |
Pattern Recognit. | 4 |
| 2022 | Revisiting ensemble adversarial attack
Ziwen He, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Signal Process. Image Commun. | 4 |
| 2021 | GAIA: A Transfer Learning System of Object Detection That Fits Your NeedsabstractTransfer learning with pre-training on large-scale datasets has played an increasingly significant role in computer vision and natural language processing recently. However, as there exist numerous application scenarios that have distinctive demands such as certain latency constraints and specialized data distributions, it is prohibitively expensive to take advantage of large-scale pre-training for per-task requirements. In this paper, we focus on the area of object detection and present a transfer learning system named GAIA, which could automatically and efficiently give birth to customized solutions according to heterogeneous downstream needs. GAIA is capable of providing powerful pre-trained weights, selecting models that conform to downstream demands such as latency constraints and specified data domains, and collecting relevant data for practitioners who have very few datapoints for their tasks. With GAIA, we achieve promising results on COCO, Objects365, Open Images, Caltech, CityPersons, and UODB which is a collection of datasets including KITTI, VOC, WiderFace, DOTA, Clipart, Comic, and more. Taking COCO as an ex-ample, GAIA is able to efficiently produce models covering a wide range of latency from 16ms to 53ms, and yields AP from 38.2 to 46.5 without whistles and bells. To benefit every practitioner in the community of object detection, GAIA is released at https://github.com/GAIA-vision. Xingyuan Bu, Junran Peng, Tieniu Tan, Zhaoxiang Zhang 0001 |
CVPR | 4 |
| 2021 | Locate Then Segment: A Strong Pipeline for Referring Image SegmentationabstractReferring image segmentation aims to segment the objects referred by a natural language expression. Previous methods usually focus on designing an implicit and recurrent feature interaction mechanism to fuse the visual-linguistic features to directly generate the final segmentation mask without explicitly modeling the localization information of the referent instances. To tackle these problems, we view this task from another perspective by decoupling it into a "Locate-Then-Segment" (LTS) scheme. Given a language expression, people generally first perform attention to the corresponding target image regions, then generate a fine segmentation mask about the object based on its context. The LTS first extracts and fuses both visual and textual features to get a cross-modal representation, then applies a cross-model interaction on the visual-textual features to locate the referred object with position prior, and finally generates the segmentation result with a light-weight segmentation network. Our LTS is simple but surprisingly effective. On three popular benchmark datasets, the LTS outperforms all the previous state-of-the-arts methods by a large margin (e.g., +3.2% on RefCOCO+ and +3.4% on RefCOCOg). In addition, our model is more interpretable with explicitly locating the object, which is also proved by visualization experiments. We believe this framework is promising to serve as a strong baseline for referring image segmentation. Ya Jing, Tao Kong, Wei Wang 0115, Liang Wang 0001, Lei Li 0005, Tieniu Tan |
CVPR | 6 |
| 2021 | Rethinking the Heatmap Regression for Bottom-Up Human Pose EstimationabstractHeatmap regression has become the most prevalent choice for nowadays human pose estimation methods. The ground-truth heatmaps are usually constructed via covering all skeletal keypoints by 2D gaussian kernels. The standard deviations of these kernels are fixed. However, for bottom-up methods, which need to handle a large variance of human scales and labeling ambiguities, the current practice seems unreasonable. To better cope with these problems, we propose the scale-adaptive heatmap regression (SAHR) method, which can adaptively adjust the standard deviation for each keypoint. In this way, SAHR is more tolerant of various human scales and labeling ambiguities. However, SAHR may aggravate the imbalance between fore-background samples, which potentially hurts the improvement of SAHR. Thus, we further introduce the weight-adaptive heatmap regression (WAHR) to help balance the fore-background samples. Extensive experiments show that SAHR together with WAHR largely improves the accuracy of bottom-up human pose estimation. As a result, we finally outperform the state-of-the-art model by +1.5AP and achieve 72.0AP on COCO test-dev2017, which is comparable with the performances of most top-down methods. Source codes are available at https://github.com/greatlog/SWAHR-HumanPose. Zhengxiong Luo 0001, Zhicheng Wang 0001, Yan Huang 0008, Liang Wang 0001, Tieniu Tan, Erjin Zhou |
CVPR | 5 |
| 2021 | Learning Instance-level Spatial-Temporal Patterns for Person Re-identificationabstractPerson re-identification (Re-ID) aims to match pedestrians under dis-joint cameras. Most Re-ID methods formulate it as visual representation learning and image search, and its accuracy is consequently affected greatly by the search space. Spatial-temporal information has been proven to be efficient to filter irrelevant negative samples and significantly improve Re-ID accuracy. However, existing spatial-temporal person Re-ID methods are still rough and do not exploit spatial-temporal information sufficiently. In this paper, we propose a novel Instance-level and Spatial-Temporal Disentangled Re-ID method (InSTD), to improve Re-ID accuracy. In our proposed framework, personalized information such as moving direction is explicitly considered to further narrow down the search space. Besides, the spatial-temporal transferring probability is disentangled from joint distribution to marginal distribution, so that outliers can also be well modeled. Abundant experimental analyses are presented, which demonstrates the superiority and provides more insights into our method. The proposed method achieves mAP of 90.8% on Market-1501 and 89.1% on DukeMTMC-reID, improving from the baseline 82.2% and 72.7%, respectively. Besides, in order to provide a better benchmark for person re-identification, we release a cleaned data list of DukeMTMC-reID with this paper: https://github.com/RenMin1991/cleaned-DukeMTMC-reID/ Lingxiao He, Xingyu Liao, Wu Liu 0005, Yunlong Wang 0003, Tieniu Tan |
ICCV | 6 |
| 2021 | Efficient Human Pose Estimation by Learning Deeply Aggregated RepresentationsabstractIn this paper, we propose an efficient human pose estimation network (DANet) by learning deeply aggregated representations. Most existing models explore multi-scale infonnation mainly from features with different spatial sizes. Powerful multi-scale representations usually rely on the cascaded pyramid framework. This framework largely boosts the performance but in the meanwhile makes networks very deep and complex. Instead, we focus on exploiting multi-scale information from layers with different receptive-field sizes and then making full of use this infonnation by improving the fusion method. Specifically, we propose an orthogonal attention block (OAB) and a second-order fusion unit (SFU). The OAB learns multi-scale infonnation from different layers and enhances them by encouraging them to be diverse. The SFU adaptively selects and fuses diverse multi-scale infonnation and suppress the redundant ones. With the help of OAB and SFU, our networks could achieve comparable or even better accuracy with much smaller model complexity. Specifically, our DANet-72 achieves 71.0 in AP score on COCO val2017 with only 1.0G FLOPS. Its speed on a CPU platfonn achieves 58 Persons-Per-Second (PPS). Zhengxiong Luo 0001, Zhicheng Wang 0001, Yuanhao Cai, Guan'an Wang, Liang Wang 0001, Yan Huang 0008, Erjin Zhou, Tieniu Tan, Jian Sun 0001 |
ICME | 8 |
| 2021 | Few-Shot Learning with Part Discovery and Augmentation from Unlabeled ImagesabstractFew-shot learning is a challenging task since only few instances are given for recognizing an unseen class. One way to alleviate this problem is to acquire a strong inductive bias via meta-learning on similar tasks. In this paper, we show that such inductive bias can be learned from a flat collection of unlabeled images, and instantiated as transferable representations among seen and unseen classes. Specifically, we propose a novel part-based self-supervised representation learning scheme to learn transferable representations by maximizing the similarity of an image to its discriminative part. To mitigate the overfitting in few-shot classification caused by data scarcity, we further propose a part augmentation strategy by retrieving extra images from a base dataset. We conduct systematic studies on miniImageNet and tieredImageNet benchmarks. Remarkably, our method yields impressive results, outperforming the previous best unsupervised methods by 7.74% and 9.24% under 5-way 1-shot and 5-way 5-shot settings, which are comparable with state-of-the-art supervised methods. Chenyang Si, Wei Wang 0115, Liang Wang 0001, Zilei Wang, Tieniu Tan |
IJCAI | 6 |
| 2021 | Neighbor-view Enhanced Model for Vision and Language NavigationabstractVision and Language Navigation (VLN) requires an agent to navigate to a target location by following natural language instructions. Most of existing works represent a navigation candidate by the feature of the corresponding single view where the candidate lies in. However, an instruction may mention landmarks out of the single view as references, which might lead to failures of textual-visual matching of existing methods. In this work, we propose a multi-module Neighbor-View Enhanced Model (NvEM) to adaptively incorporate visual contexts from neighbor views for better textual-visual matching. Specifically, our NvEM utilizes a subject module and a reference module to collect contexts from neighbor views. The subject module fuses neighbor views at a global level, and the reference module fuses neighbor objects at a local level. Subjects and references are adaptively determined via attention mechanisms. Our model also includes an action module to utilize the strong orientation guidance (e.g., "turn left'') in instructions. Each module predicts navigation action separately and their weighted sum is used for predicting the final action. Extensive experimental results demonstrate the effectiveness of the proposed method on the R2R and R4R benchmarks against several state-of-the-art navigators, and NvEM even beats some pre-training ones. Our code is available at https://github.com/MarSaKi/NvEM. Dong An 0002, Yuankai Qi, Yan Huang 0008, Qi Wu 0001, Liang Wang 0001, Tieniu Tan |
ACM Multimedia | 6 |
| 2021 | SOGAN: 3D-Aware Shadow and Occlusion Robust GAN for Makeup TransferabstractIn recent years, virtual makeup applications have become more and more popular. However, it is still challenging to propose a robust makeup transfer method in the real-world environment. Current makeup transfer methods mostly work well on good-conditioned clean makeup images, but transferring makeup that exhibits shadow and occlusion is not satisfying. To alleviate it, we propose a novel makeup transfer method, called 3D-Aware Shadow and Occlusion Robust GAN (SOGAN). Given the source and the reference faces, we first fit a 3D face model and then disentangle the faces into shape and texture. In the texture branch, we map the texture to the UV space and design a UV texture generator to transfer the makeup. Since human faces are symmetrical in the UV space, we can conveniently remove the undesired shadow and occlusion from the reference image by carefully designing a Flip Attention Module (FAM). After obtaining cleaner makeup features from the reference image, a Makeup Transfer Module (MTM) is introduced to perform accurate makeup transfer. The qualitative and quantitative experiments demonstrate that our SOGAN not only achieves superior results in shadow and occlusion situations but also performs well in large pose and expression variations. Yueming Lyu, Jing Dong 0003, Bo Peng 0002, Wei Wang 0025, Tieniu Tan |
ACM Multimedia | 5 |
| 2021 | Selective Wavelet Attention Learning for Single Image Deraining
Huaibo Huang, Aijing Yu, Zhenhua Chai, Ran He 0001, Tieniu Tan |
Int. J. Comput. Vis. | 5 |
| 2021 | Learning pose-invariant 3D object reconstruction from single-view images
Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
Neurocomputing | 4 |
| 2021 | Adaptive super-resolution for person re-identification with low-resolution images
Yan Huang 0008, Chunfeng Song, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 5 |
| 2021 | GraphAIR: Graph representation learning with neighborhood aggregation and interaction
Fenyu Hu, Yanqiao Zhu 0001, Weiran Huang 0002, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 6 |
| 2021 | A3GAN: An Attribute-Aware Attentive Generative Adversarial Network for Face AgingabstractFace aging has received significant research attention in recent years. Although great progress has been achieved with the success of Generative Adversarial Networks (GANs) in synthesizing realistic images, most existing GAN-based face aging methods have two main problems: 1) unnatural changes of high-level semantic information due to the insufficient consideration of prior knowledge of input faces, and 2) distortions of low-level image content (e.g. modifications in age-irrelevant regions). In this article, we introduce A3GAN, an Attribute-Aware Attentive face aging model to address the above issues. Facial attribute vectors are regarded as the conditional information and embedded into both the generator and discriminator, encouraging synthesized faces to be faithful to attributes of corresponding inputs. To improve the visual fidelity of generation results, we leverage the attention mechanism to restrict modifications to age-related areas and preserve image details. Unlike previous works with attention modules, we introduce face parsing maps to help the generator distinguish image regions of interest and suppress attention activation elsewhere. Moreover, the wavelet packet transform is employed to capture textural features at multiple scales in the frequency space. Extensive experimental results demonstrate the effectiveness of our model in synthesizing photo-realistic aged face images and achieving state-of-the-art performance on popular datasets. Yunfan Liu 0001, Qi Li 0005, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Learning Aligned Image-Text Representations Using Graph Attentive Relational NetworkabstractImage-text matching aims to measure the similarities between images and textual descriptions, which has made great progress recently. The key to this cross-modal matching task is to build the latent semantic alignment between visual objects and words. Due to the widespread variations of sentence structures, it is very difficult to learn the latent semantic alignment using only global cross-modal features. Many previous methods attempt to learn the aligned image-text representations by the attention mechanism but generally ignore the relationships within textual descriptions which determine whether the words belong to the same visual object. In this paper, we propose a graph attentive relational network (GARN) to learn the aligned image-text representations by modeling the relationships between noun phrases in a text for the identity-aware image-text matching. In the GARN, we first decompose images and texts into regions and noun phrases, respectively. Then a skip graph neural network (skip-GNN) is proposed to learn effective textual representations which are a mixture of textual features and relational features. Finally, a graph attention network is further proposed to obtain the probabilities that the noun phrases belong to the image regions by modeling the relationships between noun phrases. We perform extensive experiments on the CUHK Person Description dataset (CUHK-PEDES), Caltech-UCSD Birds dataset (CUB), Oxford-102 Flowers dataset and Flickr30K dataset to verify the effectiveness of each component in our model. Experimental results show that our approach achieves the state-of-the-art results on these four benchmark datasets. Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Image Process. | 4 |
| 2021 | Meta-USR: A Unified Super-Resolution Network for Multiple Degradation ParametersabstractRecent research on single image super-resolution (SISR) has achieved great success due to the development of deep convolutional neural networks. However, most existing SISR methods merely focus on super-resolution of a single fixed integer scale factor. This simplified assumption does not meet the complex conditions for real-world images which often suffer from various blur kernels or various levels of noise. More importantly, previous methods lack the ability to cope with arbitrary degradation parameters (scale factors, blur kernels, and noise levels) with a single model. A few methods can handle multiple degradation factors, e.g., noninteger scale factors, blurring, and noise, simultaneously within a single SISR model. In this work, we propose a simple yet powerful method termed meta-USR which is the first unified super-resolution network for arbitrary degradation parameters with meta-learning. In Meta-USR, a meta-restoration module (MRM) is proposed to enhance the traditional upscale module with the capability to adaptively predict the weights of the convolution filters for various combinations of degradation parameters. Thus, the MRM can not only upscale the feature maps with arbitrary scale factors but also restore the SR image with different blur kernels and noise levels. Moreover, the lightweight MRM can be placed at the end of the network, which makes it very efficient for iteratively/repeatedly searching the various degradation factors. We evaluate the proposed method through extensive experiments on several widely used benchmark data sets on SISR. The qualitative and quantitative experimental results show the superiority of our Meta-USR. Xuecai Hu, Zhang Zhang 0001, Caifeng Shan, Zilei Wang, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation with only image-level labels saves large human effort to annotate pixel-level labels. Cutting-edge approaches rely on various innovative constraints and heuristic rules to generate the masks for every single image. Although great progress has been achieved by these methods, they treat each image independently and do not take account of the relationships across different images. In this paper, however, we argue that the cross-image relationship is vital for weakly supervised segmentation. Because it connects related regions across images, where supplementary representations can be propagated to obtain more consistent and integral regions. To leverage this information, we propose an end-to-end cross-image affinity module, which exploits pixel-level cross-image relationships with only image-level labels. By means of this, our approach achieves 64.3% and 65.3% mIoU on Pascal VOC 2012 validation and test set respectively, which is a new state-of-the-art result by only using image-level labels for weakly supervised semantic segmentation, demonstrating the superiority of our approach. Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan, Chunfeng Song, Jun Xiao 0005 |
AAAI | 3 |
| 2020 | Pose-Guided Multi-Granularity Attention Network for Text-Based Person SearchabstractText-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surveillance. Extracting visual contents corresponding to the human description is the key to this cross-modal matching problem. Moreover, correlated images and descriptions involve different granularities of semantic relevance, which is usually ignored in previous methods. To exploit the multilevel corresponding visual contents, we propose a pose-guided multi-granularity attention network (PMA). Firstly, we propose a coarse alignment network (CA) to select the related image regions to the global description by a similarity-based attention. To further capture the phrase-related visual body part, a fine-grained alignment network (FA) is proposed, which employs pose information to learn latent semantic alignment between visual body part and textual noun phrase. To verify the effectiveness of our model, we perform extensive experiments on the CUHK Person Description Dataset (CUHK-PEDES) which is currently the only available dataset for text-based person search. Experimental results show that our approach outperforms the state-of-the-art methods by 15 % in terms of the top-1 metric. Ya Jing, Chenyang Si, Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
AAAI | 6 |
| 2020 | Dynamic Graph Representation for Occlusion Handling in BiometricsabstractThe generalization ability of Convolutional neural networks (CNNs) for biometrics drops greatly due to the adverse effects of various occlusions. To this end, we propose a novel unified framework integrated the merits of both CNNs and graphical models to learn dynamic graph representations for occlusion problems in biometrics, called Dynamic Graph Representation (DGR). Convolutional features onto certain regions are re-crafted by a graph generator to establish the connections among the spatial parts of biometrics and build Feature Graphs based on these node representations. Each node of Feature Graphs corresponds to a specific part of the input image and the edges express the spatial relationships between parts. By analyzing the similarities between the nodes, the framework is able to adaptively remove the nodes representing the occluded parts. During dynamic graph matching, we propose a novel strategy to measure the distances of both nodes and adjacent matrixes. In this way, the proposed method is more convincing than CNNs-based methods because the dynamic graph method implies a more illustrative and reasonable inference of the biometrics decision. Experiments conducted on iris and face demonstrate the superiority of the proposed framework, which boosts the accuracy of occluded biometrics recognition by a large margin comparing with baseline methods. Yunlong Wang 0003, Zhenan Sun, Tieniu Tan |
AAAI | 4 |
| 2020 | Instance Guided Proposal Network for Person SearchabstractPerson detection networks have been widely used in person search. These detectors discriminate persons from the background and generate proposals of all the persons from a gallery of scene images for each query. However, such a large number of proposals have a negative influence on the following identity matching process because many distractors are involved. In this paper, we propose a new detection network for person search, named Instance Guided Proposal Network (IGPN), which can learn the similarity between query persons and proposals. Thus, we can decrease proposals according to the similarity scores. To incorporate information of the query into the detection network, we introduce the Siamese region proposal network to Faster-RCNN and we propose improved cross-correlation layers to alleviate the imbalance of parameters distribution. Furthermore, we design a local relation block and a global relation branch to leverage the proposal-proposal relations and query-scene relations, respectively. Extensive experiments show that our method improves the person search performance through decreasing proposals and achieves competitive performance on two large person search benchmark datasets, CUHK-SYSU and PRW. Wenkai Dong, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan |
CVPR | 4 |
| 2020 | Bi-Directional Interaction Network for Person SearchabstractExisting works have designed end-to-end frameworks based on Faster-RCNN for person search. Due to the large receptive fields in deep networks, the feature maps of each proposal, cropped from the stem feature maps, involve redundant context information outside the bounding boxes. However, person search is a fine-grained task which needs accurate appearance information. Such context information can make the model fail to focus on persons, so the learned representations lack the capacity to discriminate various identities. To address this issue, we propose a Siamese network which owns an additional instance-aware branch, named Bi-directional Interaction Network (BINet). During the training phase, in addition to scene images, BINet also takes as inputs person patches which help the model discriminate identities based on human appearance. Moreover, two interaction losses are designed to achieve bi-directional interaction between branches at two levels. The interaction can help the model learn more discriminative features for persons in the scene. At the inference stage, only the major branch is applied, so BINet introduces no additional computation. Extensive experiments on two widely used person search benchmarks, CUHK-SYSU and PRW, have shown that our BINet achieves state-of-the-art results among end-to-end methods without loss of efficiency. Wenkai Dong, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan |
CVPR | 4 |
| 2020 | Learning Integral Objects With Intra-Class Discriminator for Weakly-Supervised Semantic SegmentationabstractImage-level weakly-supervised semantic segmentation (WSSS) aims at learning semantic segmentation by adopting only image class labels. Existing approaches generally rely on class activation maps (CAM) to generate pseudo-masks and then train segmentation models. The main difficulty is that the CAM estimate only covers partial foreground objects. In this paper, we argue that the critical factor preventing to obtain the full object mask is the classification boundary mismatch problem in applying the CAM to WSSS. Because the CAM is optimized by the classification task, it focuses on the discrimination across different image-level classes. However, the WSSS requires to distinguish pixels sharing the same image-level class to separate them into the foreground and the background. To alleviate this contradiction, we propose an efficient end-to-end Intra-Class Discriminator (ICD) framework, which learns intra-class boundaries to help separate the foreground and the background within each image-level class. Without bells and whistles, our approach achieves the state-of-the-art performance of image label based WSSS, with mIoU 68.0% on the VOC 2012 semantic segmentation benchmark, demonstrating the effectiveness of the proposed approach. Junsong Fan, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan |
CVPR | 4 |
| 2020 | Cross-Modal Cross-Domain Moment Alignment Network for Person SearchabstractText-based person search has drawn increasing attention due to its wide applications in video surveillance. However, most of the existing models depend heavily on paired image-text data, which is very expensive to acquire. Moreover, they always face huge performance drop when directly exploiting them to new domains. To overcome this problem, we make the first attempt to adapt the model to new target domains in the absence of pairwise labels, which combines the challenges from both cross-modal (text-based) person search and cross-domain person search. Specially, we propose a moment alignment network (MAN) to solve the cross-modal cross-domain person search task in this paper. The idea is to learn three effective moment alignments including domain alignment (DA), cross-modal alignment (CA) and exemplar alignment (EA), which together can learn domain-invariant and semantic aligned cross-modal representations to improve model generalization. Extensive experiments are conducted on CUHK Person Description dataset (CUHK-PEDES) and Richly Annotated Pedestrian dataset (RAP). Experimental results show that our proposed model achieves the state-of-the-art performances on five transfer tasks. Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
CVPR | 4 |
| 2020 | Large-Scale Object Detection in the Wild From Imbalanced Multi-LabelsabstractTraining with more data has always been the most stable and effective way of improving performance in deep learn-ing era. As the largest object detection dataset so far, OpenImages brings great opportunities and challenges for object detection in general and sophisticated scenarios. However, owing to its semi-automatic collecting and labeling pipeline to deal with the huge data scale, Open Images dataset suffers from label-related problems that objects may explicitly or implicitly have multiple labels and the label distribution is extremely imbalanced. In this work, we quantitatively analyze these label problems and provide a simple but effective solution. We design a concurrent softmax to handle the multi-label problems in object detection and propose a soft-sampling methods with hybrid training scheduler to deal with the label imbalance. Overall, our method yields a dramatic improvement of 3.34 points, leading to the best single model with 60.90 mAP on the public object detection test set of Open Images. And our ensembling result achieves 67.17mAP, which is 4.29 points higher than the first place method last year. Junran Peng, Xingyuan Bu, Ming Sun 0008, Zhaoxiang Zhang 0001, Tieniu Tan |
CVPR | 5 |
| 2020 | Employing Multi-estimations for Weakly-Supervised Semantic Segmentation
Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan |
ECCV (17) | 3 |
| 2020 | Prediction and Recovery for Adaptive Low-Resolution Person Re-Identification
Yan Huang 0008, Zerui Chen, Liang Wang 0001, Tieniu Tan |
ECCV (26) | 5 |
| 2020 | Adversarial Self-supervised Learning for Semi-supervised 3D Action Recognition
Chenyang Si, Xuecheng Nie, Wei Wang 0115, Liang Wang 0001, Tieniu Tan, Jiashi Feng |
ECCV (7) | 5 |
| 2020 | Efficient Super Resolution by Recursive AggregationabstractDeep neural networks have achieved remarkable results on image super resolution (SR), but the efficiency problem of deep SR networks is rarely studied. We experimentally find that many sequentially stacked convolutional blocks in nowadays SR networks are far from being fully optimized, which largely damages their overall efficiency. It indicates that comparable or even better results could be achieved with less but sufficiently optimized blocks. In this paper, we try to construct more efficient SR model via the proposed recursive aggregation network (RAN). It recursively aggregates convolutional blocks in different orders, and avoids too many sequentially stacked blocks. In this way, multiple shortcuts are introduced in RAN, and help gradients easier flow to all inner layers, even for very deep SR networks. As a result, all blocks in RAN can be better optimized, thus RAN can achieve better performance with smaller model size than existing methods. Zhengxiong Luo 0001, Yan Huang 0008, Liang Wang 0001, Tieniu Tan |
ICPR | 5 |
| 2020 | Image Inpainting with Contrastive Relation NetworkabstractImage inpainting faces the challenging issue of the requirements on structure reasonableness and texture coherence. In this paper, we propose a two-stage inpainting framework to address this issue. The basic idea is to address the two requirements in two separate stages. Completed segmentation of the corrupted image is firstly predicted through segmentation reconstruction network, while fine-grained image details are restored in the second stage through an image generator. The two stages are connected in series as the image details are generated under the guidance of completed segmentation map that predicted in the first stage. Specifically, in the second stage, we propose a novel graph-based relation network to model the relationship existed in corrupted image. In relation network, both intra-relationship for pixels in the same semantic region and inter-relationship between different semantic parts are considered, improving the consistency and compatibility of image textures. Besides, contrastive loss is designed to facilitate the relation network training. Such a framework not only simplifies the inpainting problem directly, but also exploits the relationship in corrupted image explicitly. Extensive experiments on various public datasets quantitatively and qualitatively demonstrate the superiority of our approach compared with the state-of-the-art. Xiaoqiang Zhou, Junjie Li 0002, Zilei Wang, Ran He 0001, Tieniu Tan |
ICPR | 5 |
| 2020 | Unfolding the Alternating Optimization for Blind Super ResolutionabstractPrevious methods decompose blind super resolution (SR) problem into two sequential steps: \textit{i}) estimating blur kernel from given low-resolution (LR) image and \textit{ii}) restoring SR image based on estimated kernel. This two-step solution involves two independently trained models, which may not well compatible with each other. Small estimation error of the first step could cause severe performance drop of the second one. While on the other hand, the first step can only utilize limited information from LR image, which makes it difficult to predict highly accurate blur kernel. Towards these issues, instead of considering these two steps separately, we adopt an alternating optimization algorithm, which can estimate blur kernel and restore SR image in a single model. Specifically, we design two convolutional neural modules, namely \textit{Restorer} and \textit{Estimator}. \textit{Restorer} restores SR image based on predicted kernel, and \textit{Estimator} estimates blur kernel with the help of restored SR image. We alternate these two modules repeatedly and unfold this process to form an end-to-end trainable network. In this way, \textit{Estimator} utilizes information from both LR and SR images, which makes the estimation of blur kernel easier. More importantly, \textit{Restorer} is trained with the kernel estimated by \textit{Estimator}, instead of ground-truth kernel, thus \textit{Restorer} could be more tolerant to the estimation error of \textit{Estimator}. Extensive experiments on synthetic datasets and real-world images show that our model can largely outperform state-of-the-art methods and produce more visually favorable results at much higher speed. The source code will be publicly available. Zhengxiong Luo 0001, Yan Huang 0008, Liang Wang 0001, Tieniu Tan |
NeurIPS | 5 |
| 2020 | TFNet: Multi-Semantic Feature Interaction for CTR PredictionabstractThe CTR (Click-Through Rate) prediction plays a central role in the domain of computational advertising and recommender systems. There exists several kinds of methods proposed in this field, such as Logistic Regression (LR), Factorization Machines (FM) and deep learning based methods like Wide&Deep, Neural Factorization Machines (NFM) and DeepFM. However, such approaches generally use the vector-product of each pair of features, which have ignored the different semantic spaces of the feature interactions. In this paper, we propose a novel Tensor-based Feature interaction Network (TFNet) model, which introduces an operating tensor to elaborate feature interactions via multi-slice matrices in multiple semantic spaces. Extensive offline and online experiments show that TFNet: 1) outperforms the competitive compared methods on the typical Criteo and Avazu datasets; 2) achieves large improvement of revenue and click rate in online A/B tests in the largest Chinese App recommender system, Tencent MyApp. Feng Yu 0001, Xueli Yu, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
SIGIR | 6 |
| 2020 | TAGNN: Target Attentive Graph Neural Networks for Session-based RecommendationabstractSession-based recommendation nowadays plays a vital role in many websites, which aims to predict users' actions based on anonymous sessions. There have emerged many studies that model a session as a sequence or a graph via investigating temporal transitions of items in a session. However, these methods compress a session into one fixed representation vector without considering the target items to be predicted. The fixed vector will restrict the representation ability of the recommender model, considering the diversity of target items and users' interests. In this paper, we propose a novel target attentive graph neural network (TAGNN) model for session-based recommendation. In TAGNN, target-aware attention adaptively activates different user interests with respect to varied target items. The learned interest representation vector varies with different target items, greatly improving the expressiveness of the model. Moreover, TAGNN harnesses the power of graph neural networks to capture rich item transitions in sessions. Comprehensive experiments conducted on real-world datasets demonstrate its superiority over state-of-the-art methods. Feng Yu 0001, Yanqiao Zhu 0001, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
SIGIR | 6 |
| 2020 | Disentangled Representation Learning of Makeup Portraits in the Wild
Yi Li 0018, Huaibo Huang, Jie Cao 0002, Ran He 0001, Tieniu Tan |
Int. J. Comput. Vis. | 5 |
| 2020 | A General Framework for Deep Supervised Discrete Hashing
Qi Li 0005, Zhenan Sun, Ran He 0001, Tieniu Tan |
Int. J. Comput. Vis. | 4 |
| 2020 | Kinematic skeleton graph augmented network for human parsing
Jinde Liu, Zhang Zhang 0001, Caifeng Shan, Tieniu Tan |
Neurocomputing | 4 |
| 2020 | Adversarial Cross-Spectral Face Completion for NIR-VIS Face RecognitionabstractNear infrared-visible (NIR-VIS) heterogeneous face recognition refers to the process of matching NIR to VIS face images. Current heterogeneous methods try to extend VIS face recognition methods to the NIR spectrum by synthesizing VIS images from NIR images. However, due to the self-occlusion and sensing gap, NIR face images lose some visible lighting contents so that they are always incomplete compared to VIS face images. This paper models high-resolution heterogeneous face synthesis as a complementary combination of two components: a texture inpainting component and a pose correction component. The inpainting component synthesizes and inpaints VIS image textures from NIR image textures. The correction component maps any pose in NIR images to a frontal pose in VIS images, resulting in paired NIR and VIS textures. A warping procedure is developed to integrate the two components into an end-to-end deep network. A fine-grained discriminator and a wavelet-based discriminator are designed to improve visual quality. A novel 3D-based pose correction loss, two adversarial losses, and a pixel loss are imposed to ensure synthesis results. We demonstrate that by attaching the correction component, we can simplify heterogeneous face synthesis from one-to-many unpaired image translation to one-to-one paired image translation, and minimize the spectral and pose discrepancy during heterogeneous recognition. Extensive experimental results show that our network not only generates high-resolution VIS face images but also facilitates the accuracy improvement of heterogeneous face recognition. Ran He 0001, Jie Cao 0002, Lingxiao Song, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Relational graph neural network for situation recognition
Ya Jing, Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 5 |
| 2020 | Skeleton-based action recognition with hierarchical spatial reasoning and temporal stack learning network
Chenyang Si, Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 5 |
| 2020 | Learning visual relationship and context-aware attention for image captioning
Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Zhiyong Wang 0001, David Dagan Feng, Tieniu Tan |
Pattern Recognit. | 6 |
| 2020 | Graph Sequence Recurrent Neural Network for Vision-Based Freezing of Gait DetectionabstractFreezing of gait (FoG) is one of the most common symptoms of Parkinson's disease (PD), a neurodegenerative disorder which impacts millions of people around the world. Accurate assessment of FoG is critical for the management of PD and to evaluate the efficacy of treatments. Currently, the assessment of FoG requires well-trained experts to perform time-consuming annotations via vision-based observations. Thus, automatic FoG detection algorithms are needed. In this study, we formulate vision-based FoG detection, as a fine-grained graph sequence modelling task, by representing the anatomic joints in each temporal segment with a directed graph, since FoG events can be observed through the motion patterns of joints. A novel deep learning method is proposed, namely graph sequence recurrent neural network (GS-RNN), to characterize the FoG patterns by devising graph recurrent cells, which take graph sequences of dynamic structures as inputs. For the cases of which prior edge annotations are not available, a data-driven based adjacency estimation method is further proposed. To the best of our knowledge, this is one of the first studies on vision-based FoG detection using deep neural networks designed for graph sequences of dynamic structures. Experimental results on more than 150 videos collected from 45 patients demonstrated promising performance of the proposed GS-RNN for FoG detection with an AUC value of 0.90. Kun Hu 0008, Zhiyong Wang 0001, Wei Wang 0115, Kaylena A. Ehgoetz Martens, Liang Wang 0001, Tieniu Tan, Simon J. G. Lewis, David Dagan Feng |
IEEE Trans. Image Process. | 6 |
| 2020 | Deep Unbiased Embedding Transfer for Zero-Shot LearningabstractZero-shot learning aims to recognize objects which do not appear in the training dataset. Previous prevalent mapping-based zero-shot learning methods suffer from the projection domain shift problem due to the lack of image classes in the training stage. In order to alleviate the projection domain shift problem, a deep unbiased embedding transfer (DUET) model is proposed in this paper. The DUET model is composed of a deep embedding transfer (DET) module and an unseen visual feature generation (UVG) module. In the DET module, a novel combined embedding transfer net which integrates the complementary merits of the linear and nonlinear embedding mapping functions is proposed to connect the visual space and semantic space. What's more, the end-to-end joint training process is implemented to train the visual feature extractor and the combined embedding transfer net simultaneously. In the UVG module, a visual feature generator trained with a conditional generative adversarial framework is used to synthesize the visual features of the unseen classes to ease the disturbance of the projection domain shift problem. Furthermore, a quantitative index, namely the score of resistance on domain shift (ScoreRDS), is proposed to evaluate different models regarding their resistance capability on the projection domain shift problem. The experiments on five zero-shot learning benchmarks verify the effectiveness of the proposed DUET model. As demonstrated by the qualitative and quantitative analysis, the unseen class visual feature generation, the combined embedding transfer net and the end-to-end joint training process all contribute to alleviating projection domain shift in zero-shot learning. Zhang Zhang 0001, Liang Wang 0001, Caifeng Shan, Tieniu Tan |
IEEE Trans. Image Process. | 5 |
| 2020 | Binocular Light-Field: Imaging Theory and Occlusion-Robust Depth Perception ApplicationabstractBinocular stereo vision (SV) has been widely used to reconstruct the depth information, but it is quite vulnerable to scenes with strong occlusions. As an emerging computational photography technology, light-field (LF) imaging brings about a novel solution to passive depth perception by recording multiple angular views in a single exposure. In this paper, we explore binocular SV and LF imaging to form the binocular-LF imaging system. An imaging theory is derived by modeling the imaging process and analyzing disparity properties based on the geometrical optics theory. Then an accurate occlusion-robust depth estimation algorithm is proposed by exploiting multibaseline stereo matching cues and defocus cues. The occlusions caused by binocular SV and LF imaging are detected and handled to eliminate the matching ambiguities and outliers. Finally, we develop a binocular-LF database and capture realworld scenes by our binocular-LF system to test the accuracy and robustness. The experimental results demonstrate that the proposed algorithm definitely recovers high quality depth maps with smooth surfaces and precise geometric shapes, which tackles the drawbacks of binocular SV and LF imaging simultaneously. Fei Liu 0031, Shubo Zhou, Yunlong Wang 0003, Guangqi Hou, Zhenan Sun, Tieniu Tan |
IEEE Trans. Image Process. | 6 |
| 2019 | Attention-Aware Sampling via Deep Reinforcement Learning for Action RecognitionabstractDeep learning based methods have achieved remarkable progress in action recognition. Existing works mainly focus on designing novel deep architectures to achieve video representations learning for action recognition. Most methods treat sampled frames equally and average all the frame-level predictions at the testing stage. However, within a video, discriminative actions may occur sparsely in a few frames and most other frames are irrelevant to the ground truth and may even lead to a wrong prediction. As a result, we think that the strategy of selecting relevant frames would be a further important key to enhance the existing deep learning based action recognition. In this paper, we propose an attentionaware sampling method for action recognition, which aims to discard the irrelevant and misleading frames and preserve the most discriminative frames. We formulate the process of mining key frames from videos as a Markov decision process and train the attention agent through deep reinforcement learning without extra labels. The agent takes features and predictions from the baseline model as input and generates importance scores for all frames. Moreover, our approach is extensible, which can be applied to different existing deep learning based action recognition models. We achieve very competitive action recognition performance on two widely used action recognition datasets. Wenkai Dong, Zhaoxiang Zhang 0001, Tieniu Tan |
AAAI | 3 |
| 2019 | Session-Based Recommendation with Graph Neural NetworksabstractThe problem of session-based recommendation aims to predict user actions based on anonymous sessions. Previous methods model a session as a sequence and estimate user representations besides item representations to make recommendations. Though achieved promising results, they are insufficient to obtain accurate user vectors in sessions and neglect complex transitions of items. To obtain accurate item embedding and take complex transitions of items into account, we propose a novel method, i.e. Session-based Recommendation with Graph Neural Networks, SR-GNN for brevity. In the proposed method, session sequences are modeled as graphstructured data. Based on the session graph, GNN can capture complex transitions of items, which are difficult to be revealed by previous conventional sequential methods. Each session is then represented as the composition of the global preference and the current interest of that session using an attention network. Extensive experiments conducted on two real datasets show that SR-GNN evidently outperforms the state-of-the-art session-based recommendation methods consistently. Yuyuan Tang, Yanqiao Zhu 0001, Liang Wang 0001, Xing Xie 0001, Tieniu Tan |
AAAI | 6 |
| 2019 | A Comprehensive Study on Large-Scale Person Retrieval in Real Surveillance ScenariosabstractPerson retrieval is a hot research topic due to its important application potential for public security. Though existing algorithms have achieved impressive progresses on current public datasets, it is still a challenging task in the real surveillance scenarios due to the various viewpoints, pose variations and occlusions. Moreover, few of the existing works study the problem of person retrieval on large-scale gallery set, where lots of distractions may deteriorate the retrieval results heavily. To have a deep understanding on the above challenges, we perform a comprehensive study on current state-of-the-art person retrieval algorithms with a large-scale benchmark in real surveillance scenarios. In the study, two kinds of techniques, i.e., attribute recognition and person re-identification, including eight algorithms, are evaluated at both algorithm level and system level. Here, the system-level evaluations investigate the effects of the combinations of the above algorithms with the module of person detection, where lots of distractions in person detection results pose a big challenge for person retrieval in real scenes. Extensive evaluations with large gallery sizes (up to 243k) and comprehensive analyses are presented in the study, which will guide researchers to develop more advanced algorithms in future. Da Li 0003, Zhang Zhang 0001, Caifeng Shan, Liang Wang 0001, Tieniu Tan |
AVSS | 5 |
| 2019 | Meta-SR: A Magnification-Arbitrary Network for Super-ResolutionabstractRecent research on super-resolution has achieved great success due to the development of deep convolutional neural networks (DCNNs). However, super-resolution of arbitrary scale factor has been ignored for a long time. Most previous researchers regard super-resolution of differentscale factors as independent tasks. They train a specific model for each scale factor which is inefficient in computing, and prior work only take the super-resolution of several integer scale factors into consideration. In this work,we propose a novel method called Meta-SR to firstly solve super-resolution of arbitrary scale factor (including non-integer scale factors) with a single model. In our Meta-SR,the Meta-Upscale Module is proposed to replace the traditional upscale module. For arbitrary scale factor, the Meta-Upscale Module dynamically predicts the weights of the up-scale filters by taking the scale factor as input and use these weights to generate the HR image of arbitrary size. For any low-resolution image, our Meta-SR can continuously zoomin it with arbitrary scale factor by only using a single model.We evaluated the proposed method through extensive experiments on widely used benchmark datasets on single image super-resolution. The experimental results show the superiority of our Meta-Upscale. Xuecai Hu, Haoyuan Mu, Xiangyu Zhang 0005, Zilei Wang, Tieniu Tan, Jian Sun 0001 |
CVPR | 5 |
| 2019 | Distant Supervised Centroid Shift: A Simple and Efficient Approach to Visual Domain AdaptationabstractConventional domain adaptation methods usually resort to deep neural networks or subspace learning to find invariant representations across domains. However, most deep learning methods highly rely on large-size source domains and are computationally expensive to train, while subspace learning methods always have a quadratic time complexity that suffers from the large domain size. This paper provides a simple and efficient solution, which could be regarded as a well-performing baseline for domain adaptation tasks. Our method is built upon the nearest centroid classifier, seeking a subspace where the centroids in the target domain are moderately shifted from those in the source domain. Specifically, we design a unified objective without accessing the source domain data and adopt an alternating minimization scheme to iteratively discover the pseudo target labels, invariant subspace, and target centroids. Besides its privacy-preserving property (distant supervision), the algorithm is provably convergent and has a promising linear time complexity. In addition, the proposed method can be readily extended to multi-source setting and domain generalization, and it remarkably enhances popular deep adaptation methods by borrowing the learned transferable features. Extensive experiments on several benchmarks including object, digit, and face recognition datasets validate that our methods yield state-of-the-art results in various domain adaptation tasks. Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
CVPR | 4 |
| 2019 | An Attention Enhanced Graph Convolutional LSTM Network for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition is an important task that requires the adequate understanding of movement characteristics of a human action from the given skeleton sequence. Recent studies have shown that exploring spatial and temporal features of the skeleton sequence is vital for this task. Nevertheless, how to effectively extract discriminative spatial and temporal features is still a challenging problem. In this paper, we propose a novel Attention Enhanced Graph Convolutional LSTM Network (AGC-LSTM) for human action recognition from skeleton data. The proposed AGC-LSTM can not only capture discriminative features in spatial configuration and temporal dynamics but also explore the co-occurrence relationship between spatial and temporal domains. We also present a temporal hierarchical architecture to increase temporal receptive fields of the top AGC-LSTM layer, which boosts the ability to learn the high-level semantic representation and significantly reduces the computation cost. Furthermore, to select discriminative spatial information, the attention mechanism is employed to enhance information of key joints in each AGC-LSTM layer. Experimental results on two datasets are provided: NTU RGB+D dataset and Northwestern-UCLA dataset. The comparison results demonstrate the effectiveness of our approach and show that our approach outperforms the state-of-the-art methods on both datasets. Chenyang Si, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
CVPR | 5 |
| 2019 | POD: Practical Object Detection With Scale-Sensitive NetworkabstractScale-sensitive object detection remains a challenging task, where most of the existing methods not learn it explicitly and not robust to scale variance. In addition, the most existing methods are less efficient during training or slow during inference, which are not friendly to real-time application. In this paper, we propose a practical object detection with scale-sensitive network.Our method first predicts a global continuous scale ,which shared by all position, for each convolution filter of each network stage. To effectively learn the scale, we average the spatial features and distill the scale from channels. For fast-deployment, we propose a scale decomposition method that transfers the robust fractional scale into combinations of fixed integral scales for each convolution filter, which exploit the dilated convolution. We demonstrate it on one-stage and two-stage algorithm under almost different configure. For practical application, training of our method is of efficiency and simplicity which gets rid of complex data sampling or optimize strategy. During testing, the proposed method requires no extra operation and is very friendly to hardware acceleration like TensorRT and TVM.On the COCO test-dev, our model could achieve a 41.5mAP on one-stage detector and 42.1 mAP on two-stage detectors based on ResNet-101, outperforming baselines by 2.4 and 2.1 respectively without extra FLOPS. Junran Peng, Ming Sun 0008, Zhaoxiang Zhang 0001, Tieniu Tan |
ICCV | 4 |
| 2019 | Unsupervised Cross-Domain Person Re-Identification: A New FrameworkabstractAlthough existing person Re-IDentification (ReID) methods have achieved great progress with large-scale labeled data, it is still hard to generalize to unseen scenarios without any la-beled person identities. To alleviate this problem, this paper proposes a new framework to take full advantage of the label information of source domain and the data distribution geometry of unlabeled target domain to improve the ReID performance in the unlabeled target domain. Instead of direct model transfer, the data transfer is first adopted where the identity preserving samples are generated from the labeled source domain to unlabeled target domain. Accordingly, a better initialized target domain adapted ReID model could be obtained with the generated samples. The fine-grained part-level features are then learned instead of global features to better mine new persons in the unlabeled target domain. Finally, the proposed framework iteratively updates the ReID model with the generated persons and the mined persons in last iteration, and explores new persons from the unlabeled target domain. The state-of-the-art experimental results are achieved on Market1501 and DukeMTMC-reID in terms of unsupervised cross-domain person ReID. Da Li 0003, Dangwei Li, Zhang Zhang 0001, Liang Wang 0001, Tieniu Tan |
ICIP | 5 |
| 2019 | Hierarchical Graph Convolutional Networks for Semi-supervised Node ClassificationabstractGraph convolutional networks (GCNs) have been successfully applied in node classification tasks of network mining. However, most of these models based on neighborhood aggregation are usually shallow and lack the “graph pooling” mechanism, which prevents the model from obtaining adequate global information. In order to increase the receptive field, we propose a novel deep Hierarchical Graph Convolutional Network (H-GCN) for semi-supervised node classification. H-GCN first repeatedly aggregates structurally similar nodes to hyper-nodes and then refines the coarsened graph to the original to restore the representation for each node. Instead of merely aggregating one- or two-hop neighborhood information, the proposed coarsening procedure enlarges the receptive field for each node, hence more global information can be captured. The proposed H-GCN model shows strong empirical performance on various public benchmark graph datasets, outperforming state-of-the-art methods and acquiring up to 5.9% performance improvement in terms of accuracy. In addition, when only a few labeled samples are provided, our model gains substantial improvements. Fenyu Hu, Yanqiao Zhu 0001, Liang Wang 0001, Tieniu Tan |
IJCAI | 5 |
| 2019 | Stacked Memory Network for Video SummarizationabstractIn recent years, supervised video summarization has achieved promising progress with various recurrent neural networks (RNNs) based methods, which treats video summarization as a sequence-to-sequence learning problem to exploit temporal dependency among video frames across variable ranges. However, RNN has limitations in modelling the long-term temporal dependency for summarizing videos with thousands of frames due to the restricted memory storage unit. Therefore, in this paper we propose a stacked memory network called SMN to explicitly model the long dependency among video frames so that redundancy could be minimized in the video summaries produced. Our proposed SMN consists of two key components: Long Short-Term Memory (LSTM) layer and memory layer, where each LSTM layer is augmented with an external memory layer. In particular, we stack multiple LSTM layers and memory layers hierarchically to integrate the learned representation from prior layers. By combining the hidden states of the LSTM layers and the read representations of the memory layers, our SMN is able to derive more accurate video summaries for individual video frames. Compared with the existing RNN based methods, our SMN is particularly good at capturing long temporal dependency among frames with few additional training parameters. Experimental results on two widely used public benchmark datasets: SumMe and TVsum, demonstrate that our proposed model is able to clearly outperform a number of state-of-the-art ones under various settings. Junbo Wang 0003, Wei Wang 0115, Zhiyong Wang 0001, Liang Wang 0001, David Dagan Feng, Tieniu Tan |
ACM Multimedia | 6 |
| 2019 | Efficient Neural Architecture Transformation Search in Channel-Level for Object DetectionabstractRecently, Neural Architecture Search has achieved great success in large-scale image classification. In contrast, there have been limited works focusing on architecture search for object detection, mainly because the costly ImageNet pretraining is always required for detectors. Training from scratch, as a substitute, demands more epochs to converge and brings no computation saving. To overcome this obstacle, we introduce a practical neural architecture transformation search(NATS) algorithm for object detection in this paper. Instead of searching and constructing an entire network, NATS explores the architecture space on the base of existing network and reusing its weights. We propose a novel neural architecture search strategy in channel-level instead of path-level and devise a search space specially targeting at object detection. With the combination of these two designs, an architecture transformation scheme could be discovered to adapt a network designed for image classification to task of object detection. Since our method is gradient-based and only searches for a transformation scheme, the weights of models pretrained in ImageNet could be utilized in both searching and retraining stage, which makes the whole process very efficient. The transformed network requires no extra parameters and FLOPs, and is friendly to hardware optimization, which is practical to use in real-time application. In experiments, we demonstrate the effectiveness of NATS on networks like {\em ResNet} and {\em ResNeXt}. Our transformed networks, combined with various detection frameworks, achieve significant improvements on the COCO dataset while keeping fast. Junran Peng, Ming Sun 0008, Zhaoxiang Zhang 0001, Tieniu Tan |
NeurIPS | 4 |
| 2019 | Attention-based convolutional approach for misinformation identification from massive and noisy microblog posts
Feng Yu 0001, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
Comput. Secur. | 5 |
| 2019 | Wavelet Domain Generative Adversarial Network for Multi-scale Face Hallucination
Huaibo Huang, Ran He 0001, Zhenan Sun, Tieniu Tan |
Int. J. Comput. Vis. | 4 |
| 2019 | Toward practical remote iris recognition: A boosting based framework
Man Zhang 0005, Zhaofeng He 0001, Hui Zhang 0061, Tieniu Tan, Zhenan Sun |
Neurocomputing | 4 |
| 2019 | Feedback Convolutional Neural Network for Visual Localization and SegmentationabstractFeedback is a fundamental mechanism existing in the human visual system, but has not been explored deeply in designing computer vision algorithms. In this paper, we claim that feedback plays a critical role in understanding convolutional neural networks (CNNs), e.g., how a neuron in CNNs describes an object's pattern, and how a collection of neurons form comprehensive perception to an object. To model the feedback in CNNs, we propose a novel model named Feedback CNN and develop two new processing algorithms, i.e., neural pathway pruning and pattern recovering. We mathematically prove that the proposed method can reach local optimum. Note that Feedback CNN belongs to weakly supervised methods and can be trained only using category-level labels. But it possesses a powerful capability to accurately localize and segment category-specific objects. We conduct extensive visualization analysis, and the results reveal the close relationship between neurons and object parts in Feedback CNN. Finally, we evaluate the proposed Feedback CNN over the tasks of weakly supervised object localization and segmentation, and the experimental results on ImageNet and Pascal VOC show that our method remarkably outperforms the state-of-the-art ones. Chunshui Cao, Yongzhen Huang, Yi Yang 0007, Liang Wang 0001, Zilei Wang, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2019 | Wasserstein CNN: Learning Invariant Features for NIR-VIS Face RecognitionabstractHeterogeneous face recognition (HFR) aims at matching facial images acquired from different sensing modalities with mission-critical applications in forensics, security and commercial sectors. However, HFR presents more challenging issues than traditional face recognition because of the large intra-class variation among heterogeneous face images and the limited availability of training samples of cross-modality face image pairs. This paper proposes the novel Wasserstein convolutional neural network (WCNN) approach for learning invariant features between near-infrared (NIR) and visual (VIS) face images (i.e., NIR-VIS face recognition). The low-level layers of the WCNN are trained with widely available face images in the VIS spectrum, and the high-level layer is divided into three parts: the NIR layer, the VIS layer and the NIR-VIS shared layer. The first two layers aim at learning modality-specific features, and the NIR-VIS shared layer is designed to learn a modality-invariant feature subspace. The Wasserstein distance is introduced into the NIR-VIS shared layer to measure the dissimilarity between heterogeneous feature distributions. W-CNN learning is performed to minimize the Wasserstein distance between the NIR distribution and the VIS distribution for invariant deep feature representations of heterogeneous face images. To avoid the over-fitting problem on small-scale heterogeneous face data, a correlation prior is introduced on the fully-connected WCNN layers to reduce the size of the parameter space. This prior is implemented by a low-rank constraint in an end-to-end network. The joint formulation leads to an alternating minimization for deep feature representation at the training stage and an efficient computation for heterogeneous data at the testing stage. Extensive experiments using three challenging NIR-VIS face recognition databases demonstrate the superiority of the WCNN method over state-of-the-art methods. Ran He 0001, Xiang Wu 0001, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Aggregating Randomized Clustering-Promoting Invariant Projections for Domain AdaptationabstractUnsupervised domain adaptation aims to leverage the labeled source data to learn with the unlabeled target data. Previous trandusctive methods tackle it by iteratively seeking a low-dimensional projection to extract the invariant features and obtaining the pseudo target labels via building a classifier on source data. However, they merely concentrate on minimizing the cross-domain distribution divergence, while ignoring the intra-domain structure especially for the target domain. Even after projection, possible risk factors like imbalanced data distribution may still hinder the performance of target label inference. In this paper, we propose a simple yet effective domain-invariant projection ensemble approach to tackle these two issues together. Specifically, we seek the optimal projection via a novel relaxed domain-irrelevant clustering-promoting term that jointly bridges the cross-domain semantic gap and increases the intra-class compactness in both domains. To further enhance the target label inference, we first develop a 'sampling-and-fusion' framework, under which multiple projections are independently learned based on various randomized coupled domain subsets. Subsequently, aggregating models such as majority voting are utilized to leverage multiple projections and classify unlabeled target data. Extensive experimental results on six visual benchmarks including object, face, and digit images, demonstrate that the proposed methods gain remarkable margins over state-of-the-art unsupervised domain adaptation methods. Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Learning a bi-level adversarial network with global and local perception for makeup-invariant face verification
Yi Li 0018, Lingxiao Song, Xiang Wu 0001, Ran He 0001, Tieniu Tan |
Pattern Recognit. | 5 |
| 2019 | MAPNet: Multi-modal attentive pooling network for RGB-D indoor scene classification
Yabei Li, Zhang Zhang 0001, Yanhua Cheng, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 5 |
| 2019 | Corrigendum to "MAPNet: Multi-modal attentive pooling network for RGB-D indoor scene classification" [Pattern Recognition 90 (2019) 436-449]
Yabei Li, Zhang Zhang 0001, Yanhua Cheng, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 5 |
| 2019 | Exploring uncertainty in pseudo-label guided unsupervised domain adaptation
Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
Pattern Recognit. | 4 |
| 2019 | ISEE: An Intelligent Scene Exploration and Evaluation Platform for Large-Scale Visual SurveillanceabstractIntelligent video surveillance (IVS) is always an interesting research topic to utilize visual analysis algorithms for exploring richly structured information from big surveillance data. However, existing IVS systems either struggle to utilize computing resources adequately to improve the efficiency of large-scale video analysis, or present a customized system for specific video analytic functions. It still lacks of a comprehensive computing architecture to enhance efficiency, extensibility and flexibility of IVS system. Moreover, it is also an open problem to study the effect of the combinations of multiple vision modules on the final performance of end applications of IVS system. Motivated by these challenges, we develop an Intelligent Scene Exploration and Evaluation (ISEE) platform based on a heterogeneous CPU-GPU cluster and some distributed computing tools, where Spark Streaming serves as the computing engine for efficient large-scale video processing and Kafka is adopted as a middle-ware message center to decouple different analysis modules flexibly. To validate the efficiency of the ISEE and study the evaluation problem on composable systems, we instantiate the ISEE for an end application on person retrieval with three visual analysis modules, including pedestrian detection with tracking, attribute recognition and re-identification. Extensive experiments are performed on a large-scale surveillance video dataset involving 25 camera scenes, totally 587 hours 720p synchronous videos, where a two-stage question-answering procedure is proposed to measure the performance of execution pipelines composed of multiple visual analysis algorithms based on millions of attribute-based and relationship-based queries. The case study of system-level evaluations may inspire researchers to improve visual analysis algorithms and combining strategies from the view of a scalable and composable system in the future. Da Li 0003, Zhang Zhang 0001, Kai Yu 0003, Kaiqi Huang, Tieniu Tan |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Lateral Inhibition-Inspired Convolutional Neural Network for Visual Attention and Saliency DetectionabstractLateral inhibition in top-down feedback is widely existing in visual neurobiology, but such an important mechanism has not be well explored yet in computer vision. In our recent research, we find that modeling lateral inhibition in convolutional neural network (LICNN) is very useful for visual attention and saliency detection. In this paper, we propose to formulate lateral inhibition inspired by the related studies from neurobiology, and embed it into the top-down gradient computation of a general CNN for classification, i.e. only category-level information is used. After this operation (only conducted once), the network has the ability to generate accurate category-specific attention maps. Further, we apply LICNN for weakly-supervised salient object detection.Extensive experimental studies on a set of databases, e.g., ECSSD, HKU-IS, PASCAL-S and DUT-OMRON, demonstrate the great advantage of LICNN which achieves the state-of-the-art performance. It is especially impressive that LICNN with only category-level supervised information even outperforms some recent methods with segmentation-level supervised learning. Chunshui Cao, Yongzhen Huang, Zilei Wang, Liang Wang 0001, Ninglong Xu, Tieniu Tan |
AAAI | 6 |
| 2018 | Deep Semantic Structural Constraints for Zero-Shot LearningabstractZero-shot learning aims to classify unseen image categories by learning a visual-semantic embedding space. In most cases, the traditional methods adopt a separated two-step pipeline that extracts image features are utilized to learn the embedding space. It leads to the lack of specific structural semantic information of image features for zero-shot learning task. In this paper, we propose an end-to-end trainable Deep Semantic Structural Constraints model to address this issue. The proposed model contains the Image Feature Structure constraint and the Semantic Embedding Structure constraint, which aim to learn structure-preserving image features and endue the learned embedding space with stronger generalization ability respectively. With the assistance of semantic structural information, the model gains more auxiliary clues for zero-shot learning. The state-of-the-art performance certifies the effectiveness of our proposed method. Yan Li 0043, Junge Zhang, Kaiqi Huang, Tieniu Tan |
AAAI | 5 |
| 2018 | Anti-Makeup: Learning A Bi-Level Adversarial Network for Makeup-Invariant Face VerificationabstractMakeup is widely used to improve facial attractiveness and is well accepted by the public. However, different makeup styles will result in significant facial appearance changes. It remains a challenging problem to match makeup and non-makeup face images. This paper proposes a learning from generation approach for makeup-invariant face verification by introducing a bi-level adversarial network (BLAN). To alleviate the negative effects from makeup, we first generate non-makeup images from makeup ones, and then use the synthesized non-makeup images for further verification. Two adversarial networks in BLAN are integrated in an end-to-end deep network, with the one on pixel level for reconstructing appealing facial images and the other on feature level for preserving identity information. These two networks jointly reduce the sensing gap between makeup and non-makeup images. Moreover, we make the generator well constrained by incorporating multiple perceptual losses. Experimental results on three benchmark makeup face datasets demonstrate that our method achieves state-of-the-art verification accuracy across makeup status and can produce photo-realistic non-makeup face images. Yi Li 0018, Lingxiao Song, Xiang Wu 0001, Ran He 0001, Tieniu Tan |
AAAI | 5 |
| 2018 | DF2Net: Discriminative Feature Learning and Fusion Network for RGB-D Indoor Scene ClassificationabstractThis paper focuses on the task of RGB-D indoor scene classification. It is a very challenging task due to two folds. 1) Learning robust representation for indoor scene is difficult because of various objects and layouts. 2) Fusing the complementary cues in RGB and Depth is nontrivial since there are large semantic gaps between the two modalities. Most existing works learn representation for classification by training a deep network with softmax loss and fuse the two modalities by simply concatenating the features of them. However, these pipelines do not explicitly consider intra-class and inter-class similarity as well as inter-modal intrinsic relationships. To address these problems, this paper proposes a Discriminative Feature Learning and Fusion Network (DF2Net) with two-stage training. In the first stage, to better represent scene in each modality, a deep multi-task network is constructed to simultaneously minimize the structured loss and the softmax loss. In the second stage, we design a novel discriminative fusion network which is able to learn correlative features of multiple modalities and distinctive features of each modality. Extensive analysis and experiments on SUN RGB-D Dataset and NYU Depth Dataset V2 show the superiority of DF2Net over other state-of-the-art methods in RGB-D indoor scene classification task. Yabei Li, Junge Zhang, Yanhua Cheng, Kaiqi Huang, Tieniu Tan |
AAAI | 5 |
| 2018 | Coupled Deep Learning for Heterogeneous Face RecognitionabstractHeterogeneous face matching is a challenge issue in face recognition due to large domain difference as well as insufficient pairwise images in different modalities during training. This paper proposes a coupled deep learning (CDL) approach for the heterogeneous face matching. CDL seeks a shared feature space in which the heterogeneous face matching problem can be approximately treated as a homogeneous face matching problem. The objective function of CDL mainly includes two parts. The first part contains a trace norm and a block-diagonal prior as relevance constraints, which not only make unpaired images from multiple modalities be clustered and correlated, but also regularize the parameters to alleviate overfitting. An approximate variational formulation is introduced to deal with the difficulties of optimizing low-rank constraint directly. The second part contains a cross modal ranking among triplet domain specific images to maximize the margin for different identities and increase data for a small amount of training samples. Besides, an alternating minimization method is employed to iteratively update the parameters of CDL. Experimental results show that CDL achieves better performance on the challenging CASIA NIR-VIS 2.0 face recognition database, the IIIT-D Sketch database, the CUHK Face Sketch (CUFS), and the CUHK Face Sketch FERET (CUFSF), which significantly outperforms state-of-the-art heterogeneous face recognition methods. Xiang Wu 0001, Lingxiao Song, Ran He 0001, Tieniu Tan |
AAAI | 4 |
| 2018 | Multistage Adversarial Losses for Pose-Based Human Image SynthesisabstractHuman image synthesis has extensive practical applications e.g. person re-identification and data augmentation for human pose estimation. However, it is much more challenging than rigid object synthesis, e.g. cars and chairs, due to the variability of human posture. In this paper, we propose a pose-based human image synthesis method which can keep the human posture unchanged in novel viewpoints. Furthermore, we adopt multistage adversarial losses separately for the foreground and background generation, which fully exploits the multi-modal characteristics of generative loss to generate more realistic looking images. We perform extensive experiments on the Human3.6M dataset and verify the effectiveness of each stage of our method. The generated human images not only keep the same pose as the input image, but also have clear detailed foreground and background. The quantitative comparison results illustrate that our approach achieves much better results than several state-of-the-art methods. Chenyang Si, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
CVPR | 4 |
| 2018 | M3: Multimodal Memory Modelling for Video CaptioningabstractVideo captioning which automatically translates video clips into natural language sentences is a very important task in computer vision. By virtue of recent deep learning technologies, video captioning has made great progress. However, learning an effective mapping from the visual sequence space to the language space is still a challenging problem due to the long-term multimodal dependency modelling and semantic misalignment. Inspired by the facts that memory modelling poses potential advantages to long-term sequential problems [35] and working memory is the key factor of visual attention [33], we propose a Multimodal Memory Model (M3) to describe videos, which builds a visual and textual shared memory to model the long-term visual-textual dependency and further guide visual attention on described visual targets to solve visual-textual alignments. Specifically, similar to [10], the proposed M3 attaches an external memory to store and retrieve both visual and textual contents by interacting with video and sentence with multiple read and write operations. To evaluate the proposed model, we perform experiments on two public datasets: MSVD and MSR-VTT. The experimental results demonstrate that our method outperforms most of the state-of-the-art methods in terms of BLEU and METEOR. Junbo Wang 0003, Wei Wang 0115, Yan Huang 0008, Liang Wang 0001, Tieniu Tan |
CVPR | 5 |
| 2018 | Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning
Chenyang Si, Ya Jing, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
ECCV (1) | 5 |
| 2018 | End-to-End View Synthesis for Light Field Imaging with Pseudo 4DCNN
Yunlong Wang 0003, Fei Liu 0031, Zilei Wang, Guangqi Hou, Zhenan Sun, Tieniu Tan |
ECCV (2) | 6 |
| 2018 | Inception Donut Convolution for Top-down Semantic SegmentationabstractOne of recent trends in network architecture design confirms that the inception-block convolutional group is efficient, since it can aggregate spatial context information in lower dimensions without causing significant loss in representative capabilities. We believe that not only the strong correlation between adjacent cells, multi-scale feature extraction also plays a vital role in this novel module. In this paper, we extend the profits of the block to a top-down donut convolutional network for semantic segmentation task. Our network automatically learns rich convolution kernels to capture more structure prior. In the inception-block design, it overcomes the limitations in larger kernel size and adaptively captures different object-scales contexts without chain sampling. Our experiments demonstrate that the proposed inception-block donut convolutional network is orthogonal and can further improve the performance of most off-the-shelf bottom-up based methods. He Guan, Zhaoxiang Zhang 0001, Tieniu Tan |
ICPR | 3 |
| 2018 | DeepFirearm: Learning Discriminative Feature Representation for Fine-grained Firearm RetrievalabstractThere are great demands for automatically regulating inappropriate appearance of shocking firearm images in social media or identifying firearm types in forensics. Image retrieval techniques have great potential to solve these problems. To facilitate research in this area, we introduce Firearm 14k, a large dataset consisting of over 14,000 images in 167 categories. It can be used for both fine-grained recognition and retrieval of firearm images. Recent advances in image retrieval are mainly driven by fine-tuning state-of-the-art convolutional neural networks for retrieval task. The conventional single margin contrastive loss, known for its simplicity and good performance, has been widely used. We find that it performs poorly on the Firearm 14k dataset due to: (1) Loss contributed by positive and negative image pairs is unbalanced during training process. (2) A huge domain gap exists between this dataset and ImageNet. We propose to deal with the unbalanced loss by employing a double margin contrastive loss. We tackle the domain gap issue with a two-stage training strategy, where we first fine-tune the network for classification, and then fine-tune it for retrieval. Experimental results show that our approach outperforms the conventional single margin approach by a large margin (up to 88.5% relative improvement) and even surpasses the strong triplet-loss-based approach. Jiedong Hao, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICPR | 4 |
| 2018 | Geometry Guided Adversarial Facial Expression SynthesisabstractFacial expression synthesis has drawn much attention in the field of computer graphics and pattern recognition. It has been widely used in face animation and recognition. However, it is still challenging due to the high-level semantic presence of large and non-linear face geometry variations. This paper proposes a Geometry-Guided Generative Adversarial Network (G2-GAN) for continuously-adjusting and identity-preserving facial expression synthesis. We employ facial geometry (fiducial points) as a controllable condition to guide facial texture synthesis with specific expression. A pair of generative adversarial subnetworks is jointly trained towards opposite tasks: expression removal and expression synthesis. The paired networks form a mapping cycle between neutral expression and arbitrary expressions, with which the proposed approach can be conducted among unpaired data. The proposed paired networks also facilitate other applications such as face transfer, expression interpolation and expression-invariant face recognition. Experimental results on several facial expression databases show that our method can generate compelling perceptual results on different expression editing tasks. Lingxiao Song, Zhihe Lu, Ran He 0001, Zhenan Sun, Tieniu Tan |
ACM Multimedia | 5 |
| 2018 | Hierarchical Memory Modelling for Video CaptioningabstractTranslating videos into natural language sentences has drawn much attention recently. The framework of combining visual attention with Long Short-Term Memory (LSTM) based text decoder has achieved much progress. However, the vision-language translation still remains unsolved due to the semantic gap and misalignment between video content and described semantic concept. In this paper, we propose a Hierarchical Memory Model (HMM) - a novel deep video captioning architecture which unifies a textual memory, a visual memory and an attribute memory in a hierarchical way. These memories can guide attention for efficient video representation extraction and semantic attribute selection in addition to modelling the long-term dependency for video sequence and sentences, respectively. Compared with traditional vision-based text decoder, the proposed attribute-based text decoder can largely reduce the semantic discrepancy between video and sentence. To prove the effectiveness of the proposed model, we perform extensive experiments on two public benchmark datasets: MSVD and MSR-VTT. Experiments show that our model not only can discover appropriate video representation and semantic attributes but also can achieve comparable or superior performances than state-of-the-art methods on these datasets. Junbo Wang 0003, Wei Wang 0115, Yan Huang 0008, Liang Wang 0001, Tieniu Tan |
ACM Multimedia | 5 |
| 2018 | IntroVAE: Introspective Variational Autoencoders for Photographic Image SynthesisabstractWe present a novel introspective variational autoencoder (IntroVAE) model for synthesizing high-resolution photographic images. IntroVAE is capable of self-evaluating the quality of its generated samples and improving itself accordingly. Its inference and generator models are jointly trained in an introspective way. On one hand, the generator is required to reconstruct the input images from the noisy outputs of the inference model as normal VAEs. On the other hand, the inference model is encouraged to classify between the generated and real samples while the generator tries to fool it as GANs. These two famous generative frameworks are integrated in a simple yet efficient single-stream architecture that can be trained in a single stage. IntroVAE preserves the advantages of VAEs, such as stable training and nice latent manifold. Unlike most other hybrid models of VAEs and GANs, IntroVAE requires no extra discriminators, because the inference model itself serves as a discriminator to distinguish between the generated and real samples. Experiments demonstrate that our method produces high-resolution photo-realistic images (e.g., CELEBA images at (1024^{2})), which are comparable to or better than the state-of-the-art GANs. Huaibo Huang, Zhihang Li, Ran He 0001, Zhenan Sun, Tieniu Tan |
NeurIPS | 5 |
| 2018 | Feature learning for steganalysis using convolutional neural networks
Yinlong Qian, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
Multim. Tools Appl. | 4 |
| 2018 | Fast Supervised Discrete HashingabstractLearning-based hashing algorithms are "hot topics" because they can greatly increase the scale at which existing methods operate. In this paper, we propose a new learning-based hashing method called "fast supervised discrete hashing" (FSDH) based on "supervised discrete hashing" (SDH). Regressing the training examples (or hash code) to the corresponding class labels is widely used in ordinary least squares regression. Rather than adopting this method, FSDH uses a very simple yet effective regression of the class labels of training examples to the corresponding hash code to accelerate the algorithm. To the best of our knowledge, this strategy has not previously been used for hashing. Traditional SDH decomposes the optimization into three sub-problems, with the most critical sub-problem - discrete optimization for binary hash codes - solved using iterative discrete cyclic coordinate descent (DCC), which is time-consuming. However, FSDH has a closed-form solution and only requires a single rather than iterative hash code-solving step, which is highly efficient. Furthermore, FSDH is usually faster than SDH for solving the projection matrix for least squares regression, making FSDH generally faster than SDH. For example, our results show that FSDH is about 12-times faster than SDH when the number of hashing bits is 128 on the CIFAR-10 data base, and FSDH is about 151-times faster than FastHash when the number of hashing bits is 64 on the MNIST data-base. Our experimental results show that FSDH is not only fast, but also outperforms other comparative methods. Jie Gui, Tongliang Liu, Zhenan Sun, Dacheng Tao, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Demographic Analysis from Biometric Data: Achievements, Challenges, and New FrontiersabstractBiometrics is the technique of automatically recognizing individuals based on their biological or behavioral characteristics. Various biometric traits have been introduced and widely investigated, including fingerprint, iris, face, voice, palmprint, gait and so forth. Apart from identity, biometric data may convey various other personal information, covering affect, age, gender, race, accent, handedness, height, weight, etc. Among these, analysis of demographics (age, gender, and race) has received tremendous attention owing to its wide real-world applications, with significant efforts devoted and great progress achieved. This survey first presents biometric demographic analysis from the standpoint of human perception, then provides a comprehensive overview of state-of-the-art advances in automated estimation from both academia and industry. Despite these advances, a number of challenging issues continue to inhibit its full potential. We second discuss these open problems, and finally provide an outlook into the future of this very active field of research by sharing some promising opportunities. Yunlian Sun, Man Zhang 0005, Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Learning structured ordinal measures for video based face recognition
Ran He 0001, Tieniu Tan, Larry Davis 0001, Zhenan Sun |
Pattern Recognit. | 2 |
| 2018 | Robust linear representation via exploiting structure prior
Dong Wang 0004, Ran He 0001, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. | 4 |
| 2018 | Efficient auto-refocusing for light field camera
Chi Zhang 0060, Guangqi Hou, Zhaoxiang Zhang 0001, Zhenan Sun, Tieniu Tan |
Pattern Recognit. | 5 |
| 2018 | Image Forensics Based on Planar Contact Constraints of 3D ObjectsabstractStanding objects on planar surfaces are common to see in images, e.g., people on the ground. For most objects to stay stable on the plane, planar contact is a necessary requirement. However, 2D image splicing usually disregards this physical constraint of 3D world, leading to a potential artifact of object not attached to the plane. This paper is the first attempt to use the contact constraint of standing objects as a new clue for image forensics. Accordingly, we propose a novel approach to first reconstruct the 3D poses of standing objects and their supporting plane and then measure the contact conditions for splicing detection. To tackle the problem of unknown object shape for pose estimation, we effectively employ the prior knowledge of 3D morphable model to simultaneously estimate both shape and pose parameters by fitting to image observations. The 3D normal orientation of the supporting plane is estimated given its vanishing line. Dealing with uncertainty factors in estimations, we approximate a distribution of estimates using sampling strategies and then make the final decision. Particularly, we focused our method on the important scenario of human figure splicing detection, and comprehensive experiments on multiple data sets and typical images proved the encouraging effectiveness of the new forensic clue and the proposed approach. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | A Light CNN for Deep Face Representation With Noisy LabelsabstractThe volume of convolutional neural network (CNN) models proposed for face recognition has been continuously growing larger to better fit the large amount of training data. When training data are obtained from the Internet, the labels are likely to be ambiguous and inaccurate. This paper presents a Light CNN framework to learn a compact embedding on the large-scale face data with massive noisy labels. First, we introduce a variation of maxout activation, called max-feature-map (MFM), into each convolutional layer of CNN. Different from maxout activation that uses many feature maps to linearly approximate an arbitrary convex activation function, MFM does so via a competitive relationship. MFM can not only separate noisy and informative signals but also play the role of feature selection between two feature maps. Second, three networks are carefully designed to obtain better performance, meanwhile, reducing the number of parameters and computational costs. Finally, a semantic bootstrapping method is proposed to make the prediction of the networks more consistent with noisy labels. Experimental results show that the proposed framework can utilize large-scale noisy data to learn a Light model that is efficient in computational costs and storage spaces. The learned single network with a 256-D representation achieves state-of-the-art results on various face benchmarks without fine-tuning. Xiang Wu 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | DeMeshNet: Blind Face Inpainting for Deep MeshFace VerificationabstractMeshFace photos have been widely used in many Chinese business organizations to protect ID face photos from being misused. The occlusions incurred by random meshes severely degenerate the performance of face verification systems, which raises the MeshFace verification problem between MeshFace and daily photos. Previous methods cast this problem as a typical low-level vision problem, i.e., blind inpainting. They recover perceptually pleasing clear ID photos from MeshFaces by enforcing pixel level similarity between the recovered ID images and the ground-truth clear ID images and then perform face verification on them. Essentially, face verification is conducted on a compact feature space rather than the image pixel space. Therefore, this paper argues that pixel level similarity and feature level similarity jointly offer the key to improve the verification performance. Based on this insight, we offer a novel feature oriented blind face inpainting framework. Specifically, we implement this by establishing a novel DeMeshNet, which consists of three parts. The first part addresses blind inpainting of the MeshFaces by implicitly exploiting extra supervision from the occlusion position to enforce pixel level similarity. The second part explicitly enforces a feature level similarity in the compact feature space, which can explore informative supervision from the feature space to produce better inpainting results for verification. The last part copes with face alignment within the net via a customized spatial transformer module when extracting deep facial features. All three parts are implemented within an end-to-end network that facilitates efficient optimization. Extensive experiments on two MeshFace data sets demonstrate the effectiveness of the proposed DeMeshNet as well as the insight of this paper. Shu Zhang 0015, Ran He 0001, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | Deep Feature Fusion for Iris and Periocular Biometrics on Mobile DevicesabstractThe quality of iris images on mobile devices is significantly degraded due to hardware limitations and less constrained environments. Traditional iris recognition methods cannot achieve high identification rate using these low-quality images. To enhance the performance of mobile identification, we develop a deep feature fusion network that exploits the complementary information presented in iris and periocular regions. The proposed method first applies maxout units into the convolutional neural networks (CNNs) to generate a compact representation for each modality and then fuses the discriminative features of two modalities through a weighted concatenation. The parameters of convolutional filters and fusion weights are simultaneously learned to optimize the joint representation of iris and periocular biometrics. To promote the iris recognition research on mobile devices under near-infrared (NIR) illumination, we publicly release the CASIA-Iris-Mobile-V1.0 database, which in total includes 11 000 NIR iris images of both eyes from 630 Asians. It is the largest NIR mobile iris database as far as we know. On the newly built CASIA-Iris-M1-S3 data set, the proposed method achieves 0.60% equal error rate and 2.32% false non-match rate at false match rate =10-5, which are obviously better than unimodal biometrics as well as traditional fusion methods. Moreover, the proposed model requires much fewer storage spaces and computational resources than general CNNs. Qi Zhang 0015, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | LFNet: A Novel Bidirectional Recurrent Convolutional Neural Network for Light-Field Image Super-ResolutionabstractThe low spatial resolution of light-field image poses significant difficulties in exploiting its advantage. To mitigate the dependency of accurate depth or disparity information as priors for light-field image super-resolution, we propose an implicitly multi-scale fusion scheme to accumulate contextual information from multiple scales for super-resolution reconstruction. The implicitly multi-scale fusion scheme is then incorporated into bidirectional recurrent convolutional neural network, which aims to iteratively model spatial relations between horizontally or vertically adjacent sub-aperture images of light-field data. Within the network, the recurrent convolutions are modified to be more effective and flexible in modeling the spatial correlations between neighboring views. A horizontal sub-network and a vertical sub-network of the same network structure are ensembled for final outputs via stacked generalization. Experimental results on synthetic and real-world data sets demonstrate that the proposed method outperforms other state-of-the-art methods by a large margin in peak signal-to-noise ratio and gray-scale structural similarity indexes, which also achieves superior quality for human visual systems. Furthermore, the proposed method can enhance the performance of light field applications such as depth estimation. Yunlong Wang 0003, Fei Liu 0031, Kunbo Zhang, Guangqi Hou, Zhenan Sun, Tieniu Tan |
IEEE Trans. Image Process. | 6 |
| 2018 | Supervised Discrete Hashing With RelaxationabstractData-dependent hashing has recently attracted attention due to being able to support efficient retrieval and storage of high-dimensional data, such as documents, images, and videos. In this paper, we propose a novel learning-based hashing method called "supervised discrete hashing with relaxation" (SDHR) based on "supervised discrete hashing" (SDH). SDH uses ordinary least squares regression and traditional zero-one matrix encoding of class label information as the regression target (code words), thus fixing the regression target. In SDHR, the regression target is instead optimized. The optimized regression target matrix satisfies a large margin constraint for correct classification of each example. Compared with SDH, which uses the traditional zero-one matrix, SDHR utilizes the learned regression target matrix and, therefore, more accurately measures the classification error of the regression model and is more flexible. As expected, SDHR generally outperforms SDH. Experimental results on two large-scale image data sets (CIFAR-10 and MNIST) and a large-scale and challenging face data set (FRGC) demonstrate the effectiveness and efficiency of SDHR. Jie Gui, Tongliang Liu, Zhenan Sun, Dacheng Tao, Tieniu Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Learning Invariant Deep Representation for NIR-VIS Face RecognitionabstractVisual versus near infrared (VIS-NIR) face recognition is still a challenging heterogeneous task due to large appearance difference between VIS and NIR modalities. This paper presents a deep convolutional network approach that uses only one network to map both NIR and VIS images to a compact Euclidean space. The low-level layers of this network are trained only on large-scale VIS data. Each convolutional layer is implemented by the simplest case of maxout operator. The high-level layer is divided into two orthogonal subspaces that contain modality-invariant identity information and modality-variant spectrum information respectively. Our joint formulation leads to an alternating minimization approach for deep representation at the training time and an efficient computation for heterogeneous data at the testing time. Experimental evaluations show that our method achieves 94% verification rate at FAR=0.1% on the challenging CASIA NIR-VIS 2.0 face recognition dataset. Compared with state-of-the-art methods, it reduces the error rate by 58% only with a compact 64-D representation. Ran He 0001, Xiang Wu 0001, Zhenan Sun, Tieniu Tan |
AAAI | 4 |
| 2017 | See the Forest for the Trees: Joint Spatial and Temporal Recurrent Neural Networks for Video-Based Person Re-identificationabstractSurveillance cameras have been widely used in different scenes. Accordingly, a demanding need is to recognize a person under different cameras, which is called person re-identification. This topic has gained increasing interests in computer vision recently. However, less attention has been paid to video-based approaches, compared with image-based ones. Two steps are usually involved in previous approaches, namely feature learning and metric learning. But most of the existing approaches only focus on either feature learning or metric learning. Meanwhile, many of them do not take full use of the temporal and spatial information. In this paper, we concentrate on video-based person re-identification and build an end-to-end deep neural network architecture to jointly learn features and metrics. The proposed method can automatically pick out the most discriminative frames in a given video by a temporal attention model. Moreover, it integrates the surrounding information at each location by a spatial recurrent model when measuring the similarity with another pedestrian video. That is, our method handles spatial and temporal information simultaneously in a unified manner. The carefully designed experiments on three public datasets show the effectiveness of each component of the proposed deep network, performing better in comparison with the state-of-the-art methods. Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
CVPR | 5 |
| 2017 | LivDet iris 2017 - Iris liveness detection competition 2017abstractPresentation attacks such as using a contact lens with a printed pattern or printouts of an iris can be utilized to bypass a biometric security system. The first international iris liveness competition was launched in 2013 in order to assess the performance of presentation attack detection (PAD) algorithms, with a second competition in 2015. This paper presents results of the third competition, LivDet-Iris 2017. Three software-based approaches to Presentation Attack Detection were submitted. Four datasets of live and spoof images were tested with an additional cross-sensor test. New datasets and novel situations of data have resulted in this competition being of a higher difficulty than previous competitions. Anonymous received the best results with a rate of rejected live samples of 3.36% and rate of accepted spoof samples of 14.71%. The results show that even with advances, printed iris attacks as well as patterned contacts lenses are still difficult for software-based systems to detect. Printed iris images were easier to be differentiated from live images in comparison to patterned contact lenses as was also seen in previous competitions. David Yambay, Benedict Becker, Naman Kohli, Daksha Yadav, Adam Czajka, Kevin W. Bowyer, Stephanie Schuckers, Richa Singh 0001, Mayank Vatsa, Afzel Noore, Diego Gragnaniello, Carlo Sansone, Luisa Verdoliva, Lingxiao He, Yiwei Ru, Nianfeng Liu, Zhenan Sun, Tieniu Tan |
IJCB | 19 |
| 2017 | Wavelet-SRNet: A Wavelet-Based CNN for Multi-scale Face Super ResolutionabstractMost modern face super-resolution methods resort to convolutional neural networks (CNN) to infer highresolution (HR) face images. When dealing with very low resolution (LR) images, the performance of these CNN based methods greatly degrades. Meanwhile, these methods tend to produce over-smoothed outputs and miss some textural details. To address these challenges, this paper presents a wavelet-based CNN approach that can ultra-resolve a very low resolution face image of 16 × 16 or smaller pixelsize to its larger version of multiple scaling factors (2×, 4×, 8× and even 16×) in a unified framework. Different from conventional CNN methods directly inferring HR images, our approach firstly learns to predict the LR's corresponding series of HR's wavelet coefficients before reconstructing HR images from them. To capture both global topology information and local texture details of human faces, we present a flexible and extensible convolutional neural network with three types of loss: wavelet prediction loss, texture loss and full-image loss. Extensive experiments demonstrate that the proposed approach achieves more appealing results both quantitatively and qualitatively than state-ofthe- art super-resolution methods. Huaibo Huang, Ran He 0001, Zhenan Sun, Tieniu Tan |
ICCV | 4 |
| 2017 | Encyclopedia enhanced semantic embedding for zero-shot learningabstractThere are tremendous object categories in the real world besides those in image datasets. Zero-shot learning aims to recognize image categories which are unseen in the training set. A large number of previous zero-shot learning models use word vectors of the class labels directly as category prototypes in the semantic embedding space. But word vectors cannot obtain the global knowledge of an image category sufficiently. In this paper, we propose a new encyclopedia enhanced semantic embedding model to promote the discriminative capability of word vector prototypes with the global knowledge of each image category. The proposed model extracts the TF-IDF key words from encyclopedia articles to acquire the global knowledge of each category. The convex combination of the key words' word vectors acts as the prototypes of the object categories. The prototypes of seen and unseen classes build up the embedding space where the nearest neighbour search is implemented to recognize the unseen images. The experiments show that the proposed method achieves the state-of-the-art performance on the challenging ImageNet Fall 2011 1k2hop dataset. Junge Zhang, Kaiqi Huang, Tieniu Tan |
ICIP | 4 |
| 2017 | Semantics-guided multi-level RGB-D feature fusion for indoor semantic segmentationabstractIndoor RGB-D semantic segmentation is a new and challenging problem. Traditional methods usually apply two-stream convolutional neural networks (CNNs) to represent RGB and depth images respectively, and fuse the two streams on a specific layer. In this paper, we explore several fusion strategies based on this two-stream-CNN framework and point out such a single-layer fusion method cannot exploit the complementary RGB and depth cues well for semantic segmentation. To address this problem, we propose a novel Semantics-guided Multi-level feature fusion approach, which first learns deep feature representation from bottom to up, and then gradually fuses the RGB and depth features from high level to low level under the guidance of the semantic cues. Experimental results on SUN RGB-D dataset demonstrate the advantages of the proposed method over the state of the arts. Yabei Li, Junge Zhang, Yanhua Cheng, Kaiqi Huang, Tieniu Tan |
ICIP | 5 |
| 2017 | A Convolutional Approach for Misinformation IdentificationabstractThe fast expanding of social media fuels the spreading of misinformation which disrupts people's normal lives. It is urgent to achieve goals of misinformation identification and early detection in social media. In dynamic and complicated social media scenarios, some conventional methods mainly concentrate on feature engineering which fail to cover potential features in new scenarios and have difficulty in shaping elaborate high-level interactions among significant features. Moreover, a recent Recurrent Neural Network (RNN) based method suffers from deficiencies that it is not qualified for practical early detection of misinformation and poses a bias to the latest input. In this paper, we propose a novel method, Convolutional Approach for Misinformation Identification (CAMI) based on Convolutional Neural Network (CNN). CAMI can flexibly extract key features scattered among an input sequence and shape high-level interactions among significant features, which help effectively identify misinformation and achieve practical early detection. Experiment results on two large-scale datasets validate the effectiveness of CAMI model on both misinformation identification and early detection tasks. Feng Yu 0001, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
IJCAI | 5 |
| 2017 | Deep Supervised Discrete HashingabstractWith the rapid growth of image and video data on the web, hashing has been extensively studied for image or video search in recent years. Benefiting from recent advances in deep learning, deep hashing methods have achieved promising results for image retrieval. However, there are some limitations of previous deep hashing methods (e.g., the semantic information is not fully exploited). In this paper, we develop a deep supervised discrete hashing algorithm based on the assumption that the learned binary codes should be ideal for classification. Both the pairwise label information and the classification information are used to learn the hash codes within one stream framework. We constrain the outputs of the last layer to be binary codes directly, which is rarely investigated in deep hashing algorithm. Because of the discrete nature of hash codes, an alternating minimization method is used to optimize the objective function. Experimental results have shown that our method outperforms current state-of-the-art methods on benchmark datasets. Qi Li 0005, Zhenan Sun, Ran He 0001, Tieniu Tan |
NIPS | 4 |
| 2017 | Local structured representation for generic object detection
Junge Zhang, Kaiqi Huang, Tieniu Tan, Zhaoxiang Zhang 0001 |
Frontiers Comput. Sci. | 3 |
| 2017 | High quality depth map estimation of object surface from light-field images
Fei Liu 0031, Guangqi Hou, Zhenan Sun, Tieniu Tan |
Neurocomputing | 4 |
| 2017 | Bin-based classifier fusion of iris and face biometrics
Di Miao, Man Zhang 0005, Zhenan Sun, Tieniu Tan, Zhaofeng He 0001 |
Neurocomputing | 4 |
| 2017 | A Comprehensive Study on Cross-View Gait Based Human Identification with Deep CNNsabstractThis paper studies an approach to gait based human identification via similarity learning by deep convolutional neural networks (CNNs). With a pretty small group of labeled multi-view human walking videos, we can train deep networks to recognize the most discriminative changes of gait patterns which suggest the change of human identity. To the best of our knowledge, this is the first work based on deep CNNs for gait recognition in the literature. Here, we provide an extensive empirical evaluation in terms of various scenarios, namely, cross-view and cross-walking-condition, with different preprocessing approaches and network architectures. The method is first evaluated on the challenging CASIA-B dataset in terms of cross-view gait recognition. Experimental results show that it outperforms the previous state-of-the-art methods by a significant margin. In particular, our method shows advantages when the cross-view angle is large, i.e., no less than 36 degree. And the average recognition rate can reach 94 percent, much better than the previous best result (less than 65 percent). The method is further evaluated on the OU-ISIR gait dataset to test its generalization ability to larger data. OU-ISIR is currently the largest dataset available in the literature for gait recognition, with 4,007 subjects. On this dataset, the average accuracy of our method under identical view conditions is above 98 percent, and the one for cross-view scenarios is above 91 percent. Finally, the method also performs the best on the USF gait dataset, whose gait sequences are imaged in a real outdoor scene. These results show great potential of this method for practical applications. Zifeng Wu, Yongzhen Huang, Liang Wang 0001, Xiaogang Wang 0001, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Exploring generalized shape analysis by topological representations
Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
Pattern Recognit. Lett. | 4 |
| 2017 | A Semi-Supervised Method for Surveillance-Based Visual Location RecognitionabstractIn this paper, we are devoted to solving the problem of crossing surveillance and mobile phone visual location recognition, especially for the case that the query and reference images are captured by mobile phone and surveillance camera, respectively. Besides, we also study the influence of the environmental condition variations on this problem. To explore that problem, we first build a cross-device location recognition dataset, which includes images of 22 locations taken by mobile phones and surveillance cameras under different time and weather conditions. Then based on careful analysis of the problems existing in the data, we specifically design a method which unifies an unsupervised subspace alignment method and the semi-supervised Laplacian support vector machine. Experiments are performed on our dataset. Compared with several related methods, our method shows to be more efficient on the problem of crossing surveillance and mobile phone visual location recognition. Furthermore, the influence of several factors such as feature, time, and weather is studied. Pengcheng Liu 0001, Peipei Yang, Kaiqi Huang, Tieniu Tan |
IEEE Trans. Cybern. | 5 |
| 2017 | A Code-Level Approach to Heterogeneous Iris RecognitionabstractMatching heterogeneous iris images in less constrained applications of iris biometrics is becoming a challenging task. The existing solutions try to reduce the difference between heterogeneous iris images in pixel intensities or filtered features. In contrast, this paper proposes a code-level approach in heterogeneous iris recognition. The non-linear relationship between binary feature codes of heterogeneous iris images is modeled by an adapted Markov network. This model transforms the number of iris templates in the probe into a homogenous iris template corresponding to the gallery sample. In addition, a weight map on the reliability of binary codes in the iris template can be derived from the model. The learnt iris template and weight map are jointly used in building a robust iris matcher against the variations of imaging sensors, capturing distance, and subject conditions. Extensive experimental results of matching cross-sensor, high-resolution versus low-resolution and, clear versus blurred iris images demonstrate the code-level approach can achieve the highest accuracy in compared with the existing pixel-level, feature-level, and score-level solutions. Nianfeng Liu, Jing Liu 0062, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | Optimized 3D Lighting Environment Estimation for Image Forgery DetectionabstractImage forgery is becoming a growing threat to information credibility. Among all kinds of image forgeries, photographic composites of human faces have very serious impacts. To combat this kind of forgery, some forensic methods propose to estimate the 3D lighting environments from different faces and investigate the consistency between them. Although they are very effective, existing 3D lighting-based forensic methods are limited by many simplifying assumptions about the surface reflection model, among which convexity and constant reflectance are two critical ones. In this paper, we propose an optimized 3D lighting estimation method by incorporating a more general surface reflection model. In this model, we relax the convexity and constant reflectance assumptions by taking the occlusion geometry and surface texture information into consideration. The proposed reflection model is more general and accurate; hence, it can achieve better lighting estimation accuracy and more reliable discrimination performance. Comprehensive experiments on both synthetic and real data sets validate the correctness and efficacy of the proposed method. Comparisons with two existing 3D lighting-based forensic methods also demonstrate the superiority of the proposed method for detecting face splicing. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | Conditional High-Order Boltzmann Machines for Supervised Relation LearningabstractRelation learning is a fundamental problem in many vision tasks. Recently, high-order Boltzmann machine and its variants have shown their great potentials in learning various types of data relation in a range of tasks. But most of these models are learned in an unsupervised way, i.e., without using relation class labels, which are not very discriminative for some challenging tasks, e.g., face verification. In this paper, with the goal to perform supervised relation learning, we introduce relation class labels into conventional high-order multiplicative interactions with pairwise input samples, and propose a conditional high-order Boltzmann Machine (CHBM), which can learn to classify the data relation in a binary classification way. To be able to deal with more complex data relation, we develop two improved variants of CHBM: 1) latent CHBM, which jointly performs relation feature learning and classification, by using a set of latent variables to block the pathway from pairwise input samples to output relation labels and 2) gated CHBM, which untangles factors of variation in data relation, by exploiting a set of latent variables to multiplicatively gate the classification of CHBM. To reduce the large number of model parameters generated by the multiplicative interactions, we approximately factorize high-order parameter tensors into multiple matrices. Then, we develop efficient supervised learning algorithms, by first pretraining the models using joint likelihood to provide good parameter initialization, and then finetuning them using conditional likelihood to enhance the discriminant ability. We apply the proposed models to a series of tasks including invariant recognition, face verification, and action similarity labeling. Experimental results demonstrate that by exploiting supervised relation labels, our models can greatly improve the performance. Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Image Process. | 4 |
| 2017 | Feature Selection Based on Structured Sparsity: A Comprehensive StudyabstractFeature selection (FS) is an important component of many pattern recognition tasks. In these tasks, one is often confronted with very high-dimensional data. FS algorithms are designed to identify the relevant feature subset from the original features, which can facilitate subsequent analysis, such as clustering and classification. Structured sparsity-inducing feature selection (SSFS) methods have been widely studied in the last few years, and a number of algorithms have been proposed. However, there is no comprehensive study concerning the connections between different SSFS methods, and how they have evolved. In this paper, we attempt to provide a survey on various SSFS methods, including their motivations and mathematical representations. We then explore the relationship among different formulations and propose a taxonomy to elucidate their evolution. We group the existing SSFS methods into two categories, i.e., vector-based feature selection (feature selection based on lasso) and matrix-based feature selection (feature selection based on lr,p-norm). Furthermore, FS has been combined with other machine learning algorithms for specific applications, such as multitask learning, multilabel learning, multiview learning, classification, and clustering. This paper not only compares the differences and commonalities of these methods based on regression and regularization strategies, but also provides useful guidelines to practitioners working in related fields to guide them how to do feature selection. Jie Gui, Zhenan Sun, Shuiwang Ji, Dacheng Tao, Tieniu Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Guest Editorial Introduction to the Special Issue on Large-Scale Video Analytics for Enhanced Security: Algorithms and SystemsabstractDue to the rapid increase of the number of cameras used in the video surveillance and the huge needs of the smart city and public security, video surveillance by human beings is no longer suitable. Hence, since the end of the last century, video analytics for security or visual surveillance has become one of the hottest research topics. Wide-area video surveillance systems can have extremely high data rates and high data volumes. Therefore, the challenge of video analytics is to extract meaningful information efficiently from the huge flow of video data in order to produce high-level semantic descriptions of the activities occurring in the area under surveillance. Kaiqi Huang, Tieniu Tan, Stephen J. Maybank, Rama Chellappa, Jake Aggarval |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2016 | Predicting the Next Location: A Recurrent Model with Spatial and Temporal ContextsabstractSpatial and temporal contextual information plays a key role for analyzing user behaviors, and is helpful for predicting where he or she will go next. With the growing ability of collecting information, more and more temporal and spatial contextual information is collected in systems, and the location prediction problem becomes crucial and feasible. Some works have been proposed to address this problem, but they all have their limitations. Factorizing Personalized Markov Chain (FPMC) is constructed based on a strong independence assumption among different factors, which limits its performance. Tensor Factorization (TF) faces the cold start problem in predicting future actions. Recurrent Neural Networks (RNN) model shows promising performance comparing with PFMC and TF, but all these methods have problem in modeling continuous time interval and geographical distance. In this paper, we extend RNN and propose a novel method called Spatial Temporal Recurrent Neural Networks (ST-RNN). ST-RNN can model local temporal and spatial contexts in each layer with time-specific transition matrices for different time intervals and distance-specific transition matrices for different geographical distances. Experimental results show that the proposed ST-RNN model yields significant improvements over the competitive compared methods on two typical datasets, i.e., Global Terrorism Database (GTD) and Gowalla dataset. Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
AAAI | 4 |
| 2016 | SAPE: A System for Situation-Aware Public Security EvaluationabstractPublic security events are occurring all over the world, bringing threat to personal and property safety, and homeland security. It is vital to construct an effective model to evaluate and predict the public security. In this work, we establish a Situation-Aware Public Security Evaluation (SAPE) platform. Based on conventional Recurrent Neural Networks (RNN), we develop a new variant of RNN to handle temporal contexts in public security event datasets. The proposed model can achieve better performance than the compared state-of-the-art methods. On SAPE, There are two parts of demonstrations, i.e., global public security evaluation and China public security evaluation. In the global part, based on Global Terrorism Database from UMD, for each country, SAPE can predict risk level and top-n potential terrorist organizations which might attack the country. The users can also view the actual attacking organizations and predicted results. For each province in China, SAPE can predict the risk level and the probability scores of different types of events in the next month. The users can also view the actual numbers of events and predicted risk levels of the past one year. Qiang Liu 0006, Ping Bai, Liang Wang 0001, Tieniu Tan |
AAAI | 5 |
| 2016 | Information Credibility Evaluation on Social MediaabstractWith the growing online social media, rumors are spread fast and viewed by more and more people on the Internet. Rumors bring significant harm to daily life and public security. It is crucial to evaluate the credibility of information and detect the rumors on social media automatically. In this work, we establish a Network Information Credibility Evaluation (NICE) platform, which collects a database of rumors that have been verified on Sina Weibo and automatically evaluates the information generated by users on social media but has not been verified. Users can use a query to search related information. If the according information appears in our database, users can identify it is a rumor immediately. Otherwise, NICE will show users with real-time results crawled automatically from social media and can calculate credibility of a specific result with our algorithm. Our algorithm learns dynamic representations for information on social media based on behavior information, dynamic information, user information and comment information. Then, we use an ordinary logistic regression to classify information into rumors and non-rumors. Based on our algorithm, NICE system achieves satisfactory performance on evaluating information credibility and detecting rumors on social media. Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
AAAI | 5 |
| 2016 | Simultaneous Feature and Sample Reduction for Image-Set ClassificationabstractImage-set classification is the assignment of a label to a given image set. In real-life scenarios such as surveillance videos, each image set often contains much redundancy in terms of features and samples. This paper introduces a joint learning method for image-set classification that simultaneously learns compact binary codes and removes redundant samples. The joint objective function of our model mainly includes two parts. The first part seeks a hashing function to generate binary codes that have larger inter-class and smaller intra-class distances. The second one reduces redundant samples with discrete constraints in a low-rank way. A kernel method based on anchor points is further used to reduce sample variations. The proposed discrete objective function is simplified to a series of sub-problems that admit an analytical solution, resulting in a high-quality discrete solution with a low computational cost. Experiments on three commonly used image-set datasets show that the proposed method for the tasks of face recognition from image sets is efficient and effective. Man Zhang 0005, Ran He 0001, Zhenan Sun, Tieniu Tan |
AAAI | 5 |
| 2016 | ReD-SFA: Relation Discovery Based Slow Feature Analysis for Trajectory ClusteringabstractFor spectral embedding/clustering, it is still an open problem on how to construct an relation graph to reflect the intrinsic structures in data. In this paper, we proposed an approach, named Relation Discovery based Slow Feature Analysis (ReD-SFA), for feature learning and graph construction simultaneously. Given an initial graph with only a few nearest but most reliable pairwise relations, new reliable relations are discovered by an assumption of reliability preservation, i.e., the reliable relations will preserve their reliabilities in the learnt projection subspace. We formulate the idea as a cross entropy (CE) minimization problem to reduce the discrepancy between two Bernoulli distributions parameterized by the updated distances and the existing relation graph respectively. Furthermore, to overcome the imbalanced distribution of samples, a Boosting-like strategy is proposed to balance the discovered relations over all clusters. To evaluate the proposed method, extensive experiments are performed with various trajectory clustering tasks, including motion segmentation, time series clustering and crowd detection. The results demonstrate that ReDSFA can discover reliable intra-cluster relations with high precision, and competitive clustering performance can be achieved in comparison with state-of-the-art. Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan, Peipei Yang, Jun Li 0010 |
CVPR | 3 |
| 2016 | Automatic detection of 3D lighting inconsistencies via a facial landmark based morphable modelabstractExisting 3D lighting consistency based forensic methods have some practical problems. They usually require additional images and human labor to reconstruct the 3D face model for lighting estimation, and furthermore, they cannot deal with expressional faces effectively. These drawbacks make them unusable in many practical cases. In this paper, we propose a more practical 3D lighting based forensic method by incorporating a facial landmark based 3D morphable model to efficiently fit the face shape. We also introduce a residual error based algorithm to automatically exclude outliers in lighting estimation. Our proposed method is fully automatic and very efficient compared to previous ones. Also, it does not depend on additional images and has better performance for expressional faces. Experiments on a realistic face dataset with variational lighting conditions indicate the efficacy and superiority of our method. Bo Peng 0002, Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 4 |
| 2016 | Learning and transferring representations for image steganalysis using convolutional neural networkabstractThe major challenge of machine learning based image steganalysis lies in obtaining powerful feature representations. Recently, Qian et al. have shown that Convolutional Neural Network (CNN) is effective for learning features automatically for steganalysis. In this paper, we follow up this new paradigm in steganalysis, and propose a framework based on transfer learning to help the training of CNN for steganalysis, hence to achieve a better performance. We show that feature representations learned with a pre-trained CNN for detecting a steganographic algorithm with a high payload can be efficiently transferred to improve the learning of features for detecting the same steganographic algorithm with a low pay-load. By detecting representative WOW and S-UNIWARD steganographic algorithms, we demonstrate that the proposed scheme is effective in improving the feature learning in CNN models for steganalysis. Yinlong Qian, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 4 |
| 2016 | A simple and robust super resolution method for light field imagesabstractLight field cameras generate low-resolution images due to the tradeoff between spatial and angular resolution. Traditional light field super-resolution (LFSR) methods depend on prior knowledge of depth information. This paper presents a projection-based LFSR solution without prior information based on redefinition of the mapping function between disparity and shearing shift. Moreover, simplified variational regularization is imposed in global optimization formulation to the rendered high-resolution images. Both a synthetic dataset and a real-world dataset of light field images captured by a self-developed light field camera are used to demonstrate the state-of-the-art performance of the proposed method. Yunlong Wang 0003, Guangqi Hou, Zhenan Sun, Zilei Wang, Tieniu Tan |
ICIP | 5 |
| 2016 | Group-Invariant Cross-Modal Subspace Learning
Jian Liang 0001, Ran He 0001, Zhenan Sun, Tieniu Tan |
IJCAI | 4 |
| 2016 | A Dynamic Recurrent Model for Next Basket RecommendationabstractNext basket recommendation becomes an increasing concern. Most conventional models explore either sequential transaction features or general interests of users. Further, some works treat users' general interests and sequential behaviors as two totally divided matters, and then combine them in some way for next basket recommendation. Moreover, the state-of-the-art models are based on the assumption of Markov Chains (MC), which only capture local sequential features between two adjacent baskets. In this work, we propose a novel model, Dynamic REcurrent bAsket Model (DREAM), based on Recurrent Neural Network (RNN). DREAM not only learns a dynamic representation of a user but also captures global sequential features among baskets. The dynamic representation of a specific user can reveal user's dynamic interests at different time, and the global sequential features reflect interactions of all baskets of the user over time. Experiment results on two public datasets indicate that DREAM is more effective than the state-of-the-art models for next basket recommendation. Feng Yu 0001, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
SIGIR | 5 |
| 2016 | Learning Relevance Restricted Boltzmann Machine for Unstructured Group Activity and Event Understanding
Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tao Xiang 0002, Tieniu Tan |
Int. J. Comput. Vis. | 5 |
| 2016 | Personalized ranking with pairwise Factorization Machines
Weiyu Guo, Liang Wang 0001, Tieniu Tan |
Neurocomputing | 4 |
| 2016 | Weakly Supervised Large Scale Object Localization with Multiple Instance Learning and Bag SplittingabstractLocalizing objects of interest in images when provided with only image-level labels is a challenging visual recognition task. Previous efforts have required carefully designed features and have difficulty in handling images with cluttered backgrounds. Up-scaling to large datasets also poses a challenge to applying these methods to real applications. In this paper, we propose an efficient and effective learning framework called MILinear, which is able to learn an object localization model from large-scale data without using bounding box annotations. We integrate rich general prior knowledge into a learning model using a large pre-trained convolutional network. Moreover, to reduce ambiguity in positive images, we present a bag-splitting algorithm that iteratively generates new negative bags from positive ones. We evaluate the proposed approach on the challenging Pascal VOC 2007 dataset, and our method outperforms other state-of-the-art methods by a large margin; some results are even comparable to fully supervised models trained with bounding box annotations. To further demonstrate scalability, we also present detection results on the ILSVRC 2013 detection dataset, and our method outperforms supervised deformable part-based model without using box annotations. Weiqiang Ren, Kaiqi Huang, Dacheng Tao, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Joint Feature Selection and Subspace Learning for Cross-Modal RetrievalabstractCross-modal retrieval has recently drawn much attention due to the widespread existence of multimodal data. It takes one type of data as the query to retrieve relevant data objects of another type, and generally involves two basic problems: the measure of relevance and coupled feature selection. Most previous methods just focus on solving the first problem. In this paper, we aim to deal with both problems in a novel joint learning framework. To address the first problem, we learn projection matrices to map multimodal data into a common subspace, in which the similarity between different modalities of data can be measured. In the learning procedure, the l21-norm penalties are imposed on the projection matrices separately to solve the second problem, which selects relevant and discriminative features from different feature spaces simultaneously. A multimodal graph regularization term is further imposed on the projected data,which preserves the inter-modality and intra-modality similarity relationships.An iterative algorithm is presented to solve the proposed joint learning problem, along with its convergence analysis. Experimental results on cross-modal retrieval tasks demonstrate that the proposed method outperforms the state-of-the-art subspace approaches. Kaiye Wang, Ran He 0001, Liang Wang 0001, Wei Wang 0115, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2016 | Transformation invariant subspace clustering
Qi Li 0005, Zhenan Sun, Zhouchen Lin, Ran He 0001, Tieniu Tan |
Pattern Recognit. | 5 |
| 2016 | DeepIris: Learning pairwise filter bank for heterogeneous iris verification
Nianfeng Liu, Man Zhang 0005, Zhenan Sun, Tieniu Tan |
Pattern Recognit. Lett. | 5 |
| 2016 | Representative Vector Machines: A Unified Framework for Classical ClassifiersabstractClassifier design is a fundamental problem in pattern recognition. A variety of pattern classification methods such as the nearest neighbor (NN) classifier, support vector machine (SVM), and sparse representation-based classification (SRC) have been proposed in the literature. These typical and widely used classifiers were originally developed from different theory or application motivations and they are conventionally treated as independent and specific solutions for pattern classification. This paper proposes a novel pattern classification framework, namely, representative vector machines (or RVMs for short). The basic idea of RVMs is to assign the class label of a test example according to its nearest representative vector. The contributions of RVMs are twofold. On one hand, the proposed RVMs establish a unified framework of classical classifiers because NN, SVM, and SRC can be interpreted as the special cases of RVMs with different definitions of representative vectors. Thus, the underlying relationship among a number of classical classifiers is revealed for better understanding of pattern classification. On the other hand, novel and advanced classifiers are inspired in the framework of RVMs. For example, a robust pattern classification method called discriminant vector machine (DVM) is motivated from RVMs. Given a test example, DVM first finds its k -NNs and then performs classification based on the robust M-estimator and manifold regularization. Extensive experimental evaluations on a variety of visual recognition tasks such as face recognition (Yale and face recognition grand challenge databases), object categorization (Caltech-101 dataset), and action recognition (Action Similarity LAbeliNg) demonstrate the advantages of DVM over other classifiers. Jie Gui, Tongliang Liu, Dacheng Tao, Zhenan Sun, Tieniu Tan |
IEEE Trans. Cybern. | 5 |
| 2016 | Coupled Topic Model for Collaborative Filtering With User-Generated ContentabstractThe user-generated content (UGC) is a type of dyadic information that provides description of the interaction between users and items (such as rating, purchasing, etc.). Most conventional methods incorporate either a user profile or the item description, which cannot well utilize this kind of content information. Some other works jointly consider user ratings and reviews, but they are based on the factorization technique and have difficulty in providing explanations on generated recommendations. In this study, a coupled topic model (CoTM) for recommendation with UGC is developed. By combining UGC and ratings, the method discussed in this study captures both the content-based preferences and collaborative preferences and, thus, can explain both the user and item latent spaces using the topics discovered from the UGC. The learned topics in CoTM can also serve as proper explanations for the generated recommendations. Experimental results show that the proposed CoTM model yields significant improvements over the compared competitive methods on two typical datasets, that is, MovieLens-10M and Citation-network V1. The topics discovered by CoTM can be used not only to illustrate the topic distributions of users and items, but also to explain the generated user-item recommendations. Weiyu Guo, Song Xu 0002, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2016 | Complementary Cohort Strategy for Multimodal Face Pair MatchingabstractFace pair matching is the task of determining whether two face images represent the same person. Due to the limited expressive information embedded in the two face images as well as various sources of facial variations, it becomes a quite difficult problem. Toward the issue of few available images provided to represent each face, we propose to exploit an extra cohort set (identities in the cohort set are different from those being compared) by a series of cohort list comparisons. Useful cohort coefficients are then extracted from both sorted cohort identities and sorted cohort images for complementary information. To augment its robustness to complicated facial variations, we further employ multiple face modalities owing to their complementary value to each other for the face pair matching task. The final decision is made by fusing the extracted cohort coefficients with the direct matching score for all the available face modalities. To investigate the capacity of each individual modality on matching faces, the cohort behavior, and the performance achieved using our complementary cohort strategy, we conduct a set of experiments on two recently collected multimodal face databases. It is shown that using different modalities leads to different face pair matching performance. For each modality, employing our cohort scheme significantly reduces the equal error rate. By applying the proposed multimodal complementary cohort strategy, we achieve the best performance on our face pair matching task. Yunlian Sun, Kamal Nasrollahi, Zhenan Sun, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2016 | Contextual Operation for Recommender SystemsabstractWith the rapid growth of various applications on the Internet, recommender systems become fundamental for helping users alleviate the problem of information overload. Since contextual information is a significant factor in modeling the user behavior, various context-aware recommendation methods have been proposed recently. The state-of-the-art context modeling methods usually treat contexts as certain dimensions similar to those of users and items, and capture relevances between contexts and users/items. However, such kind of relevance has much difficulty in explanation. Some works on multi-domain relation prediction can also be used for the context-aware recommendation, but they have limitations in generating recommendations under a large amount of contextual information. Motivated by recent works in natural language processing, we represent each context value with a latent vector, and model the contextual information as a semantic operation on the user and item. Besides, we use the contextual operating tensor to capture the common semantic effects of contexts. Experimental results show that the proposed Context Operating Tensor (COT) model yields significant improvements over the competitive compared methods on three typical datasets. From the experimental results of COT, we also obtain some interesting observations which follow our intuition. Qiang Liu 0006, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Multiple Attribute Aware Personalized Ranking
Weiyu Guo, Liang Wang 0001, Tieniu Tan |
APWeb | 4 |
| 2015 | Cross-Domain Object Recognition Using Object AlignmentabstractIn this paper, we focus on the problem of cross-domain object recognition [4], which has long been one of the challenging problems in computer vision. This problem typically arises when training (source domain) and test (target domain) samples are drawn from different distributions. In the problem of object recognition, this case is usually caused by the situation that training and test samples are acquired under different sets of background, lighting, view point, resolution conditions, etc. One popular solution to the problem of cross-domain object recognition is minimizing the difference between the source and target distributions. Existing methods are devoted to minimizing that domain difference in a complex image space, which makes the problem hard to solve because of background influence, as shown in Figure 1 (a). Since the object and background are twisted in that image feature space, the discrepancy caused by background is difficult to eliminate, which makes it hard to learn optimal fS and fT for minimizing D( fS(XS), fT (XT )). To discount the influence of the background, we propose to minimize that difference using object alignment. As shown in Figure 1 (b), we minimize the domain difference by transferring to the feature space of aligned objects XS and XT , but not the image feature space having background influence. The key insight of our approach is that the difference between the source and target distributions can be reduced by discounting the influence from the ambiguous background. We define the semantic object as the object that occurs in all the images of one class. To discount the background influence, our primary goal is to automatically localize the semantic object so that the irrelevant background can be eliminated. Then based on the semantic object regions, we can learn an object detector that is robust to the influence of the irrelevant background and makes the crossdomain object recognition much easier than before. In addition, since our detectors are learned in a weakly supervised way, we utilize the classificaSource domain Selective search Object alignment — Topic discovery Pengcheng Liu 0001, Peipei Yang, Kaiqi Huang, Tieniu Tan |
BMVC | 5 |
| 2015 | Social-Relational Topic Model for Social NetworksabstractSocial networking services, such as Twitter and Sina Weibo, have tremendous popularity in recent years. Mass of short texts and social links are aggregated into these service platforms. To realize personalized services on social network, topic inference from both short texts and social links plays more and more important role. Most conventional topic modeling methods focus on analyzing formal texts, e.g., papers, news and blogs, and usually assume that the links are only generated by topical factors. As a result, on social network, the learned topics of these methods are usually affected by topic-irrelevant links. Recently, a few approaches use artificial priors to recognize the links generated by the popularity factor in topic modeling. However, employing global priors, these methods can not well capture the distinct properties of each link and still suffer from the effect of topic-irrelevant links. To address the above limitations, we propose a novel Social-Relational Topic Model (SRTM), which can alleviate the effect of topic-irrelevant links by analyzing relational users' topics of each link. SRTM jointly models texts and social links for learning the topic distribution and topical influence of each user. The experimental results show that, our model outperforms the state-of-the-arts in topic modeling and social link prediction. Weiyu Guo, Liang Wang 0001, Tieniu Tan |
CIKM | 4 |
| 2015 | Multi-view Clustering via Structured Low-rank RepresentationabstractIn this paper, we present a novel solution to multi-view clustering through a structured low-rank representation. When assuming similar samples can be linearly reconstructed by each other, the resulting representational matrix reflects the cluster structure and should ideally be block diagonal. We first impose low-rank constraint on the representational matrix to encourage better grouping effect. Then representational matrices under different views are allowed to communicate with each other and share their mutual cluster structure information. We develop an effective algorithm inspired by iterative re-weighted least squares for solving our formulation. During the optimization process, the intermediate representational matrix from one view serves as a cluster structure constraint for that from another view. Such mutual structural constraint fine-tunes the cluster structures from both views and makes them more and more agreeable. Extensive empirical study manifests the superiority and efficacy of the proposed method. Dong Wang 0004, Qiyue Yin, Ran He 0001, Liang Wang 0001, Tieniu Tan |
CIKM | 5 |
| 2015 | Deep semantic ranking based hashing for multi-label image retrievalabstractWith the rapid growth of web images, hashing has received increasing interests in large scale image retrieval. Research efforts have been devoted to learning compact binary codes that preserve semantic similarity based on labels. However, most of these hashing methods are designed to handle simple binary similarity. The complex multi-level semantic structure of images associated with multiple labels have not yet been well explored. Here we propose a deep semantic ranking based method for learning hash functions that preserve multilevel semantic similarity between multi-label images. In our approach, deep convolutional neural network is incorporated into hash functions to jointly learn feature representations and mappings from them to hash codes, which avoids the limitation of semantic representation power of hand-crafted features. Meanwhile, a ranking list that encodes the multilevel similarity information is employed to guide the learning of such deep hash functions. An effective scheme based on surrogate loss is used to solve the intractable optimization problem of nonsmooth and multivariate ranking measures involved in the learning procedure. Experimental results show the superiority of our proposed approach over several state-of-the-art hashing methods in term of ranking evaluation metrics when tested on multi-label image datasets. Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
CVPR | 4 |
| 2015 | Albedo assisted high-quality shape recovery from 4D light fieldsabstractOver the past decade, shape reconstruction methods have been limited to Lambertian reflectance with uniform albedo and controlled lighting environment. In this paper, we present an approach for recovering high-quality shapes from 4D light fields, which can handle non-Lambertian and multi-albedo scenes with shadows and inter-reflections. 4D light fields represent all light rays that hit the sensor plane from different directions, and the depth map from light fields is robust to non-Lambertian objects. Specifically, we estimate the albedos by eliminating shadows and inter-reflections with the edge and chromaticity. Then the lighting environment is analyzed from albedos and shading. Finally, the high-quality surface geometry is exactly recovered through normal refinement. We evaluate the effectiveness and robustness on the public 4D light fields database with both synthetic and real-world scenes. Fei Liu 0031, Guangqi Hou, Zhenan Sun, Tieniu Tan |
ICIP | 4 |
| 2015 | Learning occlusion patterns using semantic phrases for object detectionabstractOcclusion inference in image is a classical as well as difficult problem in computer vision. Most approaches model occlusion with different occlusion patterns in complex ways. In this paper, we propose a simple and efficient way to represent occlusion patterns called `occlusion pattern phrase', for example `dog occlude person'. These phrases model the occlusion patterns between two occluded objects, which can be used in applications like object detection and object classification. Here, we focus on using the occlusion pattern phrases in object detection. DPM is used to learn appearance models and a inference procedure is introduced to infer occlusion patterns with structural outputs. Unlike other methods, our method produces not only location boxes for different objects, but also demonstrates their occlusion relations. A new dataset with well annotated occlusion pattern images collected from Pascal VOC2007 and search engines like Bing and Google is introduced in this paper. Experiments show that our method outperforms the baseline and the occlusion pattern phrases can describe the relations between objects as we expect. In the future, we will explore the use of occlusion pattern phrases in scene understanding. Jinde Liu, Kaiqi Huang, Tieniu Tan |
ICIP | 3 |
| 2015 | Robust steganalysis based on training set construction and ensemble classifiers weightingabstractThe cover source mismatch problem in steganalysis is a serious problem which keeps current steganalysis from practical use. It is mainly because of the high intra-class variation of cover and stego samples in the feature space, since current steganalytic features are inevitably affected much by the image content, size, quality and many other factors. Small training set often reflects only part of the real data distribution, hence the classifier (steganalyzer) may be undertrained and lack of robustness. In this paper, we propose a scheme to efficiently construct large representative training set for steganalysis. We also scheme out weighted ensemble classifiers which can be adaptive to testing data. Experimental results show that our method can improve the performance and robustness of ste-ganalysis under high intra-class variation. Xikai Xu, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 4 |
| 2015 | Semi-supervised learning and feature evaluation for RGB-D object recognition
Yanhua Cheng, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan |
Comput. Vis. Image Underst. | 4 |
| 2015 | Learning predictable binary codes for face indexing
Ran He 0001, Yinghao Cai, Tieniu Tan, Larry Davis 0001 |
Pattern Recognit. | 3 |
| 2014 | Deformable Object Matching via Deformation Decomposition Based 2D Label MRFabstractDeformable object matching, which is also called elastic matching or deformation matching, is an important and challenging problem in computer vision. Although numerous deformation models have been proposed in different matching tasks, not many of them investigate the intrinsic physics underlying deformation. Due to the lack of physical analysis, these models cannot describe the structure changes of deformable objects very well. Motivated by this, we analyze the deformation physically and propose a novel deformation decomposition model to represent various deformations. Based on the physical model, we formulate the matching problem as a two-mensional label Markov Random Field. The MRF energy function is derived from the deformation decomposition model. Furthermore, we propose a two-stage method to optimize the MRF energy function. To provide a quantitative benchmark, we build a deformation matching database with an evaluation criterion. Experimental results show that our method outperforms previous approaches especially on complex deformations. Junge Zhang, Kaiqi Huang, Tieniu Tan |
CVPR | 4 |
| 2014 | Weakly Supervised Object Localization with Latent Category Learning
Weiqiang Ren, Kaiqi Huang, Tieniu Tan |
ECCV (6) | 4 |
| 2014 | An optimal set of code words and correntropy for rotated least squares regressionabstractThis paper presents a robust feature extraction method for face recognition based on least squares regression (LSR). Our focus is to enhance the robustness and discriminability of the LSR. First, an optimal set of code words is introduced in LSR. Compared to the traditional set of code words, this new set uses less number of code words. Furthermore, it can make the distance of the regression targets of different classes as large as possible. Then, correntropy is integrated into the LSR model for better robustness. Furthermore, considering the commonly used distance metrics such as Euclidean distance and Cosine distance in the subspace are invariant to rotation transformation, rotation is introduced as additional freedom to promote flexibility without sacrificing accuracy. Our objective function is optimized using half-quadratic (HQ) optimization, which facilitates algorithm development and convergence study. Experimental results show that our method outperforms several subspace methods for face recognition, which indicates the validity of the proposed method. Jie Gui, Zhenan Sun, Guangqi Hou, Tieniu Tan |
IJCB | 4 |
| 2014 | Efficient auto-refocusing of iris images for light-field camerasabstractLight field photography provides a revolutionary possibility to reconstruct well-focused iris region from a 4D light-field image. However, such a “shoot and refocus” scheme is time-consuming in practice because it commonly needs to render an image sequence for finding the optimally refocused frame. This paper presents an efficient auto-refocusing iris imaging solution for lenselet-based light-field cameras. Firstly, a refocusing point spread function (R-PSF) is derived by detailed analysis of the relationship between refocusing depth and defocus blurriness. Secondly, an initial image is rendered at arbitrary depth. Thirdly, a content independent blurriness assessment method based on SVR (support vector regression) modeling is performed on the rendered image to locate depth shift from optimal focusing plane based on R-PSF. Finally, the optimally focused iris image is selected from a frontal candidate and a back candidate. Because our method only involves three times of image rendering based on precise localization of the optimal focusing plane, it is much more efficient than conventional “rendering and selection” solutions which need to render a large number of refocused images. Chi Zhang 0060, Guangqi Hou, Zhenan Sun, Tieniu Tan |
IJCB | 4 |
| 2014 | The first ICB* competition on iris recognitionabstractIris recognition becomes an important technology in our society. Visual patterns of human iris provide rich texture information for personal identification. However, it is greatly challenging to match intra-class iris images with large variations in unconstrained environments because of noises, illumination variation, heterogeneity and so on. To track current state-of-the-art algorithms in iris recognition, we organized the first ICB* Competition on Iris Recognition in 2013 (or ICIR2013 shortly). In this competition, 8 participants from 6 countries submitted 13 algorithms totally. All the algorithms were trained on a public database (e.g. CASIA-Iris-Thousand [3]) and evaluated on an unpublished database. The testing results in terms of False Non-match Rate (FNMR) when False Match Rate (FMR) is 0.0001 are taken to rank the submitted algorithms. Man Zhang 0005, Jing Liu 0062, Zhenan Sun, Tieniu Tan, Wu Su, Fernando Alonso-Fernandez, Valérian Némesin, Nadia Othman, Koichi Noda, Peihua Li, Edmundo Hoyle, Akanksha Joshi |
IJCB | 4 |
| 2014 | An effective watermarking method against valumetric distortionsabstractMost of the quantization based watermarking algorithms are very sensitive to valumetric distortions, while these distortions are regarded as common processing in audio/video analysis. In recent years, watermarking methods which can resist this kind of distortions have attracted a lot of interests. But still many proposed methods can only deal with one certain kind of valumetric distortion as amplitude scaling, and fail in other kinds of valumetric distortions like constant change attack or gamma correction. In this paper, we propose a very simple method to tackle all the three kinds of valumetric distortions. A constant change invariant domain is first constructed by spread transform, in which the watermark is embedded using a certain amplitude scaling invariant based watermarking scheme. Several typical watermarking methods and attacks have been implemented in our experiments to demonstrate the effectiveness of the proposed method. Zairan Wang, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 4 |
| 2014 | Semi-supervised subspace segmentationabstractSubspace segmentation methods usually rely on the raw explicit feature vectors in an unsupervised manner. In many applications, it is cheap to obtain some pairwise link information that tells whether two data points are in the same subspace or not. Though partially available, such link information serves as some kind of high-level semantics, which can be further used as a constraint to improve the segmentation accuracy. By constructing a link matrix and using it as a regularizer, we propose a semi-supervised subspace segmentation model where the partially observed subspace membership prior can be encoded. Specificly, under the common linear representation assumption, we enforce the representational coefficient to be consistent with the link matrix. Thus the low-level and high-level information about the data can be integrated to produce more precise segmentation results. We then develop an effective algorithm to optimize our model in an alternating minimization way. Experimental results for both motion segmentation and face clustering validate that incorporating such link information is helpful to assist and bias the unsupervised subspace segmentation methods. Dong Wang 0004, Qiyue Yin, Ran He 0001, Liang Wang 0001, Tieniu Tan |
ICIP | 5 |
| 2014 | Semi-supervised Learning for RGB-D Object RecognitionabstractConventional supervised object recognition methods have been investigated for many years. Despite their successes, there are still two suffering limitations: (1) various information of an object is represented by artificial features only derived from RGB images, (2) lots of manually labeled data is required by supervised learning. To address those limitations, we propose a new semi-supervised learning framework based on RGB and depth (RGB-D) images to improve object recognition. In particular, our framework has two modules: (1) RGB and depth images are represented by convolutional-recursive neural networks to construct high level features, respectively, (2) co-training is exploited to make full use of unlabeled RGB-D instances due to the existing two independent views. Experiments on the standard RGB-D object dataset demonstrate that our method can compete against with other state-of-the-art methods with only 20% labeled data. Yanhua Cheng, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan |
ICPR | 4 |
| 2014 | A General Nonlinear Embedding Framework Based on Deep Neural NetworkabstractRecently there has been increasing interest in deep neural network due to its powerful represent ability in several successful applications such as speech recognition and image classification. In this paper, we propose a general nonlinear embedding framework based on deep neural network which can be utilized to implement a family of dimensionality reduction algorithms. The objective function of our framework consists of two terms: 1) an embedding term which transforms the input to a low-dimensional representation with a multilayer network, and 2) a regularization term which computes the reconstruction error of the original input by unrolling the multilayer network to a deep auto encoder. We adopt a layer-by-layer pretraining procedure to obtain good initial weights for the network, and then minimize the objective function by back propagating derivatives of the two terms. To evaluate the proposed framework, we perform face recognition and digit classification experiments. The experiments demonstrate that the proposed framework achieves better results than the state-of-the-art algorithms. The success of our framework further verifies deep neural network's advantages in representation learning. Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
ICPR | 4 |
| 2014 | Semi-supervised Learning for Cross-Device Visual Location RecognitionabstractThe aim of this work is to localize a query mobile photograph by utilizing surveillance images, which naturally provide location information. We cast this cross-device visual localization problem as a classification task. By exploiting the surveillance network to collect reference images, the data acquisition process is significantly facilitated. However, the discrepancy between mobile images and surveillance images makes the training samples difficult to be used directly, and the scarcity of training samples caused by the immobility of surveillance cameras further degrades the performance. In contrast to most traditional domain adaptation problems and semi-supervised problems, the scarce labeled data and plentiful unlabeled data exist in different domains. Our location recognition method first exploits the unsupervised subspace alignment to weaken the discrepancy between the two domains, and then adopts the semi-supervised Laplacian SVM to reinforce the discriminant information utilizing the unlabeled mobile images. Experimental results show that our location recognition method significantly outperforms other related methods. Pengcheng Liu 0001, Peipei Yang, Kaiqi Huang, Tieniu Tan, Hongwei Hao |
ICPR | 4 |
| 2014 | Early Hierarchical Contexts Learned by Convolutional Networks for Image SegmentationabstractWe propose a foreground segmentation method based on convolutional networks. To predict the label of a pixel in an image, the model takes a hierarchical context as the input, which is obtained by combining multiple context patches on different scales. Short range contexts depict the local details, while long range contexts capture the object-scene relationships in an image. Early means that we combine the context patches of a pixel into a hierarchical one before any trainable layers are learned, i.e., early-combing. In contrast, late-combing means that the combination occurs later, e.g., when the convolutional feature extractor in a network has already been learned. We find that it is vital for the whole model to jointly learn the patterns of contexts on different scales in our task. Experiments show that early-combing performs better than late-combing. On the dataset1 built up by Baidu IDL2 for a latest person segmentation contest, our method beats all the competitors with a considerable margin. Qualitative results also show that the proposed method is almost ready for practical application. Zifeng Wu, Yongzhen Huang, Yinan Yu, Liang Wang 0001, Tieniu Tan |
ICPR | 5 |
| 2014 | Effects of Fragile and Semi-fragile Watermarking on Iris Recognition System
Zairan Wang, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
IWDW | 4 |
| 2014 | Slice representation of range data for head pose estimation
Yunqi Tang, Zhenan Sun, Tieniu Tan |
Comput. Vis. Image Underst. | 3 |
| 2014 | Distance metric learning for recognizing low-resolution iris images
Jing Liu 0062, Zhenan Sun, Tieniu Tan |
Neurocomputing | 3 |
| 2014 | Special issue on "Multi-biometrics and Mobile-biometrics: Recent Advances and Future Research"
Lei Zhang 0006, Tieniu Tan, Arun Ross, Stefanos Zafeiriou |
Image Vis. Comput. | 2 |
| 2014 | Robust Recovery of Corrupted Low-RankMatrix by Implicit RegularizersabstractLow-rank matrix recovery algorithms aim to recover a corrupted low-rank matrix with sparse errors. However, corrupted errors may not be sparse in real-world problems and the relationship between ℓ1 regularizer on noise and robust M-estimators is still unknown. This paper proposes a general robust framework for low-rank matrix recovery via implicit regularizers of robust M-estimators, which are derived from convex conjugacy and can be used to model arbitrarily corrupted errors. Based on the additive form of half-quadratic optimization, proximity operators of implicit regularizers are developed such that both low-rank structure and corrupted errors can be alternately recovered. In particular, the dual relationship between the absolute function in ℓ1 regularizer and Huber M-estimator is studied, which establishes a connection between robust low-rank matrix recovery methods and M-estimators based robust principal component analysis methods. Extensive experiments on synthetic and real-world data sets corroborate our claims and verify the robustness of the proposed framework. Ran He 0001, Tieniu Tan, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Half-Quadratic-Based Iterative Minimization for Robust Sparse RepresentationabstractRobust sparse representation has shown significant potential in solving challenging problems in computer vision such as biometrics and visual surveillance. Although several robust sparse models have been proposed and promising results have been obtained, they are either for error correction or for error detection, and learning a general framework that systematically unifies these two aspects and explores their relation is still an open problem. In this paper, we develop a half-quadratic (HQ) framework to solve the robust sparse representation problem. By defining different kinds of half-quadratic functions, the proposed HQ framework is applicable to performing both error correction and error detection. More specifically, by using the additive form of HQ, we propose an ℓ1-regularized error correction method by iteratively recovering corrupted data from errors incurred by noises and outliers; by using the multiplicative form of HQ, we propose an ℓ1-regularized error detection method by learning from uncorrupted data iteratively. We also show that the ℓ1-regularization solved by soft-thresholding function has a dual relationship to Huber M-estimator, which theoretically guarantees the performance of robust sparse representation in terms of M-estimation. Experiments on robust face recognition under severe occlusion and corruption validate our framework and findings. Ran He 0001, Wei-Shi Zheng 0001, Tieniu Tan, Zhenan Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Feature Coding in Image Classification: A Comprehensive StudyabstractImage classification is a hot topic in computer vision and pattern recognition. Feature coding, as a key component of image classification, has been widely studied over the past several years, and a number of coding algorithms have been proposed. However, there is no comprehensive study concerning the connections between different coding methods, especially how they have evolved. In this paper, we first make a survey on various feature coding methods, including their motivations and mathematical representations, and then exploit their relations, based on which a taxonomy is proposed to reveal their evolution. Further, we summarize the main characteristics of current algorithms, each of which is shared by several coding strategies. Finally, we choose several representatives from different kinds of coding approaches and empirically evaluate them with respect to the size of the codebook and the number of training samples on several widely used databases (15-Scenes, Caltech-256, PASCAL VOC07, and SUN397). Experimental findings firmly justify our theoretical analysis, which is expected to benefit both practical applications and future research. Yongzhen Huang, Zifeng Wu, Liang Wang 0001, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Iris Image Classification Based on Hierarchical Visual CodebookabstractIris recognition as a reliable method for personal identification has been well-studied with the objective to assign the class label of each iris image to a unique subject. In contrast, iris image classification aims to classify an iris image to an application specific category, e.g., iris liveness detection (classification of genuine and fake iris images), race classification (e.g., classification of iris images of Asian and non-Asian subjects), coarse-to-fine iris identification (classification of all iris images in the central database into multiple categories). This paper proposes a general framework for iris image classification based on texture analysis. A novel texture pattern representation method called Hierarchical Visual Codebook (HVC) is proposed to encode the texture primitives of iris images. The proposed HVC method is an integration of two existing Bag-of-Words models, namely Vocabulary Tree (VT), and Locality-constrained Linear Coding (LLC). The HVC adopts a coarse-to-fine visual coding strategy and takes advantages of both VT and LLC for accurate and sparse representation of iris texture. Extensive experimental results demonstrate that the proposed iris image classification method achieves state-of-the-art performance for iris liveness detection, race classification, and coarse-to-fine iris identification. A comprehensive fake iris image database simulating four types of iris spoof attacks is developed as the benchmark for research of iris liveness detection. Zhenan Sun, Hui Zhang 0061, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Object tracking across non-overlapping views by learning inter-camera transfer models
Xiaotang Chen, Kaiqi Huang, Tieniu Tan |
Pattern Recognit. | 3 |
| 2014 | Gabor Ordinal Measures for Face RecognitionabstractGreat progress has been achieved in face recognition in the last three decades. However, it is still challenging to characterize the identity related features in face images. This paper proposes a novel facial feature extraction method named Gabor ordinal measures (GOM), which integrates the distinctiveness of Gabor features and the robustness of ordinal measures as a promising solution to jointly handle inter-person similarity and intra-person variations in face images. In the proposal, different kinds of ordinal measures are derived from magnitude, phase, real, and imaginary components of Gabor images, respectively, and then are jointly encoded as visual primitives in local regions. The statistical distributions of these visual primitives in face image blocks are concatenated into a feature vector and linear discriminant analysis is further used to obtain a compact and discriminative feature representation. Finally, a two-stage cascade learning method and a greedy block selection method are used to train a strong classifier for face recognition. Extensive experiments on publicly available face image databases, such as FERET, AR, and large scale FRGC v2.0, demonstrate state-of-the-art face recognition performance of GOM. Zhenhua Chai, Zhenan Sun, Heydi Mendez Vazquez, Ran He 0001, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2014 | Exploring DCT Coefficient Quantization Effects for Local Tampering DetectionabstractIn this paper, we focus on local image tampering detection. For a JPEG image, the probability distributions of its DCT coefficients will be disturbed by tampering operation. The tampered region and the unchanged region have different distributions, which is an important clue for locating tampering. Based on the assumption of Laplacian distribution of unquantized ac DCT coefficients, these two distributions as well as the size of tampered region can be estimated so that the probability of each DCT block being tampered is obtained. More accurate localization results could be got when we consider the prior knowledge of common tampered regions. We also design three kinds of features that can distinguish truly tampered regions from the false ones to reduce false alarm. For a tampered image which is saved in lossless compressed format, we also propose the specialized approach, which employs the quantization noise of high-frequency DCT coefficient, to improve the tampering localization performance. Extensive experiments on large scale databases prove the effectiveness of our proposed method and demonstrate that our method is suitable for locating tampered regions with different scales. Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Ordinal Feature Selection for Iris and Palmprint RecognitionabstractOrdinal measures have been demonstrated as an effective feature representation model for iris and palmprint recognition. However, ordinal measures are a general concept of image analysis and numerous variants with different parameter settings, such as location, scale, orientation, and so on, can be derived to construct a huge feature space. This paper proposes a novel optimization formulation for ordinal feature selection with successful applications to both iris and palmprint recognition. The objective function of the proposed feature selection method has two parts, i.e., misclassification error of intra and interclass matching samples and weighted sparsity of ordinal feature descriptors. Therefore, the feature selection aims to achieve an accurate and sparse representation of ordinal measures. And, the optimization subjects to a number of linear inequality constraints, which require that all intra and interclass matching pairs are well separated with a large margin. Ordinal feature selection is formulated as a linear programming (LP) problem so that a solution can be efficiently obtained even on a large-scale feature pool and training database. Extensive experimental results demonstrate that the proposed LP formulation is advantageous over existing feature selection methods, such as mRMR, ReliefF, Boosting, and Lasso for biometric recognition, reporting state-of-the-art accuracy on CASIA and PolyU databases. Zhenan Sun, Tieniu Tan |
IEEE Trans. Image Process. | 3 |
| 2014 | Introduction to the Special Section on Biometric Systems and ApplicationsabstractNowadays, biometrics is an important technological area receiving continuously growing interest from academia, industry, government, and the general public, due to the criticality and the social impact of its applications. Biometric systems are in fact rapidly being adopted in a wide variety of applications such as security, ambient intelligence, electronic and physical access control, digital rights management, background checking and defense, medical diagnosis as well as for adaptive environments. Michele Nappi, Vincenzo Piuri, Tieniu Tan, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2013 | Auto-encoder Based Data Clustering
Chunfeng Song, Feng Liu 0036, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
CIARP (1) | 5 |
| 2013 | Recent Progress on Object Classification and Detection
Tieniu Tan, Yongzhen Huang, Junge Zhang |
CIARP (2) | 1 |
| 2013 | Learning Coupled Feature Spaces for Cross-Modal MatchingabstractCross-modal matching has recently drawn much attention due to the widespread existence of multimodal data. It aims to match data from different modalities, and generally involves two basic problems: the measure of relevance and coupled feature selection. Most previous works mainly focus on solving the first problem. In this paper, we propose a novel coupled linear regression framework to deal with both problems. Our method learns two projection matrices to map multimodal data into a common feature space, in which cross-modal data matching can be performed. And in the learning procedure, the ell_21-norm penalties are imposed on the two projection matrices separately, which leads to select relevant and discriminative features from coupled feature spaces simultaneously. A trace norm is further imposed on the projected data as a low-rank constraint, which enhances the relevance of different modal data with connections. We also present an iterative algorithm based on half-quadratic minimization to solve the proposed regularized linear regression problem. The experimental results on two challenging cross-modal datasets demonstrate that the proposed method outperforms the state-of-the-art approaches. Kaiye Wang, Ran He 0001, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
ICCV | 5 |
| 2013 | Robust Subspace Clustering via Half-Quadratic MinimizationabstractSubspace clustering has important and wide applications in computer vision and pattern recognition. It is a challenging task to learn low-dimensional subspace structures due to the possible errors (e.g., noise and corruptions) existing in high-dimensional data. Recent subspace clustering methods usually assume a sparse representation of corrupted errors and correct the errors iteratively. However large corruptions in real-world applications can not be well addressed by these methods. A novel optimization model for robust subspace clustering is proposed in this paper. The objective function of our model mainly includes two parts. The first part aims to achieve a sparse representation of each high-dimensional data point with other data points. The second part aims to maximize the correntropy between a given data point and its low-dimensional representation with other points. Correntropy is a robust measure so that the influence of large corruptions on subspace clustering can be greatly suppressed. An extension of our method with explicit introduction of representation error terms into the model is also proposed. Half-quadratic minimization is provided as an efficient solution to the proposed robust subspace clustering formulations. Experimental results on Hopkins 155 dataset and Extended Yale Database B demonstrate that our method outperforms state-of-the-art subspace clustering methods. Yingya Zhang, Zhenan Sun, Ran He 0001, Tieniu Tan |
ICCV | 4 |
| 2013 | Two Notes from Experimental Study on Image Steganalysis
Qingxiao Guan, Jing Dong 0003, Tieniu Tan |
ICIC (1) | 3 |
| 2013 | Multi-task deep neural network for multi-label learningabstractThis paper proposes a multi-task deep neural network (MT-DNN) architecture to handle the multi-label learning problem, in which each label learning is defined as a binary classification task, i.e., a positive class for “an instance owns this label” and a negative class for “an instance does not own this label”. Multi-label learning is accordingly transformed to multiple binary-class classification tasks. Considering that a deep neural nets (DNN) architecture can learn good intermediate representations shared across tasks, we generalize one classification task of traditional DNN into multiple binary classification tasks through defining the output layer with a negative class node and a positive class node for each label. After a similar pretraining process to deep belief nets, we redefine the label assignment error of MT-DNN and perform the back-propagation algorithm to fine-tune the network. To evaluate the proposed model, we carry out image annotation experiments on two public image datasets, with 2000 images and 30,000 images respectively. The experiments demonstrate that the proposed model achieves the state-of-the-art performance. Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
ICIP | 4 |
| 2013 | Video steganalysis based on the constraints of motion vectorsabstractIn this paper, we focus on detecting data hiding in motion vectors of compressed video and propose a new steganalytic algorithm based on the mutual constraints of motion vectors. The constraints of motion vectors from multiple frames are analyzed and formulized by three functions, then statistical features are extracted based on these functions. Moreover, we also incorporate calibration method to improve the detection accuracy. Experimental results demonstrate that the proposed method can effectively attack typical motion-vector-based video steganography. Xikai Xu, Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
ICIP | 4 |
| 2013 | Discovering compact topical descriptors for web video retrievalabstractDescribing videos efficiently is an important task for content based web video retrieval. To solve this problem, we propose an unsupervised approach based on an undirected topic model to learn a compact topical descriptor upon the bag-of-words (BoW) video representation. In our method, words in a BoW are assumed to have different topic features, and the topical descriptor of an entire video is obtained by aggregating those features, which makes the descriptor contain information about relative strength of topics. To improve the descriptor interpretability, an L1penalty is used to control the topical sparsity. Furthermore, efficient learning and inference algorithms are presented. We evaluate the proposed descriptor on the Columbia Consumer Video dataset. Experimental results demonstrate that compared with the BoW and other topical representations, the proposed compact descriptor has better performance in web video retrieval. Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ICIP | 4 |
| 2013 | Depth-embedded multiple pooling for image classificationabstractMost existing methods of image classification ignore the role of depth information hidden in 2-D images. However, the depth information is important for visual perception, especially when the appearance information does not perform well. In this paper, we propose to embed depth information within multiple pooling into the classic platform of image classification, namely bag-of-features. The proposed method quantifies depth diversity by projecting objects to their nearby depth planes, resulting pooling features in the 3-D space indirectly. Experimental results on the MIT Indoor Scene database demonstrate that our proposed depth-embedded multiple pooling is effective to enhance the accuracy of image classification, especially when the appearance features alone are not so discriminative. Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ICIP | 4 |
| 2013 | Relevance Topic Model for Unstructured Social Group Activity RecognitionabstractUnstructured social group activity recognition in web videos is a challenging task due to 1) the semantic gap between class labels and low-level visual features and 2) the lack of labeled training data. To tackle this problem, we propose a relevance topic model" for jointly learning meaningful mid-level representations upon bag-of-words (BoW) video representations and a classifier with sparse weights. In our approach, sparse Bayesian learning is incorporated into an undirected topic model (i.e., Replicated Softmax) to discover topics which are relevant to video classes and suitable for prediction. Rectified linear units are utilized to increase the expressive power of topics so as to explain better video data containing complex contents and make variational inference tractable for the proposed model. An efficient variational EM algorithm is presented for model parameter estimation and inference. Experimental results on the Unstructured Social Activity Attribute dataset show that our model achieves state of the art performance and outperforms other supervised topic model in terms of classification accuracy, particularly in the case of a very small number of labeled training videos." Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
NIPS | 4 |
| 2013 | Practical Camera Calibration From Moving Objects for Traffic Scene SurveillanceabstractWe address the problem of camera calibration for traffic scene surveillance, which supplies a connection between 2-D image features and 3-D measurement. It is helpful to deal with appearance distortion related to view angles, establish multiview correspondences, and make use of 3-D object models as prior information to enhance surveillance performance. A convenient and practical camera calibration method is proposed in this paper. With the camera heightHmeasured as the only user input, we can recover both intrinsic and extrinsic parameters of the camera based on redundant information supplied by moving objects in monocular videos. All cases of traffic scene layouts are considered and corresponding solutions are given to make our method applicable to almost all kinds of traffic scenes in reality. Numerous experiments are conducted in different scenes, and experimental results demonstrate the accuracy and practicability of our approach. It is shown that our approach can be effectively adopted in all kinds of traffic scene surveillance applications. Zhaoxiang Zhang 0001, Tieniu Tan, Kaiqi Huang, Yunhong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Semantic Pixel Sets Based Local Binary Patterns for Face Recognition
Zhenhua Chai, Heydi Mendez Vazquez, Ran He 0001, Zhenan Sun, Tieniu Tan |
ACCV (2) | 5 |
| 2012 | Local Hypersphere Coding Based on Edges between Visual Words
Weiqiang Ren, Yongzhen Huang, Xin Zhao 0012, Kaiqi Huang, Tieniu Tan |
ACCV (1) | 5 |
| 2012 | Contextual Pooling in Image Classification
Zifeng Wu, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ACCV (1) | 4 |
| 2012 | Spatial Graph for Image Classification
Zifeng Wu, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ACCV (1) | 4 |
| 2012 | Data Decomposition and Spatial Mixture Modeling for Part Based Model
Junge Zhang, Yongzhen Huang, Kaiqi Huang, Zifeng Wu, Tieniu Tan |
ACCV (1) | 5 |
| 2012 | Tracking Blurred Object with Data-Driven TrackerabstractMotion blur is very common in the low quality of image sequences and videos captured by low speed of cameras. Object tracking without accounting for the motion blur would easily fail in these kinds of videos. We propose a new data-driven tracker in the particle filter framework to address this problem without deblurring the image sequences. The motion blur is detected by exploring the property of the blurred input image through Fourier analysis. The appearance model is integrated with a set of motion blur kernels which could reflect different blur effects in real scenes. The motion model is improved to be more robust to sudden motion of the target object. To evaluate the proposed algorithm, several challenging videos with significant motion blur are used in the experiments. The experimental results demonstrate the robustness and accuracy of our algorithm. Jianwei Ding, Kaiqi Huang, Tieniu Tan |
AVSS | 3 |
| 2012 | Baseline Results for Violence Detection in Still ImagesabstractRecognizing objectionable content draws more and more attention nowadays given the rapid proliferation of images and videos on the Internet. Although there are some investigations about violence video detection and pornographic information filtering, very few existing methods touch on the problem of violence detection in still images. However, given its potential use in violence webpage filtering, online public opinion monitoring and some other aspects, recognizing violence in still images is worth being deeply investigated. To this end, we first establish a new database containing 500 violence images and 1500 non-violence images. And we use the Bag-of-Words (BoW) model which is frequently adopted in image classification domain to discriminate violence images and non-violence images. The effectiveness of four different feature representations are tested within the BoW framework. Finally the baseline results for violence image detection on our newly built database are reported. Dong Wang 0004, Zhang Zhang 0001, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
AVSS | 5 |
| 2012 | CLUMOC: Multiple Motion Estimation by Cluster Motion ConsensusabstractIn this paper, we present techniques for robust multiple motions estimation based on dual consensus via clustering in both the image spatial space and the motion parameter space. Starting from traditional Random Samples Consensus algorithm, we novelly propose the CLUster MOtion Consensus (CLUMOC) to extract robust motions. The proposed algorithm has two advantages: (1), instead of random samples, the CLUMOC employs clustering in initial sample selection, which can remove outliers from correct pairs of motion, (2), CLUMOC automatically decides the number of motions, by employing competition among motion and samples, that each motion needs to compete for matching pairs and each pair of matching competes for motions. The experimental results show that the proposed method is effective and efficient under various situations. Yinan Yu, Weiqiang Ren, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
AVSS | 5 |
| 2012 | l2, 1 Regularized correntropy for robust feature selectionabstractIn this paper, we study the problem of robust feature extraction based on l2,1regularized correntropy in both theoretical and algorithmic manner. In theoretical part, we point out that an l2,1-norm minimization can be justified from the viewpoint of half-quadratic (HQ) optimization, which facilitates convergence study and algorithmic development. In particular, a general formulation is accordingly proposed to unify l1-norm and l2,1-norm minimization within a common framework. In algorithmic part, we propose an l2,1regularized correntropy algorithm to extract informative features meanwhile to remove outliers from training data. A new alternate minimization algorithm is also developed to optimize the non-convex correntropy objective. In terms of face recognition, we apply the proposed method to obtain an appearance-based model, called Sparse-Fisherfaces. Extensive experiments show that our method can select robust and sparse features, and outperforms several state-of-the-art subspace methods on largescale and open face recognition datasets. Ran He 0001, Tieniu Tan, Liang Wang 0001, Wei-Shi Zheng 0001 |
CVPR | 2 |
| 2012 | Universal spatial feature set for video steganalysisabstractIn this paper, we propose a universal spatial feature set for video steganalysis. This feature set comprehensively exploits the correlation of adjacent pixels and can be viewed as a generalized extension of most correlation based features. We also develop a new approach to extract inter-frame features for video steganalysis. Our method can be universally applied to detect different video steganographic algorithms regardless of video format. The experimental results show that it outperforms current correlation based methods. Xikai Xu, Jing Dong 0003, Tieniu Tan |
ICIP | 3 |
| 2012 | Feature coding via vector difference for image classificationabstractAn effective image representation is important to an image classification task. The most popular image representation framework utilizes a feature coding algorithm to encode the extracted low-level feature descriptors into a vector representation. In this paper, we analyze the recently developed feature coding methods in a general way. According to their common characteristics, we propose a new coding scheme to perform feature coding based on the vector difference in a high-dimensional space which is obtained by explicit feature maps. As we illustrate, our method has promising results with small codebook sizes and generalizes most existing coding methods in a unified form. Xin Zhao 0012, Yinan Yu, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICIP | 5 |
| 2012 | Regularization parameter estimation for spectral regression discriminant analysis based on perturbation theory
Jie Gui, Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2012 | An effective regional saliency model based on extended site entropy rate
Yan Huang 0008, Wei Wang 0115, Liang Wang 0001, Tieniu Tan |
ICPR | 4 |
| 2012 | Accurate iris localization using contour segments
Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2012 | Robust regularized feature selection for iris recognition via linear programming
Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2012 | Group encoding of local features in image classification
Zifeng Wu, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ICPR | 4 |
| 2012 | Iris image classification based on color information
Hui Zhang 0061, Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2012 | Semantic windows mining in sliding window based object detection
Junge Zhang, Xin Zhao 0012, Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICPR | 5 |
| 2012 | A cascade fusion scheme for gait and cumulative foot pressure image recognition
Shuai Zheng 0001, Kaiqi Huang, Tieniu Tan, Dacheng Tao |
Pattern Recognit. | 3 |
| 2012 | Noisy iris image matching by using multiple cues
Tieniu Tan, Zhenan Sun, Hui Zhang 0061 |
Pattern Recognit. Lett. | 1 |
| 2012 | Cast Shadow Removal in a Hierarchical Manner Using MRFabstractIn this paper, we present a novel method for shadow removal using Markov random fields (MRF). In our method, we first construct the shadow model in a hierarchical manner. At the pixel level, we use the Gaussian mixture model to model the behavior of cast shadows for every pixel in the HSV color space. The samples which are used to update the shadow model should satisfy a pre-classifier. This pre-classifier indicates the color feature of shadow in current frame. At the global level, we exploit the statistical features of shadow in the whole scene over several consecutive frames to make this pre-classifier accurate and adaptive to the change of shadow. Then, based on the shadow model, an MRF model is constructed for shadow removal. The main contribution of this paper is twofold. First, although our method is a chroma-based method, we make the pre-classifier accurate and adaptive to the change of shadow by using the statistical features of shadow at the global level. Moreover, tracking information can make this global-level statistical information more robust. Second, we construct an MRF model to represent the dependencies between the label of a pixel and the shadow models of its neighbors. Experimental results show that the proposed method is efficient and robust. Kaiqi Huang, Tieniu Tan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | A Discriminative Model of Motion and Cross Ratio for View-Invariant Action RecognitionabstractAction recognition is very important for many applications such as video surveillance, human-computer interaction, and so on; view-invariant action recognition is hot and difficult as well in this field. In this paper, a new discriminative model is proposed for video-based view-invariant action recognition. In the discriminative model, motion pattern and view invariants are perfectly fused together to make a better combination of invariance and distinctiveness. We address a series of issues, including interest point detection in image sequence, motion feature extraction and description, and view-invariant calculation. First, motion detection is used to extract motion information from videos, which is much more efficient than traditional background modeling and tracking-based methods. Second, as for feature representation, we exact variety of statistical information from motion and view-invariant feature based on cross ratio. Last, in the action modeling, we apply a discriminative probabilistic model-hidden conditional random field to model motion patterns and view invariants, by which we could fuse the statistics of motion and projective invariability of cross ratio in one framework. Experimental results demonstrate that our method can improve the ability to distinguish different categories of actions with high robustness to view change in real circumstances. Kaiqi Huang, Yeying Zhang, Tieniu Tan |
IEEE Trans. Image Process. | 3 |
| 2012 | Efficient Object Tracking by Incremental Self-Tuning Particle Filtering on the Affine GroupabstractWe propose an incremental self-tuning particle filtering (ISPF) framework for visual tracking on the affine group, which can find the optimal state in a chainlike way with a very small number of particles. Unlike traditional particle filtering, which only relies on random sampling for state optimization, ISPF incrementally draws particles and utilizes an online-learned pose estimator (PE) to iteratively tune them to their neighboring best states according to some feedback appearance-similarity scores. Sampling is terminated if the maximum similarity of all tuned particles satisfies a target-patch similarity distribution modeled online or if the permitted maximum number of particles is reached. With the help of the learned PE and some appearance-similarity feedback scores, particles in ISPF become "smart" and can automatically move toward the correct directions; thus, sparse sampling is possible. The optimal state can be efficiently found in a step-by-step way in which some particles serve as bridge nodes to help others to reach the optimal state. In addition to the single-target scenario, the "smart" particle idea is also extended into a multitarget tracking problem. Experimental results demonstrate that our ISPF can achieve great robustness and very high accuracy with only a very small number of particles. Min Li 0022, Tieniu Tan, Wei Chen 0012, Kaiqi Huang |
IEEE Trans. Image Process. | 2 |
| 2012 | Foreground Object Detection Using Top-Down Information Based on EM FrameworkabstractIn this paper, we present a novel foreground object detection scheme that integrates the top-down information based on the expectation maximization (EM) framework. In this generalized EM framework, the top-down information is incorporated in an object model. Based on the object model and the state of each target, a foreground model is constructed. This foreground model can augment the foreground detection for the camouflage problem. Thus, an object's state-specific Markov random field (MRF) model is constructed for detection based on the foreground model and the background model. This MRF model depends on the latent variables that describe each object's state. The maximization of the MRF model is the M-step in the EM framework. Besides fusing spatial information, this MRF model can also adjust the contribution of the top-down information for detection. To obtain detection result using this MRF model, sampling importance resampling is used to sample the latent variable and the EM framework refines the detection iteratively. Besides the proposed generalized EM framework, our method does not need any prior information of the moving object, because we use the detection result of moving object to incorporate the domain knowledge of the object shapes into the construction of top-down information. Moreover, in our method, a kernel density estimation (KDE)-Gaussian mixture model (GMM) hybrid model is proposed to construct the probability density function of background and moving object model. For the background model, it has some advantages over GMM- and KDE-based methods. Experimental results demonstrate the capability of our method, particularly in handling the camouflage problem. Kaiqi Huang, Tieniu Tan |
IEEE Trans. Image Process. | 3 |
| 2012 | A Novel Algorithm for View and Illumination Invariant Image MatchingabstractThe challenges in local-feature-based image matching are variations of view and illumination. Many methods have been recently proposed to address these problems by using invariant feature detectors and distinctive descriptors. However, the matching performance is still unstable and inaccurate, particularly when large variation in view or illumination occurs. In this paper, we propose a view and illumination invariant image-matching method. We iteratively estimate the relationship of the relative view and illumination of the images, transform the view of one image to the other, and normalize their illumination for accurate matching. Our method does not aim to increase the invariance of the detector but to improve the accuracy, stability, and reliability of the matching results. The performance of matching is significantly improved and is not affected by the changes of view and illumination in a valid range. The proposed method would fail when the initial view and illumination method fails, which gives us a new sight to evaluate the traditional detectors. We propose two novel indicators for detector evaluation, namely, valid angle and valid illumination, which reflect the maximum allowable change in view and illumination, respectively. Extensive experimental results show that our method improves the traditional detector significantly, even in large variations, and the two indicators are much more distinctive. Yinan Yu, Kaiqi Huang, Wei Chen 0012, Tieniu Tan |
IEEE Trans. Image Process. | 4 |
| 2012 | Three-Dimensional Deformable-Model-Based Localization and Recognition of Road VehiclesabstractWe address the problem of model-based object recognition. Our aim is to localize and recognize road vehicles from monocular images or videos in calibrated traffic scenes. A 3-D deformable vehicle model with 12 shape parameters is set up as prior information, and its pose is determined by three parameters, which are its position on the ground plane and its orientation about the vertical axis under ground-plane constraints. An efficient local gradient-based method is proposed to evaluate the fitness between the projection of the vehicle model and image data, which is combined into a novel evolutionary computing framework to estimate the 12 shape parameters and three pose parameters by iterative evolution. The recovery of pose parameters achieves vehicle localization, whereas the shape parameters are used for vehicle recognition. Numerous experiments are conducted in this paper to demonstrate the performance of our approach. It is shown that the local gradient-based method can evaluate accurately and efficiently the fitness between the projection of the vehicle model and the image data. The evolutionary computing framework is effective for vehicles of different types and poses is robust to all kinds of occlusion. Zhaoxiang Zhang 0001, Tieniu Tan, Kaiqi Huang, Yunhong Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2011 | Recovery of corrupted low-rank matrices via half-quadratic based nonconvex minimizationabstractRecovering arbitrarily corrupted low-rank matrices arises in computer vision applications, including bioinformatic data analysis and visual tracking. The methods used involve minimizing a combination of nuclear norm and l1norm. We show that by replacing the l1norm on error items with nonconvex M-estimators, exact recovery of densely corrupted low-rank matrices is possible. The robustness of the proposed method is guaranteed by the M-estimator theory. The multiplicative form of half-quadratic optimization is used to simplify the nonconvex optimization problem so that it can be efficiently solved by iterative regularization scheme. Simulation results corroborate our claims and demonstrate the efficiency of our proposed method under tough conditions. Ran He 0001, Zhenan Sun, Tieniu Tan, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2011 | Exploring relations of visual codes for image classificationabstractThe classic Bag-of-Features (BOF) model and its extensional work use a single value to represent a visual code. This strategy ignores the relation of visual codes. In this paper, we explore this relation and propose a new algorithm for image classification. It consists of two main parts: 1) construct the codebook graph wherein a visual code is linked with other codes; 2) describe each local feature using a pair of related codes, corresponding to an edge of the graph. Our approach contains richer information than previous BOF models. Moreover, we demonstrate that these models are special cases of ours. Various coding and pooling algorithms can be embedded into our framework to obtain better performance. Experiments on different kinds of image classification databases demonstrate that our approach can stably achieve excellent performance compared with various BOF models. Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
CVPR | 4 |
| 2011 | Salient coding for image classificationabstractThe codebook based (bag-of-words) model is a widely applied model for image classification. We analyze recent coding strategies in this model, and find that saliency is the fundamental characteristic of coding. The saliency in coding means that if a visual code is much closer to a descriptor than other codes, it will obtain a very strong response. The salient representation under maximum pooling operation leads to the state-of-the-art performance on many databases and competitions. However, most current coding schemes do not recognize the role of salient representation, so that they may lead to large deviations in representing local descriptors. In this paper, we propose “salient coding”, which employs the ratio between descriptors' nearest code and other codes to describe descriptors. This approach can guarantee salient representation without deviations. We study salient coding on two sets of image classification databases (15-Scenes and PASCAL VOC2007). The experimental results demonstrate that our approach outperforms all other coding methods in image classification. Yongzhen Huang, Kaiqi Huang, Yinan Yu, Tieniu Tan |
CVPR | 4 |
| 2011 | Boosted local structured HOG-LBP for object localizationabstractObject localization is a challenging problem due to variations in object's structure and illumination. Although existing part based models have achieved impressive progress in the past several years, their improvement is still limited by low-level feature representation. Therefore, this paper mainly studies the description of object structure from both feature level and topology level. Following the bottom-up paradigm, we propose a boosted Local Structured HOG-LBP based object detector. Firstly, at feature level, we propose Local Structured Descriptor to capture the object's local structure, and develop the descriptors from shape and texture information, respectively. Secondly, at topology level, we present a boosted feature selection and fusion scheme for part based object detector. All experiments are conducted on the challenging PASCAL VOC2007 datasets. Experimental results show that our method achieves the state-of-the-art performance. Junge Zhang, Kaiqi Huang, Yinan Yu, Tieniu Tan |
CVPR | 4 |
| 2011 | Graph modeling based local descriptor selection via a hierarchical structure for biometric recognitionabstractLocal descriptor based image representation is widely used in biometrics and has achieved promising results. We usually extract the most distinctive local descriptors for image sparse representation due to the large feature space and the redundancy among local descriptors. In this paper, we describe the local descriptor based image representation via a graph model, in which each node is a local descriptor (we call it “atom”) and the edges denote the relationship between atoms. Based on this model, a hierarchical structure is constructed to select the most distinctive local descriptors. Two-layer structure is adopted in our work, including local selection and global selection. In the first layer, L1/Lqregularized least square regression is adopted to reduce the redundancy of local descriptors in local regions. In the second layer, AdaBoost learning is performed for local descriptor selection based on the results of the first layer. We apply this method to long-range personal identification by using binocular regions. Our method can select the distinctive local descriptors and reduce the redundancy among them, and achieve encouraging results on the collected binocular database and CASIA-Iris-Distance. Particularly, our method is about 50 times faster than the traditional AdaBoost learning based method in the experiments. Zhenan Sun, Tieniu Tan |
IJCB | 3 |
| 2011 | Direction-based stochastic matching for pedestrian recognition in non-overlapping camerasabstractPedestrian recognition is a challenging problem in non-overlapping multi-camera object tracking. In this paper, we present a novel approach for matching pedestrians across non-overlapping multiple cameras without the need of a training phase or spatio-temporal cues across cameras. To deal with viewpoint changes, we introduce the concept of directional angles estimated using the spatio-temporal continuity in the single camera tracking. To deal with pose changes, a stochastic matching strategy is performed, where the similarity of two blobs belonging to different viewpoints is calculated by a novel similarity measurement algorithm. The experiments are performed on different multi-view datasets. Experimental results demonstrate the effectiveness and robustness of the proposed method. Xiaotang Chen, Kaiqi Huang, Tieniu Tan |
ICIP | 3 |
| 2011 | An effective image steganalysis method based on neighborhood information of pixelsabstractThis paper focuses on image steganalysis. We use higher order image statistics based on neighborhood information of pixels (NIP) to detect the stego images from original ones. We use subtracting gray values of adjacent pixels to capture neighborhood information, and also make use of “rotation invariant” property to reduce the dimensionality for the whole feature sets. We tested two kinds of NIP feature, the experimental results illustrates that our proposed feature sets are with good performance and even outperform the state-of-art in certain aspect. Qingxiao Guan, Jing Dong 0003, Tieniu Tan |
ICIP | 3 |
| 2011 | Comprehensive assessment of iris image qualityabstractIris image quality critically determines iris recognition performance and the quality metrics of iris images are also useful prior information for adaptive selection of optimal recognition strategy. Iris image quality is jointly determined by multiple factors such as focus, occlusion, off-angle, deformation, etc. So it is a complex problem to assess the overall quality score of an iris image. This paper proposes a novel framework for comprehensive assessment of iris image quality. The contributions of the paper include three aspects: (i) Three novel approaches are proposed to estimate the quality metrics (QM) of defocus, motion blur and off-angle in an iris image respectively, (ii) A fusion method based on likelihood ratio is proposed to combine six quality factors of an iris image into an unified quality score. (iii) A statistical quantization method based on t-test is proposed to adaptively classify the iris images in a database into a number of quality levels. Extensive experiments demonstrate the proposed framework can effectively assess the overall quality of iris images. And the relationship between iris recognition results and the quality level of iris images can be explicitly formulated. Zhenan Sun, Tieniu Tan |
ICIP | 3 |
| 2011 | Partial Least Squares based subwindow search for pedestrian detectionabstractIn this paper, we propose a Partial Least Squares based sub- window search method for pedestrian detection, by which the detection speed can be improved effectively while maintaining high detection accuracy. Firstly, a sparse search is implemented to find all the possible locations containing parts of a pedestrian. Then a pre-learned Partial Least Squares regression model is applied to estimate the displacements of the subwindows to guide them towards the approximate locations of the pedestrians. Finally, we conduct a dense search around the approximate locations to obtain the exact locations of the pedestrians. Experiments on the INRIA dataset demonstrate that our method greatly reduces the number of search windows, which leads to much fewer feature extraction in the detection phase. Thus, it is about 10 times faster than the sliding window method with a jump step of 8 x 8. Jinchen Wu, Wei Chen 0012, Kaiqi Huang, Tieniu Tan |
ICIP | 4 |
| 2011 | Deformable DAISY Matcher for robust iris recognitionabstractIris is rich of texture information for reliable personal identification. However, nonlinear deformation of iris pattern caused by pupil dilation or contraction raises a grand challenge to iris recognition. This paper proposes a novel iris recognition method namely Deformable DAISY Matcher (DDM) for robust iris feature matching. Firstly, dense DAISY descriptors are extracted to represent regional iris features, which are robust against intra-class variations of iris images. Then a set of iris key points are localized on the feature map. Finally deformation tolerant matching strategy is proposed to match corresponding key points of iris images. Experimental results on two iris image databases demonstrate DDM is better than state-of-the-art iris recognition methods. Man Zhang 0005, Zhenan Sun, Tieniu Tan |
ICIP | 3 |
| 2011 | Evaluation framework on translation-invariant representation for cumulative foot pressure imageabstractGround reaction force can be used to distinguish different human gait like limb movement. Cumulative foot pressure image is a 2-D data that recorded the spatial and temporal change of ground reaction force during one gait cycle. However, when putting it into practice as a new biometric for gait recognition, it suffers from the problem of large translation variation within class caused by wearing different shoes and walking in different speed. In this paper, an evaluation framework is proposed to address the problem. The framework consists of a database containing cumulative foot pressure images and a well designed benchmark. The data are collected from 118 subjects. A locality-constrained sparse coding scheme is developed to be compared with the benchmark PCA approach. Experimental results show the potential of evaluation framework for evaluating the translation-invariant power of image representation algorithms. Shuai Zheng 0001, Kaiqi Huang, Tieniu Tan |
ICIP | 3 |
| 2011 | Robust view transformation model for gait recognitionabstractRecent gait recognition systems often suffer from the challenges including viewing angle variation and large intra-class variations. In order to address these challenges, this paper presents a robust View Transformation Model for gait recognition. Based on the gait energy image, the proposed method establishes a robust view transformation model via robust principal component analysis. Partial least square is used as feature selection method. Compared with the existing methods, the proposed method finds out a shared linear correlated low rank subspace, which brings the advantages that the view transformation model is robust to viewing angle variation, clothing and carrying condition changes. Conducted on the CASIA gait dataset, experimental results show that the proposed method outperforms the other existing methods. Shuai Zheng 0001, Junge Zhang, Kaiqi Huang, Ran He 0001, Tieniu Tan |
ICIP | 5 |
| 2011 | Iris Matching Based on Personalized Weight MapabstractIris recognition typically involves three steps, namely, iris image preprocessing, feature extraction, and feature matching. The first two steps of iris recognition have been well studied, but the last step is less addressed. Each human iris has its unique visual pattern and local image features also vary from region to region, which leads to significant differences in robustness and distinctiveness among the feature codes derived from different iris regions. However, most state-of-the-art iris recognition methods use a uniform matching strategy, where features extracted from different regions of the same person or the same region for different individuals are considered to be equally important. This paper proposes a personalized iris matching strategy using a class-specific weight map learned from the training images of the same iris class. The weight map can be updated online during the iris recognition procedure when the successfully recognized iris images are regarded as the new training data. The weight map reflects the robustness of an encoding algorithm on different iris regions by assigning an appropriate weight to each feature code for iris matching. Such a weight map trained by sufficient iris templates is convergent and robust against various noise. Extensive and comprehensive experiments demonstrate that the proposed personalized iris matching strategy achieves much better iris recognition performance than uniform strategies, especially for poor quality iris images. Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | An Extended Grammar System for Learning and Recognizing Complex Visual EventsabstractFor a grammar-based approach to the recognition of visual events, there are two major limitations that prevent it from real application. One is that the event rules are predefined by domain experts, which means huge manual cost. The other is that the commonly used grammar can only handle sequential relations between subevents, which is inadequate to recognize more complex events involving parallel subevents. To solve these problems, we propose an extended grammar approach to modeling and recognizing complex visual events. First, motion trajectories as original features are transformed into a set of basic motion patterns of a single moving object, namely, primitives (terminals) in the grammar system. Then, a Minimum Description Length (MDL) based rule induction algorithm is performed to discover the hidden temporal structures in primitive stream, where Stochastic Context-Free Grammar (SCFG) is extended by Allen's temporal logic to model the complex temporal relations between subevents. Finally, a Multithread Parsing (MTP) algorithm is adopted to recognize interesting complex events in a given primitive stream, where a Viterbi-like error recovery strategy is also proposed to handle large-scale errors, e.g., insertion and deletion errors. Extensive experiments, including gymnastic exercises, traffic light events, and multi-agent interactions, have been executed to validate the effectiveness of the proposed approach. Zhang Zhang 0001, Tieniu Tan, Kaiqi Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Enhanced Biologically Inspired Model for Object RecognitionabstractThe biologically inspired model (BIM) proposed by Serre presents a promising solution to object categorization. It emulates the process of object recognition in primates' visual cortex by constructing a set of scale- and position-tolerant features whose properties are similar to those of the cells along the ventral stream of visual cortex. However, BIM has potential to be further improved in two aspects: mismatch by dense input and randomly feature selection due to the feedforward framework. To solve or alleviate these limitations, we develop an enhanced BIM (EBIM) in terms of the following two aspects: 1) removing uninformative inputs by imposing sparsity constraints, 2) apply a feedback loop to middle level feature selection. Each aspect is motivated by relevant psychophysical research findings. To show the effectiveness of the EBIM, we apply it to object categorization and conduct empirical studies on four computer vision data sets. Experimental results demonstrate that the EBIM outperforms the BIM and is comparable to state-of-the-art approaches in terms of accuracy. Moreover, the new system is about 20 times faster than the BIM. Yongzhen Huang, Kaiqi Huang, Dacheng Tao, Tieniu Tan, Xuelong Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2011 | Biologically Inspired Features for Scene Classification in Video SurveillanceabstractInspired by human visual cognition mechanism, this paper first presents a scene classification method based on an improved standard model feature. Compared with state-of-the-art efforts in scene classification, the newly proposed method is more robust, more selective , and of lower complexity. These advantages are demonstrated by two sets of experiments on both our own database and standard public ones. Furthermore, occlusion and disorder problems in scene classification in video surveillance are also first studied in this paper. Kaiqi Huang, Dacheng Tao, Yuan Yan Tang, Xuelong Li 0001, Tieniu Tan |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2010 | Modeling Complex Scenes for Accurate Moving Objects Segmentation
Jianwei Ding, Min Li 0022, Kaiqi Huang, Tieniu Tan |
ACCV (2) | 4 |
| 2010 | A Heuristic Deformable Pedestrian Detection Method
Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ACCV (2) | 3 |
| 2010 | Multi-Target Tracking by Learning Class-Specific and Instance-Specific Cues
Min Li 0022, Wei Chen 0012, Kaiqi Huang, Tieniu Tan |
ACCV (2) | 4 |
| 2010 | Visual tracking via incremental self-tuning particle filtering on the affine groupabstractWe propose an incremental self-tuning particle filtering (ISPF) framework for visual tracking on the affine group. SIFT (Scale Invariant Feature Transform) like descriptors are used as basic features, and IPCA (Incremental Principle Component Analysis) is utilized to learn an adaptive appearance subspace for similarity measurement. ISPF tries to find the optimal target position in a step-by-step way: particles are incrementally drawn and intelligently tuned to their best states by an online LWPR (Local Weighted Projection Regression) pose estimator; searching is terminated if the maximum similarity of all tuned particles satisfies a target similarity distribution (TSD) modeled online or the permitted maximum number of particles is reached. Experimental results demonstrate that our ISPF can achieve great robustness and very high accuracy with only a very small number of random particles. Min Li 0022, Wei Chen 0012, Kaiqi Huang, Tieniu Tan |
CVPR | 4 |
| 2010 | Image tampering detection based on stationary distribution of Markov chainabstractIn this paper, we propose a passive image tampering detection method based on modeling edge information. We model the edge image of image chroma component as a finite-state Markov chain and extract low dimensional feature vector from its stationary distribution for tampering detection. The support vector machine (SVM) is utilized as classifier to evaluate the effectiveness of the proposed algorithm. The experimental results in a large scale of evaluation database illustrates that our proposed method is promising. Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 3 |
| 2010 | Texture removal for adaptive level set based iris segmentationabstractLevel set based active contour method has been proposed for iris segmentation in recent years, but it can not converge to iris contours in real applications because of its sensitivity to local gradient extremes due to the complex iris texture. In this paper, a novel scheme is proposed to remove local gradient extremes before using level set directly. Firstly, we use two orthogonal ordinal filters to obtain robust gradient map. Then we localize the iris region on the gradient map by an improved Hough transform. After that, a Semantic Iris Contour Map is generated by combining the spatial information of coarse iris location and the gradient map as the edge indicator for level set segmentation. For robust and accurate segmentation, we propose a convergence criterion and a means of updating the parameters for level set. Finally, the accurate segmentation is obtained by the robust adaptive level set method. Encouraging results on ICE 2005 database and CASIA v3 database show the efficiency and effectiveness of our method. Zhenan Sun, Tieniu Tan |
ICIP | 3 |
| 2010 | Statistics of local surface curvatures for mis-localized iris detectionabstractEye detection is a hot research topic in computer vision for its wide applications in human-computer interaction, face and iris recognition, etc. However, robust eye detection is still a grand challenge due to the numerous appearance variations of eye images in real-world applications. In this paper, we present a novel local surface curvature analysis method to deal with this problem. Firstly, by regarding an eye image as a 2D surface in 3D space, we propose to use the histogram of local surface curvatures as the general representation of eye pattern. Then, a SVM classifier is employed for eye detection using the histogram vectors of eye and non-eye samples. Extensive experiments are performed and the results show that the proposed method achieves state-of-the-art performance in eye detection. In particular, it is more efficient in mistakenly localized iris detection. Hui Zhang 0061, Zhenan Sun, Tieniu Tan |
ICIP | 3 |
| 2010 | Recovering the Topology of Multiple Cameras by Finding Continuous Paths in a TrellisabstractIn this paper, we propose an unsupervised method for recovering the topology of multiple cameras with non-overlapping fields of view. The nodes in the topology graph are defined as entry/exit zones in each camera while the connectivity between nodes is inferred through finding continuous paths in a trellis where appearance information and temporal information of moving objects are encoded. Unlike previous methods which assume a single mode transition distribution between nodes, our method is capable of dealing with multi-modal transition situations when both cars and pedestrians are in the scene. Results on simulated and real-life datasets demonstrate the effectiveness of the proposed method. Yinghao Cai, Kaiqi Huang, Tieniu Tan, Matti Pietikäinen |
ICPR | 3 |
| 2010 | 3D Model Based Vehicle Tracking Using Gradient Based Fitness Evaluation under Particle Filter FrameworkabstractWe address the problem of 3D model based vehicle tracking from monocular videos of calibrated traffic scenes. A 3D wire-frame model is set up as prior information and an efficient fitness evaluation method based on image gradients is introduced to estimate the fitness score between the projection of vehicle model and image data, which is then combined into a particle filter based framework for robust vehicle tracking. Numerous experiments are conducted and experimental results demonstrate the effectiveness of our approach for accurate vehicle tracking and robustness to noise and occlusions. Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan, Yunhong Wang 0001 |
ICPR | 3 |
| 2010 | Hierarchical Fusion of Face and Iris for Personal IdentificationabstractMost existing face and iris fusion schemes are concerned about improving performance on good quality images under controlled environments. In this paper, we propose a hierarchical fusion scheme for low quality images under uncontrolled situations. In the training stage, canonical correlation analysis (CCA) is adopted to construct a statistical mapping from face to iris in pixel level. In the testing stage, firstly the probe face image is used to obtain a subset of candidate gallery samples via regression between the probe face and gallery irises, then ordinal representation and sparse representation are performed on these candidate samples for iris recognition and face recognition respectively. Finally, score level fusion via min-max normalization is performed to make final decision. Experimental results on our low quality database show the outperforming performance of proposed method. Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2010 | Contact Lens Detection Based on Weighted LBPabstractSpoof detection is a critical function for iris recognition because it reduces the risk of iris recognition systems being forged. Despite various counterfeit artifacts, cosmetic contact lens is one of the most common and difficult to detect. In this paper, we proposed a novel fake iris detection algorithm based on improved LBP and statistical features. Firstly, a simplified SIFT descriptor is extracted at each pixel of the image. Secondly, the SIFT descriptor is used to rank the LBP encoding sequence. Then, statistical features are extracted from the weighted LBP map. Lastly, SVM classifier is employed to classify the genuine and counterfeit iris images. Extensive experiments are conducted on a database containing more than 5000 fake iris images by wearing 70 kinds of contact lens, and captured by four iris devices. Experimental results show that the proposed method achieves state-of-the-art performance in contact lens spoof detection. Hui Zhang 0061, Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2010 | New developments in color image tampering detectionabstractIn this paper, an efficient framework for passive-blind color image tampering detection is presented. Statistical features are extracted from a given test image and a set of 2-D arrays derived by applying multi-size block discrete cosine transform to the given test image. Image features are extracted from Cr channel, a chroma channel in YCbCr color space, because of its observed sensitivity to color image tampering. A support vector machine is employed to evaluate the effectiveness of image features over a color image dataset recently established for tampering detection. Boosting feature selection is applied to having feature dimensionality reduced so as to make detection accuracy generalizable and computational complexity decreased. Experimental results have demonstrated that the proposed framework applied to the aforementioned dataset outperforms the state of the arts by distinct margins. Patchara Sutthiwan, Yun Q. Shi 0001, Jing Dong 0003, Tieniu Tan, Tian-Tsong Ng |
ISCAS | 4 |
| 2010 | Blind Quantitative Steganalysis Based on Feature Fusion and Gradient Boosting
Qingxiao Guan, Jing Dong 0003, Tieniu Tan |
IWDW | 3 |
| 2010 | Tampered Region Localization of Digital Color Images Based on JPEG Compression Noise
Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IWDW | 3 |
| 2010 | Efficient and robust segmentation of noisy iris images for non-cooperative iris recognition
Tieniu Tan, Zhaofeng He 0001, Zhenan Sun |
Image Vis. Comput. | 1 |
| 2010 | Topology modeling for Adaboost-cascade based object detection
Zhaofeng He 0001, Tieniu Tan, Zhenan Sun |
Pattern Recognit. Lett. | 2 |
| 2010 | Vs-star: A visual interpretation system for visual surveillance
Kaiqi Huang, Tieniu Tan |
Pattern Recognit. Lett. | 2 |
| 2009 | A Novel Visual Organization Based on Topological Perception
Yongzhen Huang, Kaiqi Huang, Tieniu Tan, Dacheng Tao |
ACCV (1) | 3 |
| 2009 | A Harris-Like Scale Invariant Feature Detector
Yinan Yu, Kaiqi Huang, Tieniu Tan |
ACCV (2) | 3 |
| 2009 | Hierarchical Shape Primitive Features for Online Text-independent Writer IdentificationabstractThis paper proposes a novel method to text independent writer identification from online handwriting. The main contributions of our method include two parts: shape primitive representation and hierarchical structure. Both shape primitive's features are developed to represent the robust and distinctive characteristics of handwriting in two hierarchies. In first hierarchy, the shape primitives probability distribution function (SPPDF)is defined as the static features, to characterize orientation information of writing style. For each shape primitive, the statistics of pressure is defined as the dynamic shape primitives probability distribution function (DSPPDF) and the second hierarchy we build Gaussian model in dynamic attributes (DA) according to curvature of shape primitives. Experiments were conducted on the NLPR handwriting database collected from 242 persons. The results show that the new method achieves high accuracy, fast speed and low requirement of the amount of characters in handwriting samples. We achieve a writer identification rate of 91.5% with datasets in Chinese text and 93.6% in English text. Bangy Li, Zhenan Sun, Tieniu Tan |
ICDAR | 3 |
| 2009 | Online Text-independent Writer Identification Based on Temporal Sequence and Shape CodesabstractIn this paper we present a novel method for online text-independent writer identification. Most of the existing writer identification techniques require the data to be from a specific text which is not applicable to cases where such text is not available, such as in criminal justice systems when text documents with different content need to be compared. Text-independent approaches often require a large amount of data to be confident of good results. We propose temporal sequence and shape codes to encode online handwriting. Temporal sequence codes (TSC) are to characterize trajectory in speed and pressure change in writing, and shape codes (SC) are to characterize direction of trajectory in writing handwriting. For TSC, we use two different codes to encode speed and pressure to codebook: stroke temporal sequence codes (STSC) and neighbor temporal sequence codes (NTSC). At identification stage, we implement decision and fusion strategy to identify writer. Experimental results show that our proposed method can improve the identification accuracy with a small number of characters. Moreover, we find that the proposed method is even effective for cross-language (English & Chinese) writer identification. Bangy Li, Tieniu Tan |
ICDAR | 2 |
| 2009 | A convergent solution to two dimensional linear discriminant analysisabstractThe matrix based data representation has been recognized to be effective for face recognition because it can deal with the undersampled problem. One of the most popular algorithms, the two dimensional linear discriminant analysis (2DLDA), has been identified to be effective to encode the discriminative information for training matrix represented samples. However, 2DLDA does not converge in the training stage. This paper presents an evolutionary computation based solution, referred to as E-2DLDA, to provide a convergent training stage for 2DLDA. In E-2DLDA, every randomly generated candidate projection matrices are first normalized. The evolutionary computation method optimizes the projection matrices to best separate different classes. Experimental results show E-2DLDA is convergent and outperforms 2DLDA. Wei Chen 0012, Kaiqi Huang, Tieniu Tan, Dacheng Tao |
ICIP | 3 |
| 2009 | Quality-based dynamic threshold for iris matchingabstractCurrent iris recognition systems usually regard poor quality iris images useless since defocused or partially occluded iris images may cause false acceptance. However, such a strategy may lose an opportunity to correctly report a genuine match with poor-quality samples. This paper proposes an adaptive iris matching method to improve the throughput of iris recognition systems. The core idea of the method is to dynamically adjust the decision threshold of iris matching module based on the quality measure of input iris image. So that the poor quality iris images also have a chance to match template database under the controlled false accept rate. Experiment results on the real system demonstrate the effectiveness of the proposed method and the recognition time is expected to be greatly reduced. Zhenan Sun, Tieniu Tan, Zhuoshi Wei |
ICIP | 3 |
| 2009 | Palmprint recognition using coarse-to-fine statistical image representationabstractRecent literatures have revealed that statistics of local texture measures can provide accurate descriptions of palmprint appearances. In this framework, one palmprint image is divided into local blocks with multiple spatial resolutions. The statistical texture descriptions of each block are then concatenated to form a multi-scale image representation. However, resultant high-dimensional statistical features lead to increasing of computational cost. In this paper, we tackle this problem by performing a coarse-to-fine cascade scheme, which makes use of information redundancy of statistical texture descriptions between different spatial scales. In contrast with non-cascade strategies, the proposed method reduces most of computational burden and achieves accurate classification simultaneously. Zhenan Sun, Tieniu Tan |
ICIP | 3 |
| 2009 | Computational primitives of visual perceptionabstractGreat stride has been made in psychological research about primitives of visual perception, which is important to computer vision and image processing. In this paper, we propose a computational model to imitate the primitives of visual perception based on the pyschological theory of topological perceptual organization. First, we adopt geodesic distance based descriptor to describe an independent topological structure. Then, we consider the spatial relationship of two independent structures. Experiments on structures classification demonstrates that the propose model is consistent with the psychological theory. Further experiments on patches clustering prove that our approach can be used to enhance other algorithms. Yongzhen Huang, Kaiqi Huang, Tieniu Tan |
ICIP | 3 |
| 2009 | Rapid and robust human detection and tracking based on omega-shape featuresabstractThis paper proposes a novel method for rapid and robust human detection and tracking based on the omega-shape features of people's head-shoulder parts. There are two modules in this method. In the first module, a Viola-Jones type classifier and a local HOG (Histograms of Oriented Gradients) feature based AdaBoost classifier are combined to detect head-shoulders rapidly and effectively. Then, in the second module, each detected head-shoulder is tracked by a particle filter tracker using local HOG features to model target's appearance, which shows great robustness in scenarios of crowding, background distractors and partial occlusions. Experimental results demonstrate the effectiveness and efficiency of the proposed approach. Min Li 0022, Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan |
ICIP | 4 |
| 2009 | Robust visual tracking based on simplified biologically inspired featuresabstractWe address the problem of robust appearance-based visual tracking. First, a set of simplified biologically inspired features (SBIF) is proposed for object representation and the Bhattacharyya coefficient is used to measure the similarity between the target model and candidate targets. Then, the proposed appearance model is combined into a Bayesian state inference tracking framework utilizing the SIR (sampling importance resampling) particle filter to propagate sample distributions over time. Numerous experiments are conducted and experimental results demonstrate that our algorithm is robust to partial occlusions and variations of illumination and pose, resistent to nearby distractors, as well as possesses the state-of-the-art tracking accuracy. Min Li 0022, Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan |
ICIP | 4 |
| 2009 | Effective image splicing detection based on image chromaabstractA color image splicing detection method based on gray level co-occurrence matrix (GLCM) of thresholded edge image of image chroma is proposed in this paper. Edge images are generated by subtracting horizontal, vertical, main and minor diagonal pixel values from current pixel values respectively and then thresholded with a predefined threshold T. The GLCMs of edge images along the four directions serve as features for image splicing detection. Boosting feature selection is applied to select optimal features and Support Vector Machine (SVM) is utilized as classifier in our approach. The effectiveness of the proposed method has been demonstrated by our experimental results. Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
ICIP | 3 |
| 2009 | Object detection and tracking for night surveillance based on salient contrast analysisabstractNight surveillance is a challenging task because of low brightness, low contrast, low Signal to Noise Ratio (SNR) and low appearance information. Most existing models for night surveillance share the following problems: a lack of adaptability for different scenes and separation between detection and tracking. To solve these problems we propose a model based on Salient Contrast Change (SCC) feature, which applies learning process to enhance adaptability and analyzes trajectories to improve the effectiveness of detection. Empirical studies on several real night videos show that the proposed model is more effective than the original CC model and other traditional models. Liangsheng Wang, Kaiqi Huang, Yongzhen Huang, Tieniu Tan |
ICIP | 4 |
| 2009 | A compact optical flowbased motion representation for real-time action recognition in surveillance scenesabstractWe address the problem of action recognition. Our aim is to recognize single person activities in surveillance scenes. To meet the requirements of real scene action recognition, we present a compact motion representation for human activity recognition. With the employment of efficient features extracted from optical flow as the main part, together with global information, our motion representation is compact and discriminative. We also build a novel human action dataset(CASIA) in surveillance scene with three vertically different viewpoints and distant people. Experiments on CASIA dataset and WEIZMANN dataset show that our method can achieve satisfying recognition performance with low computational cost as well as robustness against both horizontal(panning) and vertical(tilting) viewpoint changes. Shiquan Wang, Kaiqi Huang, Tieniu Tan |
ICIP | 3 |
| 2009 | View-invariant action recognition using cross ratios across framesabstractWe present a new method of computing invariants in videos captured from different views to achieve view-invariant action recognition. To avoid the constraints of collinearity or coplanarity of image points for constructing invariants, we consider several neighboring frames to compute cross ratios, namely cross ratios across frames (CRAF), as our invariant representation of action. For every five points sampled with different intervals from the trajectories of action, we construct a pair of cross ratios (CRs). Afterwards, we transform the CRs to histograms as the feature vectors for classification. Experimental results demonstrate that the proposed method outperforms the state-of-the-art methods in effectiveness and stability. Yeying Zhang, Kaiqi Huang, Yongzhen Huang, Tieniu Tan |
ICIP | 4 |
| 2009 | Multi-class Blind Steganalysis Based on Image Run-Length Analysis
Jing Dong 0003, Wei Wang 0025, Tieniu Tan |
IWDW | 3 |
| 2009 | A Survey of Passive Image Tampering Detection
Wei Wang 0025, Jing Dong 0003, Tieniu Tan |
IWDW | 3 |
| 2009 | Toward Accurate and Fast Iris Segmentation for Iris BiometricsabstractIris segmentation is an essential module in iris recognition because it defines the effective image region used for subsequent processing such as feature extraction. Traditional iris segmentation methods often involve an exhaustive search of a large parameter space, which is time consuming and sensitive to noise. To address these problems, this paper presents a novel algorithm for accurate and fast iris segmentation. After efficient reflection removal, an Adaboost-cascade iris detector is first built to extract a rough position of the iris center. Edge points of iris boundaries are then detected, and an elastic model named pulling and pushing is established. Under this model, the center and radius of the circular iris boundaries are iteratively refined in a way driven by the restoring forces of Hooke's law. Furthermore, a smoothing spline-based edge fitting scheme is presented to deal with noncircular iris boundaries. After that, eyelids are localized via edge detection followed by curve fitting. The novelty here is the adoption of a rank filter for noise elimination and a histogram filter for tackling the shape irregularity of eyelids. Finally, eyelashes and shadows are detected via a learned prediction model. This model provides an adaptive threshold for eyelash and shadow detection by analyzing the intensity distributions of different iris regions. Experimental results on three challenging iris image databases demonstrate that the proposed algorithm outperforms state-of-the-art methods in both accuracy and speed. Zhaofeng He 0001, Tieniu Tan, Zhenan Sun, Xianchao Qiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Ordinal Measures for Iris RecognitionabstractImages of a human iris contain rich texture information useful for identity authentication. A key and still open issue in iris recognition is how best to represent such textural information using a compact set of features (iris features). In this paper, we propose using ordinal measures for iris feature representation with the objective of characterizing qualitative relationships between iris regions rather than precise measurements of iris image structures. Such a representation may lose some image-specific information, but it achieves a good trade-off between distinctiveness and robustness. We show that ordinal measures are intrinsic features of iris patterns and largely invariant to illumination changes. Moreover, compactness and low computational complexity of ordinal measures enable highly efficient iris recognition. Ordinal measures are a general concept useful for image analysis and many variants can be derived for ordinal feature extraction. In this paper, we develop multilobe differential filters to compute ordinal measures with flexible intralobe and interlobe parameters such as location, scale, orientation, and distance. Experimental results on three public iris image databases demonstrate the effectiveness of the proposed ordinal feature models. Zhenan Sun, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Automatic 3D face recognition from depth and intensity Gabor features
Chenghua Xu, Stan Z. Li, Tieniu Tan, Long Quan |
Pattern Recognit. | 3 |
| 2009 | Human Behavior Analysis Based on a New Motion DescriptorabstractHuman behavior analysis is an important area of research in computer vision and is also driven by a wide spectrum of applications, such as smart video surveillance and human-computer interface. In this paper, we present a novel approach for human behavior analysis. Two research challenges, motion representation and behavior recognition, are addressed. A novel motion descriptor, which is an improved feature based on optical flow, is proposed for motion representation. Optical flow is improved with a motion filter, and feature fusion with the shape and trajectory information. To recognize the behavior, the support vector machine is employed to train the classifier where the concatenation of histograms is formed as the input features. Experimental results on the Weizmann behavior database and the Institute of Automation, Chinese Academy of Science real-world multiview behavior database demonstrate the robustness and effectiveness of our method. Kaiqi Huang, Shiquan Wang, Tieniu Tan, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | A Study on Gait-Based Gender ClassificationabstractGender is an important cue in social activities. In this correspondence, we present a study and analysis of gender classification based on human gait. Psychological experiments were carried out. These experiments showed that humans can recognize gender based on gait information, and that contributions of different body components vary. The prior knowledge extracted from the psychological experiments can be combined with an automatic method to further improve classification accuracy. The proposed method which combines human knowledge achieves higher performance than some other methods, and is even more accurate than human observers. We also present a numerical analysis of the contributions of different human components, which shows that head and hair, back, chest and thigh are more discriminative than other components. We also did challenging cross-race experiments that used Asian gait data to classify the gender of Europeans, and vice versa. Encouraging results were obtained. All the above prove that gait-based gender classification is feasible in controlled environments. In real applications, it still suffers from many difficulties, such as view variation, clothing and shoes changes, or carrying objects. We analyze the difficulties and suggest some possible solutions. Shiqi Yu 0001, Tieniu Tan, Kaiqi Huang, Kui Jia, Xinyu Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2009 | View-Independent Behavior AnalysisabstractThe motion analysis of the human body is an important topic of research in computer vision devoted to detecting, tracking, and understanding people's physical behavior. This strong interest is driven by a wide spectrum of applications in various areas such as smart video surveillance. Most research in behavior (or gesture) representation focusses on view-dependent representation, and some research on view invariance considers only information from 3-D models, which is effective under considerable changes of viewpoint. This paper introduces a view-independent behavior-analysis framework based on decision fusion in which distance and view angle factors are analyzed. This is a first effort to tackle the problem of behaviors under significant changes in view angle, and a first corresponding video database is built. Kaiqi Huang, Dacheng Tao, Yuan Yuan 0001, Xuelong Li 0001, Tieniu Tan |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2008 | Boosting ordinal features for accurate and fast iris recognitionabstractIn this paper, we present a novel iris recognition method based on learned ordinal features.Firstly, taking full advantages of the properties of iris textures, a new iris representation method based on regional ordinal measure encoding is presented, which provides an over-complete iris feature set for learning. Secondly, a novel Similarity Oriented Boosting (SOBoost) algorithm is proposed to train an efficient and stable classifier with a small set of features. Compared with Adaboost, SOBoost is advantageous in that it operates on similarity oriented training samples, and therefore provides a better way for boosting strong classifiers. Finally, the well-known cascade architecture is adopted to reorganize the learned SOBoost classifier into a dasiacascadepsila, by which the searching ability of iris recognition towards large-scale deployments is greatly enhanced. Extensive experiments on two challenging iris image databases demonstrate that the proposed method achieves state-of-the-art iris recognition accuracy and speed. In addition, SOBoost outperforms Adaboost (Gentle-Adaboost, JS-Adaboost, etc.) in terms of both accuracy and generalization capability across different iris databases. Zhaofeng He 0001, Zhenan Sun, Tieniu Tan, Xianchao Qiu |
CVPR | 3 |
| 2008 | Enhanced biologically inspired modelabstractIt has been demonstrated by Serre et al. that the biologically inspired model (BIM) is effective for object recognition. It outperforms many state-of-the-art methods in challenging databases. However, BIM has the following three problems: a very heavy computational cost due to dense input, a disputable pooling operation in modeling relations of the visual cortex, and blind feature selection in a feed-forward framework. To solve these problems, we develop an enhanced BIM (EBIM), which removes uninformative input by imposing sparsity constraints, utilizes a novel local weighted pooling operation with stronger physiological motivations, and applies a feedback procedure that selects effective features for combination. Empirical studies on the CalTech5 database and CalTech101 database show that EBIM is more effective and efficient than BIM. We also apply EBIM to the MIT-CBCL street scene database to show it achieves comparable performance in comparison with the current best performance. Moreover, the new system can process images with resolution 128 times 128 at a rate of 50 frames per second and enhances the speed 20 times at least in comparison with BIM in common applications. Yongzhen Huang, Kaiqi Huang, Liangsheng Wang, Dacheng Tao, Tieniu Tan, Xuelong Li 0001 |
CVPR | 5 |
| 2008 | Practical camera auto-calibration based on object appearance and motion for traffic scene visual surveillanceabstractCamera calibration, as a fundamental issue in computer vision, is indispensable in many visual surveillance applications. Firstly, calibrated camera can help to deal with perspective distortion of object appearance on image plane. Secondly, calibrated camera makes it possible to recover metrics from images which are robust to scene or view an gle changes. In addition, with calibrated cameras, we can make use of prior information of 3D models to estimate 3D pose of objects and make object detection or tracking more robust to noise and occlusions. In this paper, we propose an automatic method to recover camera models from traffic scene surveillance videos. With only the camera height H measured, we can completely recover both intrinsic and extrinsic parameters of cameras based on appearance and motion of objects in videos. Experiments are conducted in different scenes and experimental results demonstrate the effectiveness and practicability of our approach, which can be adopted in many traffic scene surveillance applications. Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan |
CVPR | 4 |
| 2008 | Robust 3D face recognition in uncontrolled environmentsabstractMost current 3D face recognition algorithms are designed based on the data collected in controlled situations, which leads to the un-guaranteed performance in practical systems. In this paper, we propose a Robust Local Log-Gabor Histograms (RLLGH) method to handle the uncontrolled problems encountered in 3D face recognition. In this challenging topic, large expressions and data noises are two main obstacles. To overcome the large expressions, we choose Log-Gabor features (LGF) to extract the distinctive and robust information embedded in 3D faces, which will be represented as 3D Log-Gabor faces. Data noises are summarized as distorted meshes, hair occlusions and misalignments. To overcome these problems, we introduce a robust local histogram (RLH) strategy, which takes advantage of the robustness of the accurate local statistical information. The combination of LGF and RLH leads to RLLGH. The novelties of this paper come from 1) Our work aims at studying 3D face recognition performance in uncontrolled environments; 2) We find that embedding LGF into the LVC framework leads to robustness in handling large expression variations; 3) The RLH strategy gives a promising way to solve the data noises problem. Our experiments are based on the large expression subset in FRGC2.0 3D face database and the expression subset in CASIA 3D face database. Experimental results show the efficiency, robustness and generalization of our proposed method. Zhenan Sun, Tieniu Tan, Zhaofeng He 0001 |
CVPR | 3 |
| 2008 | Multi-thread Parsing for Recognizing Complex Events in Videos
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan |
ECCV (3) | 3 |
| 2008 | Effects of watermarking on iris recognition performanceabstractProtection of biometric data and templates is a crucial issue for the security of biometric systems, and biometric watermarking is introduced for this purpose. However, watermarking introduces extra information into the biometric data (biometric images or biometric feature templates) which leads to certain distortion. In addition, watermarked images are always subject to the risk of being attacked. Hence, whether and how biometric recognition performance will be affected by biometric watermarking deserves investigation. In this paper, we make a first attempt in such investigations by studying two application scenarios in the context of iris recognition, namely protection of iris templates by hiding them in cover images as watermarks (iris watermarks), and protection of iris images by watermarking them. Experimental results suggest that watermark embedding in iris images does not introduce detectable decreases on iris recognition performance whereas recognition performance drops significantly if iris watermarks suffer from severe attacks. Jing Dong 0003, Tieniu Tan |
ICARCV | 2 |
| 2008 | Matching tracking sequences across widely separated camerasabstractIn this paper, we present a new solution to the problem of matching tracking sequences across different cameras. Unlike snapshot-based appearance matching which matches objects by a single image, we focus on sequence matching to alleviate the uncertainties brought by segmentation errors and partial occlusions. By incorporating multiple snapshots of the same object, the influence of the variation is alleviated. At the training stage, given the sequence of a queried person under one camera, the appearance model is formulated by concatenating feature vectors with the majority of votes over the sequence. At the testing stage, Bayesian inference is incorporated into the identification framework to accumulate the temporal information in the sequence. Experimental results demonstrate the effectiveness of the proposed method. Yinghao Cai, Kaiqi Huang, Tieniu Tan |
ICIP | 3 |
| 2008 | Blind image steganalysis based on run-length histogram analysisabstractIn this paper, a new, simple but effective method is proposed for blind image steganalysis, which is based on run-length histogram analysis. Higher-order statistics of characteristic functions of three types of image run-length histograms are selected as features. Support vector machine is used as classifier. Experimental results demonstrate that the proposed scheme significantly outperforms prior arts in detection accuracy and generality. Jing Dong 0003, Tieniu Tan |
ICIP | 2 |
| 2008 | Multispectral palm image fusion for accurate contact-free palmprint recognitionabstractIn this paper, we propose to improve the verification performance of a contract-free palmprint recognition system by means of feature- level image registration and pixel-level fusion of multi-spectral palm images. Our method involves image acquisition via a dedicated device under contact-free and multi-spectral environment, preprocessing to locate region of interest (ROI) from each individual hand images, feature-level registration to align ROIs from different spectral images in one sequence and fusion to combine images from multiple spectra. The advantages of the proposed method include better hygiene and higher verification performance. Given a database composed of images from 330 hands, two out of four state of the art fusion strategies offer significant performance gain and the best equal error rate (EER) is 0.5%. Ying Hao, Zhenan Sun, Tieniu Tan |
ICIP | 3 |
| 2008 | Enhanced usability of iris recognition via efficient user interface and iris image restorationabstractIn this paper, we investigate the possibility of enhancing the usability of iris recognition via exploration of the specular spots in iris images. Firstly, the spatial configuration of the specular spots in iris images is utilized to estimate the distance between the user and the camera. Based on this a friendly user interface is established to assist users for their range adjustment. Furthermore, the estimated distance is used by an adaptive image restoration scheme to restore the blurred iris image, thereby increasing the depth of field of the iris camera. Experimental results show that the proposed method significantly enhances the usability of iris recognition without noticeable computation cost. Zhaofeng He 0001, Zhenan Sun, Tieniu Tan, Xianchao Qiu |
ICIP | 3 |
| 2008 | Robust eyelid, eyelash and shadow localization for iris recognitionabstractEyelids, eyelashes and shadows are three major challenges for effective iris segmentation, which have not been adequately addressed in the current literature. In this paper, we present a novel method to localize each of them. First, a novel coarse-line to fine-parabola eyelid fitting scheme is developed for accurate and fast eyelid localization. Then, a smart prediction model is established to determine an appropriate threshold for eyelash and shadow detection. Experimental results on the challenging CASIA-IrisV3-Lamp iris image database demonstrate that the proposed method outperforms state-of-the-art methods in both accuracy and speed. Zhaofeng He 0001, Tieniu Tan, Zhenan Sun, Xianchao Qiu |
ICIP | 2 |
| 2008 | Palmprint image synthesis: A preliminary studyabstractIn this paper we present a preliminary study of palmprint image synthesis and propose a framework for synthesizing palmprint texture. We first extract principal lines of real palmprints using edge detection and synthesize wrinkles and ridges of palm using patch-based sampling. Then we incorporate principal lines, wrinkles and ridges to obtain the final synthetic image. After that multiple images are derived from each artificial palm to simulate the intra-class images. Our approach can generate large palmprint databases which preserve inter-class and intra-class variations. Experimental results demonstrate that the synthetic images bear a close resemblance to real palmprints in terms of appearance as well as statistical properties, showing a promising usage in algorithms evaluation and comparison. Zhuoshi Wei, Zhenan Sun, Tieniu Tan |
ICIP | 4 |
| 2008 | Robust automated ground plane rectification based on moving vehicles for traffic scene surveillanceabstractMost outdoor visual surveillance scenes involve objects of interest moving on the ground plane. However, perspective distortion introduces many difficulties to various applications like object classification and activity recognition. In this paper, we propose a robust automated method for both affine and metric rectification of the ground plane based on appearance and motion of vehicles in traffic scene surveillance videos. This rectification enables normalization of object properties like size, length and velocity. Various useful applications are presented and experimental results demonstrate the effectiveness and robustness of the proposed method. Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan |
ICIP | 4 |
| 2008 | Learning efficient codes for 3D face recognitionabstractFace representation based on the visual codebook becomes popular because of its excellent recognition performance, in which the critical problem is how to learn the most efficient codes to represent the facial characteristics. In this paper, we introduce the quadtree clustering algorithm to learn the facial-codes to boost 3D face recognition performance. The merits of quadtree clustering come from: (1) It is robust to data noises; (2) It can adaptively assign clustering centers according to the density of data distribution. We make a comparison between quadtree and some widely used clustering methods, such as g-means, k-means, normalized-cut and mean-shift. Experimental results show that using the facial- codes learned by quadtree clustering gives the best performance for 3D face recognition. Zhenan Sun, Tieniu Tan |
ICIP | 3 |
| 2008 | Human appearance matching across multiple non-overlapping camerasabstractIn this paper, we present a new solution to the problem of appearance matching across multiple non-overlapping cameras. Objects of interest, pedestrians are represented by a set of region signatures centered at points sampled from edges. The problem of frame-to-frame appearance matching is formulated as finding corresponding points in two images as minimization of a cost function over the space of correspondence. The correspondence problem is solved under integer optimization framework where the cost function is determined by similarity of region signatures as well as geometric constraints between points. Experimental results demonstrate the effectiveness of the proposed method. Yinghao Cai, Kaiqi Huang, Tieniu Tan |
ICPR | 3 |
| 2008 | How to make iris recognition easier?abstractIris recognition is regarded as the most reliable biometrics and has been widely applied in both public and personal security areas. However users have to highly cooperate with the iris cameras to make his iris images well captured. In this paper, we aim to discuss whether and how we can make iris recognition easier. Firstly the restricting factors of iris image acquisition are analyzed and the optical formulas are derived. Then the solutions of state-of-the-art iris recognition systems are reviewed and summarized. Finally, we propose two novel iris recognition systems with good human-computer-interface but with two different strategies which respectively meet the requirements of low-end and high-end market. Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2008 | Combine hierarchical appearance statistics for accurate palmprint recognitionabstractPalmprint recognition is an active member of biometrics in recent years. State-of-the-art algorithms of palmprint recognition describe appearances of palmprints efficiently through local texture analysis. Following this framework, we propose a novel approach of palmprint recognition in this paper, which represents palmprint images based on statistics and spatial arrangement of appearance descriptors within local image areas. In this method, we firstly design a robust descriptor to encode properties of palmprint appearances of local regions. The whole image is divided into non-overlapped blocks at increasingly fine resolutions successively, so as to describe the spatial layout in hierarchical scales. For a specific spatial resolution, local distributions of the proposed descriptors in the blocks are concatenated to represent structures of palmprint structures. Finally, distribution information of different resolutions is combined to provide complementary descriptive power. Promising experimental results demonstrate that the proposed method achieves even better performances than the state-of-the-art approaches. Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2008 | A hierarchical model for the evaluation of biometric sample qualityabstractThe evaluation of biometric sample quality is of great importance in the evaluation of biometric algorithms. In this paper, we propose a novel hierarchical model to compute the sample quality at three levels. This model is developed on the basis of three types of influencing factors: global factors, subjective factors and variable factors. We adopt different strategies to compute the corresponding three level qualities: database level quality, class level quality and image level quality. The database level quality is estimated by experience. Then, we compute the mean value of variable number of normalized genuine scores, the quantiles of which are used to determine the class level quality. On the image level quality evaluation, a novel concept of subset frequency is proposed. Zhenan Sun, Tieniu Tan |
ICPR | 3 |
| 2008 | Estimating the number of people in crowded scenes by MID based foreground segmentation and head-shoulder detectionabstractThis paper proposes a novel method to address the problem of estimating the number of people in surveillance scenes with people gathering and waiting. The proposed method combines a MID (mosaic image difference) based foreground segmentation algorithm and a HOG (histograms of oriented gradients) based head-shoulder detection algorithm to provide an accurate estimation of people counts in the observed area. In our framework, the MID-based foreground segmentation module provides active areas for the head-shoulder detection module to detect heads and count the number of people. Numerous experiments are conducted and convincing results demonstrate the effectiveness of our method. Min Li 0022, Zhaoxiang Zhang 0001, Kaiqi Huang, Tieniu Tan |
ICPR | 4 |
| 2008 | Counterfeit iris detection based on texture analysisabstractThis paper addresses the issue of counterfeit iris detection, which is a liveness detection problem in biometrics. Fake iris mentioned here refers to iris wearing color contact lens with textures printed onto them. We propose three measures to detect fake iris: measuring iris edge sharpness, applying Iris-Texton feature for characterizing the visual primitives of iris textures and using selected features based on co-occurrence matrix (CM). Extensive testing is carried out on two datasets containing different types of contact lens with totally 640 fake iris images, which demonstrates that Iris-Texton and CM features are effective and robust in anticounterfeit iris. Detailed comparisons with two state-of-the-art methods are also presented, showing that the proposed iris edge sharpness measure acquires a comparable performance with these two methods, while Iris-Texton and CM features outperform the state-of-the-art. Zhuoshi Wei, Xianchao Qiu, Zhenan Sun, Tieniu Tan |
ICPR | 4 |
| 2008 | Synthesis of large realistic iris databases using patch-based samplingabstractThis paper presents a framework to synthesize large realistic iris databases, providing an alternative to iris database collection. Firstly, iris patch is used as a basic element to characterize visual primitive of iris texture, and patch-based sampling is applied to create an iris prototype. Then a set of pseudo irises with intra-class variations are derived from the prototype. Qualitative and quantitative studies reveal that synthetic databases are well suited for evaluating iris recognition systems by achieving three goals: (1) the synthetic iris images bear a close resemblance to real iris images in terms of visual appearance; (2) the proposed framework is able to generate databases with large capacity; (3) statistical performance shows that the synthetic iris images hold all the major characteristics of real iris images. Zhuoshi Wei, Tieniu Tan, Zhenan Sun |
ICPR | 2 |
| 2008 | Boosting local feature descriptors for automatic objects classification in traffic scene surveillanceabstractWe address the problem of automatic object classification for traffic scene surveillance, which is very challenging for the low resolution videos, large intra-class variations and real-time requirement. In this paper, we propose a new strategy for object classification by boosting different local feature descriptors in motion blobs. We not only evaluate the performance of each local feature descriptor, but also fuse these descriptors to achieve better performance. Numerous experiments are conducted and experimental results demonstrate the effectiveness and efficiency of our approach with robustness to noise and variance of view angles, lighting conditions and environments. Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan |
ICPR | 4 |
| 2008 | 3D model based vehicle localization by optimizing local gradient based fitness evaluationabstractWe address the problem of 3D model based vehicle localization in calibrated traffic scenes. A wire-frame vehicle model is set up as prior information and an efficient local gradient based method is proposed to evaluate the fitness between the projection of 3D model and image data, which illustrates smooth optimization surface and more conspicuous peak with low computational cost. Gradient decent is then applied to optimize the evaluation score for localization. Experimental results demonstrate the accuracy, efficiency and robustness of the proposed method for model based vehicle localization. Zhaoxiang Zhang 0001, Min Li 0022, Kaiqi Huang, Tieniu Tan |
ICPR | 4 |
| 2008 | Run-Length and Edge Statistics Based Approach for Image Splicing Detection
Jing Dong 0003, Wei Wang 0025, Tieniu Tan, Yun Q. Shi 0001 |
IWDW | 3 |
| 2008 | A real-time object detecting and tracking system for outdoor night surveillance
Kaiqi Huang, Liangsheng Wang, Tieniu Tan, Stephen J. Maybank |
Pattern Recognit. | 3 |
| 2007 | Continuously Tracking Objects Across Multiple Widely Separated Cameras
Yinghao Cai, Wei Chen 0012, Kaiqi Huang, Tieniu Tan |
ACCV (1) | 4 |
| 2007 | Palmprint Recognition Under Unconstrained Scenes
Zhenan Sun, Tieniu Tan |
ACCV (2) | 4 |
| 2007 | Comparative Studies on Multispectral Palm Image Fusion for Biometrics
Ying Hao, Zhenan Sun, Tieniu Tan |
ACCV (2) | 3 |
| 2007 | Discriminating 3D Faces by Statistics of Depth Differences
Yunhong Wang 0001, Tieniu Tan |
ACCV (2) | 3 |
| 2007 | Multi-view Gymnastic Activity Recognition with Fused HMM
Ying Wang 0003, Kaiqi Huang, Tieniu Tan |
ACCV (1) | 3 |
| 2007 | Fusion of Face and Palmprint for Personal Identification Based on Ordinal FeaturesabstractIn this paper, we present a face and palmprint multimodal biometric identification method and system to improve the identification performance. Effective classifiers based on ordinal features are constructed for faces and palmprints, respectively. Then, the matching scores from the two classifiers are combined using several fusion strategies. Experimental results on a middle-scale data set have demonstrated the effectiveness of the proposed system. Rufeng Chu, Shengcai Liao, Zhenan Sun, Stan Z. Li, Tieniu Tan |
CVPR | 6 |
| 2007 | Cast Shadow Removal Combining Local and Global FeaturesabstractIn this paper, we present a method using pixel-level information, local region-level information and global-level information to remove shadow. At the pixel-level, we employ GMM to model the behavior of cast shadow for every pixel in the HSV color space, as it can deal with complex illumination conditions. However, unlike the GMM for background which can obtain sample every frame, this model for shadow needs more frames to get the same number of sample, because shadow may not appear at the same pixel for each frame. Therefore, it will take a long time to converge. To overcome this drawback, we use the local region-level information to get more samples and global-level information to improve a preclassifier and then, by using it, we get samples which are more likely to be shadow. Also, at the local region-level, we use Markov random fields to represent dependencies between the label of single pixel and labels of its neighborhood. Moreover, to make global level information more robust, tracking information is used. Experimental results show that the proposed method is efficient and robust. Kaiqi Huang, Tieniu Tan, Liangsheng Wang |
CVPR | 3 |
| 2007 | Online Appearance Model Learning for Video-Based Face RecognitionabstractIn this paper, we propose a novel online learning method which can learn appearance models incrementally from a given video stream. The data of each frame in the video can be discarded as soon as it has been processed. We only need to maintain a few linear eigenspace models and a transition matrix to approximately construct face appearance manifolds. It is convenient to use these learnt models for video-based face recognition. There are mainly two contributions in this paper. First, we propose an algorithm which can learn appearance models online without using a pre-trained model. Second, we propose a method for eigenspace splitting to prevent that most samples cluster into the same eigenspace. This is useful for clustering and classification. Experimental results show that the proposed method can both learn appearance models online and achieve high recognition rate. Yunhong Wang 0001, Tieniu Tan |
CVPR | 3 |
| 2007 | Recognizing Night Walkers Based on One Pseudoshape Representation of GaitabstractGait is a promising biometric cue which can facilitate the recognition of human beings, particularly when other biometrics are unavailable. Existing work for gait recognition, however, lays more emphasis on the problem of daytime walker recognition and overlooks the significance of walker recognition at night. This paper deals with the problem of recognizing nighttime walkers. We take advantage of infrared gait patterns to accomplish this task: 1) Walker detection is improved using intensity compensation-based background subtraction; 2) pseudoshape-based features are proposed to describe gait patterns; 3) the dimension of gait features is reduced through the principal component analysis (PCA) and linear discriminant analysis (LDA) techniques; 4) temporal cues are exploited in the form of the relevant component analysis (RCA) learning; 5) the nearest neighbor classifier is used to recognize unknown gait. Experimental results justify the effectiveness of our method and show that our method has an encouraging potential for the application in surveillance systems. Daoliang Tan, Kaiqi Huang, Shiqi Yu 0001, Tieniu Tan |
CVPR | 4 |
| 2007 | Human Activity Recognition Based on R TransformabstractThis paper addresses human activity recognition based on a new feature descriptor. For a binary human silhouette, an extended radon transform, R transform, is employed to represent low-level features. The advantage of the R transform lies in its low computational complexity and geometric invariance. Then a set of HMMs based on the extracted features are trained to recognize activities. Compared with other commonly-used feature descriptors, R transform is robust to frame loss in video, disjoint silhouettes and holes in the shape, and thus achieves better performance in recognizing similar activities. Rich experiments have proved the efficiency of the proposed method. Ying Wang 0003, Kaiqi Huang, Tieniu Tan |
CVPR | 3 |
| 2007 | Semi-supervised Learning on Semantic Manifold for Event Analysis in Dynamic ScenesabstractEvents can be considered as obvious changes of important properties with semantic meanings. Usually, all these properties are measurable and continual in complex formats and higher dimensions. It is hard to define and measure semantic events on the original observed data. However, according to the perception process of human being, these spatial-temporal continuous data can be mapped onto corresponding smooth manifolds, and different appearances on manifolds can indicate different semantic meanings. In this paper, we propose a semi-supervised learning method, which is based on partially labeled data, to map original observed data onto semantic manifolds for events definition and analysis in dynamic scenes. Furthermore we also perform semantic representations for various events in real world scenes. Finally, we present experimental results to evaluate the performance of our method. Lun Xin, Tieniu Tan |
CVPR | 2 |
| 2007 | EDA Approach for Model Based Localization and Recognition of VehiclesabstractWe address the problem of model based recognition. Our aim is to localize and recognize road vehicles from monocular images in calibrated scenes. A deformable 3D geometric vehicle model with 12 parameters is set up as prior information and Bayesian Classification Error is adopted for evaluation of fitness between the model and images. Using a novel evolutionary computing method called EDA (Estimation of Distribution Algorithm), we can not only determine the 3D pose of the vehicle, but also obtain a 12 dimensional vector which corresponds to the 12 shape parameters of the model. By clustering obtained vectors in the parameter space, we can recognize different types of vehicles. Experimental results demonstrate the effectiveness of the approach to vehicles of different types and poses. Thanks to EDA, we can not only localize and recognize vehicles, but also show the whole evolution procedure of the deformable model which gradually fits the image better and better. Zhaoxiang Zhang 0001, Weishan Dong, Kaiqi Huang, Tieniu Tan |
CVPR | 4 |
| 2007 | Trajectory Series Analysis based Event Rule Induction for Visual SurveillanceabstractIn this paper, a generic rule induction framework based on trajectory series analysis is proposed to learn the event rules. First the trajectories acquired by a tracking system are mapped into a set of primitive events that represent some basic motion patterns of moving object. Then a minimum description length (MDL) principle based grammar induction algorithm is adopted to infer the meaningful rules from the primitive event series. Compared with previous grammar rule based work on event recognition where the rules are all defined manually, our work aims to learn the event rules automatically. Experiments in a traffic crossroad have demonstrated the effectiveness of our methods. Shown in the experimental results, most of the grammar rules obtained by our algorithm are consistent with the actual traffic events in the crossroad. Furthermore the traffic lights rule in the crossroad can also be leaned correctly with the help of eliminating the irrelevant trajectories. Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan, Liangsheng Wang |
CVPR | 3 |
| 2007 | Robust 3D Face Recognition Using Learned Visual CodebookabstractIn this paper, we propose a novel learned visual code-book (LVC) for 3D face recognition. In our method, we first extract intrinsic discriminative information embedded in 3D faces using Gabor filters, then K-means clustering is adopted to learn the centers from the filter response vectors. We construct LVC by these learned centers. Finally we represent 3D faces based on LVC and achieve recognition using a nearest neighbor (NN) classifier. The novelty of this paper comes from 1) We first apply textons based methods into 3D face recognition; 2) We encompass the efficiency of Gabor features for face recognition and the robustness of texton strategy for texture classification simultaneously. Our experiments are based on two challenging databases, CASIA 3D face database and FRGC2.0 3D face database. Experimental results show LVC performs better than many commonly used methods. Zhenan Sun, Tieniu Tan |
CVPR | 3 |
| 2007 | Fast Principal Component Analysis using Eigenspace MergingabstractIn this paper, we propose a fast algorithm for principal component analysis (PCA) dealing with large high-dimensional data sets. A large data set is firstly divided into several small data sets. Then, the traditional PCA method is applied on each small data set and several eigenspace models are obtained, where each eigenspace model is computed from a small data set. At last, these eigenspace models are merged into one eigenspace model which contains the PCA result of the original data set. Experiments on the FERET data set show that this algorithm is much faster than the traditional PCA method, while the principal components and the reconstruction errors are almost the same as that given by the traditional method. Yunhong Wang 0001, Tieniu Tan |
ICIP (6) | 4 |
| 2007 | Learning Appearance Primitives of Iris Images for Ethnic ClassificationabstractIris pattern is commonly regarded as a kind of phenotypic feature without relation to genes. In our previous work, we argued that iris texture is race related, and its genetic information is illustrated in coarse scale texture features, rather than preserved in the minute local features of state-of-the-art iris recognition algorithms. In this paper, we propose a novel ethnic classification method based on learning appearance primitives of iris images. So we not only confirm that iris texture is race related, but also try to find out which kinds of iris visual primitives make iris images look different between Asian and non-Asian. In our scheme, we learned a small finite vocabulary of micro-structures, which are called iris-textons, to represent visual primitives of iris images. Then we use iris-texton histogram to capture the difference between iris textures. Finally iris images are grouped into two race categories, Asian and non-Asian, by support vector machine (SVM). Based on the proposed method, we get a higher correct classification rate (CCR) of 91.02% than our previous method on a database containing 2400 iris samples. Xianchao Qiu, Zhenan Sun, Tieniu Tan |
ICIP (2) | 3 |
| 2007 | SAR and SPOT Image Registration Based on Mutual Information with Contrast MeasureabstractIn this study, we propose a novel robust mutual information (MI) based method to register SAR and SPOT images. Traditional MI based method can register SAR and SPOT images well. However, its robustness is not satisfying for local spatial information is absent. In our approach, first, local contrast of 5*5 windows centered at each point in both images is calculated, then the contrast value is assigned to each pixel and two contrast images are obtained. Finally, the SAR and SPOT images are registered by maximizing the MI between their contrast images. Experimental results show that compared with traditional MI, our approach is much more robust and acquires comparable or even higher accuracy. Meanwhile, compared with the MI with orientation information based registration (MIOI), another robust MI based method, our algorithm works much faster and more accurately. Lixia Shu, Tieniu Tan |
ICIP (5) | 2 |
| 2007 | Orthogonal Diagonal Projections for Gait RecognitionabstractGait has received much attention from researchers in the vision field due to its utility in walker identification. One of the key issues in gait recognition is how to extract discriminative shape features from 2D human silhouette images. This paper deals with the problem of gait-based walker recognition using statistical shape features. First, we normalize walkers' silhouettes (to facilitate gait feature comparison) into a square form and use the orthogonal projections in the positive and negative diagonal directions to draw personal signatures contained in gait patterns. Then principal component analysis (PCA) and linear discriminant analysis (LDA) are applied to reduce the dimensionality of original gait features and to improve the topological structure in the feature space. Finally, this paper accomplishes the recognition of unknown gait features based on the nearest neighbor rule, with the discussion of the effect of distance metrics and scales on discriminating performance. Experimental results justify the potential of our method. Daoliang Tan, Kaiqi Huang, Shiqi Yu 0001, Tieniu Tan |
ICIP (1) | 4 |
| 2007 | Abnormal Activity Recognition in Office Based on R TransformabstractThis paper introduces an abnormal activity recognition method based on a new feature descriptor for human silhouette. For a binary human silhouette, an extended radon transform, R transform, is employed to represent low-level features. The information that the initial silhouette carries is transformed in a compact way preserving important spatial information of the activities. Then a set of HMMs based on the features extracted by our method are trained to recognize abnormal activities. Experiments have proved the accuracy and efficiency of the proposed method, and the comparison with Fourier descriptor illustrates its robustness to disjoint shapes and shapes with holes. Ying Wang 0003, Kaiqi Huang, Tieniu Tan |
ICIP (1) | 3 |
| 2007 | Group Activity Recognition Based on ARMA Shape Sequence ModelingabstractIn this paper, we propose a system identification approach for group activity recognition in traffic surveillance. Statistical shape theory is used to extract features, and then ARMA (autoregressive and moving average) is adopted for feature learning and activity identification. Here only a few points, instead of the complete trajectory of each object are used to describe the dynamic information of group activity. And ARMA is employed to learn activity sequences. The performance of the proposed method is proved by experiments on 570 video sequences, with the average recognition rate of 88% (compared with 81% of HMM). The extracted features are invariant to zoom, pan and tilt, which is also proved in the experiments. Ying Wang 0003, Kaiqi Huang, Tieniu Tan |
ICIP (3) | 3 |
| 2007 | Dynamic Audio-Visual Mapping using Fused Hidden Markov Model Inversion MethodabstractRealistic audio-visual mapping remains a very challenging problem. Having short time delay between inputs and outputs is also of great importance. In this paper, we present a new dynamic audio-visual mapping approach based on the Fused Hidden Markov Model Inversion method. In our work, the Fused HMM is used to model the loose synchronization nature of the two tightly coupled audio speech and visual speech streams explicitly. Given novel audio inputs, the inversion algorithm is derived to synthesize visual counterparts by maximizing the joint probabilistic distribution of the Fused HMM. When it is implemented in the subsets built from the training corpus, realistic synthesized facial animation having relative short time delay is obtained. Experiments on a 3D motion capture bimodal database show that the synthetic results are comparable with the ground truth. Le Xin, Jianhua Tao 0001, Tieniu Tan |
ICIP (3) | 3 |
| 2007 | Real-Time Moving Object Classification with Automatic Scene DivisionabstractWe address the problem of moving object classification. Our aim is to classify moving objects of traffic scene videos into pedestrians, bicycles and vehicles. Instead of supervised learning and manual labeling of large training samples, our classifiers are initialized and refined online automatically. With efficient features extracted and organized, the approach can be real-time and achieve high classification accuracy. Once the view or scene changes detected, the algorithm can automatically refine the classifiers and adapt them to new environments. Experimental results demonstrate the effectiveness and robustness of the proposed approach. Zhaoxiang Zhang 0001, Yinghao Cai, Kaiqi Huang, Tieniu Tan |
ICIP (5) | 4 |
| 2007 | Fusion Based Blind Image Steganalysis by Boosting Feature Selection
Jing Dong 0003, Xiaochuan Chen, Tieniu Tan |
IWDW | 4 |
| 2007 | Introduction to the Special Issue on Biometrics: Progress and DirectionsabstractThe guest editors provide an overview of the articles selected for this special issue. The issue's goal is to document the current state-of-the-art, acknowledge the latest breakthroughs achieved by scientists working in the area of biometric recognition, and identify future promising research areas. It is thought the selection of papers discussed should give readers a good idea of where researchers have been focusing, both on long- studied problems still needing more work and on newer challenges. A fundamental of the field of biometrics is an ever-increasing need for better recognition and stronger security. But, as public and commercial biometric deployments increase in number, there is also more need to understand privacy issues and to provide greater ease-of-use. The volume and quality of papers in this special issue indicate that much progress has been made in many aspects of the biometrics field and that there are challenging and promising future directions still to follow. Salil Prabhakar, Josef Kittler, Davide Maltoni, Lawrence O'Gorman, Tieniu Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2007 | Real-time hand tracking using a mean shift embedded particle filter
Caifeng Shan, Tieniu Tan, Yucheng Wei |
Pattern Recognit. | 2 |
| 2006 | Detecting and Tracking Distant Objects at Night Based on Human Visual System
Kaiqi Huang, Liangsheng Wang, Tieniu Tan |
ACCV (2) | 3 |
| 2006 | Recognize Multi-people Interaction Activity by PCA-HMMs
Ying Wang 0003, Xinwen Hou, Tieniu Tan |
ACCV (1) | 3 |
| 2006 | From Motion Patterns to Visual Concepts for Event Analysis in Dynamic Scenes
Lun Xin, Tieniu Tan |
ACCV (1) | 2 |
| 2006 | Modelling the Effect of View Angle Variation on Appearance-Based Gait Recognition
Shiqi Yu 0001, Daoliang Tan, Tieniu Tan |
ACCV (1) | 3 |
| 2006 | Complex Activity Representation and Recognition by Extended Stochastic Grammar
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan |
ACCV (1) | 3 |
| 2006 | Combining Statistics of Geometrical and Correlative Features for 3D Face RecognitionabstractIn this paper, we present a new method for face recognition using range data. The proposed method is based on both global statistics of geometrical features and local statistics of correlative features of facial surfaces. Firstly, we analyze the performances of common geometrical representations by using global histograms for matching. Secondly, we propose a new method to encode the relationships between points and their neighbors, which are demonstrated to own great power to represent the intrinsic structure of facial surfaces. Finally, the two kinds of features are supposed to be complementary to some extent, and the combination of them is proven to be able to improve the recognition performance. All the experiments are performed on the full 3D face dataset of FRGC 2.0 which is the largest 3D face database so far. Promising results have demonstrated the effectiveness of our proposed method. 1 Yunhong Wang 0001, Tieniu Tan |
BMVC | 3 |
| 2006 | Learning Boosted Asymmetric Classifiers for Object DetectionabstractObject detection can be posted as those classification tasks where the rare positive patterns are to be distinguished from the enormous negative patterns. To avoid the danger of missing positive patterns, more attention should be payed on them. Therefore there should be different requirements for False Reject Rate (FRR) and False Accept Rate (FAR) , and learning a classifier should use an asymmetric factor to balance between FRR and FAR. In this paper, a normalized asymmetric classification error is proposed for the task of rejecting negative patterns. Minimizing it not only controls the ratio of FRR and FAR, but more importantly limits the upper-bound of FRR. The latter characteristic is advantageous for those tasks where there is a requirement for low FRR. Based on this normalized asymmetric classification error, we develop an asymmetric AdaBoost algorithm with variable asymmetric factor and apply it to the learning of cascade classifiers for face detection. Experiments demonstrate that the proposed method achieves less complex classifiers and better performance than some previous AdaBoost methods. Xinwen Hou, Cheng-Lin Liu 0001, Tieniu Tan |
CVPR (1) | 3 |
| 2006 | Learning Effective Intrinsic Features to Boost 3D-Based Face Recognition
Chenghua Xu, Tieniu Tan, Stan Z. Li, Yunhong Wang 0001 |
ECCV (2) | 2 |
| 2006 | Principal Axis-Based Correspondence between Multiple Cameras for People TrackingabstractVisual surveillance using multiple cameras has attracted increasing interest in recent years. Correspondence between multiple cameras is one of the most important and basic problems which visual surveillance using multiple cameras brings. In this paper, we propose a simple and robust method, based on principal axes of people, to match people across multiple cameras. The correspondence likelihood reflecting the similarity of pairs of principal axes of people is constructed according to the relationship between "ground-points" of people detected in each camera view and the intersections of principal axes detected in different camera views and transformed to the same view. Our method has the following desirable properties: 1) Camera calibration is not needed. 2) Accurate motion detection and segmentation are less critical due to the robustness of the principal axis-based feature to noise. 3) Based on the fused data derived from correspondence results, positions of people in each camera view can be accurately located even when the people are partially occluded in all views. The experimental results on several real video sequences from outdoor environments have demonstrated the effectiveness, efficiency, and robustness of our method. Weiming Hu 0004, Tieniu Tan, Jianguang Lou, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2006 | A System for Learning Statistical Motion PatternsabstractAnalysis of motion patterns is an effective approach for anomaly detection and behavior prediction. Current approaches for the analysis of motion patterns depend on known scenes, where objects move in predefined ways. It is highly desirable to automatically construct object motion patterns which reflect the knowledge of the scene. In this paper, we present a system for automatically learning motion patterns for anomaly detection and behavior prediction based on a proposed algorithm for robustly tracking multiple objects. In the tracking algorithm, foreground pixels are clustered using a fast accurate fuzzy K-means algorithm. Growing and prediction of the cluster centroids of foreground pixels ensure that each cluster centroid is associated with a moving object in the scene. In the algorithm for learning motion patterns, trajectories are clustered hierarchically using spatial and temporal information and then each motion pattern is represented with a chain of Gaussian distributions. Based on the learned statistical motion patterns, statistical methods are used to detect anomalies and predict behaviors. Our system is tested using image sequences acquired, respectively, from a crowded real traffic scene and a model traffic scene. Experimental results show the robustness of the tracking algorithm, the efficiency of the algorithm for learning motion patterns, and the encouraging performance of algorithms for anomaly detection and behavior prediction. Weiming Hu 0004, Xuejuan Xiao, Zhouyu Fu, Tieniu Tan, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2006 | Combining local features for robust nose location in 3D facial data
Chenghua Xu, Tieniu Tan, Yunhong Wang 0001, Long Quan |
Pattern Recognit. Lett. | 2 |
| 2005 | Affective Computing: A Review
Jianhua Tao 0001, Tieniu Tan |
ACII | 2 |
| 2005 | Ordinal Palmprint Represention for Personal IdentificationabstractPalmprint-based personal identification, as a new member in the biometrics family, has become an active research topic in recent years. Although great progress has been made, how to represent palmprint for effective classification is still an open problem. In this paper, we present a novel palmprint representation - ordinal measure, which unifies several major existing palmprint algorithms into a general framework. In this framework, a novel palmprint representation method, namely orthogonal line ordinal features, is proposed. The basic idea of this method is to qualitatively compare two elongated, line-like image regions, which are orthogonal in orientation and generate one bit feature code. A palmprint pattern is represented by thousands of ordinal feature codes. In contrast to the state-of-the-art algorithm reported in the literature, our method achieves higher accuracy, with the equal error rate reduced by 42% for a difficult set, while the complexity of feature extraction is halved. Zhenan Sun, Tieniu Tan, Yunhong Wang 0001, Stan Z. Li |
CVPR (1) | 2 |
| 2005 | Automatic 3D Face Modeling from VideoabstractIn this paper, we develop an efficient technique for fully automatic recovery of accurate 3D face shape from videos captured by a low cost camera. The method is designed to work with a short video containing a face rotating from frontal view to profile view. The whole approach consists of three components. First, automatic initialization is performed in the first frame with approximately frontal face. Then, to handle the case of low quality image captured by low cost camera, the 2D feature matching, head poses and underlying 3D face shape are estimated and refined iteratively in an efficient way based on image sequence segmentation. Finally, to take advantage of the sparse structure of the proposed algorithm, sparse bundle adjustment technique is further employed to speed up the computation. We demonstrate the accuracy and robustness of the algorithm using a set of experiments Le Xin, Qiang Wang 0023, Jianhua Tao 0001, Xiaoou Tang, Tieniu Tan, Harry Shum |
ICCV | 5 |
| 2005 | Similarity based vehicle trajectory clustering and anomaly detectionabstractIn this paper, we proposed a hierarchical clustering framework to classify vehicle motion trajectories in real traffic video based on their pairwise similarities. First raw trajectories are pre-processed and resampled at equal space intervals. Then spectral clustering is used to group trajectories with similar spatial patterns. Dominant paths and lanes can be distinguished as a result of two-layer hierarchical clustering. Detection of novel trajectories is also possible based on the clustering results. Experimental results demonstrate the superior performance of spectral clustering compared with conventional fuzzy K-means clustering and some results of anomaly detection are presented. Zhouyu Fu, Weiming Hu 0004, Tieniu Tan |
ICIP (2) | 3 |
| 2005 | Local manifold matching for face recognitionabstractIn this paper, we propose a novel classification method, called local manifold matching (LMM), for face recognition. LMM has great representational capacity of available prototypes and is based on the local linearity assumption that each data point and its k nearest neighbors from the same class lie on a linear manifold locally embedded in the image space. We present a supervised local manifold learning algorithm for learning all locally linear manifold structures. Then we propose the nearest manifold criterion for the classification in which the query feature point is assigned to the most matching face manifold. Experimental results show that kernel PCA incorporated with the LMM classifier achieves the best face recognition performance. Wei Liu 0035, Yunhong Wang 0001, Tieniu Tan |
ICIP (2) | 4 |
| 2005 | A novel registration method for SAR and SPOT imagesabstractIn this paper, we propose a novel mutual information (MI) based method to register SAR and SPOT images. The traditional MI can register SAR and SPOT images well. However, its robustness is weakened by the absence of orientation information. In our approach, we first extract orientation information at four directions by Gabor filters, then MI of each corresponding image pair is calculated and the average value of MI is used as an improved measure for MI. Experiments show that our method is more robust than the traditional MI method. Meanwhile our method maintains comparable accuracy to the traditional MI, which is much better than coarse manual registration. Lixia Shu, Tieniu Tan, Chunhong Pan |
ICIP (2) | 2 |
| 2005 | Extended optimization method of LSB steganalysisabstractImage steganalysis has attracted increasing attention recently. LSB steganalysis is one of the most active research topics. The paper proposes a method for LSB steganalysis of images, where the secret message is embedded in a given number L of the least significant bits. The proposed estimation method is an extension of Fridrich's method from the case L = 1 to arbitrary L > 0. A weighted stego image is defined first and then estimation formula is derived. To evaluate the proposed steganalytic method, two experiments of detection and estimation are performed. It is shown that the accuracy of detecting the existence of secret messages in images and of estimating the embedding ratio of secret messages is relatively high. Estimation errors and further studies are also discussed. Experimental results and theoretical verification show that this method is an effective method of LSB steganalysis. Xiaoyi Yu, Tieniu Tan, Yunhong Wang 0001 |
ICIP (2) | 2 |
| 2005 | Phase Correlation Based Iris Image Registration Model
Junzhou Huang, Tieniu Tan, Li Ma 0001, Yunhong Wang 0001 |
J. Comput. Sci. Technol. | 2 |
| 2005 | 3-D Model-Based Vehicle TrackingabstractThis paper aims at tracking vehicles from monocular intensity image sequences and presents an efficient and robust approach to three-dimensional (3-D) model-based vehicle tracking. Under the weak perspective assumption and the ground-plane constraint, the movements of model projection in the two-dimensional image plane can be decomposed into two motions: translation and rotation. They are the results of the corresponding movements of 3-D translation on the ground plane (GP) and rotation around the normal of the GP, which can be determined separately. A new metric based on point-to-line segment distance is proposed to evaluate the similarity between an image region and an instantiation of a 3-D vehicle model under a given pose. Based on this, we provide an efficient pose refinement method to refine the vehicle's pose parameters. An improved EKF is also proposed to track and to predict vehicle motion with a precise kinematics model. Experimental results with both indoor and outdoor data show that the algorithm obtains desirable performance even under severe occlusion and clutter. Jianguang Lou, Tieniu Tan, Weiming Hu 0004, Hao Yang 0010, Stephen J. Maybank |
IEEE Trans. Image Process. | 2 |
| 2005 | Improving iris recognition accuracy via cascaded classifiersabstractAs a reliable approach to human identification, iris recognition has received increasing attention in recent years. The most distinguishing feature of an iris image comes from the fine spatial changes of the image structure. So iris pattern representation must characterize the local intensity variations in iris signals. However, the measurements from minutiae are easily affected by noise, such as occlusions by eyelids and eyelashes, iris localization error, nonlinear iris deformations, etc. This greatly limits the accuracy of iris recognition systems. In this paper, an elastic iris blob matching algorithm is proposed to overcome the limitations of local feature based classifiers (LFC). In addition, in order to recognize various iris images efficiently a novel cascading scheme is proposed to combine the LFC and an iris blob matcher. When the LFC is uncertain of its decision, poor quality iris images are usually involved in intra-class comparison. Then the iris blob matcher is resorted to determine the input iris' identity because it is capable of recognizing noisy images. Extensive experimental results demonstrate that the cascaded classifiers significantly improve the system's accuracy with negligible extra computational cost. Zhenan Sun, Yunhong Wang 0001, Tieniu Tan, Jiali Cui |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2004 | Adaptive Multi-Resolution Fitting and its Application to Realistic Head ModelingabstractThe general approach for object modeling is to construct the surface from the high-quality range points obtained from laser scanners. In this paper, we face the noise point cloud obtained from image sequences by a common camera and develop a novel algorithm of adaptive multi-resolution fitting (AMRF) for object modeling. This algorithm combines the adaptive subdivision scheme with multi-resolution fitting so that the control model is subdivided locally and adaptively according to the local complexity of the point cloud and approximates the 3D data level by level. The proposed method can conquer the holes and outliers efficiently and create full compatibility between the complexity of the mesh model and the representation of the local details. We apply the proposed method to the complete head modeling with the real data, and the results seem very promising. Chenghua Xu, Long Quan, Yunhong Wang 0001, Tieniu Tan, Maxime Lhuillier |
GMP | 4 |
| 2004 | Fast recursive mathematical morphological transformsabstractSince many mathematical morphology operations are recursive transforms of dilation and erosion, this paper proposes fast recursive transforms to reduce computational complexity. The basic idea of the method is to compute the temporary results within a series of adaptive windows and the computing is performed on specific pixels. Each step of the recursive process consists of two parts: 1) computation is limited to the specific pixels (foreground or background pixels) within a window; 2) update the window adoptively and delete those varied pixels. Extensive results show that the time complexity of the method is proportional to the number of the specific pixels. Jiali Cui, Yunhong Wang 0001, Tieniu Tan, Zhenan Sun |
ICIG | 3 |
| 2004 | Detection of circular oil tanks based on the fusion of SAR and optical imagesabstractThis paper aims to detect circular oil tanks automatically by using both SAR and optical images. The proposed method consists of two main steps: first, the bright areas are extracted from the SAR image to obtain the candidate regions for oil tanks; second, shape detection is performed in the optical image on these corresponding candidate areas, before we finally determine whether they are oil tanks or not. Prior knowledge, such as the validate dimensions of oil tanks and the group emerging property, is taken into account in this method. The experimental results demonstrate the effectiveness of the proposed method. Tieniu Tan, Xianqing Tai |
ICIG | 3 |
| 2004 | Reliable detection of BPCS-steganography in natural imagesabstractImage steganalysis has attracted increasing attention recently. BPCS steganalysis is currently a hot research topic. Although some BPCS steganalysis methods have been proposed, these methods have their limitations. In this paper, we propose a new approach to detect BPCS-steganography. The approach, based on statistical features, can detect the existence of secret message not only in spatial domain, but also in transform domain. To evaluate the proposed steganalytic framework, two examples of detection are performed. It is shown that the accuracy of detection of the existence of secret message in images is relatively high. Experimental results and theoretical analysis show that the proposed method is an effective steganlytic method of BPCS-steganography. Xiaoyi Yu, Tieniu Tan, Yunhong Wang 0001 |
ICIG | 2 |
| 2004 | Gait analysis for human identification in frequency domainabstractIn this paper, we analyze the spatio-temporal human characteristic of moving silhouettes in frequency domain, and find key Fourier descriptors that have better discriminatory capability for recognition than the other Fourier descriptors. A large number of experimental results and analysis show that the proposed algorithm based on the key Fourier descriptors can not only greatly reduce the gait data dimensionality, but also lighten the computation cost, with a satisfactory CCR. Besides that, classification performance can be further improved using feature fusion. Shiqi Yu 0001, Liang Wang 0001, Weiming Hu 0004, Tieniu Tan |
ICIG | 4 |
| 2004 | Multi-camera correspondence based on principal axis of human bodyabstractMulticamera correspondence of moving people is a relatively new issue in computer vision. To cope with it, we propose a simple but effective method based on the principal axis of human body. We apply the method to real video sequences in outdoor environments. The experimental results have demonstrated the efficiency of the proposed method. Jianguang Lou, Weiming Hu 0004, Tieniu Tan |
ICIP | 4 |
| 2004 | Noise removal and impainting model for IRIS imageabstractNoise removal is an important problem for iris recognition. If the iris regions were not correctly segmented in iris images, segmented iris regions possibly include noises, namely eyelashes, eyelids, reflections and pupil. Noises influence the features of both noise regions and their neighboring regions, which will result in poor recognition performance. To solve this problem, this paper proposes a method for removing noises and impainting iris images. The whole procedure includes three steps: 1) localization and normalization, 2) noise removal based on phase congruency and 3) iris image impainting. A series of experiments show that the proposed method has encouraging performance for improving the recognition accuracy. Junzhou Huang, Yunhong Wang 0001, Jiali Cui, Tieniu Tan |
ICIP | 4 |
| 2004 | Cascading statistical and structural classifiers for iris recognitionabstractReliable human identification using iris pattern has recently gained growing interests from pattern recognition researchers. In literature of iris recognition, almost all algorithms are based on statistical information. In this paper, a structural iris image analysis method is proposed, which provides complementary information to statistical classifier. In order to save computational cost, the structural matcher is not consulted unless the statistical classifier is uncertain of its decision. At the second stage, the structural classifier may be combined with statistical classifier with different fusion strategies. The experimental results of decision-level classifiers combination are reported, which demonstrate that the cascaded classification system significantly outperforms single classifier. Zhenan Sun, Yunhong Wang 0001, Tieniu Tan, Jiali Cui |
ICIP | 3 |
| 2004 | Semantic-based traffic video retrieval using activity pattern analysisabstractA semantic based retrieval framework for traffic video sequences is proposed. In order to estimate the low-level motion data, a cluster tracking algorithm is developed. A novel hierarchical self-organizing map is applied to learn the activity patterns. By using activity pattern analysis and semantic concepts assignment, a set of activity models is generated, which is used as the indexing key for accessing video clips and individual vehicles in the semantic level. The proposed retrieval framework supports various queries including query by keywords, query by sketch and multiple object queries. Weiming Hu 0004, Tieniu Tan, Junyi Peng |
ICIP | 3 |
| 2004 | Robust nose detection in 3d facial data using local characteristics
Chenghua Xu, Yunhong Wang 0001, Tieniu Tan, Long Quan |
ICIP | 3 |
| 2004 | Adaptive skin detection using multiple cuesabstractThis paper presents an adaptive approach to skin detection. First, we propose a nonlinear relationship among R, G and B components and use a closed curve to identify the skin cluster region. Then, a split machine is designed that aids the extraction of the pixels with similar low-level features from images. Finally, a nonlinear skin color classifier with an adaptive threshold is developed by analyzing the properties of the extracted pixels in the HSL, YCbCr, YUV and YIQ color spaces. Experimental results show that our proposed method works very well in skin detection. Jinfeng Yang, Zhouyu Fu, Tieniu Tan, Weiming Hu 0004 |
ICIP | 3 |
| 2004 | Model based steganalysisabstractIn this paper, we consider a new method for performing steganalysis using a statistical model of the cover medium. Using model based methodology, examples of detecting secret message and estimating the secret message length of bit-streams embedded using JSteg-like steganography and quantization index modulation are proposed. This steganalysis technique is based on the model of statistical distribution of quantized DCT coefficients. The histogram of cover image and "shrinkage histogram" are estimated from stego image using the statistical model. Then the secret message is detected and the secret message length is estimated. The methodology described in this paper is a framework which can also be applied to virtually any type of media such as JPEG2000 file format embedding. Xiaoyi Yu, Yunhong Wang 0001, Tieniu Tan |
ICIP | 3 |
| 2004 | Emotional Chinese talking head systemabstractNatural Human-Computer Interface requires integration of realistic audio and visual information for perception and display. In this paper, a lifelike talking head system is proposed. The system converts text to speech with synchronized animation of mouth movements and emotion expression. The talking head is based on a generic 3D human head model. The personalized model is incorporated into the system. With texture mapping, the personalized model offers a more natural and realistic look than the generic model. To express emotion, both emotional speech synthesis and emotional facial animation are integrated and Chinese viseme models are also created in the paper. Finally, the emotional talking head system is created to generate the natural and vivid audio-visual output. Jianhua Tao 0001, Tieniu Tan |
ICMI | 2 |
| 2004 | Kinematics-based tracking of human walking in monocular video sequences
Huazhong Ning, Tieniu Tan, Liang Wang 0001, Weiming Hu 0004 |
Image Vis. Comput. | 2 |
| 2004 | Local intensity variation analysis for iris recognition
Li Ma 0001, Tieniu Tan, Yunhong Wang 0001 |
Pattern Recognit. | 2 |
| 2004 | People tracking based on motion model and motion constraints with automatic initialization
Huazhong Ning, Tieniu Tan, Liang Wang 0001, Weiming Hu 0004 |
Pattern Recognit. | 2 |
| 2004 | Fusion of static and dynamic body biometrics for gait recognitionabstractVision-based human identification at a distance has recently gained growing interest from computer vision researchers. This paper describes a human recognition algorithm by combining static and dynamic body biometrics. For each sequence involving a walker, temporal pose changes of the segmented moving silhouettes are represented as an associated sequence of complex vector configurations and are then analyzed using the Procrustes shape analysis method to obtain a compact appearance representation, called static information of body. In addition, a model-based approach is presented under a Condensation framework to track the walker and to further recover joint-angle trajectories of lower limbs, called dynamic information of gait. Both static and dynamic cues obtained from walking video may be independently used for recognition using the nearest exemplar classifier. They are fused on the decision level using different combinations of rules to improve the performance of both identification and verification. Experimental results of a dataset including 20 subjects demonstrate the feasibility of the proposed algorithm. Liang Wang 0001, Huazhong Ning, Tieniu Tan, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | Efficient iris recognition by characterizing key local variationsabstractUnlike other biometrics such as fingerprints and face, the distinct aspect of iris comes from randomly distributed features. This leads to its high reliability for personal identification, and at the same time, the difficulty in effectively representing such details in an image. This paper describes an efficient algorithm for iris recognition by characterizing key local variations. The basic idea is that local sharp variation points, denoting the appearing or vanishing of an important image structure, are utilized to represent the characteristics of the iris. The whole procedure of feature extraction includes two steps: 1) a set of one-dimensional intensity signals is constructed to effectively characterize the most important information of the original two-dimensional image; 2) using a particular class of wavelets, a position sequence of local sharp variation points in such signals is recorded as features. We also present a fast matching scheme based on exclusive OR operation to compute the similarity between a pair of position sequences. Experimental results on 2255 iris images show that the performance of the proposed method is encouraging and comparable to the best iris recognition algorithm found in the current literature. Li Ma 0001, Tieniu Tan, Yunhong Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2004 | A hierarchical self-organizing approach for learning the patterns of motion trajectoriesabstractThe understanding and description of object behaviors is a hot topic in computer vision. Trajectory analysis is one of the basic problems in behavior understanding, and the learning of trajectory patterns that can be used to detect anomalies and predict object trajectories is an interesting and important problem in trajectory analysis. In this paper, we present a hierarchical self-organizing neural network model and its application to the learning of trajectory distribution patterns for event recognition. The distribution patterns of trajectories are learnt using a hierarchical self-organizing neural network. Using the learned patterns, we consider anomaly detection as well as object behavior prediction. Compared with the existing neural network structures that are used to learn patterns of trajectories, our network structure has smaller scale and faster learning speed, and is thus more effective. Experimental results using two different sets of data demonstrate the accuracy and speed of our hierarchical self-organizing neural network in learning the distribution patterns of object trajectories. Weiming Hu 0004, Tieniu Tan |
IEEE Trans. Neural Networks | 3 |
| 2004 | A survey on visual surveillance of object motion and behaviorsabstractVisual surveillance in dynamic scenes, especially for humans and vehicles, is currently one of the most active research topics in computer vision. It has a wide spectrum of promising applications, including access control in special areas, human identification at a distance, crowd flux statistics and congestion analysis, detection of anomalous behaviors, and interactive surveillance using multiple cameras, etc. In general, the processing framework of visual surveillance in dynamic scenes includes the following stages: modeling of environments, detection of motion, classification of moving objects, tracking, understanding and description of behaviors, human identification, and fusion of data from multiple cameras. We review recent developments and general strategies of all these stages. Finally, we analyze possible research directions, e.g., occlusion handling, a combination of twoand three-dimensional tracking, a combination of motion analysis and biometrics, anomaly detection and behavior prediction, content-based retrieval of surveillance videos, behavior understanding and natural language description, fusion of information from multiple sensors, and remote surveillance. Weiming Hu 0004, Tieniu Tan, Liang Wang 0001, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2004 | Learning activity patterns using fuzzy self-organizing neural networkabstractActivity understanding in visual surveillance has attracted much attention in recent years. In this paper, we present a new method for learning patterns of object activities in image sequences for anomaly detection and activity prediction. The activity patterns are constructed using unsupervised learning of motion trajectories and object features. Based on the learned activity patterns, anomaly detection and activity prediction can be achieved. Unlike existing neural network based methods, our method uses a whole trajectory as an input to the network. This makes the network structure much simpler. Furthermore, the fuzzy set theory based method and the batch learning method are introduced into the network learning process, and make the learning process much more efficient. Two sets of data acquired, respectively, from a model scene and a campus scene are both used to test the proposed algorithms. Experimental results show that the fuzzy self-organizing neural network (fuzzy SOM) is much more efficient than the Kohonen self-organizing feature map (SOFM) and vector quantization in both speed and accuracy, and the anomaly detection and activity prediction algorithms have encouraging performances. Weiming Hu 0004, Tieniu Tan, Stephen J. Maybank |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2003 | Learning Based Resolution Enhancement of Iris ImagesabstractIris recognition is one of the most reliable personal identification methods. The potential requirement of obtaining high accuracy is that users supply iris images with good quality. It is thus necessary for an iris recognition system to operate the possibly blurred iris images due to less cooperation of users and camera with low resolution. This paper proposes a new algorithm for resolution enhancement of iris images captured by the low resolution camera in less cooperative situations. The prior probability relation between the information of different frequency bands of iris features useful for recognition is firstly learned. Then, it is incorporated into resolution enhancement algorithms to recover the lost information for the seriously blurred images. A large number of experiments on the CASIA iris database demonstrate the validity of the proposed approach. Junzhou Huang, Li Ma 0001, Tieniu Tan, Yunhong Wang 0001 |
BMVC | 3 |