EDBT 2026 Demo / reviewers in the wild / expert
Deepu Rajan
dblp:95/3115
· DBLP profile ↗
112ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0001-7788-8368ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 88 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 42 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task KnowledgeabstractLarge-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowledge of the VFMs to launch potent attacks. This paper investigates a novel and practical adversarial threat scenario: attacking downstream models or MLLMs fine-tuned from open-source VFMs, without requiring access to the victim task, training data, model query, and architecture. In contrast to conventional transfer-based attacks that rely on task-aligned surrogate models, we demonstrate that adversarial vulnerabilities can be exploited directly from the VFMs. To this end, we propose the Transferable Video Attack (TVA), a temporal-aware adversarial attack method that leverages the temporal representation dynamics of VFMs to craft effective perturbations. TVA integrates a bidirectional contrastive learning mechanism to maximize the discrepancy between the clean and adversarial features, and introduces a temporal consistency loss that exploits motion cues to enhance the sequential impact of perturbations. TVA avoids the need to train expensive surrogate models or access to domain-specific data, thereby offering a more practical and efficient attack strategy. Extensive experiments across 24 video-related tasks demonstrate the efficacy of TVA against downstream models and MLLMs, revealing a previously underexplored security vulnerability in the deployment of video models. Yi Yu 0011, Song Xia, Deepu Rajan, Boon Poh Ng, Alex Chichung Kot, Xudong Jiang 0001 |
AAAI | 5 |
| 2026 | A-V Representation Learning via Audio Shift Prediction for Multimodal Deepfake Detection and Temporal LocalizationabstractRecent multimodal deepfake detection methods typically rely on single-stage training, which can cause the model to focus on dataset-specific multimodal cues while missing important features that are helpful to detect unseen manipulations, thereby limiting generalization. While some approaches attempt to address this using self-supervised audio-visual pretraining, they may not fully exploit cross-modal temporal information. Also, they often assume that manipulations affect the entire video, ignoring more realistic cases where only short segments are altered. To overcome these limitations, we propose a two-stage training framework that first learns audio-visual temporal alignment in real videos and then uses this information to detect and localize potential deepfakes by identifying temporal inconsistencies. We propose a self-supervised shift-prediction pretraining objective to fully understand cross-modal temporal alignment across multiple temporal shifts applied to the audio input. The pretrained features enable the model to identify manipulations across entire videos as well as accurately localize deepfake segments in partially tampered content. Moreover, the pretrained components do not require task-specific fine-tuning, improving the model’s flexibility for both classification and localization. Experiments on benchmark datasets demonstrate strong within-dataset performance, robust generalization to new manipulations and datasets, and accurate temporal localization.1 Ashutosh Anshul, Chng Eng Siong, Deepu Rajan |
WACV | 3 |
| 2025 | Chebyshev Attention Depth Permutation Texture Network with Latent Texture Attribute LossabstractDespite recent advances in deep texture recognition, existing methods still lack representational diversity and struggle to capture and preserve discriminative cues across stages of representation hierarchies. Moreover, many rely on loss formulations that prioritize recognition accuracy while overlooking spatial coherence and statistical consistency in the feature space. To address these issues, we propose three key innovations: (1) Stochastic Local Texture Masking (SLTM), a regularization strategy that randomly occludes small texture patches to promote the learning of broader spatial and contextual dependencies; (2) the Chebyshev Attention Depth Permutation Texture Network (CAPTN), a novel architecture that learns expressive and persistent Latent Texture Attribute (LTA) representations. CAPTN integrates a Texture Frequency Attention (TFA) module that generates Latent Texture Attributes (LTAs) and enables frequency-aware interpretability, a Dual Depth Permutation (D2P) module to expose complementary channel adjacency patterns, and Learnable Chebyshev Polynomials (LCPs) to model high-order orderless LTA transformations via recursive Chebyshev basis expansion; and (3) a Latent Texture Attribute Loss that jointly optimizes classification accuracy, statistical alignment, and spatial fidelity. CAPTN supports end-to-end training without relying on fine-tuned CNN backbones and achieves state-of-the-art performance on several texture and material recognition benchmarks. (Code: https://github.com/RavishankarEvani/CAPTN) Ravishankar Evani, Deepu Rajan, Shangbo Mao |
CVPR | 2 |
| 2025 | Intra-Modal and Cross-Modal Synchronization for Audio-Visual Deepfake Detection and Temporal Localization
Ashutosh Anshul, Shreyas Gopal, Deepu Rajan, Chng Eng Siong |
ICCV | 3 |
| 2025 | Imore: Implicit Program-Guided Reasoning for Human Motion QA
Chinthani Sugandhika, Ee Yeo Keat, Eric P. Xing, Deepu Rajan, Basura Fernando |
ICCV | 7 |
| 2025 | Action Sequence Augmentation for Action AnticipationabstractAction anticipation models require an understanding of temporal action patterns and dependencies to predict future actions from previous events. The key challenges arise from the vast number of possible action sequences, given the flexibility in action ordering and the interleaving of multiple goals. Since only a subset of such action sequences are present in action anticipation datasets, there is an inherent ordering bias in them. Another challenge is the presence of noisy input to the models due to erroneous action recognition or other upstream tasks. This paper addresses these challenges by introducing a novel data augmentation strategy that separately augments observed action sequences and next actions. To address biased action ordering, we introduce a grammar induction algorithm that derives a powerful context-free grammar from action sequence data. We also develop an efficient parser to generate plausible next-action candidates beyond the ground truth. For noisy input, we enhance model robustness by randomly deleting or replacing actions in observed sequences. Our experiments on the 50Salads, EGTEA Gaze+, and Epic-Kitchens-100 datasets demonstrate significant performance improvements over existing state-of-the-art methods. Yihui Qiu, Deepu Rajan |
ICLR | 2 |
| 2025 | Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across LanguagesabstractMultilingual speech emotion recognition aims to estimate a speaker’s emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses significant challenges for zero-shot speech emotion recognition, especially with multilingual datasets. In this paper, we propose leveraging contrastive learning to refine multilingual speech features and extend large language models for zero-shot multilingual speech emotion estimation. Specifically, we employ a novel two-stage training framework to align speech signals with linguistic features in the emotional space, capturing both emotion-aware and language-agnostic speech representations. To advance research in this field, we introduce a large-scale synthetic multilingual speech emotion dataset, M5SER. Our experiments demonstrate the effectiveness of the proposed method in both speech emotion recognition and zero-shot multilingual speech emotion recognition, including previously unseen datasets and languages. Our introduced dataset and related code will be available on GitHub1. Heqing Zou, Fengmao Lv, Desheng Zheng, Chng Eng Siong, Deepu Rajan |
ICME | 5 |
| 2025 | An Encoder-Agnostic Weakly Supervised Method For Describing TexturesabstractRecent advances in Large Language Models (LLMs) have enabled the semantic description of textures in natural language, aiming to capture them in richer detail. However, most methods are confined to either depending on supervised training with pairs of images and manually annotated visual attributes that most texture datasets lack or using Vision-Language Models (VLMs) such as CLIP. In this paper, we develop an encoder-agnostic Weakly supervised Texture Description Generator (WTDG) that employs a novel Scaled Ranked Kullback-Leibler divergence (SR-KL) loss between image and text modalities. Within the SR-KL loss formulation, we leverage category information, which is always available as ground-truths for all benchmark texture recognition datasets. We further extend our proposed WTDG to assist in texture recognition by using its generated texture descriptions. Thus, we develop a multimodal framework, called$T e x^2$, which is adept at simultaneous generation of texture description and recognition. Our approach exhibits promising performance in describing and recognizing textures on benchmark datasets. Shangbo Mao, Deepu Rajan |
WACV | 2 |
| 2025 | Situational Scene Graph for Structured Human-Centric Situation Understanding
Chinthani Sugandhika, Deepu Rajan, Basura Fernando |
WACV | 3 |
| 2024 | Multiscale Graph Texture Network
Ravishankar Evani, Deepu Rajan, Shangbo Mao |
ECCV (32) | 2 |
| 2024 | Cross-Modality and Within-Modality Regularization for Audio-Visual Deepfake DetectionabstractAudio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal targets face challenges due to uncertainties and inconsistencies in learned representations caused by independent modality manipulations in deepfake videos. To address this, we propose cross-modality and within-modality regularization to preserve modality distinctions during multimodal representation learning. Our approach includes an audio-visual transformer module for modality correspondence and a cross-modality regularization module to align paired audio-visual signals, preserving modality distinctions. Simultaneously, a within-modality regularization module refines unimodal representations with modality-specific targets to retain modal-specific details. Experimental results on the public audio-visual dataset, FakeAVCeleb, demonstrate the effectiveness and competitiveness of our approach. Heqing Zou, Meng Shen 0002, Chen Chen 0075, Chng Eng Siong, Deepu Rajan |
ICASSP | 6 |
| 2024 | Improving Temporal Action Segmentation and Detection with Hierarchical Task Grammar
Yihui Qiu, Deepu Rajan |
ICPR (29) | 2 |
| 2024 | Enhancing Modality Representation and Alignment for Multimodal Cold-start Active LearningabstractTraining multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on sufficient labeled data to train a well-calibrated model that can assess the uncertainty and diversity of unlabeled data. However, when assembling a dataset, labeled data are often scarce initially, leading to a cold-start problem. Additionally, most AL methods seldom address multimodal data, highlighting a research gap in this field. Our research addresses these issues by developing a two-stage method for Multi-Modal Cold-Start Active Learning (MMCSAL). Firstly, we observe the modality gap, a significant distance between the centroids of representations from different modalities, when only using cross-modal pairing information as self-supervision signals. This modality gap affects data selection process, as we calculate both uni-modal and cross-modal distances. To address this, we introduce uni-modal prototypes to bridge the modality gap. Secondly, conventional AL methods often falter in multimodal scenarios where alignment between modalities is overlooked. Therefore, we propose enhancing cross-modal alignment through regularization, thereby improving the quality of selected multimodal data pairs in AL. Finally, our experiments demonstrate MMCSAL's efficacy in selecting multimodal data pairs across three multimodal datasets. Meng Shen 0002, Yake Wei, Jianxiong Yin, Deepu Rajan, Di Hu 0001, Simon See |
MMAsia | 4 |
| 2024 | A multivariate Markov chain model for interpretable dense action anticipation
Yihui Qiu, Deepu Rajan |
Neurocomputing | 2 |
| 2024 | A Unified Framework for Guiding Generative AI With Wireless Perception in Resource Constrained Mobile Edge NetworksabstractWith the significant advancements in artificial intelligence (AI) technologies and computational capabilities, generative AI (GAI) has become a pivotal digital content generation technique for offering superior digital services. However, due to the inherent instability of AI models, directing GAI towards the desired output remains a challenging task. Therefore, in this paper, we design a novel framework that utilizeswirelessperception to guideGAI(WiPe-GAI) in delivering AI-generated content (AIGC) service, within resource-constrained mobile edge networks. Specifically, we first propose a new sequential multi-scale perception (SMSP) algorithm to predict user skeleton based on the channel state information (CSI) extracted from wireless signals. This prediction then guides GAI to provide users with AIGC, i.e., virtual character generation. To ensure the efficient operation of the proposed framework in resource constrained networks, we further design a pricing-based incentive mechanism and propose a diffusion model based approach to generate an optimal pricing strategy for the service provisioning. The strategy maximizes the user's utility while incentivizing the participation of the virtual service provider (VSP) in AIGC provision. The experimental results demonstrate the effectiveness of the designed framework in terms of skeleton prediction and optimal pricing strategy generation, outperforming other existing solutions. Jiacheng Wang 0001, Hongyang Du 0001, Dusit Niyato, Jiawen Kang 0001, Zehui Xiong, Deepu Rajan, Shiwen Mao, Xuemin Shen |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Controllable Video Generation With Text-Based InstructionsabstractMost of the existing studies on controllable video generation either transfer disentangled motion to an appearance without detailed control over motion or generate videos of simple actions such as the movement of arbitrary objects conditioned on a control signal from users. In this study, we introduce Controllable Video Generation with text-based Instructions (CVGI) framework that allows text-based control over action performed on a video. CVGI generates videos where hands interact with objects to perform the desired action by generating hand motions with detailed control through text-based instruction from users. By incorporating the motion estimation layer, we divide the task into two sub-tasks: (1) control signal estimation and (2) action generation. In control signal estimation, an encoder models actions as a set of simple motions by estimating low-level control signals for text-based instructions with given initial frames. In action generation, generative adversarial networks (GANs) generate realistic hand-based action videos as a combination of hand motions conditioned on the estimated low control level signal. Evaluations on several datasets (EPIC-Kitchens-55, BAIR robot pushing, and Atari Breakout) show the effectiveness of CVGI in generating realistic videos and in the control over actions. Ali Koksal, Kenan E. Ak, Ying Sun 0001, Deepu Rajan, Joo-Hwee Lim |
IEEE Trans. Multim. | 4 |
| 2023 | Towards Balanced Active Learning for Multimodal ClassificationabstractTraining multimodal networks requires a vast amount of data due to their larger parameter space compared to unimodal networks. Active learning is a widely used technique for reducing data annotation costs by selecting only those samples that could contribute to improving model performance. However, current active learning strategies are mostly designed for unimodal tasks, and when applied to multimodal data, they often result in biased sample selection from the dominant modality. This unfairness hinders balanced multimodal learning, which is crucial for achieving optimal performance. To address this issue, we propose three guidelines for designing a more balanced multimodal active learning strategy. Following these guidelines, a novel approach is proposed to achieve more fair data selection by modulating the gradient embedding with the dominance degree among modalities. Our studies demonstrate that the proposed method achieves more balanced multimodal learning by avoiding greedy sample selection from the dominant modality. Our approach outperforms existing active learning strategies on a variety of multimodal classification tasks. Overall, our work highlights the importance of balancing sample selection in multimodal active learning and provides a practical solution for achieving more balanced active learning for multimodal classification. Meng Shen 0002, Yizheng Huang 0001, Jianxiong Yin, Heqing Zou, Deepu Rajan, Simon See |
ACM Multimedia | 5 |
| 2022 | Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic InformationabstractSpeech Emotion Recognition (SER) aims to help the machine to understand human’s subjective emotion from only audio in-formation. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task. In this paper, we propose an end-to-end speech emotion recognition system using multi-level acoustic information with a newly designed co-attention module. We firstly extract multi-level acoustic information, including MFCC, spectrogram, and the embedded high-level acoustic information with CNN, BiL-STM and wav2vec2, respectively. Then these extracted features are treated as multimodal inputs and fused by the pro-posed co-attention mechanism. Experiments are carried on the IEMOCAP dataset, and our model achieves competitive performance with two different speaker-independent cross-validation strategies. Our code is available on GitHub. Heqing Zou, Yuke Si, Chen Chen 0075, Deepu Rajan, Chng Eng Siong |
ICASSP | 4 |
| 2022 | Smart interpretable model (SIM) enabling subject matter experts in rule generation
Hotman Christianto, Gary Kee Khoon Lee, Weigui Jair Zhou, Henry Kasim, Deepu Rajan |
Expert Syst. Appl. | 5 |
| 2021 | An interpretable Neural Fuzzy Hammerstein-Wiener network for stock price prediction
Xie Chen 0002, Deepu Rajan, Hiok Chai Quek |
Inf. Sci. | 2 |
| 2021 | Deep residual pooling network for texture recognition
Shangbo Mao, Deepu Rajan, Liang-Tien Chia |
Pattern Recognit. | 2 |
| 2020 | An image similarity descriptor for classification tasks
Deepu Rajan |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | Saliency-based bit plane detection for network applications
Maryam Asadzadeh Kaljahi, Palaiahnakote Shivakumara, Saqib Hakak, Mohd Yamani Idna Bin Idris, Mohammad Hossein Anisi, Deepu Rajan |
Multim. Tools Appl. | 6 |
| 2020 | Are Object Detection Assessment Criteria Ready for Maritime Computer Vision?abstractMaritime vessels equipped with visible and infrared cameras can complement other conventional sensors for object detection. However, application of computer vision techniques in maritime domain received attention only recently. The maritime environment offers its own unique requirements and challenges. Assessment of the quality of detections is a fundamental need in computer vision. However, the conventional assessment metrics suitable for usual object detection are deficient in the maritime setting. Thus, a large body of related work in computer vision appears inapplicable to the maritime setting at the first sight. We discuss the problem of defining assessment metrics suitable for maritime computer vision. We consider new bottom edge proximity metrics as assessment metrics for maritime computer vision. These metrics indicate that existing computer vision approaches are indeed promising for maritime computer vision and can play a foundational role in the emerging field of maritime computer vision. Dilip K. Prasad, Huixu Dong, Deepu Rajan, Hiok Chai Quek |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Comparative Convolutional Neural Network for Younger Face IdentificationabstractWe consider the problem of determining whether a pair of face images can be distinguishable in terms of age and if so, which is the younger of the two. We also determine the degree of distinguishability in which age differences are categorized into large, medium, small and tiny. We propose a comparative convolutional neural network combining two parallel deep architectures. Based on the two deep learnt face features, we introduce a comparative layer to represent their mutual relationships, followed by a concatenatation implementation. Softmax is adopted to complete the classification task. To demonstrate our approach, we construct a very large dataset consisting of over 1.7 million face image pairs with young/old labels. Deepu Rajan |
VCIP | 2 |
| 2019 | CamType: assistive text entry using gaze with an off-the-shelf webcam
Yi Liu 0040, Bu-Sung Lee, Deepu Rajan, Andrzej Stefan Sluzek, Martin J. McKeown |
Mach. Vis. Appl. | 3 |
| 2019 | CNN-based gender classification in near-infrared periocular images
Anirudh Manyala, Hisham Cholakkal, Vijay Anand, Vivek Kanhangad, Deepu Rajan |
Pattern Anal. Appl. | 5 |
| 2019 | Object Detection in a Maritime Environment: Performance Evaluation of Background Subtraction MethodsabstractThis paper provides a benchmark of the performance of 23 classical and state-of-the-art background subtraction (BS) algorithms on visible range and near infrared range videos in the Singapore Maritime dataset. Importantly, our study indicates the limitations of the conventional performance evaluation criteria for maritime vision and proposes new performance evaluation criteria that is better suited to this problem. This paper provides insight into the specific challenges of BS in maritime vision. We identify four open challenges that plague BS methods in maritime scenario. These include spurious dynamics of water, wakes, ghost effect, and multiple detections. Poor recall and extremely poor precision of all the 23 methods, which have been otherwise successful for other challenging BS situations, allude to the need for new BS methods custom designed for maritime vision. Dilip K. Prasad, Chandrashekar Krishna Prasath, Deepu Rajan, Lily Rachmawati, Eshan Rajabally, Hiok Chai Quek |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Locality and context-aware top-down saliencyabstractIn this study, the authors propose a novel framework for top‐down (TD) saliency detection, which is well suited to locate category‐specific objects in natural images. Saliency value is defined as the probability of a target based on its visual feature. They introduce an effective coding strategy called locality constrained contextual coding (LCCC) that enforces locality and contextual constraints. Furthermore, a contextual pooling operation is presented to take advantages of feature contextual information. Benefiting from LCCC and contextual pooling, the obtained feature representation has high discriminative power, which makes the authors' saliency detection method achieving competitive results with existing saliency detection algorithms. They also include bottom‐up cues into their framework to supplement the proposed TD saliency algorithm. Experimental results on three datasets (Graz‐02, Weizmann Horse and PASCAL VOC 2007) show that the proposed framework outperforms state‐of‐the‐art methods in terms of visual quality and accuracy. Junxia Li, Deepu Rajan, Jian Yang 0003 |
IET Image Process. | 2 |
| 2018 | Backtracking Spatial Pyramid Pooling-Based Image Classifier for Weakly Supervised Top-Down Salient Object DetectionabstractTop-down saliency models produce a probability map that peaks at target locations specified by a task/goal such as object detection. They are usually trained in a fully supervised setting involving pixel-level annotations of objects. We propose a weakly supervised top-down saliency framework using only binary labels that indicate the presence/absence of an object in an image. First, the probabilistic contribution of each image region to the confidence of a CNN-based image classifier is computed through a backtracking strategy to produce top-down saliency. From a set of saliency maps of an image produced by fast bottom-up saliency approaches, we select the best saliency map suitable for the top-down task. The selected bottom-up saliency map is combined with the top-down saliency map. Features having high combined saliency are used to train a linear SVM classifier to estimate feature saliency. This is integrated with combined saliency and further refined through a multi-scale superpixel-averaging of saliency map. We evaluate the performance of the proposed weakly supervised topdown saliency and achieve comparable performance with fully supervised approaches. Experiments are carried out on seven challenging datasets and quantitative results are compared with 40 closely related approaches across 4 different applications. Code will be made publicly available. Hisham Cholakkal, Jubin Johnson, Deepu Rajan |
IEEE Trans. Image Process. | 3 |
| 2017 | L1-Regularized Reconstruction Error as Alpha MatteabstractSampling-based alpha matting methods have traditionally followed the compositing equation to estimate the alpha value at a pixel from a pair of foreground (F) and background (B) samples. The (F,B) pair that produces the least reconstruction error is selected, followed by alpha estimation. The significance of that residual error has been left unexamined. In this letter, we propose a video matting algorithm that uses L1-regularized reconstruction error of F and B samples as a measure of the alpha matte. A multiframe nonlocal means framework using coherency sensitive hashing is utilized to ensure temporal coherency in the video mattes. Qualitative and quantitative evaluations on a dataset exclusively for video matting demonstrate the effectiveness of the proposed matting algorithm. Jubin Johnson, Hisham Cholakkal, Deepu Rajan |
IEEE Signal Process. Lett. | 3 |
| 2017 | Video Processing From Electro-Optical Sensors for Object Detection and Tracking in a Maritime Environment: A SurveyabstractWe present a survey on maritime object detection and tracking approaches, which are essential for the development of a navigational system for autonomous ships. The electro-optical (EO) sensor considered here is a video camera that operates in the visible or the infrared spectra, which conventionally complements radar and sonar for situational awareness at sea and has demonstrated its effectiveness over the last few years. This paper provides a comprehensive overview of various approaches of video processing for object detection and tracking in the maritime environment. We follow an approach-based taxonomy wherein the advantages and limitations of each approach are compared. The object detection system consists of the following modules: horizon detection, static background subtraction, and foreground segmentation. Each of these has been studied extensively in maritime situations and has been shown to be challenging due to the presence of background motion especially due to waves and wakes. The key processes involved in object tracking include video frame registration, dynamic background subtraction, and the object tracking algorithm itself. The challenges for robust tracking arise due to camera motion, dynamic background, and low contrast of tracked object, possibly due to environmental degradation. The survey also discusses multisensor approaches and commercial maritime systems that use EO sensors. The survey also highlights methods from computer vision research, which hold promise to perform well in maritime EO data processing. Performance of several maritime and computer vision techniques is evaluated on Singapore Maritime Dataset. Dilip K. Prasad, Deepu Rajan, Lily Rachmawati, Eshan Rajabally, Hiok Chai Quek |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | Backtracking ScSPM Image Classifier for Weakly Supervised Top-Down SaliencyabstractTop-down saliency models produce a probability map that peaks at target locations specified by a task/goal such as object detection. They are usually trained in a supervised setting involving annotations of objects. We propose a weakly supervised top-down saliency framework using only binary labels that indicate the presence/absence of an object in an image. First, the probabilistic contribution of each image patch to the confidence of an ScSPM-based classifier produces a Reverse-ScSPM (R-ScSPM) saliency map. Neighborhood information is then incorporated through a contextual saliency map which is estimated using logistic regression learnt on patches having high R-ScSPM saliency. Both the saliency maps are combined to obtain the final saliency map. We evaluate the performance of the proposed weakly supervised top-down saliency and achieves comparable performance with fully supervised approaches. Experiments are carried out on 5 challenging datasets across 3 different applications. Hisham Cholakkal, Jubin Johnson, Deepu Rajan |
CVPR | 3 |
| 2016 | Salient object detection in tracking shotsabstractTracking shots have posed a significant challenge for salient region detection due to the presence of highly competing background motion. In this paper, we propose a computationally efficient technique to detect salient objects in a tracking shot. We first separate the tracked foreground pixels from the background by accounting for the variability of the pixels in a set of frames. The focus of the tracked foreground pixels is utilized as a measure of saliency of objects in the scene. We evaluate the performance of this method by comparing the salient region detection with ground truth data of the location of the salient object that are manually generated. The results of the evaluation show that the proposed method is able to achieve superior salient object detection performance with very low computational load. Karthik Muthuswamy, Deepu Rajan |
ICPR | 2 |
| 2016 | Efficient 2D viewpoint combination for human action recognition
Behrouz Saghafi, Deepu Rajan, Wanqing Li 0001 |
Pattern Anal. Appl. | 2 |
| 2016 | A classifier-guided approach for top-down salient object detection
Hisham Cholakkal, Jubin Johnson, Deepu Rajan |
Signal Process. Image Commun. | 3 |
| 2016 | Sparse Coding for Alpha MattingabstractExisting color sampling-based alpha matting methods use the compositing equation to estimate alpha at a pixel from the pairs of foreground ( F ) and background ( B ) samples. The quality of the matte depends on the selected ( F,B ) pairs. In this paper, the matting problem is reinterpreted as a sparse coding of pixel features, wherein the sum of the codes gives the estimate of the alpha matte from a set of unpaired F and B samples. A non-parametric probabilistic segmentation provides a certainty measure on the pixel belonging to foreground or background, based on which a dictionary is formed for use in sparse coding. By removing the restriction to conform to ( F,B ) pairs, this method allows for better alpha estimation from multiple F and B samples. The same framework is extended to videos, where the requirement of temporal coherence is handled effectively. Here, the dictionary is formed by samples from multiple frames. A multi-frame graph model, as opposed to a single image as for image matting, is proposed that can be solved efficiently in closed form. Quantitative and qualitative evaluations on a benchmark dataset are provided to show that the proposed method outperforms the current stateoftheart in image and video matting. Jubin Johnson, Ehsan Shahrian, Hisham Cholakkal, Deepu Rajan |
IEEE Trans. Image Process. | 4 |
| 2016 | Double Low Rank Matrix Recovery for Saliency FusionabstractIn this paper, we address the problem of fusing various saliency detection methods such that the fusion result outperforms each of the individual methods. We observe that the saliency regions shown in different saliency maps are with high probability covering parts of the salient object. With image regions being represented by the saliency values of multiple saliency maps, the object regions have strong correlation and thus lie in a low-dimensional subspace. Meanwhile, most of background regions tend to have lower saliency values in various saliency maps. They are also strongly correlated and lie in a lowdimensional subspace that is independent of the object subspace. Therefore, an image can be represented as the combination of two low rank matrices. To obtain a unified low rank matrix that represents the salient object, this paper presents a double low rank matrix recovery model for saliency fusion. The inference process is formulated as a constrained nuclear norm minimization problem, which is convex and can be solved efficiently with the alternating direction method of multipliers (ADMM). Furthermore, to reduce the computational complexity of the proposed saliency fusion method, a saliency model selection strategy based on the sparse representation is proposed. Experiments on five datasets show that our method consistently outperforms each individual saliency detection approach and other state-of-the-art saliency fusion methods. Junxia Li, Lei Luo 0001, Fanlong Zhang, Jian Yang 0003, Deepu Rajan |
IEEE Trans. Image Process. | 5 |
| 2015 | Top-down saliency with Locality-constrained Contextual Sparse CodingabstractWe propose a sparse coding based framework for top-down salient object detection in which three locality constraints are integrated. First is the spatial or contextual lo- cality constraint in which features from adjacent regions have similar code, second is the feature-domain locality constraint in which similar features have similar code, and third is the category-domain locality constraint in which features are coded using similar atoms from each partition of the dictionary, where each partition corresponds to an object category. This faster coding strategy produces better saliency maps compared to conven- tional sparse coding. Proposed codes are max-pooled over a spatial neighborhood for saliency estimation. In spite of its simplicity, the proposed top-down saliency achieves state-of-the-art results at patch-level on two challenging datasets-Graz-02 and PASCAL VOC-07. A novel Gaussian-weighted interpolation further improves pixel-level saliency map derived from the patch-level map. Hisham Cholakkal, Deepu Rajan, Jubin Johnson |
BMVC | 2 |
| 2015 | Local feature embedding for supervised image classificationabstractLocal feature embedding considers two constraints: intra-image spatial and inter-image feature affinity in the embedding process. However, it does not work well for the image classification task when the images are with intra-class variation, background clutter, etc. In this paper, we enhance the manifold structure by adding the class label of images into the embedding process. Since class labels are used in the training, our method can be considered as supervised. Four constituents are included in our model: feature consistency, spatial consistency, intra-class compactness and inter-class separability. With the defined Hausdorff distance between two images, different classifiers are exploited for classification. Extensive experiments on seven datasets demonstrate the effectiveness of our proposed image classification model. Junxia Li, Deepu Rajan, Jian Yang 0003 |
ICIP | 2 |
| 2015 | Temporal trimap propagation using motion-assisted shape blendingabstractIn digital matting, the availability of an accurate trimap is essential for pulling an accurate matte. This requirement becomes all the more critical in video matting where temporal coherence of the trimap is an added requirement. However, the task of manually drawing a trimap for every frame is not feasible. This paper proposes an adaptive trimap propagation framework that alleviates this task by automatically generating trimaps between two key-frames. 2-D shape blending is coupled with motion flow using a novel strategy based on angle criteria to generate accurate in-between foreground contours for the entire video sequence. A robust trimap generation step ensures temporal coherence between the contours. Quantitative and qualitative comparisons on complex video sequences demonstrates the effectiveness of the proposed method. Jubin Johnson, Deepu Rajan, Hisham Cholakkal |
VCIP | 2 |
| 2015 | Particle filter framework for salient object detection in videosabstractSalient object detection in videos is challenging because of the competing motion in the background, resulting from camera tracking an object of interest, or motion of objects in the foreground. The authors present a fast method to detect salient video objects using particle filters, which are guided by spatio‐temporal saliency maps and colour feature with the ability to quickly recover from false detections. The proposed method for generating spatial and motion saliency maps is based on comparing local features with dominant features present in the frame. A region is marked salient if there is a large difference between local and dominant features. For spatial saliency, hue and saturation features are used, while for motion saliency, optical flow vectors are used as features. Experimental results on standard datasets for video segmentation and for saliency detection show superior performance over state‐of‐the‐art methods. Karthik Muthuswamy, Deepu Rajan |
IET Comput. Vis. | 2 |
| 2015 | Multi-modal fusion for associated news story retrieval
Ehsan Younessian, Deepu Rajan |
Multim. Tools Appl. | 2 |
| 2014 | Sparse codes as Alpha Matte
Jubin Johnson, Deepu Rajan, Hisham Cholakkal |
BMVC | 2 |
| 2013 | Depth really Matters: Improving Visual Salient Region Detection with DepthabstractDepth information has been shown to affect identification of visually salient regions in images. In this paper, we investigate the role of depth in saliency detection in the presence of (i) competing saliencies due to appearance, (ii) depth-induced blur and (iii) centre-bias. Having established through experiments that depth continues to be a significant contributor to saliency in the presence of these cues, we propose a 3D-saliency formulation that takes into account structural features of objects in an indoor setting to identify regions at salient depth levels. Computed 3D-saliency is used in conjunction with 2D-saliency models through non-linear regression using SVM to improve saliency maps. Experiments on benchmark datasets containing depth information show that the proposed fusion of 3D-saliency with 2D-saliency models results in an average improvement in ROC scores of about 9% over state-of-the-art 2D saliency models. The main contributions of this paper are: (i) The development of a 3D-saliency model that integrates depth and geometric features of object surfaces in indoor scenes. (ii) Fusion of appearance (RGB) saliency with depth saliency through non-linear regression using SVM. (iii) Experiments to support the hypothesis that depth improves saliency detection in the presence of blur and centre-bias. The effectiveness of the 3D-saliency model and its fusion with RGB-saliency is illustrated through experiments on two benchmark datasets that contain depth information. Current stateof-the-art saliency detection algorithms perform poorly on these datasets that depict indoor scenes due to the presence of competing saliencies in the form of color contrast. For example in Fig. 1, saliency maps of [1] is shown for different scenes, along with its human eye fixations and our proposed saliency map after fusion. It is seen from the first scene of Fig. 1, that illumination plays spoiler role in RGB-saliency map. In second scene of Fig. 1, the RGB-saliency is focused on the cap though multiple salient objects are present in the scene. Last scene at the bottom of Fig. 1, shows the limitation of the RGB-saliency when the object is similar in appearance with the background. Effect of depth on Saliency: In [4], it is shown that depth is an important cue for saliency. In this paper we go further and verify if the depth alone influences the saliency. Different scenes were captured for experimentation using Kinect sensor. Observations resulted out of these experiments are (i) Humans fixate on the objects at closer depth, in the presence of visually competing salient objects in the background, (ii) Early attention happens on the objects at closer depth, (iii) Effective fixations are high at the low contrast foreground compared to the high contrast objects in the background which are blurred, (iv) Low contrast object placed at the center of the field of view, gets more attention compared to other locations. As a result of all these observations, we develop a 3D-saliency that captures the depth information of the regions in the scene. 3D-Saliency: We adapt the region based contrast method from Cheng et al. [1] in computing contrast strengths for the segmented 3D surfaces or regions. Each segmented region is assigned a contrast score using surface normals as the feature. Structure of the surface can be described based on the distribution of normals in the region. We compute a histogram of angular distances formed by every pair of normals in the region. Every region Rk is associated with a histogram Hk. Contrast score Ck of a region Rk is computed as the sum of the dot products of its histogram with histograms of other regions in the scene. Since the depth of the region is influencing the visual attention, the contrast score is scaled by a value Zk, which is the depth of the region Rk from the sensor. In order to define the saliency, sizes of the regions i.e. the number of the points in the region, have to be considered. We find the ratio of the region dimension to the half of the scene dimension. Considering nk as the number of 3D points in the region Rk, the constrast score becomes Figure 1: Four different scenes and their saliency maps; For each scene from top left (i) Original Image, (ii) RGB-Saliency map using RC [1], (iii) Human fixations from eye-tracker and (iv) Fused RGBD-saliency map Karthik Desingh, K. Madhava Krishna, Deepu Rajan, C. V. Jawahar |
BMVC | 3 |
| 2013 | Improving Image Matting Using Comprehensive Sampling SetsabstractIn this paper, we present a new image matting algorithm that achieves state-of-the-art performance on a benchmark dataset of images. This is achieved by solving two major problems encountered by current sampling based algorithms. The first is that the range in which the foreground and background are sampled is often limited to such an extent that the true foreground and background colors are not present. Here, we describe a method by which a more comprehensive and representative set of samples is collected so as not to miss out on the true samples. This is accomplished by expanding the sampling range for pixels farther from the foreground or background boundary and ensuring that samples from each color distribution are included. The second problem is the overlap in color distributions of foreground and background regions. This causes sampling based methods to fail to pick the correct samples for foreground and background. Our design of an objective function forces those foreground and background samples to be picked that are generated from well-separated distributions. Comparison on the dataset at and evaluation by www.alphamatting.com shows that the proposed method ranks first in terms of error measures used in the website. Ehsan Shahrian, Deepu Rajan, Brian L. Price, Scott Cohen |
CVPR | 2 |
| 2013 | Human activities recognition using depth imagesabstractWe present a new method to classify human activities by leveraging on the cues available from depth images alone. Towards this end, we propose a descriptor which couples depth and spatial information of the segmented body to describe a human pose. Unique poses (i.e. codewords) are then identified by a spatial-based clustering step. Given a video sequence of depth images, we segment humans from the depth images and represent these segmented bodies as a sequence of codewords. We exploit unique poses of an activity and the temporal ordering of these poses to learn subsequences of codewords which are strongly discriminative for the activity. Each discriminative subsequence acts as a classifier and we learn a boosted ensemble of discriminative subsequences to assign a confidence score for the activity label of the test sequence. Unlike existing methods which demand accurate tracking of 3D joint locations or couple depth with color image information as recognition cues, our method requires only the segmentation masks from depth images to recognize an activity. Experimental results on the publicly available Human Activity Dataset (which comprises 12 challenging activities) demonstrate the validity of our method, where we attain a precision/recall of 78.1%/75.4% when the person was not seen before in the training set, and 94.6%/93.1% when the person was seen before. Raj Kumar Gupta, Alex Yong Sang Chia, Deepu Rajan |
ACM Multimedia | 3 |
| 2013 | Background subtraction via coherent trajectory decompositionabstractBackground subtraction, the task to detect moving objects in a scene, is an important step in video analysis. In this paper, we propose an efficient background subtraction method based on coherent trajectory decomposition. We assume that the trajectories from background lie in a low-rank subspace, and foreground trajectories are sparse outliers in this background subspace. Meanwhile, the Markov Random Field (MRF) is used to encode the spatial coherency and trajectory consistency. With the low-rank decomposition and the MRF, our method can better handle videos with moving camera and obtain coherent foreground. Experimental results on a video dataset show our method achieves very competitive performance. Zhixiang Ren, Liang-Tien Chia, Deepu Rajan, Shenghua Gao |
ACM Multimedia | 3 |
| 2013 | Using texture to complement color in image matting
Ehsan Shahrian, Deepu Rajan |
Image Vis. Comput. | 2 |
| 2013 | Dynamic distortion maps for image retargeting
Yiqun Hu, Deepu Rajan |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Salient Motion Detection in Compressed DomainabstractWe propose a novel motion saliency framework using DCT coefficients and motion vectors as features from MPEG-2 videos. The DCT coefficients of the luma and chroma components are used to calculate the spatial saliency of a frame while the motion vectors are utilized to refine it. Spatio-temporal similarity maps are calculated separately in order to cater to tracking shots. Results show that the proposed motion saliency framework is computationally nearly an 8-fold improvement over the best performing state-of-the-art method while nearly equalling its performance in identifying salient regions. Karthik Muthuswamy, Deepu Rajan |
IEEE Signal Process. Lett. | 2 |
| 2013 | Regularized Feature Reconstruction for Spatio-Temporal Saliency DetectionabstractMultimedia applications such as image or video retrieval, copy detection, and so forth can benefit from saliency detection, which is essentially a method to identify areas in images and videos that capture the attention of the human visual system. In this paper, we propose a new spatio-temporal saliency detection framework on the basis of regularized feature reconstruction. Specifically, for video saliency detection, both the temporal and spatial saliency detection are considered. For temporal saliency, we model the movement of the target patch as a reconstruction process using the patches in neighboring frames. A Laplacian smoothing term is introduced to model the coherent motion trajectories. With psychological findings that abrupt stimulus could cause a rapid and involuntary deployment of attention, our temporal model combines the reconstruction error, regularizer, and local trajectory contrast to measure the temporal saliency. For spatial saliency, a similar sparse reconstruction process is adopted to capture the regions with high center-surround contrast. Finally, the temporal saliency and spatial saliency are combined together to favor salient regions with high confidence for video saliency detection. We also apply the spatial saliency part of the spatio-temporal model to image saliency detection. Experimental results on a human fixation video dataset and an image saliency detection dataset show that our method achieves the best performance over several state-of-the-art approaches. Zhixiang Ren, Shenghua Gao, Liang-Tien Chia, Deepu Rajan |
IEEE Trans. Image Process. | 4 |
| 2013 | Weighted Color and Texture Sample Selection for Image MattingabstractColor sampling based matting methods find the best known samples for foreground and background colors of unknown pixels. Such methods do not perform well if there is an overlap in the color distribution of foreground and background regions because color cannot distinguish between these regions and hence, the selected samples cannot reliably estimate the matte. Furthermore, current sampling based matting methods choose samples that are located around the boundaries of foreground and background regions. In this paper, we overcome these two problems. First, we propose texture as a feature that can complement color to improve matting by discriminating between known regions with similar colors. The contribution of texture and color is automatically estimated by analyzing the content of the image. Second, we combine local sampling with a global sampling scheme that prevents true foreground or background samples to be missed during the sample collection stage. An objective function containing color and texture components is optimized to choose the best foreground and background pair among a set of candidate pairs. Experiments are carried out on a benchmark data set and an independent evaluation of the results shows that the proposed method is ranked first among all other image matting methods. Ehsan Shahrian, Deepu Rajan |
IEEE Trans. Image Process. | 2 |
| 2012 | Weighted color and texture sample selection for image mattingabstractColor information is leveraged by color sampling-based matting methods to find the best known samples for foreground and background color of unknown pixels. Such methods do not perform well if there is an overlap in the color distribution of foreground and background regions because color cannot distinguish between these regions and hence, the selected samples cannot reliably estimate the matte. Similarly, alpha propagation based matting methods may fail when the affinity among neighboring pixels is reduced by strong edges. In this paper, we overcome these two problems by considering texture as a feature that can complement color to improve matting. The contribution of texture and color is automatically estimated by analyzing the content of the image. An objective function containing color and texture components is optimized to choose the best foreground and background pair among a set of candidate pairs. Experiments are carried out on a benchmark data set and an independent evaluation of the results show that the proposed method is ranked first among all other image matting methods. Ehsan Shahrian, Deepu Rajan |
CVPR | 2 |
| 2012 | Sparse likelihood saliency detectionabstractThis paper addresses the problem of detection salient regions in images by exploiting the redundancy in image patches. We assume that redundant patches are more likely to be sparsely represented by other patches in the image while salient patches are not. Such sparse likelihood can be measured via L1-minimization by finding the sparse representation of an image patch based on a dictionary constructed using all other patches from the input image. We show that this approach leads to a robust saliency algorithm and the evaluation based on a database of 1000 images demonstrates that our algorithm achieves significant improvement over existing methods. Minh-Chau Hoang, Deepu Rajan |
ICASSP | 2 |
| 2012 | Salient motion detection through state controllabilityabstractSalient motion detection is a challenging task especially when the motion is obscured by dynamic background motion. Salient motion is characterized by its consistency while the non-salient background motion typically consists of dynamic motion such as fog, waves, fire etc. In this paper, we present a novel framework for identifying salient motion by modelling the video sequence as a linear dynamic system and using controllability of states to estimate salient motion. The proposed saliency detection algorithm is tested on a challenging benchmark video dataset and the performance is compared with other state-of-the-art algorithms. The results of the comparison indicate that the proposed algorithm demonstrates superior performance when compared to other state-of-the-art methods and with higher computational efficiency. Karthik Muthuswamy, Deepu Rajan |
ICASSP | 2 |
| 2012 | Spatiotemporal Saliency Detection via Sparse RepresentationabstractMultimedia applications like retrieval, copy detection etc. can gain from saliency detection, which is essentially a method to identify areas in images and videos that capture the attention of the human visual system. In this paper, we propose a new spatiotemporal saliency framework for videos based on sparse representation. For temporal saliency, we model the movement of the target patch as a reconstruction process, and the overlapping patches in neighboring frames are used to reconstruct the target patch. The learned coefficients encode the positions of the matched patches, which are able to represent the motion trajectory of the target patch. We also introduce a smoothing term into our sparse coding framework to learn coherent motion trajectories. Based on the psychological findings that abrupt stimulus could cause a rapid and involuntary deployment of attention, our temporal model combines the reconstruction error, sparsity regularizer, and local trajectory contrast to measure the motion saliency. For spatial saliency, a similar sparse reconstruction process is adopted to capture the regions with high center-surround contrast. Finally, the temporal saliency and spatial saliency are combined by agreement to favor the salient regions with high confidence. Experimental results on a human fixation video dataset show our method achieved the best performance over five state-of-the-art approaches. Zhixiang Ren, Shenghua Gao, Deepu Rajan, Liang-Tien Chia |
ICME | 3 |
| 2012 | Salient Object Detection through Over-SegmentationabstractIn this paper we present a salient object detection model from an over-segmented image. The input image is initially segmented by the mean-shift segmentation algorithm and then over-segmented by a quad mesh to even smaller segments. Such segmented regions overcome the disadvantage of using patches or single pixels to compute saliency. Segments that are similar and spread over the image receive low saliency and a segment which is distinct in the whole image or in a local region receives high saliency. We express this as a color compactness measure which is used to derive saliency level directly. Our method is shown to outperform six existing methods in the literature using a saliency detection database containing images with human-labeled object contour ground truth. The proposed saliency model has been shown to be useful for an image retargeting application. Zhixiang Ren, Deepu Rajan, Yiqun Hu |
ICME | 3 |
| 2012 | Video saliency detection with robust temporal alignment and local-global spatial contrastabstractVideo saliency detection, the task to detect attractive content in a video, has broad applications in multimedia understanding and retrieval. In this paper, we propose a new framework for spatiotemporal saliency detection. To better estimate the salient motion in temporal domain, we take advantage of robust alignment by sparse and low-rank decomposition to jointly estimate the salient foreground motion and the camera motion. Consecutive frames are transformed and aligned, and then decomposed to a low-rank matrix representing the background and a sparse matrix indicating the objects with salient motion. In the spatial domain, we address several problems of local center-surround contrast based model, and demonstrate how to utilize global information and prior knowledge to improve spatial saliency detection. Individual component evaluation demonstrates the effectiveness of our temporal and spatial methods. Final experimental results show that the combination of our spatial and temporal saliency maps achieve the best overall performance compared to several state-of-the-art methods. Zhixiang Ren, Liang-Tien Chia, Deepu Rajan |
ICMR | 3 |
| 2012 | Image colorization using similar imagesabstractWe present a new example-based method to colorize a gray image. As input, the user needs only to supply a reference color image which is semantically similar to the target image. We extract features from these images at the resolution of superpixels, and exploit these features to guide the colorization process. Our use of a superpixel representation speeds up the colorization process. More importantly, it also empowers the colorizations to exhibit a much higher extent of spatial consistency in the colorization as compared to that using independent pixels. We adopt a fast cascade feature matching scheme to automatically find correspondences between superpixels of the reference and target images. Each correspondence is assigned a confidence based on the feature matching costs computed at different steps in the cascade, and high confidence correspondences are used to assign an initial set of chromatic values to the target superpixels. To further enforce the spatial coherence of these initial color assignments, we develop an image space voting framework which draws evidence from neighboring superpixels to identify and to correct invalid color assignments. Experimental results and user study on a broad range of images demonstrate that our method with a fixed set of parameters yields better colorization results as compared to existing methods. Raj Kumar Gupta, Alex Yong Sang Chia, Deepu Rajan, Ee Sin Ng, Zhiyong Huang 0001 |
ACM Multimedia | 3 |
| 2012 | Scene Signatures for Unconstrained News Video Stories
Ehsan Younessian, Deepu Rajan |
MMM | 2 |
| 2012 | Multi-modal Solution for Unconstrained News Story Retrieval
Ehsan Younessian, Deepu Rajan |
MMM | 2 |
| 2012 | New Edge Characteristics for Scene and Object ClassificationabstractIn this paper, we show that simple edge characteristics in images, when judiciously combined, can result in improved scene and object classification. Unlike existing methods that require a large number of training samples and complex learning schemes, our method discovers simple edge properties. We introduce three sets of edge properties, namely, centroid, compactness and aspect ratio of edges in the image. The combinations of these edge properties are used to discriminate among images in each class. A class representative is calculated for each class according to the average percentage of edges that satisfy the property of a particular class. This percentage for an unknown image is compared to the class representative to assign a label to it. It is shown that this simple edge properties-based method outperforms some of the state-of-the-art results on scene and object classification on standard databases. Palaiahnakote Shivakumara, Deepu Rajan, Suresh Anand Sadananthan |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2012 | Object Recognition by Discriminative Combinations of Line Segments, Ellipses, and Appearance FeaturesabstractWe present a novel contour-based approach that recognizes object classes in real-world scenes using simple and generic shape primitives of line segments and ellipses. Compared to commonly used contour fragment features, these primitives support more efficient representation since their storage requirements are independent of object size. Additionally, these primitives are readily described by their geometrical properties and hence afford very efficient feature comparison. We pair these primitives as shape-tokens and learn discriminative combinations of shape-tokens. Here, we allow each combination to have a variable number of shape-tokens. This, coupled with the generic nature of primitives, enables a variety of class-specific shape structures to be learned. Building on the contour-based method, we propose a new hybrid recognition method that combines shape and appearance features. Each discriminative combination can vary in the number and the types of features, where these two degrees of variability empower the hybrid method with even more flexibility and discriminative potential. We evaluate our methods across a large number of challenging classes, and obtain very competitive results against other methods. These results show the proposed shape primitives are indeed sufficiently powerful to recognize object classes in complex real-world scenes. Alex Yong Sang Chia, Deepu Rajan, Maylor Karhang Leung, Susanto Rahardja |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Human action recognition using Pose-based discriminant embedding
Behrouz Saghafi, Deepu Rajan |
Signal Process. Image Commun. | 2 |
| 2012 | A Linear Dynamical System Framework for Salient Motion DetectionabstractDetection of salient motion in a video involves determining which motion is attended to by the human visual system in the presence of background motion that consists of complex visuals that are constantly changing. Salient motion is marked by its predictability compared to the more complex unpredictable motion of the background such as fluttering of leaves, ripples in water, dispersion of smoke, and others. We introduce a novel approach to detect salient motion based on the concept of “observability” from the output pixels, when the video sequence is represented as a linear dynamical system. The group of output pixels with maximum saliency is further used to model the holistic dynamics of the salient region. The pixel saliency map is bolstered by two region-based saliency maps, which are computed based on the similarity of dynamics of the different spatiotemporal patches in the video with the salient region dynamics, in a global as well as a local sense. The resulting algorithm is tested on a set of challenging sequences and compared to state-of-the-art methods to showcase its superior performance on grounds of its computational efficiency and ability to detect salient motion. Viswanath Gopalakrishnan, Deepu Rajan, Yiqun Hu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | A Split and Merge Based Ellipse Detector With Self-Correcting CapabilityabstractA novel ellipse detector based upon edge following is proposed in this paper. The detector models edge connectivity by line segments and exploits these line segments to construct a set of elliptical-arcs. Disconnected elliptical-arcs which describe the same ellipse are identified and grouped together by incrementally finding optimal pairings of elliptical-arcs. We extract hypothetical ellipses of an image by fitting an ellipse to the elliptical-arcs of each group. Finally, a feedback loop is developed to sieve out low confidence hypothetical ellipses and to regenerate a better set of hypothetical ellipses. In this aspect, the proposed algorithm performs self-correction and homes in on "difficult" ellipses. Detailed evaluation on synthetic images shows that the algorithm outperforms existing methods substantially in terms of recall and precision scores under the scenarios of image cluttering, salt-and-pepper noise and partial occlusion. Additionally, we apply the detector on a set of challenging real-world images. Successful detection of ellipses present in these images is demonstrated. We are not aware of any other work that can detect ellipses from such difficult images. Therefore, this work presents a significant contribution towards ellipse detection. Alex Yong Sang Chia, Susanto Rahardja, Deepu Rajan, Maylor Karhang Leung |
IEEE Trans. Image Process. | 3 |
| 2010 | Sustained Observability for Salient Motion Detection
Viswanath Gopalakrishnan, Yiqun Hu, Deepu Rajan |
ACCV (3) | 3 |
| 2010 | Unsupervised Feature Selection for Salient Object Detection
Viswanath Gopalakrishnan, Yiqun Hu, Deepu Rajan |
ACCV (2) | 3 |
| 2010 | Salient Region Detection by Jointly Modeling Distinctness and Redundancy of Image Content
Yiqun Hu, Zhixiang Ren, Deepu Rajan, Liang-Tien Chia |
ACCV (2) | 3 |
| 2010 | Embedding Visual Words into Concept Space for Action and Scene RecognitionabstractIn this paper we propose a novel approach to introducing semantic relations into the bag-of-words framework. We use the latent semantic models, such as LSA and pLSA, in order to define semantically-rich features and embed the visual features into a semantic space. The semantic features used in LSA technique are derived from the low-rank approximation of word-document occurrence matrix by SVD. Similarly, by using the pLSA approach, the topic-specific distributions of words can be considered dimensions of a concept space. In the proposed space, the distances between words represent the semantic distances which are used for constructing a discriminative and semantically meaningful vocabulary. We have tested our approach on the KTH action database and on the Fifteen Scene database and have achieved very promising results on both. Behrouz Khadem, Elahe Farahzadeh, Deepu Rajan, Andrzej Stefan Sluzek |
BMVC | 3 |
| 2010 | Object recognition by discriminative combinations of line segments and ellipsesabstractWe present a contour based approach to object recognition in real-world images. Contours are represented by generic shape primitives of line segments and ellipses. These primitives offer substantial flexibility to model complex shapes. We pair connected primitives as shape tokens, and learn category specific combinations of shape tokens. We do not restrict combinations to have a fixed number of tokens, but allow each combination to flexibly evolve to best represent a category. This, coupled with the generic nature of primitives, enables a variety of discriminative shape structures of a category to be learned. We compare our approach with related methods and state-of-the-art contour based approaches on two demanding datasets across 17 categories. Highly competitive results are obtained. In particular, on the challenging Weizmann horse dataset, we attain improved image classification and object detection results over the best contour based results published so far. Alex Yong Sang Chia, Susanto Rahardja, Deepu Rajan, Maylor K. H. Leung |
CVPR | 3 |
| 2010 | Hybrid shift map for video retargetingabstractWe propose a new method for video retargeting, which can generate spatial-temporal consistent video. The new measure called spatial-temporal naturality preserves the motion in the source video without any motion analysis in contrast to other methods that need motion estimation. This advantage prevents the retargeted video from degenerating due to the propagation of the errors in motion analysis. It allows the proposed method to be applied on challenging videos with complex camera and object motion. To improve the efficiency of the retargeting process, we retarget video using a 3D shift map in low resolution and refine it using an incremental 2D shift map in higher resolution. This new hierarchical framework, denoted as hybrid shift map, can produce satisfactory retargeting results while significantly improving the computational efficiency. Yiqun Hu, Deepu Rajan |
CVPR | 2 |
| 2010 | DCT domain watermarking scheme using Chinese Remainder Theorem for image authenticationabstractIn this paper, we propose a novel Chinese Remainder Theorem (CRT) based watermarking scheme that works in the Discrete Cosine Transform (DCT) domain. We first reviewed a Singular Value Decomposition (SVD) based watermarking scheme [11] followed by an existing CRT based watermarking scheme [7] that works in the spatial domain and their shortcomings are highlighted. The proposed CRT based scheme is more resistant to different types of attacks, particularly to JPEG compression; in addition, it improves the security feature of the watermarking scheme. Experimental results have shown that the proposed scheme makes the watermark perceptually invisible and has better robustness to common image manipulation techniques such as JPEG compression, brightening and sharpening effects compared to the spatial domain based CRT scheme; in addition, its computational complexity is much lower than the SVD-based scheme. Jagdish C. Patra, Jiliang E. Phua, Deepu Rajan |
ICME | 3 |
| 2010 | Multi-view Clustering of Visual Words Using Canonical Correlation Analysis for Human Action RecognitionabstractIn this paper we propose a novel approach for introducing semantic relations into the bag-of-words framework for recognizing human actions. We represent visual words in two different views: the original features and the document co-occurrence representation. The latter view conveys semantic relations but is large, sparse and noisy. We use canonical correlation analysis between the two views to find a subspace in which the words are more semantically distributed. We apply k-means clustering in the computed space to find semantically meaningful clusters and use them as the semantic visual vocabulary. Incorporating the semantic visual vocabulary the features are quantized to form more discriminative histograms. Eventually the histograms are classified using an SVM classifier. We have tested our approach on KTH action dataset and achieved promising results. Behrouz Saghafi, Deepu Rajan |
ICMLA | 2 |
| 2010 | Image Retargeting in Compressed DomainabstractA simple algorithm for image retargeting in the compressed domain is proposed. Most existing retargeting algorithms work directly in the spatial domain of the raw image. Here, we work on the DCT coefficients of a JPEG-compressed image to generate a gradient map that serves as an importance map to help identify those parts in the image that need to be retained during the retargeting process. Each 8×8 block of DCT coefficients is scaled based on the least importance value. Retargeting can be done both in the horizontal and vertical directions with the same framework. We also illustrate image enlargement using the same method. Experimental results show that the proposed algorithm produces less distortion in the retargeted image compared to some other algorithms reported recently. O. V. Ramana Murthy, Karthik Muthuswamy, Deepu Rajan, Liang-Tien Chia |
ICPR | 3 |
| 2010 | Improved saliency detection based on superpixel clustering and saliency propagationabstractSaliency detection is useful for high level applications such as adaptive compression, image retargeting, object recognition, etc. In this paper, we introduce an effective region-based solution for saliency detection. We first use the adaptive mean shift algorithm to extract superpixels from the input image, then apply Gaussian Mixture Model (GMM) to cluster superpixels based on their color similarity, and finally calculate the saliency value for each cluster using compactness metric together with modified PageRank propagation. This solution is able to represent the image in a perceptually meaningful way and is robust to over-segmentation. It highlights salient regions with full resolution, well-defined boundary. Experimental results show that both the adaptive mean shift and the modified PageRank algorithm contribute substantially to the saliency detection result. In addition, the ROC analysis demonstrates that our approach significantly outperforms five existing popular methods. Zhixiang Ren, Yiqun Hu, Liang-Tien Chia, Deepu Rajan |
ACM Multimedia | 4 |
| 2010 | Scene classification using multiple features in a two-stage probabilistic classification framework
Deepu Rajan, Liang-Tien Chia |
Neurocomputing | 2 |
| 2010 | Random Walks on Graphs for Salient Object Detection in ImagesabstractWe formulate the problem of salient object detection in images as an automatic labeling problem on the vertices of a weighted graph. The seed (labeled) nodes are first detected using Markov random walks performed on two different graphs that represent the image. While the global properties of the image are computed from the random walk on a complete graph, the local properties are computed from a sparse k-regular graph. The most salient node is selected as the one which is globally most isolated but falls on a locally compact object. A few background nodes and salient nodes are further identified based upon the random walk based hitting time to the most salient node. The salient nodes and the background nodes will constitute the labeled nodes. A new graph representation of the image that represents the saliency between nodes more accurately, the "pop-out graph" model, is computed further based upon the knowledge of the labeled salient and background nodes. A semisupervised learning technique is used to determine the labels of the unlabeled nodes by optimizing a smoothness objective label function on the newly created "pop-out graph" model along with some weighted soft constraints on the labeled nodes. Viswanath Gopalakrishnan, Yiqun Hu, Deepu Rajan |
IEEE Trans. Image Process. | 3 |
| 2009 | Random walks on graphs to model saliency in imagesabstractWe formulate the problem of salient region detection in images as Markov random walks performed on images represented as graphs. While the global properties of the image are extracted from the random walk on a complete graph, the local properties are extracted from a k-regular graph. The most salient node is selected as the one which is globally most isolated but falls on a compact object. The equilibrium hitting times of the ergodic Markov chain holds the key for identifying the most salient node. The background nodes which are farthest from the most salient node are also identified based on the hitting times calculated from the random walk. Finally, a seeded salient region identification mechanism is developed to identify the salient parts of the image. The robustness of the proposed algorithm is objectively demonstrated with experiments carried out on a large image database annotated with “ground-truth” salient regions. Viswanath Gopalakrishnan, Yiqun Hu, Deepu Rajan |
CVPR | 3 |
| 2009 | Improved Keypoint Matching Method for Near-Duplicate Keyframe RetrievalabstractWe propose a Near-Duplicate Keyframe (NDK) retrieval method that can handle extreme zooming and significant object motion. The first stage consists of eliminating false keypoint matches using symmetric property and a ratio of nearest and second-nearest neighbor distances. Then, a pattern coherency score is assigned to each pair of keyframes. These two features are combined through linear discriminant analysis (LDA) and the separating boundary is trained using SVM. Experiments are carried out for NDK retrieval on the Columbia and NTU datasets. The promising results confirm the effectiveness of our keypoint matching algorithm and show distinguishing power of our proposed features and feature weighting role in NDK retrieval. Ehsan Younessian, Deepu Rajan, Chng Eng Siong |
ISM | 2 |
| 2009 | Attention-from-motion: A factorization approach for detecting attention objects in motion
Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
Comput. Vis. Image Underst. | 2 |
| 2009 | Structural Descriptors for Category Level Object DetectionabstractWe propose a new class of descriptors which exhibits the ability to yield meaningful structural descriptions of objects. These descriptors are constructed from two types of image primitives: quadrangles and ellipses. The primitives are extracted from an image based on human cognitive psychology and model local parts of objects. Experiments reveal that these primitives densely cover objects in images. In this regard, structural information of an object can be comprehensively described by these primitives. It is found that a combination of simple spatial relationships between primitives plus a small set of geometrical attributes provide rich and accurate local structural descriptions of objects. Category level object detection of four-legged animals, bicycles, and cars images is demonstrated under scaling, moderate viewpoint variations, and background clutter. Promising results are achieved. Alex Yong Sang Chia, Susanto Rahardja, Deepu Rajan, Maylor K. H. Leung |
IEEE Trans. Multim. | 3 |
| 2009 | Salient Region Detection by Modeling Distributions of Color and OrientationabstractWe present a robust salient region detection framework based on the color and orientation distribution in images. The proposed framework consists of a color saliency framework and an orientation saliency framework. The color saliency framework detects salient regions based on the spatial distribution of the component colors in the image space and their remoteness in the color space. The dominant hues in the image are used to initialize an expectation-maximization (EM) algorithm to fit a Gaussian mixture model in the hue-saturation (H-S) space. The mixture of Gaussians framework in H-S space is used to compute the inter-cluster distance in the H-S domain as well as the relative spread among the corresponding colors in the spatial domain. Orientation saliency framework detects salient regions in images based on the global and local behavior of different orientations in the image. The oriented spectral information from the Fourier transform of the local patches in the image is used to obtain the local orientation histogram of the image. Salient regions are further detected by identifying spatially confined orientations and with the local patches that possess high orientation entropy contrast. The final saliency map is selected as either color saliency map or orientation saliency map by automatically identifying which of the maps leads to the correct identification of the salient region. The experiments are carried out on a large image database annotated with ldquoground-truthrdquo salient regions, provided by Microsoft Research Asia, which enables us to conduct robust objective level comparisons with other salient region detection algorithms. Viswanath Gopalakrishnan, Yiqun Hu, Deepu Rajan |
IEEE Trans. Multim. | 3 |
| 2009 | Coherent Phrase Model for Efficient Image Near-Duplicate RetrievalabstractThis paper presents an efficient and effective solution for retrieving image near-duplicate (IND) from image database. We introduce the coherent phrase model which incorporates the coherency of local regions to reduce the quantization error of the bag-of-words (BoW) model. In this model, local regions are characterized byvisual phraseof multiple descriptors instead of visual word of single descriptor. We propose two types of visual phrase to encode the coherency in feature and spatial domain, respectively. The proposed model reduces the number of false matches by using this coherency and generates sparse representations of images. Compared to other method, the local coherencies among multiple descriptors of every region improve the performance and preserve the efficiency for IND retrieval. The proposed method is evaluated on several benchmark datasets for IND retrieval. Compared to the state-of-the-art methods, our proposed model has been shown to significantly improve the accuracy of IND retrieval while maintaining the efficiency of the standard bag-of-words model. The proposed method can be integrated with other extensions of BoW. Yiqun Hu, Xiangang Cheng, Liang-Tien Chia, Xing Xie 0001, Deepu Rajan, Ah-Hwee Tan |
IEEE Trans. Multim. | 5 |
| 2008 | A split and merge based ellipse detectorabstractWe present an ellipse detector that continually pools lower level information of the edge pixels together to achieve robust detection of the ellipses present in the image. In addition, the parameters of the detected ellipses are continually refined using a close loop system driven by Gestalt psychology. We highlight that we do not rely on the geometrical properties of the ellipses to detect the ellipses. In this aspect, our algorithm is well suited to detect partially occluded ellipses in the image. Experiments on real and synthetic images demonstrate the robustness of our algorithm in which both complete and incomplete ellipses can be detected. In particular, experimental results show that the mean detection accuracy of our algorithm surpasses 92% even with around 90% outliers in the images. This detection performance is superior to that achieved by the robust regression, least squares and the hough transform based ellipse detectors. Alex Yong Sang Chia, Deepu Rajan, Maylor K. H. Leung, Susanto Rahardja |
ICIP | 2 |
| 2008 | Image classification: Are rule-based systems effective when classes are fixed and known?abstractIn this paper, we investigate if rule-based systems are useful for image classification problems when the number of classes is fixed. The rules are derived from simple edge features such as width and straightness. A class representative is calculated for each class according to the average percentage of edges that satisfy the rule for a particular class. This percentage for an unknown image is compared to the class representative to assign a label to it. The proposed system does not require extensive feature extraction and classification techniques. It is shown that the rule based system outperforms some of the reported results on scene classification. Palaiahnakote Shivakumara, Deepu Rajan, Suresh Anand Sadananthan |
ICPR | 2 |
| 2008 | Detection of visual attention regions in images using robust subspace analysis
Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
J. Vis. Commun. Image Represent. | 2 |
| 2007 | A ZGPCA Algorithm for Subspace EstimationabstractWe propose a new algorithm called the ZGPCA algorithm for subspace estimation based on the GPCA (Generalized Principal Component Analysis) algorithm. It is formulated within an FIR filter framework so that the norm vectors of the subspaces correspond to filter coefficients. It is shown that such an approach leads to a more accurate and computationally efficient method compared to the GPCA algorithm. We extend the ZGPCA algorithm to make it recursive so that subspaces with possibly different dimensions can be obtained. We also propose a new distance measure that can be used for k-means clustering of sample points within a subspace. Experimental results on synthetic data and applications on face clustering and sports video clustering show good performance of the proposed algorithm. Haoran Yi, Deepu Rajan, Liang-Tien Chia |
ICME | 2 |
| 2007 | Scale adaptive visual attention detection by subspace analysisabstractWe describe a method to extract visual attention regions in images by robust subspace analysis from simple feature like intensity endowed with scale adaptivity in order to represent textured areas in an image. The scale adaptive descriptor is mapped onto clusters in linear spaces. A new subspace estimation algorithm based on the Generalized Principal Component Analysis (GPCA) is proposed to estimate multiple linear subspaces. The visual attention of each region is calculated using a new region attention measure that considers feature contrast and spatial geometric properties. Compared with existing visual attention detection methods, the proposed method directly measures global visual attention at the region level as opposed to pixel level. Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
ACM Multimedia | 2 |
| 2006 | An Event-Driven Sports Video Adaptation for the MPEG-21 DIA FrameworkabstractWe present an event-driven video adaptation system in this paper. Events are detected by audio/video analysis and annotated by the description schemes (DSs) provided by MPEG-7 multimedia description schemes (MDSs). And then, adaptation take account of users' preference of events and network characteristic to adapt video by event selection and frame dropping as following three steps: 1) the event information is parsed from MPEG-7 annotation XML file together with bitstream to generate generic bitstream syntax description (gBSD), 2) users' preference, network characteristic and adaptation QoS (AQoS) are considered for making adaptation decision, 3) adaptation engine automatically parses adaptation decisions and gBSD to achieve adaptation. Different from most existing adaptation work, the system adapts video by interesting events according to users' preference. To achieve a generic adaptation solution, the system is developed following MPEG-7 and MPEG-21 standards. gBSD based adaptation avoids complex video computation. 30 students from various departments test the system with satisfaction. Although, the system is tested on basketball video adaptation so far, it is easy to extend to other video domains Min Xu 0001, Jiaming Li 0003, Yiqun Hu, Liang-Tien Chia, Bu-Sung Lee, Deepu Rajan, Jianfei Cai 0001 |
ICME | 6 |
| 2006 | Event on demand with MPEG-21 video adaptation systemabstractIn this paper, we present an event-on-demand (EoD)video adaptation system. The proposed system supports users in deciding their events of interest and considers network conditions to adapt video source by event selection and frame dropping.Firstly, events are detected by audio/video analysis and annotated by the description schemes (DSs)provided by MPEG-7 Multimedia Description Schemes (MDSs). And then, to achieve a generic adaptation solution, the adaptation is developed following MPEG-21 Digital Item Adaptation (DIA)framework. We look at early release of the MPEG-21 Reference Software on XML generation and develop our own system for EoD video adaptation in three steps:1) the event information is parsed from MPEG-7 annotation XML file together with bitstream to generate generic Bitstream Syntax Description (gBSD). 2) Users' preference, Network Characteristic and Adaptation QoS (AQoS) are considered for making adaptation decision. 3) adaptation engine automatically parses adaptation decisions and gBSD to achieve adaptation.Unlike most existing adaptation work, the system adapts video of events with interest according to users' preference. Implementation following MPEG-7 and MPEG-21 standards provides a generic video adaptation solution. gBSD based adaptation avoids complex video computation. 30 students from various departments were invited to test the system and their responses has been positive. Min Xu 0001, Jiaming Li 0003, Liang-Tien Chia, Yiqun Hu, Bu-Sung Lee, Deepu Rajan, Jesse S. Jin |
ACM Multimedia | 6 |
| 2006 | Affective content detection in sitcom using subtitle and audioabstractFrom a personalized media point of view, many users favor a flexible tool to quickly browse the affective content in a video. Such affective content may cause audiences' strong reactions or special emotional experiences, such as anger, sadness, fear, joy and love. This paper attempts to extract affective content for digital videos by analyzing the subtitle files of DVD/DivX videos and utilize audio event to assist affective content detection. Firstly, videos are segmented by dialogue script partition. Compared to traditional video shot, video segmented by scripts is not affected by camera changes and shooting angles and easy to include video segments with compact content. Secondly, emotion-related vocabularies in video script are detected to locate affective video content. Using script to directly access video content avoids complex video analysis. Thirdly, audio event detection is utilized to assist affective content detection. Compared with traditional video semantic analysis, affective content analysis puts much more emphasis on the audience's reactions and emotions. Initial experiments are carried on sitcom videos because its simple video structure provides useful domain knowledge. The experimental results demonstrate that subtitle file analysis and audio event detection provides effective and efficient clues to determine the emotional content of the videos. Min Xu 0001, Liang-Tien Chia, Haoran Yi, Deepu Rajan |
MMM | 4 |
| 2006 | A motion-based scene tree for browsing and retrieval of compressed videos
Haoran Yi, Deepu Rajan, Liang-Tien Chia |
Inf. Syst. | 2 |
| 2006 | A motion-based scene tree for compressed video content management
Haoran Yi, Deepu Rajan, Liang-Tien Chia |
Image Vis. Comput. | 2 |
| 2006 | Dynamic Programming-Based Reverse Frame Selection for VBR Video Delivery Under Constrained ResourcesabstractIn this paper, we investigate optimal frame-selection algorithms based on dynamic programming for delivering stored variable bit rate (VBR) video under both bandwidth and buffer size constraints. Our objective is to find a feasible set of frames that can maximize the video's accumulated motion values without violating any constraint. It is well known that dynamic programming has high complexity. In this research, we propose to eliminate nonoptimal intermediate frame states, which can effectively reduce the complexity of dynamic programming. Moreover, we propose a reverse frame selection (RFS) algorithm, where the selection starts from the last frame and ends at the first frame. Compared with the conventional dynamic programming-based forward frame selection, the RFS is able to find all of the optimal results for different preloads in one round. We further extend the RFS scheme to solve the problem of frame selection for VBR channels. In particular, we first perform the RFS algorithm offline, and the complexity is modest and scalable with the aids of frame stuffing and nonoptimal state elimination. During online streaming, we only need to retrieve the optimal frame-selection path from the pregenerated offline results, and it can be applied to any VBR channels as long as the VBR channels can be modeled as piecewise CBR channels. Experimental results show good performance of our proposed algorithms Dayong Tao, Jianfei Cai 0001, Haoran Yi, Deepu Rajan, Liang-Tien Chia, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2005 | Global Motion Compensated Key Frame Extraction from Compressed VideosabstractA key frame extraction approach, based on change detection of DC images extracted from compressed video, is proposed in this paper. We define a simple pixel change map that captures additional information in a frame with respect to its adjacent frames. Since global motion contributes to pixel changes, falsely indicating the presence of key frames, it is compensated by adaptively filtering the pixel change map using a modified version of the least mean square (LMS) algorithm. The prediction errors thus obtained are used to subsequently select the key frames. The key frames are selected so that the cumulative prediction error is partitioned into equal amounts in each segment. The entire procedure is computationally simple and flexible. Experimental results illustrate the good performance of the proposed algorithm. Haoran Yi, Deepu Rajan, Liang-Tien Chia |
ICASSP (2) | 2 |
| 2005 | Adaptive local context suppression of multiple cues for salient visual attention detectionabstractVisual attention is obtained through determination of contrasts of low level features or attention cues like intensity, color etc. We propose a new texture attention cue that is shown to be more effective for images where the salient object regions and background have similar visual characteristics. Current visual attention models do not consider local contextual information to highlight attention regions. We also propose a feature combination strategy by suppressing saliency based on context information that is effective in determining the true attention region. We compare our approach with other visual attention models using a novel average discrimination ratio measure. Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
ICME | 2 |
| 2005 | Adaptive hierarchical multi-class SVM classifier for texture-based image classificationabstractIn this paper, we present a new classification scheme based on support vector machines (SVM) and a new texture feature, called texture correlogram, for high-level image classification. Originally, SVM classifier is designed for solving only binary classification problem. In order to deal with multiple classes, we present a new method to dynamically build up a hierarchical structure from the training dataset. The texture correlogram is designed to capture spatial distribution information. Experimental results demonstrate that the proposed classification scheme and texture feature are effective for high-level image classification task and the proposed classification scheme is more efficient than the other schemes while achieving almost the same classification accuracy. Another advantage of the proposed scheme is that the underlying hierarchical structure of the SVM classification tree manifests the interclass relationships among different classes. Song Liu 0001, Haoran Yi, Liang-Tien Chia, Deepu Rajan |
ICME | 4 |
| 2005 | Robust subspace analysis for detecting visual attention regions in imagesabstractDetecting visually attentive regions of an image is a challenging but useful issue in many multimedia applications. In this paper, we describe a method to extract visual attentive regions in images using subspace estimation and analysis techniques. The image is represented in a 2D space using polar transformation of its features so that each region in the image lies in a 1D linear subspace. A new subspace estimation algorithm based on Generalized Principal Component Analysis (GPCA) is proposed. The robustness of subspace estimation is improved by using weighted least square approximation where weights are calculated from the distribution of K nearest neighbors to reduce the sensitivity of outliers. Then a new region attention measure is defined to calculate the visual attention of each region by considering both feature contrast and geometric properties of the regions. The method has been shown to be effective through experiments to be able to overcome the scale dependency of other methods. Compared with existing visual attention detection methods, it directly measures the global visual contrast at the region level as opposed to pixel level contrast and can correctly extract the attentive region. Yiqun Hu, Deepu Rajan, Liang-Tien Chia |
ACM Multimedia | 2 |
| 2005 | Attention region selection with information from professional digital cameraabstractThe attentive region extraction is a challenging issue for semantic interpretation of image and video content. The successful attentive region extraction greatly facilitates image classification, adaptation, compression and retrieval. Different from the traditional visual attention detection models, we propose a new attentive region extraction method based on out-of-focus blurring (OFB) technique used by professional photographers. Firstly, we combine metadata in Exchangeable Image File Format (EXIF) with visual features to quickly select professional photographs from image database. After that, an algorithm is implemented to automatically extract the attentive region from these photographs. This algorithm measures the saliency for individual pixels based on edge distribution of the images. The experimental results on OFB images have proved that our approach is able to overcome the contrast map selection problem of traditional visual attention methods and extract the attentive region using OFB information. The attentive region generated by our algorithm has similar shape and size with the subject of photographs which is a useful information for searching and retrieving the high-level semantic meaningful objects. Song Liu 0001, Liang-Tien Chia, Deepu Rajan |
ACM Multimedia | 3 |
| 2005 | Automatic Generation of MPEG-7 Compliant XML Document for Motion Trajectory Descriptor in Sports Video
Haoran Yi, Deepu Rajan, Liang-Tien Chia |
Multim. Tools Appl. | 2 |
| 2005 | A new motion histogram to index motion content in video segments
Haoran Yi, Deepu Rajan, Liang-Tien Chia |
Pattern Recognit. Lett. | 2 |
| 2004 | Region-of-interest based image resolution adaptation for MPEG-21 digital itemabstractThe upcoming MPEG-21 standard proposes a general framework for augmented use of multimedia services in different network environments, for various users with various terminal devices. In the context of image adaptation, terminals with different screen size limitation require the multimedia adaptation engine to adapt image resources intelligently. Saliency map based visual attention analysis provides some intelligence for finding the attention area within the image. In this paper, we improved the standard MPEG-21 metadata driven adaptation engine by using enhanced saliency map based visual attention model which provides a mean to intelligently adapt JPEG2000 image resolution for different terminal devices with varying screen size according to human visual attention. Yiqun Hu, Liang-Tien Chia, Deepu Rajan |
ACM Multimedia | 3 |
| 2004 | Automatic extraction of motion trajectories in compressed sports videosabstractThis paper presents an algorithm for automatically extracting significant motion trajectories in sports videos. Our approach consists of four stages: global motion estimation, motion blob detection, trajectory evolution and trajectory refinement. Global motion is estimated from the motion vectors in the compressed video using an iterative algorithm with robust outlier rejection. A statistical hypothesis test is carried out within the Block Rejection Map(BRM), which is the by-product of the global motion estimation, for the detection of motion blobs. Trajectory evolution is the process in which the motion blobs are either appended to an existing trajectory or are considered to be the beginning of a new trajectory based on its distance to an adaptive trajectory description. Finally, the extracted motion trajectories are refined using a Kalman filter. Experimental results on both indoor and outdoor sports videos demonstrate the effectiveness and efficiency of the proposed method. Haoran Yi, Deepu Rajan, Liang-Tien Chia |
ACM Multimedia | 2 |
| 2003 | Efficient image retrieval using MPEG-7 descriptorsabstractIn this paper, a new method to calculate the similarity among images using dominant color descriptor is discussed. Using earth mover's distance (EMD), better retrieval results can be obtained compared with those obtained from the original MPEG-7 reference software (XM) [Text of ISO/IEC 15938-6/FDIS Information Technology-Multimedia content description interface-Part 6: Reference Software]. To further improve the retrieval accuracy, texture information from edge histogram descriptor is added. In order to reduce the retrieval time, two different methods which can prune the images far from the query image are discussed. One is the lower bound of EMD, while the other is the M-tree index based on EMD distance. Experiments show that the lower bound is easier to implement and more efficient than the M-tree. Surong Wang, Liang-Tien Chia, Deepu Rajan |
ICIP (3) | 3 |
| 2003 | A unified approach to detection of shot boundaries and subshots in compressed videoabstractThis paper describes a method to partition a video sequence into shots and subshots. By subshots, we mean one or a combination of the three camera motions of pan, tilt and zoom. The proposed technique detects both hard cuts and gradual transitions in MPEG compressed video using a single technique. We also present a motion estimation algorithm to compute the dominant motion represented by an affine model. The motion information is used to refine the location of dissolves as well as to subdivide the shot into subshots, thus providing a characterization of camera motion. We consider the dissimilarity between the I-, P- and B-frames with respect to the type of macroblocks used for encoding. Unlike previous algorithms reported, our method requires minimal decompression of the video sequence and uses very loose thresholds. The algorithm is evaluated on several types of video sequences to demonstrate its effectiveness. Haoran Yi, Deepu Rajan, Liang-Tien Chia |
ICIP (2) | 2 |
| 2003 | Simultaneous Estimation of Super-Resolved Scene and Depth Map from Low Resolution Defocused ObservationsabstractThis paper presents a novel technique to simultaneously estimate the depth map and the focused image of a scene, both at a super-resolution, from its defocused observations. Super-resolution refers to the generation of high spatial resolution images from a sequence of low resolution images. Hitherto, the super-resolution technique has been restricted mostly to the intensity domain. In this paper, we extend the scope of super-resolution imaging to acquire depth estimates at high spatial resolution simultaneously. Given a sequence of low resolution, blurred, and noisy observations of a static scene, the problem is to generate a dense depth map at a resolution higher than one that can be generated from the observations as well as to estimate the true high resolution focused image. Both the depth and the image are modeled as separate Markov random fields (MRF) and a maximum a posteriori estimation method is used to recover the high resolution fields. Since there is no relative motion between the scene and the camera, as is the case with most of the super-resolution and structure recovery techniques, we do away with the correspondence problem. Deepu Rajan, Subhasis Chaudhuri |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Generation of super-resolution images from blurred observations using Markov random fieldsabstractThis paper presents a new technique for generating a high resolution image from a blurred image sequence; this is also referred to as super-resolution restoration of images. The image sequence consists of decimated, blurred and noisy versions of the high resolution image. The high resolution image is modeled as a Markov random field (MRF) and a maximum a posteriori (MAP) estimation technique is used. A simple gradient descent method is used to optimize the functional. Further, line fields are introduced in the cost function and optimization using Graduated Non-Convexity (GNC) is shown to yield improved results. Lastly, we present results of optimization using Simulated Annealing (SA). Deepu Rajan, Subhasis Chaudhuri |
ICASSP | 1 |
| 2001 | Simultaneous Estimation of Super-Resolved Intensity and Depth Maps from Low Resolution Defocused Observations of a Scene
Deepu Rajan, Subhasis Chaudhuri |
ICCV | 1 |
| 2001 | Generalized interpolation and its application in super-resolution imaging
Deepu Rajan, Subhasis Chaudhuri |
Image Vis. Comput. | 1 |
| 2000 | A Perceptually Organized Method for Image InterpolationabstractPerception of an image depends on its visual representation. In this paper we present a perceptually organized scheme for image expansion or scattered data interpolation. This is done by decomposing the image (or data) into appropriate perceptual groups, carrying out the interpolation in individual groups as per the perceptual necessity and subsequently transforming the interpolated values back to the image domain. Various perceptual properties of the image, such as the 3D shape of an object, textural homogeneity, local variations in scene reflectivity, visual transparency can be better preserved during the interpolation process. Deepu Rajan, Subhasis Chaudhuri |
ICPR | 1 |