Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sujoy Roy

dblp:99/903 · DBLP profile ↗
← Back
41ranked-venue papers
15as first author
1since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 9 first-authorArtificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 39% Segmentation and scene understanding · 26% Face, body and person analysis · 24%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 100%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis › facial attribute analysis
facial attribute recognition
0.312018
Landmark Free Face Attribute Prediction · IEEE Trans. Image Process. 2018
Machine learning › Generative modeling
generative adversarial network
0.312018
Multi-Human Parsing Machines · ACM Multimedia 2018
Computer vision › Segmentation and scene understanding
human parsing
0.312018
Multi-Human Parsing Machines · ACM Multimedia 2018
Machine learning › Generative modeling
image generation
0.312018
Multi-Human Parsing Machines · ACM Multimedia 2018
Computer vision › Segmentation and scene understanding › human parsing
multi-human parsing
0.312018
Multi-Human Parsing Machines · ACM Multimedia 2018
Machine learning › Generative modeling › image generation
person image synthesis
0.312018
Multi-Human Parsing Machines · ACM Multimedia 2018
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.212016
Understanding Deep Representations Learned in Modeling Users Likes · IEEE Trans. Image Process. 2016
Recommender systems › content recommendation
image recommendation
0.212016
Understanding Deep Representations Learned in Modeling Users Likes · IEEE Trans. Image Process. 2016
Multimedia analysis and retrieval › video retrieval
news video retrieval
0.212013
Metadata enrichment for news video retrieval: a graph-based propagation approach · ACM Multimedia 2013
Multimedia analysis and retrieval
video retrieval
0.212013
Metadata enrichment for news video retrieval: a graph-based propagation approach · ACM Multimedia 2013
Computer vision › Face, body and person analysis › personality assessment
multimodal personality assessment
0.112012
Don't ask me what i'm like, just watch and listen · ACM Multimedia 2012
Computer vision › Face, body and person analysis
personality assessment
0.112012
Don't ask me what i'm like, just watch and listen · ACM Multimedia 2012
Multimedia analysis and retrieval › near-duplicate detection
image copy detection
0.112005
A unified framework for resolving ambiguity in copy detection · ACM Multimedia 2005
Multimedia analysis and retrieval
near-duplicate detection
0.112005
A unified framework for resolving ambiguity in copy detection · ACM Multimedia 2005
Digital forensics and information hiding
watermarking
0.112005
A unified framework for resolving ambiguity in copy detection · ACM Multimedia 2005
Mathematical optimization
nonconvex optimization
0.012005
A unified framework for resolving ambiguity in copy detection · ACM Multimedia 2005
Mathematical optimization › continuous optimization › convex optimization › conic optimization
second-order cone programming
0.012005
A unified framework for resolving ambiguity in copy detection · ACM Multimedia 2005

Methods — techniques the papers use, named apart from their topics

knowledge transfer · 0.5feature selection · 0.5deep learning · 0.5spatial transformation · 0.3group-individual push and pull loss · 0.3end-to-end learning · 0.3data generation · 0.3competitive learning · 0.3graph-based propagation · 0.2explicit semantic analysis · 0.2automatic speech recognition · 0.2second-order cone programming · 0.2AFMT feature representation · 0.2multimodal feature extraction · 0.1kernel regression · 0.1
YearPublicationVenuePosition
2025 FTBAC: fuzzy trust based access control for healthcare cross-domain environment
Sujoy Roy, Alok Kumar 0003, Udai Pratap Rao
Soft Comput.1
2019 Task Relation Networks
abstract
Multi-task learning is popular in machine learning and computer vision. In multitask learning, properly modeling task relations is important for boosting the performance of jointly learned tasks. Task covariance modeling has been successfully used to model the relations of tasks but is limited to homogeneous multi-task learning. In this paper, we propose a feature based task relation modeling approach, suitable for both homogeneous and heterogeneous multi-task learning. First, we propose a new metric to quantify the relations between tasks. Based on the quantitative metric, we then develop the task relation layer, which can be combined with any deep learning architecture to form task relation networks to fully exploit the relations of different tasks in an online fashion. Benefiting from the task relation layer, the task relation networks can better leverage the mutual information from the data. We demonstrate our proposed task relation networks are effective in improving the performance in both homogeneous and heterogeneous multi-task learning settings through extensive experiments on computer vision tasks.
Jianshu Li, Pan Zhou 0002, Yunpeng Chen, Jian Zhao 0006, Sujoy Roy, Shuicheng Yan, Jiashi Feng, Terence Sim
WACV5
2019 Evaluation of Sirtuin-3 probe quality and co-expressed genes using literature cohesion
abstract
BACKGROUND: Gene co-expression studies can provide important insights into molecular and cellular signaling pathways. The GeneNetwork database is a unique resource for co-expression analysis using data from a variety of tissues across genetically distinct inbred mice. However, extraction of biologically meaningful co-expressed gene sets is challenging due to variability in microarray platforms, probe quality, normalization methods, and confounding biological factors. In this study, we tested whether literature derived functional cohesion could be used as an objective metric in lieu of 'ground truth' to evaluate the quality of probes and microarray datasets. RESULTS: We examined Sirtuin-3 (Sirt3) co-expressed gene sets extracted from either liver or brain tissues of BXD recombinant inbred mice in the GeneNetwork database. Depending on the microarray platform, there were as many as 26 probes that targeted different regions of Sirt3 primary transcript. Co-expressed gene sets (ranging from 100-1000 genes) associated with each Sirt3 probe were evaluated using the previously developed literature-derived cohesion p-value (LPv) and benchmarked against 'gold standards' derived from proteomic studies or Gene Ontology classifications. We found that the maximal F-measure was obtained at an average window size of 535 genes. Using set size of 500 genes, the Pearson correlations between LPv and F-measure as well as between LPv and mitochondrial gene enrichment p-values were 0.90 and 0.93, respectively. Importantly, we found that the LPv approach can distinguish high quality Sirt3 probes. Analysis of the most functionally cohesive Sirt3 co-expressed gene set revealed core metabolic pathways that were shared between hippocampus and liver as well as distinct pathways which were unique to each tissue. These results are consistent with other studies that suggest Sirt3 is a key metabolic regulator and has distinct functions in energy-producing vs. energy-demanding tissues. CONCLUSIONS: Our results provide proof-of-concept that literature cohesion analysis is useful for evaluating the quality of probes and microarray datasets, particularly when experimentally derived gold standards are unavailable. Our approach would enable researchers to rapidly identify biologically meaningful co-expressed gene sets and facilitate discovery from high throughput genomic data.
Sujoy Roy, Kazi I. Zaman, Robert W. Williams, Ramin Homayouni
BMC Bioinform.1
2018 Multi-Human Parsing Machines
abstract
Human parsing is an important task in human-centric analysis. Despite the remarkable progress in single-human parsing, the more realistic case of multi-human parsing remains challenging in terms of the data and the model. Compared with the considerable number of available single-human parsing datasets, the datasets for multi-human parsing are very limited in number mainly due to the huge annotation effort required. Besides the data challenge to multi-human parsing, the persons in real-world scenarios are often entangled with each other due to close interaction and body occlusion, making it difficult to distinguish body parts from different person instances. In this paper we propose the Multi-Human Parsing Machines (MHPM) system, which contains an MHP Montage model and an MHP Solver, to address both challenges in multi-human parsing. Specifically, the MHP Montage model in MHPM generates realistic images with multiple persons together with the parsing labels. It intelligently composes single persons onto background scene images while maintaining the structural information between persons and the scene. The generated images can be used to train better multi-human parsing algorithms. On the other hand, the MHP Solver in MHPM solves the bottleneck of distinguishing multiple entangled persons with close interaction. It employs a Group-Individual Push and Pull (GIPP) loss function, which can effectively separate persons with close interaction. We experimentally show that the proposed MHPM can achieve state-of-the-art performance on the multi-human parsing benchmark and the person individualization benchmark, which distinguishes closely entangled person instances.
Jianshu Li, Jian Zhao 0006, Yunpeng Chen, Sujoy Roy, Shuicheng Yan, Jiashi Feng, Terence Sim
ACM Multimedia4
2018 'Who Likes What and, Why?' Insights into Modeling Users' Personality Based on Image 'Likes'
abstract
The increased proliferation of data production technologies (e.g., cameras) and consumption avenues (e.g., social media) has led to images and videos being utilized by users to convey innate preferences and tastes. This has opened up the possibility of using multimedia as a source for user-modeling. This work attempts to model personality traits (based on the Five Factor Theory) of users using a collection of images they tag as `favorite' (or like) on Flickr. First, a set of semantic features are proposed to be used for representing different concepts in images which influence users to like them. The addition of the proposed features led to improvement over state-of-the-art by 12 percent. Second, a novel machine learning approach is developed to model users' personality based on the image features (resulting in upto 15 percent improvement). Third, efficacy of the semantic features and the modeling approach is shown in recommending images based on personality modeling. Using the modeling approach, recommendations are made regarding the factors that might influence users with different personality traits to like an image.
Sharath Chandra Guntuku, Joey Tianyi Zhou, Sujoy Roy, Weisi Lin, Ivor W. Tsang
IEEE Trans. Affect. Comput.3
2018 Landmark Free Face Attribute Prediction
abstract
Face attribute prediction in the wild is important for many facial analysis applications yet it is very challenging due to ubiquitous face variations. In this paper, we address face attribute prediction in the wild by proposing a novel method, lAndmark Free Face AttrIbute pRediction (AFFAIR). Unlike traditional face attribute prediction methods that require facial landmark detection and face alignment, AFFAIR uses an endto- end learning pipeline to jointly learn a hierarchy of spatial transformations that optimize facial attribute prediction with no reliance on landmark annotations or pre-trained landmark detectors. AFFAIR achieves this through simultaneously 1) learning a global transformation which effectively alleviates negative effect of global face variation for the following attribute prediction tailored for each face, 2) locating the most relevant facial part for attribute prediction and 3) aggregating the global and local features for robust attribute prediction. Within AFFAIR, a new competitive learning strategy is developed that effectively enhances global transformation learning for better attribute prediction. We show that with zero information about landmarks, AFFAIR achieves state-of-the-art performance on three face attribute prediction benchmarks, which simultaneously learns the face-level transformation and attribute-level localization within a unified framework.
Jianshu Li, Fang Zhao 0006, Jiashi Feng, Sujoy Roy, Shuicheng Yan, Terence Sim
IEEE Trans. Image Process.4
2016 Happiness level prediction with sequential inputs via multiple regressions
abstract
This paper presents our solution submitted to the Emotion Recognition in the Wild (EmotiW 2016) group-level happiness intensity prediction sub-challenge. The objective of this sub-challenge is to predict the overall happiness level given an image of a group of people in a natural setting. We note that both the global setting and the faces of the individuals in the image influence the group-level happiness intensity of the image. Hence the challenge lies in building a solution that incorporates both these factors and also considers their right combination. Our proposed solution incorporates both these factors as a combination of global and local information. We use a convolutional neural network to extract discriminative face features, and a recurrent neural network to selectively memorize the important features to perform the group-level happiness prediction task. Experimental evaluations show promising performance improvements, resulting in Root Mean Square Error (RMSE) reduction of about 0.5 units on the test set compared to the baseline algorithm that uses only global information.
Jianshu Li, Sujoy Roy, Jiashi Feng, Terence Sim
ICMI2
2016 Personalizing User Interfaces for improving quality of experience in VoD recommender systems
abstract
Recommending content to users involves understanding a) what to present and b) how to present them, so as to increase quality of experience (QoE) and thereby, content consumption. This work attempts to address the question of how to present contents in a way so that the user finds it easy to get to desired content. While the process of User Interface (UI) design is dependent on several human factors, there are basic design components and their combination that have to be common to any recommender system user interface. Personalization of the UI design process involves picking the right components and their combination, and presenting a UI to suit the usage behavior of an individual user, so as to enhance the QoE. This work proposes a system that learns from a user's content consumption patterns and makes some recommendations regarding how to present the content for the user (in the context of Video-On-Demand/Live-TV services on Computer displays), so as to enhance the QoE of the recommender system.
Sharath Chandra Guntuku, Sujoy Roy, Weisi Lin, Kelvin Ng, Wee Keong Ng, Vinit Jakhetiya
QoMEX2
2016 Latent Factor Representations for Cold-Start Video Recommendation
abstract
Recommending items that have rarely/never been viewed by users is a bottleneck for collaborative filtering (CF) based recommendation algorithms. To alleviate this problem, item content representation (mostly in textual form) has been used as auxiliary information for learning latent factor representations. In this work we present a novel method for learning latent factor representation for videos based on modelling the emotional connection between user and item. First of all we present a comparative analysis of state-of-the art emotion modelling approaches that brings out a surprising finding regarding the efficacy of latent factor representations in modelling emotion in video content. Based on this finding we present a method visual-CLiMF for learning latent factor representations for cold start videos based on implicit feedback. Visual-CLiMF is based on the popular collaborative less-is-more approach but demonstrates how emotional aspects of items could be used as auxiliary information to improve MRR performance. Experiments on a new data set and the Amazon products data set demonstrate the effectiveness of visual-CLiMF which outperforms existing CF methods with or without content information.
Sujoy Roy, Sharath Chandra Guntuku
RecSys1
2016 Prioritization, clustering and functional annotation of MicroRNAs using latent semantic indexing of MEDLINE abstracts
abstract
BACKGROUND: The amount of scientific information about MicroRNAs (miRNAs) is growing exponentially, making it difficult for researchers to interpret experimental results. In this study, we present an automated text mining approach using Latent Semantic Indexing (LSI) for prioritization, clustering and functional annotation of miRNAs. RESULTS: For approximately 900 human miRNAs indexed in miRBase, text documents were created by concatenating titles and abstracts of MEDLINE citations which refer to the miRNAs. The documents were parsed and a weighted term-by-miRNA frequency matrix was created, which was subsequently factorized via singular value decomposition to extract pair-wise cosine values between the term (keyword) and miRNA vectors in reduced rank semantic space. LSI enables derivation of both explicit and implicit associations between entities based on word usage patterns. Using miR2Disease as a gold standard, we found that LSI identified keyword-to-miRNA relationships with high accuracy. In addition, we demonstrate that pair-wise associations between miRNAs can be used to group them into categories which are functionally aligned. Finally, term ranking by querying the LSI space with a group of miRNAs enabled annotation of the clusters with functionally related terms. CONCLUSIONS: LSI modeling of MEDLINE abstracts provides a robust and automated method for miRNA related knowledge discovery. The latest collection of miRNA abstracts and LSI model can be accessed through the web tool miRNA Literature Network (miRLiN) at http://bioinfo.memphis.edu/mirlin .
Sujoy Roy, Brandon C. Curry, Behrouz Madahian, Ramin Homayouni
BMC Bioinform.1
2016 Understanding Deep Representations Learned in Modeling Users Likes
abstract
Automatically understanding and discriminating different users' liking for an image is a challenging problem. This is because the relationship between image features (even semantic ones extracted by existing tools, viz., faces, objects, and so on) and users' likes is non-linear, influenced by several subtle factors. This paper presents a deep bi-modal knowledge representation of images based on their visual content and associated tags (text). A mapping step between the different levels of visual and textual representations allows for the transfer of semantic knowledge between the two modalities. Feature selection is applied before learning deep representation to identify the important features for a user to like an image. The proposed representation is shown to be effective in discriminating users based on images they like and also in recommending images that a given user likes, outperforming the state-of-the-art feature representations by ∼ 15 %-20%. Beyond this test-set performance, an attempt is made to qualitatively understand the representations learned by the deep architecture used to model user likes.
Sharath Chandra Guntuku, Joey Tianyi Zhou, Sujoy Roy, Weisi Lin, Ivor W. Tsang
IEEE Trans. Image Process.3
2015 Evaluating visual and textual features for predicting user 'likes'
abstract
Computationally modeling users `liking' for image(s) requires understanding how to effectively represent the image so that different factors influencing user `likes' are considered. In this work, an evaluation of the state-of-the-art visual features in multimedia understanding at the task of predicting user `likes' is presented, based on a collection of images crawled from Flickr. Secondly, a probabilistic approach for modeling `likes' based only on tags is proposed. The approach of using both visual and text-based features is shown to improve the state-of-the-art performance by 12%. Analysis of the results indicate that more human-interpretable and semantic representations are important for the task of predicting very subtle response of `likes'.
Sharath Chandra Guntuku, Sujoy Roy, Weisi Lin
ICME2
2015 Personality Modeling Based Image Recommendation
Sharath Chandra Guntuku, Sujoy Roy, Weisi Lin
MMM (2)2
2015 A Bayesian approach for inducing sparsity in generalized linear models with multi-category response
abstract
BACKGROUND: The dimension and complexity of high-throughput gene expression data create many challenges for downstream analysis. Several approaches exist to reduce the number of variables with respect to small sample sizes. In this study, we utilized the Generalized Double Pareto (GDP) prior to induce sparsity in a Bayesian Generalized Linear Model (GLM) setting. The approach was evaluated using a publicly available microarray dataset containing 99 samples corresponding to four different prostate cancer subtypes. RESULTS: A hierarchical Sparse Bayesian GLM using GDP prior (SBGG) was developed to take into account the progressive nature of the response variable. We obtained an average overall classification accuracy between 82.5% and 94%, which was higher than Support Vector Machine, Random Forest or a Sparse Bayesian GLM using double exponential priors. Additionally, SBGG outperforms the other 3 methods in correctly identifying pre-metastatic stages of cancer progression, which can prove extremely valuable for therapeutic and diagnostic purposes. Importantly, using Geneset Cohesion Analysis Tool, we found that the top 100 genes produced by SBGG had an average functional cohesion p-value of 2.0E-4 compared to 0.007 to 0.131 produced by the other methods. CONCLUSIONS: Using GDP in a Bayesian GLM model applied to cancer progression data results in better subclass prediction. In particular, the method identifies pre-metastatic stages of prostate cancer with substantially better accuracy and produces more functionally relevant gene sets.
Behrouz Madahian, Sujoy Roy, Dale Bowman, Lih-Yuan Deng, Ramin Homayouni
BMC Bioinform.2
2014 Deep Representations to Model User 'Likes'
Sharath Chandra Guntuku, Joey Tianyi Zhou, Sujoy Roy, Weisi Lin, Ivor W. Tsang
ACCV (1)3
2014 Utilizing 3D flow of points for facial expression recognition
Ruchir Srivastava, Sujoy Roy
Multim. Tools Appl.2
2013 RGB-D video content identification
abstract
This paper proposes the first content identification (ID) system for depth video as well as a first hybrid content ID system for synchronized RGB and depth (RGB-D) video. The proposed systems are tested on a public RGB-D dataset. The hybrid system demonstrates significant performance gains over RGB-alone or depth-alone systems, while depth and RGB perform comparably. Moreover, a statistical interpretation of the hybrid system's superior performance is provided.
Honghai Yu, Pierre Moulin, Sujoy Roy
ICASSP3
2013 Devanagari Character Recognition in Scene Images
abstract
Character recognition in scene images is an extremely challenging task. Although several techniques are reported performing well, they pertain to English only. This paper focuses on Devanagari character recognition from scene images. Devanagari script is very popular language and has very typical characteristics different from other scripts, particularly English. Combination of basic Devanagari consonants and vowels in multi-variegated ways can yield as many as 100s of characters. Building a classifier to recognize all these classes will be a difficult task. To alleviate this problem, a novel part-based model technique is proposed. 40 basic classes were identified from the Devanagari script for the same purpose. The technique was proposed so as to classify an instance of one these classes in any given test sample. Procuring a large dataset for training is not feasible in the case of scene images. To simultaneously solve this problem, we developed our technique that can use either the machine printed or the handwritten dataset for training. We present our results on the publicly available dataset (DSIW2K) containing images of street scenes taken in New Delhi, India.
Vipin Narang, Sujoy Roy, O. V. Ramana Murthy, Madasu Hanmandlu
ICDAR2
2013 Metadata enrichment for news video retrieval: a graph-based propagation approach
abstract
This paper summarizes our contribution to the Technicolor Rich Multimedia Retrieval from Input Videos Grand Challenge. We hold the view that semantic analysis of a given news video is best performed in the text domain. Starting with a noisy text obtained from applying Automatic Speech Recognition (ASR), a graph-based approach is then used to enrich the text by propagating labels from visually similar videos culled from parallel (YouTube) News sources. From the enriched text, we next extract salient keywords to form a query to a news video search engine, retrieving a larger corpus of related news video. Compared to a baseline method that only uses the ASR text, significant improvement in precision has been obtained, indicating that retrieval has benefited from the ingestion of the external labels. Capitalizing on the enriched metadata, we find that videos are more amenable to the Wikipedia-based Explicit Semantic Analysis (ESA), resulting in better support for subtopic news video retrieval. We apply our methods to an in-house live news search portal, and report on several best practices.
Kong-Wah Wan, Weiyun Yau, Sujoy Roy
ACM Multimedia3
2012 Recognizing emotions of characters in movies
abstract
This work presents an investigation into recognizing emotions of people in near real life scenarios. Most existing studies on recognizing emotions of people have been conducted under controlled environments where the emotions are not spontaneous, rather highly exaggerated, and the number of modalities considered and their interactions is limited. The proposed bimodal approach fuses facial expression recognition (FER) with the “semantic orientation” of dialogs of actors to identify emotions under difficult illumination conditions, pose variations and occlusions in scenes. Experiments conducted on a dataset of 700 video clips from 17 movies demonstrate that the proposed fusion approach improves emotion recognition performance over unimodal approaches.
Ruchir Srivastava, Shuicheng Yan, Terence Sim, Sujoy Roy
ICASSP4
2012 Don't ask me what i'm like, just watch and listen
abstract
Traditional (based on psychology) approaches for personality assessment of an individual require him/her to fill up a questionnaire. This paper presents a novel way of utilizing multimodal cues to automatically fill up the questionnaire. The contributions of this work are three-fold. (1) Novel psychology-based audio/visual/lexical features are proposed and shown to be effective in predicting answers to a personality questionnaire, Big-Five Inventory-10 (BFI- 10). (2) Extracted features are used to learn linear and kernel versions of a novel regression model, 'SLoT', to automatically predict BFI-10 answers. The model is based on Sparse and Low-rank Transformation (SLoT). (3) Predicted answers are used to compute personality scores using standard BFI-10 scoring scheme. We evaluated our approach on a dataset of 3907 clips (for 50 characters from movies of diverse genres) manually labeled with BFI-10 answers and personality scores as ground-truth. Experiments indicate that the proposed 'SLoT' model effectively automates the answering process by emulating human understanding. We also conclude that predicting personality scores through predicting answers first is better than directly predicting scores based on audio/visual features (as studied in state-of-the art methods).
Ruchir Srivastava, Jiashi Feng, Sujoy Roy, Shuicheng Yan, Terence Sim
ACM Multimedia3
2012 OS-Guard: on-site signature based framework for multimedia surveillance data management
Praveen Kumar 0005, Sujoy Roy, Ankush Mittal
Multim. Tools Appl.2
2011 Accumulated motion images for facial expression recognition in videos
abstract
This paper details the method and experiments conducted towards our submission to the FERA 2011 facial expression recognition benchmarking evaluations. The benchmarking evaluation task involves recognizing 5 emotion classes in videos. Our method for detecting facial expressions is a fusion of the decisions of two FER approaches based on two different feature representations, namely using motion information from facial regions and facial feature point displacement information. The main observation motivating the approach we took is that different feature representations are discriminative in detecting different facial expressions. Hence a fusion approach could complement each other to improve recognition performance. Experiments were conducted on the GEMEP-FERA data set provided by the organizers.
Ruchir Srivastava, Sujoy Roy, Shuicheng Yan, Terence Sim
FG2
2011 Wikipedia Based News Video Topic Modeling for Information Extraction
Sujoy Roy, Mun-Thye Mak, Kong-Wah Wan
MMM (2)1
2011 Multi-actor Emotion Recognition in Movies Using a Bimodal Approach
Ruchir Srivastava, Sujoy Roy, Shuicheng Yan, Terence Sim
MMM (2)2
2011 Latent Semantic Indexing of PubMed abstracts for identification of transcription factor candidates from microarray derived gene sets
abstract
BACKGROUND: Identification of transcription factors (TFs) responsible for modulation of differentially expressed genes is a key step in deducing gene regulatory pathways. Most current methods identify TFs by searching for presence of DNA binding motifs in the promoter regions of co-regulated genes. However, this strategy may not always be useful as presence of a motif does not necessarily imply a regulatory role. Conversely, motif presence may not be required for a TF to regulate a set of genes. Therefore, it is imperative to include functional (biochemical and molecular) associations, such as those found in the biomedical literature, into algorithms for identification of putative regulatory TFs that might be explicitly or implicitly linked to the genes under investigation. RESULTS: In this study, we present a Latent Semantic Indexing (LSI) based text mining approach for identification and ranking of putative regulatory TFs from microarray derived differentially expressed genes (DEGs). Two LSI models were built using different term weighting schemes to devise pair-wise similarities between 21,027 mouse genes annotated in the Entrez Gene repository. Amongst these genes, 433 were designated TFs in the TRANSFAC database. The LSI derived TF-to-gene similarities were used to calculate TF literature enrichment p-values and rank the TFs for a given set of genes. We evaluated our approach using five different publicly available microarray datasets focusing on TFs Rel, Stat6, Ddit3, Stat5 and Nfic. In addition, for each of the datasets, we constructed gold standard TFs known to be functionally relevant to the study in question. Receiver Operating Characteristics (ROC) curves showed that the log-entropy LSI model outperformed the tf-normal LSI model and a benchmark co-occurrence based method for four out of five datasets, as well as motif searching approaches, in identifying putative TFs. CONCLUSIONS: Our results suggest that our LSI based text mining approach can complement existing approaches used in systems biology research to decipher gene regulatory networks by providing putative lists of ranked TFs that might be explicitly or implicitly associated with sets of DEGs derived from microarray experiments. In addition, unlike motif searching approaches, LSI based approaches can reveal TFs that may indirectly regulate genes.
Sujoy Roy, Kevin Heinrich, Vinhthuy T. Phan, Michael W. Berry, Ramin Homayouni
BMC Bioinform.1
2010 Identifying and learning visual attributes for object recognition
abstract
We propose an attribute centric approach for visual object recognition. The attributes of an object are the observable visual properties that help to uniquely describe it. We present methods for identifying and learning these object attributes. To identify suitable object attributes, we process the corresponding Wikipedia pages to select terms that not only have high occurrence frequency, the images of these concepts must also be visually consistent. To learn object attributes, we assume prior knowledge of the object class-specific distributions of patches over the attributes, and introduce a novel algorithm that iteratively refines these distributions by a nearest-neighbor attribute classifier. Given an unseen image, its attribute vector is first formed by the distribution of patches over the attributes, and its final class is then determined by the attribute representation. We report efficacy of the proposed framework on an animal data set of ten classes, where the test set consists of images collected from the web.
Kong-Wah Wan, Sujoy Roy
ICIP2
2010 Rotation invariant Facial Expression Recognition in image sequences
abstract
Facial Expression Recognition has mostly been done on frontal or near frontal faces. However, most of the faces in real life are non-frontal. This paper deals with in-plane rotation of faces in image sequences and considers the six universal facial expressions. The proposed approach does not need to rotate the image to frontal position. FER by rotating images to frontal is sensitive to determination of rotation angle and can involve errors in tracking facial points. Directions of motion of Facial Feature Points (FFPs) is used for feature extraction. In training for six expressions, Gaussian Mixture Models are fit to the distribution of angles representing these directions of motion. These models are used for further classification of test sequences using SVM. Gaussian Mixture Modeling is experimentally found to be robust to errors in position of FFPs. For dimensionality reduction, feature selection is performed using Fisher ratio test.
Ruchir Srivastava, Sujoy Roy, Terence Sim
ICME2
2009 A Latent Model for Visual Disambiguation of Keyword-based Image Search
abstract
The problem of polysemy in keyword-based image search arises mainly from the inherent ambiguity in user queries. We propose a latent model based approach that resolves user search ambiguity by allowing sense specific diversity in search results. Given a query keyword and the images retrieved by issuing the query to an image search engine, we first learn a latent visual sense model of these polysemous images. Next, we use Wikipedia to disambiguate the word sense of the original query, and issue these Wiki-senses as new queries to retrieve sense specific images. A sense-specific image classifier is then learnt by combining information from the latent visual sense model, and used to cluster and re-rank the polysemous images from the original query keyword into its specific senses. Results on a ground truth of 17K image set returned by 10 keyword searches and their 62 word senses provides empirical indications that our method can improve upon existing keyword based search engines. Our method learns the visual word sense models in a totally unsupervised manner, effectively filters out irrelevant images, and is able to mine the long tail of image search.
Kong-Wah Wan, Ah-Hwee Tan, Joo-Hwee Lim, Liang-Tien Chia, Sujoy Roy
BMVC5
2009 LSI based framework to predict gene regulatory information
Sujoy Roy, Lijing Xu, Ramin Homayouni
BMC Bioinform.1
2008 On the security of non-forgeable robust hash functions
abstract
In many applications, it is often desirable to extract a consistent key from a multimedia object (e.g., an image), even when the object has gone through a noisy channel. For example, the extracted key can be used to generate content dependent watermarks to mitigate copy attacks, or for two or more parties to establish a session key from their noisy versions of the same object. Robust hash functions are useful in extracting such consistent keys. It differs from cryptographic hash functions in that small noise in the messages would yield the same hash value with high probability. However, the security of robust hash functions is not well understood. In this paper, we study different security notions of robust hash functions w.r.t. forgery attacks, where the goal of the attacker is to estimate the key (hash value) extracted from a given message. We show that information- theoretical security against forgery under chosen message attacks is not possible, in the sense that given enough number of observations of message/hash pairs, the entropy of the hash value of another message can be reduced arbitrarily. We further give a construction that is computationally secure, where computing the hash value can still be computationally infeasible even its entropy may not be high.
Sujoy Roy
ICIP2
2008 Performance analysis of locality preserving image hash
abstract
Bit extraction is an essential component of an image hashing system. A good bit extraction scheme should preserve the performance achieved at the feature representation level. In other words the robustness discrimination tradeoff measured by ROC analysis should be preserved. This is dependent on several factors such as, the encountered noise and the number of bits that can be extracted per sample. This paper investigates the relationship between these parameters and proposes some theoretical bounds in achieving a good tradeoff. The analysis primarily focuses on the bit extraction method proposed in [1] and its performance is compared with a scalar quantization based hashing method.
Sujoy Roy, Qibin Sun, Ton Kalker
ICIP1
2007 Robust Hash for Detecting and Localizing Image Tampering
abstract
An image hash should be (1) robust to allowable operations and (2) sensitive to illegal manipulations and distinct queries. Some applications also require the hash to be able to localize image tampering. This requires the hash to contain both robust content and alignment information to meet the above criterion. Fulfilling this is difficult because of two contradictory requirements. First, the hash should be small and second, to verify authenticity and then localize tampering, the amount of information in the hash about the original required would be large. Hence a tradeoff between these requirements needs to be found. This paper presents an image hashing method that addresses this concern, to not only detect but also localize tampering using a small signature (< 1kB). Illustrative experiments bring out the efficacy of the proposed method compared to existing methods.
Sujoy Roy, Qibin Sun
ICIP (6)1
2007 On preserving robustness-false alarm tradeoff in media hashing
abstract
This paper discusses one of the important issues in generating a robust media hash. Robustness of a media hashing algorithm is primarily determined by three factors, (1) robustness-false alarm tradeoff achieved by the chosen feature representation, (2) accuracy of the bit extraction step and (3) the distance measure used to measure similarity (dissimilarity) between two hashes. The robustness-false alarm tradeoff in feature space is measured by a similarity (dissimilarity) measure and it defines a limit on the performance of the hashing algorithm. The distance measure used to compute the distance between the hashes determines how far this tradeoff in the feature space is preserved through the bit extraction step. Hence the bit extraction step is crucial, in defining the robustness of a hashing algorithm. Although this is recognized as an important requirement by all, to our knowledge there is no work in the existing literature that elucidates the effcacy of their algorithm based on their effectiveness in improving this tradeoff compared to other methods. This paper specifically demonstrates the kind of robustness false alarm tradeoff achieved by existing methods and proposes a method for hashing that clearly improves this tradeoff.
Sujoy Roy, J. Yuan, E.-C. Chang
VCIP1
2007 A novel framework for improving bandwidth utilization for VBR video delivery over wide-area networks
abstract
There are great challenges in streaming variable-bit-rate video over wide-area networks due to the significant variation of network conditions. The utilization of the precious bandwidth of wide-area networks is often low in such streaming systems. In this paper, we propose a novel framework to improve the bandwidth utilization from a new perspective. Instead of focusing on the performance of each single media stream, we aim to improve the overall bandwidth utilization for video streaming systems. We try to exploit the unoccupied bandwidth in ongoing streams and using it to deliver some prefetched data which can be used to facilitate future streaming. Preliminary results show that our mechanism has great potential to improve both the overall bandwidth utilization and the caching performance of the proxy servers in the streaming systems.
Junli Yuan, Sujoy Roy, Qibin Sun
VCIP2
2006 In Search of Optimal Codes for DNA Computing
Max H. Garzon, Vinhthuy T. Phan, Sujoy Roy, Andrew Neel
DNA3
2005 A unified framework for resolving ambiguity in copy detection
abstract
Copy detection is an important component of digital rights management and can be implemented using a retrieval-based approach. Under this approach, a query image, suspected to be a copy, is compared against all the images in the owner database. The comparison is done based on a distance metric in feature space. The performance of such a system depends on the mutual separation of the feature representation of the images in the database. In this paper we propose a framework that increases this mutual separation by literally shifting them away from each other. The idea of modifying the features derives its inspiration from the field of watermarking. It is also important to make sure that the semantics of the images do not change after modification. Thus the focus of this paper is on how to modify the images in the database, so that the mutual separation between the images in feature space is above a certain threshold and the distortion induced is minimized. This problem can be formulated as a non-convex optimization problem which is difficult to solve. We propose a restriction of the problem and solve it using second-order cone programming. We present a practical implementation of our framework, named RAM, which uses AFMT as the feature representation. We conduct experiments to test the performance of RAM.
Sujoy Roy, Ee-Chien Chang, K. Natarajan
ACM Multimedia1
2004 Watermarking color histograms
abstract
In this paper we give a method for watermarking color histograms. Color histograms have been known M. J. Swain et al., (1991) to be robust to rotations and other geometric transformations. If the watermark can be embedded in such geometry invariant representations it should survive geometric transformations. The difficulty in watermarking color histograms is that they have a nonlinear relationship with the pixel representation. Therefore it is not clear how to get a watermarked image given its watermarked histogram. We give a method for watermarking color histograms that uses earth mover distance (EMD) to modify an image to a target histogram. We conduct extensive experiments to test our method.
Sujoy Roy, Ee-Chien Chang
ICIP1
2004 Watermarking with retrieval systems
Sujoy Roy, Ee-Chien Chang
Multim. Syst.1
2003 Watermarking with knowledge of image database
abstract
The goal of this paper is to study how a-prior knowledge of the image database could be exploited for better watermarking performance. Unlike most formulations, where the encoder and detector only know the distribution of the images, under our formulation, the actual set of images to be watermarked are known, either in a static or dynamic setting. To achieve better performance, instead of choosing a random watermarking key or predefined code-book as is the usual practice, we derive the watermarking keys from the database. We study two settings, static and dynamic. In the dynamic setting, the image database starts from a single image and grows as more images arrive. Thus the watermarking keys have to be updated frequently. This setting can be applied to applications where the detector has access to the Internet. To demonstrate the main idea, we extend a variant of spread-spectrum method to a few schemes, and analyze their performance. Interestingly, the requirements on false-alarm, robustness and distortion can be traded-off with the size of the watermarking keys. We perform our experiments on both natural images and Gaussian source. Our analysis and experiments show promising improvement in performance by exploiting the a-prior knowledge of the image database, specifically for fixed robustness and false alarm we achieve significant reduction of distortion. Similar idea can be incorporated into other watermarking methods.
Sujoy Roy, Ee-Chien Chang
ICIP (2)1
2002 Region-based image registration for wide-baseline stereo
abstract
In this paper, we present a novel system for image registration between disparate views (wide baseline stereo), using 2D regions as correspondence features. Our system consists of three stages. First, we use a region segmentation technique to find reasonably good regions in image pairs. Second, an adjacency relationship based region matching technique of time complexity 0(min(mk/sup 2//p,nk/sup 2//p)) is proposed to match regions across image pairs, where m and n denote the number of regions in each image respectively, k denotes the maximum number of neighbors of a region and p denotes the number of matches out of those k neighbors. Lastly, a new technique for finding the affine transformation between regions based on minimum enclosing ellipses is presented. The efficacy of our approach on real image pairs subject to large perspective distortions is demonstrated.
Sujoy Roy, Sanjeev Kapoor
ICIP (1)1