VLDB 2026 Research / reviewers in the wild / expert
Tanveer A. Faruquie
dblp:53/1182 · also Tanveer Afzal Faruquie
· DBLP profile ↗
36ranked-venue papers
10as first author
4since 2021 · last 2024
0009-0008-9474-7928ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 15 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Machine Learning in FinanceabstractThis workshop aims to explore the intersection of Generative AI with the rich tapestry of financial data types, seeking to uncover new methodologies and techniques that can enhance predictive analytics, fraud detection, and customer insights across the sector. By harnessing these advancements in AI, we can pave the way to not only understand customer behavior but also anticipate their needs more effectively, leading to superior customer outcomes and more personalized services. Our objective is to shed light on the challenges and opportunities presented by the diverse data formats in finance. We aim to bridge the gap between the dominance of traditional models for tabular data analysis and the emerging potential of Generative AI to revolutionize the treatment of time series, click streams, and other unstructured data forms. Leman Akoglu, Nitesh V. Chawla, Josep Domingo-Ferrer, Eren Kurshan, Senthil Kumar, Vidyut M. Naware, José A. Rodríguez-Serrano, Isha Chaturvedi, Saurabh Nagrecha, Mahashweta Das, Tanveer A. Faruquie |
KDD | 11 |
| 2023 | KDD Workshop on Machine Learning in FinanceabstractThe finance industry is constantly faced with an ever evolving set of challenges including credit card fraud, identity theft, network intrusion, money laundering, human trafficking, and illegal sales of firearms. There is also the newly emerging threat of fake news in financial media that can lead to distortions in trading strategies and investment decisions. In addition, traditional problems such as customer analytics, forecasting, and recommendations take on a unique flavor when applied to financial data. A number of new ideas are emerging to tackle all these problems including self-supervised learning methods, deep learning algorithms, network/graph based solutions as well as linguistic approaches. These methods must often be able to work in real-time and be able handle large volumes of data. The purpose of this workshop is to bring together researchers and practitioners to discuss both the problems faced by the financial industry and potential solutions. We plan to invite regular papers, positional papers and extended abstracts of work in progress. We will also encourage short papers from financial industry practitioners that introduce domain specific problems and challenges to academic researchers. Leman Akoglu, Nitesh V. Chawla, Senthil Kumar, Saurabh Nagrecha, Mahashweta Das, Vidyut M. Naware, Tanveer A. Faruquie |
KDD | 7 |
| 2022 | KDD Workshop on Machine Learning in FinanceabstractThe finance industry is constantly faced with an ever evolving set of challenges including credit card fraud, identity theft, network intrusion, money laundering, human trafficking, and illegal sales of firearms. There is also the newly emerging threat of fake news in financial media that can lead to distortions in trading strategies and investment decisions. In addition, traditional problems such as customer analytics, forecasting, and recommendations take on a unique flavor when applied to financial data. A number of new ideas are emerging to tackle all these problems including semi-supervised learning methods, deep learning algorithms, network/graph based solutions as well as linguistic approaches. These methods must often be able to work in real-time and be able handle large volumes of data. The purpose of this workshop is to bring together researchers and practitioners to discuss both the problems faced by the financial industry and potential solutions. We plan to invite regular papers, positional papers and extended abstracts of work in progress. We will also encourage short papers from financial industry practitioners that introduce domain specific problems and challenges to academic researchers. Senthil Kumar, Leman Akoglu, Nitesh V. Chawla, Saurabh Nagrecha, Vidyut M. Naware, Tanveer A. Faruquie, Hays 'Skip' McCormick |
KDD | 6 |
| 2021 | Machine Learning in FinanceabstractThe finance industry is constantly faced with an ever evolving set of challenges including credit card fraud, identity theft, network intrusion, money laundering, human trafficking, and illegal sales of firearms. There are also newly emerging threats such as fake news in financial media that can lead to distortions in trading strategies and investment decisions. In addition, traditional problems such as customer analytics, forecasting, and recommendations take on a unique flavor when applied to financial data. A number of new ideas are emerging to tackle all these problems including semi-supervised learning methods, deep learning algorithms, network/graph based solutions as well as linguistic approaches. These methods must often be able to work in real-time and be able handle large volumes of data. The purpose of this workshop is to bring together researchers and practitioners to discuss both the problems faced by the financial industry and potential solutions. We have invited regular papers, positional papers and extended abstracts of work in progress. We have also encouraged short papers from financial industry practitioners that introduce domain specific problems and challenges to academic researchers. This event is the fourth in a sequence of finance related workshops we have organized at KDD since 2017. Senthil Kumar, Leman Akoglu, Nitesh V. Chawla, José A. Rodríguez-Serrano, Tanveer A. Faruquie, Saurabh Nagrecha |
KDD | 5 |
| 2015 | Tracking Political Elections on Social Media: Applications and Experience
Danish Contractor, Bhupesh Chawda, Sameep Mehta, L. Venkata Subramaniam, Tanveer A. Faruquie |
IJCAI | 5 |
| 2014 | Processing Interval Joins On Map-ReduceabstractIn this paper we investigate the problem of processing multi-way interval joins on map-reduce platform. We look at join queries formed by interval predicates as defined by Allen’s interval algebra. These predicates can be classified in two groups: colocation based predicates and sequence based predicates. A colocation predicate requires two intervals to share at least one common point while a sequence predi-cate requires two intervals to be disjoint. An interval join query can therefore be thought of as belonging to one of the three classes: (a) queries containing only colocation based predicates, (b) queries containing only sequence based pred-icates and (c) queries containing both classes of predicates. We address these three classes of join queries, discuss the challenges and present novel approaches for processing these queries on map-reduce platform. We also discuss why the current approaches developed for handling join queries on real-valued data can not be directly used to handle inter-val joins. We finally extend the approaches developed to handle join queries containing multiple interval attributes as well as join queries containing both interval as well as non-interval attributes. Through experimental evaluations both on synthetic and real life datasets, we demonstrate that the proposed approaches comfortably outperform naive ap-proaches. 1. Bhupesh Chawda, Sumit Negi, Tanveer A. Faruquie, L. Venkata Subramaniam, Mukesh K. Mohania |
EDBT | 4 |
| 2013 | Processing multi-way spatial joins on map-reduceabstractIn this paper we investigate the problem of processing multi-way spatial joins on map-reduce platform. We look at two common spatial predicates - overlap and range. We address these two classes of join queries, discuss the challenges and outline novel approaches for executing these queries on a map-reduce framework. We then discuss how we can process join queries involving both overlap and range predicates. Specifically we present a Controlled-Replicate framework using which we design the approaches presented in this paper. The Controlled-Replicate framework is carefully engineered to minimize the communication among cluster nodes. Through experimental evaluations we discuss the complexity of the problem under investigation, details of Controlled-Replicate framework and demonstrate that the proposed approaches comfortably outperform naive approaches. Bhupesh Chawda, Sumit Negi, Tanveer A. Faruquie, L. Venkata Subramaniam, Mukesh K. Mohania |
EDBT | 4 |
| 2013 | Automating pattern discovery for rule based data standardization systemsabstractData quality is a perennial problem for many enterprise data assets. To improve data quality, businesses often employ rule based data standardization systems in which domain experts code rules for handling important and prevalent patterns. Finding these patterns is laborious and time consuming, particularly for noisy or highly specialized data sets. It is also subjective to the persons determining these patterns. In this paper we present a tool to automatically mine patterns that can help in improving the efficiency and effectiveness of these data standardization systems. The automatically extracted patterns are used by the domain and knowledge experts for rule writing. We use a greedy algorithm to extract patterns that result in a maximal coverage of data. We further group the extracted patterns such that each group represents patterns that capture similar domain knowledge. We propose a similarity measure that uses input pattern semantics to group these patterns. We demonstrate the effectiveness of our method for standardization tasks on three real world datasets. Snigdha Chaturvedi, K. Hima Prasad, Tanveer A. Faruquie, Bhupesh Chawda, L. Venkata Subramaniam, Raghu Krishnapuram |
ICDE | 3 |
| 2012 | Discovering Activities and Their Temporal SignificanceabstractIn this paper, we address the problem of discovering activities and their temporal significance in an area under surveillance. Discovering activities along with its expectation of occurrence at a particular time plays an important role in many surveillance applications. We propose an unsupervised model, called Time pLSA model, that extends the probabilistic Latent Semantic Analysis (pLSA) model to jointly capture the activities and their behaviour over time. We use adaptive background subtraction to detect spatio-temporal patches, which are used as feature representation for activity patterns. Each of these patches are associated with the time slot in which they occur. Multinomial distributions are used to model both activities as distribution over spatio-temporal patches and time significance as distribution over the time-line. We demonstrate the effectiveness of our approach on a real life surveillance feed of an outdoor scene. Ayesha Choudhary, Tanveer A. Faruquie, Subhashis Banerjee, Santanu Chaudhury |
AVSS | 2 |
| 2012 | Unsupervised Discovery of Activities and Their Temporal BehaviourabstractThis paper addresses the problem of discovering activities and their temporal significance in surveillance videos in an unsupervised manner. We propose a generative model that can jointly capture the activities and their behaviour over time. We use multinomial distribution over local motion features to model activities and a mixture distribution over their time stamps to capture the multi-modal temporal distribution of these activities. We give a Gibbs sampling algorithm to infer the parameters of the model. We demonstrate the effectiveness of our approach on real life surveillance feed of outdoor scenes. Tanveer A. Faruquie, Subhashis Banerjee, Prem Kumar Kalra |
AVSS | 1 |
| 2012 | Fusing biographical and biometric classifiers for improved person identification
Vivek Tyagi, Hima P. Karanam, Tanveer A. Faruquie, L. Venkata Subramaniam, Nalini K. Ratha |
ICPR | 3 |
| 2012 | Topic Models with Relational Features for Drug Design
Tanveer A. Faruquie, Ashwin Srinivasan 0001, Ross D. King |
ILP | 1 |
| 2012 | Using content and interactions for discovering communities in social networksabstractIn recent years, social networking sites have not only enabled people to connect with each other using social links but have also allowed them to share, communicate and interact over diverse geographical regions. Social network provide a rich source of heterogeneous data which can be exploited to discover previously unknown relationships and interests among groups of people. In this paper, we address the problem of discovering topically meaningful communities from a social network. We assume that a persons' membership in a community is conditioned on its social relationship, the type of interaction and the information communicated with other members of that community. We propose generative models that can discover communities based on the discussed topics, interaction types and the social connections among people. In our models a person can belong to multiple communities and a community can participate in multiple topics. This allows us to discover both community interests and user interests based on the information and linked associations. We demonstrate the effectiveness of our model on two real word data sets and show that it performs better than existing community discovery models. Mrinmaya Sachan, Danish Contractor, Tanveer A. Faruquie, L. Venkata Subramaniam |
WWW | 3 |
| 2012 | Data and task parallelism in ILP using MapReduce
Ashwin Srinivasan 0001, Tanveer A. Faruquie, Sachindra Joshi |
Mach. Learn. | 2 |
| 2011 | Discovering customer intent in real-time for streamlining service desk conversationsabstractBusinesses require the contact center agents to meet pre-specified customer satisfaction levels while keeping the cost of operations low or meeting sales targets, objectives that end up being complementary and difficult to achieve in real-time. In this paper, we describe a speech enabled real-time conversation management system that tracks customer-agent conversations to detect user intent (e.g. gathering information, likely to buy, etc.) that can help agents to then decide the best sequence of actions for that call. We present an entropy based decision support system that parses a text stream generated in real-time during a audio conversation and identifies the first instance at which the intent becomes distinct enough for the agent to then take subsequent actions. We provide evaluation results displaying the efficiency and effectiveness of our system. Ullas Nambiar, Tanveer A. Faruquie, L. Venkata Subramaniam, Sumit Negi, Ganesh Ramakrishnan |
CIKM | 2 |
| 2011 | Probabilistic model for discovering topic based communities in social networksabstractSocial graphs have received renewed interest as a research topic with the advent of social networking websites. These online networks provide a rich source of data to study user relationships and interaction patterns on a large scale. In this paper, we propose a generative Bayesian model for extracting latent communities from a social graph. We assume that community memberships depend on topics of interest between users and the link relationships between them in the social graph topology. In addition, we make use of the nature of interaction to gauge user interests. Our model allows communities to be related to multiple topics and each user in the graph can be a member of multiple communities. This gives an insight into user interests and topical distribution in communities. We show the effectiveness of our model using a real world data set and also compare our model with existing community discovery methods. Mrinmaya Sachan, Danish Contractor, Tanveer A. Faruquie, L. Venkata Subramaniam |
CIKM | 3 |
| 2011 | Learning Dirichlet Processes from Partially Observed GroupsabstractMotivated by the task of vernacular news analysis using known news topics from national news-papers, we study the task of topic analysis, where given source datasets with observed topics, data items from a target dataset need to be assigned either to observed source topics or to new ones. Using Hierarchical Dirichlet Processes for addressing this task imposes unnecessary and often inappropriate generative assumptions on the observed source topics. In this paper, we explore Dirichlet Processes with partially observed groups (POG-DP). POG-DP avoids modeling the given source topics. Instead, it directly models the conditional distribution of the target data as a mixture of a Dirichlet Process and the posterior distribution of a Hierarchical Dirichlet Process with known groups and topics. This introduces coupling between selection probabilities of all topics within a source, leading to effective identification of source topics. We further improve on this with a Combinatorial Dirichlet Process with partially observed groups (POG-CDP) that captures finer grained coupling between related topics by choosing intersections between sources. We evaluate our models in three different real-world applications. Using extensive experimentation, we compare against several baselines to show that our model performs significantly better in all three applications. Avinava Dubey, Indrajit Bhattacharya, Mrinal Kanti Das, Tanveer A. Faruquie, Chiranjib Bhattacharyya |
ICDM | 4 |
| 2011 | Using Text Reviews for Product Entity Completion
Mrinmaya Sachan, Tanveer A. Faruquie, L. Venkata Subramaniam, Mukesh K. Mohania |
IJCNLP | 2 |
| 2011 | Latent topic model-based group activity discovery
Tanveer A. Faruquie, Subhashis Banerjee, Prem Kumar Kalra |
Vis. Comput. | 1 |
| 2010 | Estimating accuracy for text classification tasks on large unlabeled dataabstractRule based systems for processing text data encode the knowledge of a human expert into a rule base to take decisions based on interactions of the input data and the rule base. Similarly, supervised learning based systems can learn patterns present in a given dataset to make decisions on similar and other related data. Performances of both these classes of models are largely dependent on the training examples seen by them, based on which the learning was performed. Even though trained models might fit well on training data, the accuracies they yield on a new test data may be considerably different. Computing the accuracy of the learnt models on new unlabeled datasets is a challenging problem requiring costly labeling, and which is still likely to only cover a subset of the new data because of the large sizes of datasets involved. In this paper, we present a method to estimate the accuracy of a given model on a new dataset without manually labeling the data. We verify our method on large datasets for two shallow text processing tasks: document classification and postal address segmentation, and using both supervised machine learning methods and human generated rule based models. Snigdha Chaturvedi, Tanveer A. Faruquie, L. Venkata Subramaniam, Mukesh K. Mohania |
CIKM | 2 |
| 2010 | Handling Noisy Queries in Cross Language FAQ Retrieval
Danish Contractor, Govind Kothari, Tanveer A. Faruquie, L. Venkata Subramaniam, Sumit Negi |
EMNLP | 3 |
| 2010 | Data cleansing as a transient serviceabstractThere is often a transient need within enterprises for data cleansing which can be satisfied by offering data cleansing as a transient service. Every time a data cleansing need arises it should be possible to provision hardware, software and staff for accomplishing the task and then dismantling the set up. In this paper we present such a system that uses virtualized hardware and software for data cleansing. We share actual experiences gained from building such a system.We use a cloud infrastructure to offer virtualized data cleansing instances that can be accessed as a service. We build a system that is scalable, elastic and configurable. Each enterprise has unique needs which makes it necessary to customize both the infrastructure and the cleansing algorithms to address these needs. In this paper we will present a system that is easily configurable to suit the data cleansing needs of an enterprise. Tanveer A. Faruquie, K. Hima Prasad, L. Venkata Subramaniam, Mukesh K. Mohania, Girish Venkatachaliah, Shrinivas Kulkarni, Pramit Basu |
ICDE | 1 |
| 2010 | Transfer of Supervision for Improved Address StandardizationabstractAddress Cleansing is very challenging, particularly for geographies with variability in writing addresses. Supervised learners can be easily trained for different data sources. However, training requires labeling large corpora for each data source which is time consuming and labor intensive to create. We propose a method to automatically transfer supervision from a given labeled source to a target unlabeled source using a hierarchical dirichlet process. Each dirichlet process models data from one source. The shared component distribution across these dirichlet processes captures the semantic relation between data sources. A feature projection on the component distributions from multiple sources is used to transfer supervision. Govind Kothari, Tanveer A. Faruquie, L. Venkata Subramaniam, K. Hima Prasad, Mukesh K. Mohania |
ICPR | 2 |
| 2009 | SMS based Interface for FAQ Retrieval
Govind Kothari, Sumit Negi, Tanveer A. Faruquie, Venkatesan T. Chakaravarthy, L. Venkata Subramaniam |
ACL/IJCNLP | 3 |
| 2009 | Time based Activity Inference using Latent Dirichlet AllocationabstractIn this paper we address the problem of time based activity inference in unsupervised manner for an area under surveillance. We use a Latent Dirichlet Allocation based model that captures the activities and how they change over time. We use agglomerative cluster-ing on optical flow vectors to code direction and spatial information. In this model each activity is associated with not only a mixture distribution over these cluster occurrences but also on the distribution over timestamps of their occurrences. Our method thus helps in determining the prominence and the correlation of activities over a period of time. 1 Tanveer A. Faruquie, Prem Kumar Kalra, Subhashis Banerjee |
BMVC | 1 |
| 2009 | Business Intelligence from Voice of CustomerabstractIn this paper, we present a first of a kind system, called business intelligence from voice of customer (BIVoC), that can: 1) combine unstructured information and structured information in an information intensive enterprise and 2) derive richer business insights from the combined data. Unstructured information, in this paper, refers to voice of customer (VoC) obtained from interaction of customer with enterprise namely, conversation with call-center agents, email, and sms. Structured database reflect only those business variables that are static over (a longer window of) time such as, educational qualification, age group, and employment details. In contrast, a combination of unstructured and structured data provide access to business variables that reflect up to date dynamic requirements of the customers and more importantly indicate trends that are difficult to derive from a larger population of customers through any other means. For example, some of the variables reflected in unstructured data are problem/interest in a certain product, expression of dissatisfaction with the business provided, and some unexplored category of people showing certain interest/problem. This gives the BIVoC system the ability to derive business insights that are richer, more valuable and crucial to the enterprises than the traditional business intelligence systems which utilize only structured information. We demonstrate the effectiveness of BIVoC system through one of our real-life engagements where the problem is to determine how to improve agent productivity in a call center scenario. We also highlight major challenges faced while dealing with unstructured information such as handling noise and linking with structured data. L. Venkata Subramaniam, Tanveer A. Faruquie, Shajith Ikbal, Shantanu Godbole, Mukesh K. Mohania |
ICDE | 2 |
| 2008 | Exploiting context to detect sensitive information in call center conversationsabstractProtecting sensitive information while preserving the share-ability and usability of data is becoming increasingly important. In call-centers a lot of customer related sensitive information is stored in audio recordings. In this work, we address the problem of protecting sensitive information in audio recordings and speech transcripts. We present a semi-supervised method to model sensitive information as a directed graph. Effectiveness of this approach is demonstrated by applying it to the problem of detecting and locating credit card transaction in real life conversations between agents and customers in a call center. Tanveer A. Faruquie, Sumit Negi, Anup Chalamalla, L. Venkata Subramaniam |
CIKM | 1 |
| 2008 | HMM based event detection in audio conversationabstractIn this paper, we address the problem of detecting sensitive events in speech signal such as exchange of credit card information. Although close in nature to the word spotting problem, variability in the linguistic content constituting an event and their composition makes event detection a harder task, especially in the context where it is applied such as call-center interaction. In this work we extend the hidden Markov model (HMM) based framework as used in word spotting to event detection, by constructing a network composed of HMM based acoustic models for event and garbage (non-event). Vocabularies specific to the event and non-event are used respectively to build the event and garbage models along with length constraints based on prior knowledge. Effectiveness of this approach is demonstrated by applying it to the problem of detecting credit card transaction event in real life conversations between agents and customers in call center. Our approach yield a false alarm rate of 17.0% and false miss rate of 12.5%. Shajith Ikbal, Tanveer A. Faruquie |
ICME | 2 |
| 2007 | Improving Automatic Call Classification using Machine TranslationabstractUtterance classification is an important task in spoken-dialog systems. The response of the system is dependent on category assigned to the speaker's utterance by the classifier. However, often the input speech is spontaneous and noisy which results in high word error rates. This results in unsatisfactory system performance. In this paper we describe a method to improve the natural language call classification task using statistical machine translation (SMT). We utilize the translation model in SMT to capture the relation between truth and the ASR transcribed text. The model is trained using the human transcribed text and the ASR transcribed text. During deployment SMT is used to sanitize the ASR transcribed text. Our experiments with IBM model 2 shows significant improvement in call classification accuracy. Tanveer A. Faruquie, Nitendra Rajput, Vimal Raj |
ICASSP (4) | 1 |
| 2005 | An architecture for pluggable disambiguation mechanism for RDC based voice applications
Tanveer A. Faruquie, Pankaj Kankar, Nitendra Rajput |
INTERSPEECH | 1 |
| 2004 | An Algorithmic Framework for Solving the Decoding Problem in Statistical Machine Translation
Raghavendra Udupa, Tanveer A. Faruquie, Hemanta K. Maji |
COLING | 2 |
| 2004 | An English-Hindi Statistical Machine Translation System
Raghavendra Udupa, Tanveer A. Faruquie |
IJCNLP | 2 |
| 2004 | Animating expressive faces across languagesabstractThis paper describes a morphing-based audio driven facial animation system. Based on an incoming audio stream, a face image is animated with full lip synchronization and synthesized expressions. A novel scheme to implement a language independent system for audio-driven facial animation given a speech recognition system for just one language, in our case, English, is presented. The method presented here can also be used for text to audio-visual speech synthesis. Visemes in new expressions are synthesized to be able to generate animations with different facial expressions. An animation sequence using optical flow between visemes is constructed, given an incoming audio stream and still pictures of a face representing different visemes. The presented techniques give improved lip synchronization and naturalness to the animated video. Ashish Verma 0001, L. Venkata Subramaniam, Nitendra Rajput, Chalapathy Neti, Tanveer A. Faruquie |
IEEE Trans. Multim. | 5 |
| 2001 | Audio Driven Facial Animation For Audio-Visual RealityabstractIn this paper, we demonstrate a morphing based automated audio driven facial animation system. Based on an incoming audio stream, a face image is animated with full lip synchronization and expression. An animation sequence using optical flow between visemes is constructed, given an incoming audio stream and still pictures of a face speaking different visemes. Rules are formulated based on coarticulation and the duration of a viseme to control the continuity in terms of shape and extent of lip opening. In addition to this new viseme-expression combinations are synthesized to be able to generate animations with new facial expressions. Finally various applications of this system are discussed in the context of creating audio-visual reality. 1. Tanveer A. Faruquie, Ashish Kapoor, Rohit J. Kate, Nitendra Rajput, L. Venkata Subramaniam |
ICME | 1 |
| 2001 | Robust detection of visual ROI for automatic speechreadingabstractWe present our work on visual pruning in an audio-visual (AV) speech recognition scenario. Visual speech information has been successfully used in circumstances where audio-only recognition suffers (e.g. noisy environments). Tracking and extraction of region-of-interest (ROI) (e.g., speaker's mouth region) from video is an essential component of such systems. It is important for the visual front-end to handle tracking errors that result in noisy visual data and hamper performance. We present our robust visual front-end, investigate methods to prune visual noise and its effect on the performance of the AV speech recognition systems. Specifically, we estimate the "goodness of ROI" using Gaussian mixture models and our experiments indicate that significant performance gains are achieved with good quality visual data. Giridharan Iyengar, Gerasimos Potamianos, Chalapathy Neti, Tanveer A. Faruquie, Ashish Verma 0001 |
MMSP | 4 |
| 2000 | Large Vocabulary Audio-Visual Speech Recognition Using Active Shape ModelsabstractOrthogonal information present in the video signal associated with the audio helps in improving the accuracy of a speech recognition system. Audio-visual speech recognition involves extraction of both the audio as well as visual features from the input signal. Extraction of visual parameters is done by the recognition of speech dependent features from the video sequence. The paper uses geometrical features to describe the lip shapes. Curve-based active shape models are used to extract the geometry. These geometrically represented visual parameters are used along with the audio cepstral features to perform an audio-visual classification. It is shown that the bimodal system presented gives an improvement in the classification results over classification using only the audio features. Tanveer A. Faruquie, Abhik Majumdar, Nitendra Rajput, L. Venkata Subramaniam |
ICPR | 1 |