VLDB 2026 Research / reviewers in the wild / expert
Ramesh Jain 0001
dblp:j/RameshJain · also Ramesh C. Jain, Ramesh Chandra Jain
· DBLP profile ↗
250ranked-venue papers
29as first author
12since 2021 · last 2025
0000-0003-2373-4966ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 151 · 11 first-author · 7 since 2021Artificial intelligence and machine learning · 82 · 15 first-author · 4 since 2021Databases, data management, data science and information retrieval · 35 · 3 first-author · 1 since 2021Computer networks · 13 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 1 since 2021Systems, architecture and hardware · 6Human-computer interaction and ubiquitous computing · 6 · 2 first-authorSecurity and privacy · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AVES: An Audio-Visual Emotion Stream Dataset for Temporal Emotion DetectionabstractHuman emotions vary over time, which can be vividly described as a stream of emotions. Observing the emotion stream in daily life provides valuable insights into an individual's mental state. However, existing research in emotion understanding has mainly focused on classification tasks, assigning an emotion category to a well-trimmed segment or each frame within a continuous signal. In contrast, the task of temporal emotion detection, which involveslocatingthe boundaries of emotion segments andrecognizingtheir categories in untrimmed signals, has not been fully explored. To advance research in this area, this paper introduces an in-the-wild Audio-Visual Emotion Stream (AVES) dataset, which is reliably annotated with the time boundaries and emotion category for each emotion segment in the videos. Thus, AVES can serve as a solid benchmark for temporal emotion detection tasks. Moreover, considering the flexible boundaries and varying durations of emotion segments, we propose a Boundary Combination Network (BoCoNet) for temporal emotion detection, which leverages short-term temporal context information to first predict the boundaries of emotion segments and then locate the entire emotion segments. Extensive experiments conducted on various representative unimodal and multimodal representations demonstrate that BoCoNet achieves state-of-the-art results. The AVES dataset will be released to the research community. We expect that this paper can advance the research on emotion stream and temporal emotion detection. Yan Li 0121, Ke Lu 0002, Dongmei Jiang, Ramesh Jain 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | Causal Inference Hashing for Long-Tailed Image RetrievalabstractIn hashing-based long-tailed image retrieval, the dominance of data-rich head classes often hinders the learning of effective hash codes for data-poor tail classes due to inherent long-tailed bias. Interestingly, this bias also contains valuable prior knowledge by revealing inter-class dependencies, which can be beneficial for hash learning. However, previous methods have not thoroughly analyzed this tangled negative and positive effects of long-tailed bias from a causal inference perspective. In this paper, we propose a novel hash framework that employs causal inference to disentangle detrimental bias effects from beneficial ones. To capture good bias in long-tailed datasets, we construct hash mediators that conserve valuable prior knowledge from class centers. Furthermore, we propose a de-biased hash loss To enhance the beneficial bias effects while mitigating adverse ones, leading to more discriminative hash codes. Specifically, this loss function leverages the beneficial bias captured by hash mediators to support accurate class label prediction, while mitigating harmful bias by blocking its causal path to the hash codes and refining predictions through backdoor adjustment. Extensive experimental results on four widely used datasets demonstrate that the proposed method improves retrieval performance against the state-of-the-art methods by large margins. The source code is available at https://github.com/IMAG-LuJin/CIH. Lu Jin 0001, Zhengyun Lu, Zechao Li, Yonghua Pan, Longquan Dai, Jinhui Tang 0001, Ramesh Jain 0001 |
IEEE Trans. Image Process. | 7 |
| 2024 | A Comprehensive Picture of Factors Affecting User Willingness to Use Mobile Health ApplicationsabstractMobile health (mHealth) applications have become increasingly valuable in preventive healthcare and in reducing the burden on healthcare organizations. The aim of this article is to investigate the factors that influence user acceptance of mHealth apps and identify the underlying structure that shapes users’ behavioral intention. An online study that employed factorial survey design with vignettes was conducted, and a total of 1,669 participants from eight countries across four continents were included in the study. Structural equation modeling was employed to quantitatively assess how various factors collectively contribute to users’ willingness to use mHealth apps. The results indicate that users’ digital literacy has the strongest impact on their willingness to use them, followed by their online habit of sharing personal information. Users’ concerns about personal privacy only had a weak impact. Furthermore, users’ demographic background, such as their country of residence, age, ethnicity, and education, has a significant moderating effect. Our findings have implications for app designers, healthcare practitioners, and policymakers. Efforts are needed to regulate data collection and sharing and promote digital literacy among the general population to facilitate the widespread adoption of mHealth apps. Shaojing Fan, Ramesh Jain 0001, Mohan Kankanhalli |
ACM Trans. Comput. Heal. | 2 |
| 2024 | Alleviating Over-Fitting in Hashing-Based Fine-Grained Image Retrieval: From Causal Feature Learning to Binary-Injected Hash LearningabstractHashing-based fine-grained image retrieval pursues learning diverse local features to generate inter-class discriminative hash codes. However, existing fine-grained hash methods with attention mechanisms usually tend to just focus on a few obvious areas, which misguides the network to over-fit some salient features. Such a problem raises two main limitations. 1) It overlooks some subtle local features, degrading the generalization capability of learned embedding. 2) It causes the over-activation of some hash bits correlated to salient features, which breaks the binary code balance and further weakens the discrimination abilities of hash codes. To address these limitations of the over-fitting problem, we propose a novel hash framework fromCausalFeature learning toBinary-injectedHash learning (CFBH), which captures various local information and suppresses over-activated hash bits simultaneously. For causal feature learning, we adopt causal inference theory to alleviate the bias towards the salient regions in fine-grained images. In detail, we obtain local features from the feature map and combine this local information with original image information followed by this theory. Theoretically, these fused embeddings help the network to re-weight the retrieval effort of each local feature and exploit more subtle variations without observational bias. For binary-injected hash learning, we propose a Binary Noise Injection (BNI) module inspired by Dropout. The BNI module not only mitigates over-activation to particular bits, but also makes hash codes uncorrelated and balanced in the Hamming space. Extensive experimental results on six popular fine-grained image datasets demonstrate the superiority of CFBH over several State-of-the-Art methods. Xinguang Xiang, Xinhao Ding, Lu Jin 0001, Zechao Li, Jinhui Tang 0001, Ramesh Jain 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | Integrative Multi-Modal Computing for Personal Health NavigationabstractAn individual’s health trajectory is most influenced by personal lifestyle choices made regularly and frequently. It is now possible to measure, store, and analyze both multimodal lifestyle signals as well as multimodal physiological and behavioral health signals perpetually. Moreover, it is possible to model the effects of lifestyle precisely on health signals by collecting a variety of longitudinal data. These actions provide the inputs that affect changes in the health state of an individual based on their lifestyle decisions. With the advent of modern relatively inexpensive and common multi-modal data streams and sensing technologies, individuals now have large amounts of data and information about themselves that have the potential to transform decision-making in day-to-day life for health improvement. Multimodal analytics and contextual prediction and retrieval allow this data to make the best decisions to keep the health state optimal while making good lifestyle choices. This critical problem requires making a large variety of data constantly useful, contextually relevant, and most importantly useful in personalized decision-making. Nitish Nag, Hyungik Oh, Mengfan Tang, Mingshu Shi, Ramesh Jain 0001 |
ICMR | 5 |
| 2023 | Towards Deep Personal Lifestyle Models Using Multimodal N-of-1 Data
Nitish Nagesh, Iman Azimi, Tom Andriola, Amir-Mohammad Rahmani, Ramesh Jain 0001 |
MMM (1) | 5 |
| 2023 | ScopeSense: An 8.5-Month Sport, Nutrition, and Lifestyle Lifelogging DatasetabstractNowadays, most people have a smartphone that can track their everyday activities. Furthermore, a significant number of people wear advanced smartwatches to track several vital biomarkers in addition to activity data. However, it is still unclear how these data can actually be used to improve certain aspects of people’s lives. One of the key challenges is that the collected data is often massive and unstructured. Therefore, a link to other important information (e.g., when, what, and how much food was consumed) is required. It is widely believed that such detailed and structured longitudinal data about a person is essential to model and provide personalized and precise guidance. Despite the strong belief of researchers about the power of such a data-driven approach, respective datasets have been difficult to collect. In this study, we present a unique dataset from two individuals performing a structured data collection over eight and a half months. In addition to the sensor data, we collected their nutrition, training, and well-being data. The availability of nutrition data with many other important objectives and subjective longitudinal data streams may facilitate research related to food for a healthy lifestyle. Thus, we present a sport, nutrition, and lifestyle logging dataset called ScopeSense from two individuals and discuss its potential use. The dataset is fully open for researchers, and we consider this study as a potential starting point for developing methods to collect and create knowledge for a larger cohort of people. Michael Riegler 0001, Vajira Thambawita, Binh T. Nguyen 0001, Steven Alexander Hicks, Vibeke Telle-Hansen, Svein Arne Pettersen, Dag Johansen, Ramesh Jain 0001, Pål Halvorsen |
MMM (1) | 9 |
| 2023 | Multimodal driver distraction detection using dual-channel network of CNN and Transformer
Luntian Mou, Jiali Chang, Yiyuan Zhao, Nan Ma 0008, Ramesh Jain 0001, Wen Gao 0001 |
Expert Syst. Appl. | 7 |
| 2023 | Driver Emotion Recognition With a Hybrid Attentional Multimodal Fusion FrameworkabstractNegative emotions may induce dangerous driving behaviors leading to extremely serious traffic accidents. Therefore, it is necessary to establish a system that can automatically recognize driver emotions so that some actions can be taken to avoid traffic accidents. Existing studies on driver emotion recognition have mainly used facial data and physiological data. However, there are fewer studies on multimodal data with contextual characteristics of driving. In addition, fully fusing multimodal data in the feature fusion layer to improve the performance of emotion recognition is still a challenge. To this end, we propose to recognize driver emotion using a novel multimodal fusion framework based on convolutional long-short term memory network (ConvLSTM), and hybrid attention mechanism to fuse non-invasive multimodal data of eye, vehicle, and environment. In order to verify the effectiveness of the proposed method, extensive experiments have been carried out on a dataset collected using an advanced driving simulator. The experimental results demonstrate the effectiveness of the proposed method. Finally, a preliminary exploration on the correlation between driver emotion and stress is performed. Luntian Mou, Yiyuan Zhao, Bahareh Nakisa, Mohammad Naim Rastgoo, Lei Ma 0008, Tiejun Huang 0001, Ramesh Jain 0001, Wen Gao 0001 |
IEEE Trans. Affect. Comput. | 9 |
| 2023 | Isotropic Self-Supervised Learning for Driver Drowsiness Detection With Attention-Based Multimodal FusionabstractDriverdrowsiness is an important cause of traffic accidents. Many studies using computer vision techniques to detect driver drowsiness states, such as slow blinking, yawning, and nodding, have demonstrated excellent potential. Although existing studies have made significant progress, the number of samples in the training corpora is small, which makes it difficult for a model to learn effective drowsiness representations from images or videos. To address this issue, we develop an isotropic self-supervised learning (IsoSSL) approach to learn powerful representations of images without relying on human-provided annotations and propose an IsoSSL-MoCo model by combining IsoSSL with momentum contrast (MoCo). To exploit the complementarity of multimodal data, an attention-based multimodal fusion model is also proposed to fuse features from the eye, mouth, and optical flow of the head. Specifically, we first use the IsoSSL-MoCo model to pretrain the image encoders for the three modalities in other datasets. Then, these encoders are fine-tuned and integrated into the proposed fusion model. The feature vectors generated by the image encoders of the three modalities are fed into the recursive layer to extract temporal information. To capture the importance degrees of the effects of temporal features from the three modalities on drowsiness detection, an attention mechanism is introduced to automatically weigh the feature vectors from the recursive layer to improve detection accuracy. Finally, a vector representation is generated by the attention layer and is used to detect driver drowsiness states. Experimental results based on two challenging datasets show that our method outperforms the baseline methods and the latest existing methods. Luntian Mou, Pengtao Xie, Pengfei Zhao 0008, Ramesh Jain 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Guest Editorial: Learning From Noisy Multimedia DataabstractThis special issue provides a premier forum for researchers in multimedia big data to share challenges and recent advancements in learning from noisy multimedia data. The multimedia age and its proliferation of devices and platforms is fueling exponential data growth. As computational power and deep learning algorithms rapidly evolve, the web has become a rich source of potential training data for robust machine learning, with search engines such as Google and Bing, Twitter, TikTok, Instagram, and short video sharing platforms offering large-scale data points in the hundreds of millions. The concurrent shift in the Internet to richer web data modalities such as text, audio, image, and video reveal further opportunities to leverage large-scale data for the automatic construction of a variety of datasets for model training and testing. However, the ubiquity of multimedia data means noise is a fundamental challenge, with ‘label noise’ and ‘domain mismatch’ the most critical issues in automatically collected datasets. Learning from noisy multimedia data tends towards poor performance, making it increasingly essential to address these challenges. Jian Zhang 0002, Alan Hanjalic, Ramesh Jain 0001, Xian-Sheng Hua 0001, Shin'ichi Satoh 0001, Yazhou Yao, Dan Zeng 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Driver stress detection via multimodal fusion using attention-based CNN-LSTM
Luntian Mou, Pengfei Zhao 0008, Bahareh Nakisa, Mohammad Naim Rastgoo, Ramesh Jain 0001, Wen Gao 0001 |
Expert Syst. Appl. | 6 |
| 2020 | What Should I Do?abstractI find myself asking, "What should I do?" in many situations such as when I want to go out to eat; I want to decide about my vacation; decide on spending some free time on the weekend; and numerous other decisions. All these decisions are really personalized contextual decisions that may be addressed by a contextual recommendation engine that knows me. For knowing me well, the engine should prepare my model based on all events in my life. By retrieving and mining events of various types computed using different multimodal data streams, such a personal model may be prepared and then used to help in making decisions ranging from trivial to critical. We discuss important challenges in organizing life events that may be used for building personal models and for accessing characteristics of such events as may be needed in various applications. We will demonstrate our ideas using some applications related to lifestyle and health. Ramesh Jain 0001 |
ICMR | 1 |
| 2020 | Continuous Health Interface Event RetrievalabstractKnowing the state of our health at every moment in time is critical for advances in health science. Using data obtained outside an episodic clinical setting is the first step towards building a continuous health estimation system. In this paper, we explore a system that allows users to combine events and data streams from different sources and retrieve complex biological events, such as cardiovascular volume overload, using measured lifestyle events. These complex events, which have been explored in biomedical literature and which we call interface events, have a direct causal impact on the relevant biological systems; they are the interface through which the lifestyle events influence our health. We retrieve the interface events from existing events and data streams by encoding domain knowledge using the event operator language. The interface events can then be utilized to provide a continuous estimate of the biological variables relevant to the user's health state. The event-based framework also makes it easier to estimate which event is causally responsible for a particular change in the individual's health state. Vaibhav Pandey, Nitish Nag, Ramesh Jain 0001 |
ICMR | 3 |
| 2020 | Coping with Pandemics: Opportunities and Challenges for AI Multimedia in the "New Normal"abstractTheworld iswelcoming the newnormal - the coronavirus pandemic has significantly changed the way people live, work, communicate and learn. Almost everyone now is wearing a face mask when they go in public. People are working from home, some taking care of children at the same time. Bars and restaurants are limited to carry-out and delivery only. Meetings and conferences go online. Schools are closed and educators are instead holding video conference classes regularly. All these become the new normal as our ways of life. The panel thus provides a valuable opportunity for people from a variety of backgrounds to exchange views on opportunities and challenges for AI multimedia in the current and post pandemics era. Jiaying Liu 0001, Wen-Huang Cheng, Klara Nahrstedt, Ramesh Jain 0001, Elisa Ricci 0001, Hyeran Byun |
ACM Multimedia | 4 |
| 2020 | Personal Food ModelabstractFood is central to life. Food provides us with energy and foundational building blocks for our body and is also a major source of joy and new experiences. A significant part of the overall economy is related to food. Food science, distribution, processing, and consumption have been addressed by different communities using silos of computational approaches. In this paper, we adopt a person-centric multimedia and multimodal perspective on food computing and show how multimedia and food computing are synergistic and complementary. Enjoying food is a truly multimedia experience involving sight, taste, smell, and even sound, that can be captured using a multimedia food logger. The biological response to food can be captured using multimodal data streams using available wearable devices. Central to this approach is the Personal Food Model. Personal Food Model is the digitized representation of the food-related characteristics of an individual. It is designed to be used in food recommendation systems to provide eating-related recommendations that improve the user's quality of life. To model the food-related characteristics of each person, it is essential to capture their food-related enjoyment using a Preferential Personal Food Model and their biological response to food using their Biological Personal Food Model. Inspired by the power of 3-dimensional color models for visual processing, we introduce a 6-dimensional taste-space for capturing culinary characteristics as well as personal preferences. We use event mining approaches to relate food with other life and biological events to build a predictive model that could also be used effectively in emerging food recommendation systems. Ali Rostami 0004, Vaibhav Pandey, Nitish Nag, Vesper Wang, Ramesh Jain 0001 |
ACM Multimedia | 5 |
| 2020 | Multimedia Food LoggerabstractLogging what we eat is important for individuals and the aggregated information in these logs are important for businesses as well as public health. Food logging has received very little attention and has been mostly limited only to the recognition of food items ignoring context, situation, and health variable completely. In this demo we let the audience interact with our multimedia food logger system which is described in the following. We also describe how this system captures the major food-related information that could be used by all stakeholders in the food ecosystem. We will demonstrate the complete functionality of such a system in this demo. Ali Rostami 0004, Bihao Xu, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2020 | PMData: a sports logging datasetabstractIn this paper, we present PMData: a dataset that combines traditional lifelogging data with sports-activity data. Our dataset enables the development of novel data analysis and machine-learning applications where, for instance, additional sports data is used to predict and analyze everyday developments, like a person's weight and sleep patterns; and applications where traditional lifelog data is used in a sports context to predict athletes' performance. PMData combines input from Fitbit Versa 2 smartwatch wristbands, the PMSys sports logging smartphone application, and Google forms. Logging data has been collected from 16 persons for five months. Our initial experiments show that novel analyses are possible, but there is still room for improvement. Vajira Thambawita, Steven Alexander Hicks, Hanna Borgli, Håkon Kvale Stensland, Debesh Jha, Martin Kristoffer Svensen, Svein Arne Pettersen, Dag Johansen, Håvard D. Johansen, Susann Dahl Pettersen, Simon Nordvang, Sigurd Pedersen, Anders T. Gjerdrum, Tor-Morten Grønli, Per Morten Fredriksen, Ragnhild Eg, Kjeld Hansen, Siri Fagernes, Christine Claudi, Andreas Biørn-Hansen, Duc-Tien Dang-Nguyen, Tomas Kupka, Hugo Hammer, Ramesh Jain 0001, Michael Riegler 0001, Pål Halvorsen |
MMSys | 24 |
| 2020 | Food Recommendation: Framework, Existing Solutions, and ChallengesabstractA growing proportion of the global population is becoming overweight or obese, leading to various diseases (e.g., diabetes, ischemic heart disease and even cancer) due to unhealthy eating patterns, such as increased intake of food with high energy and high fat. Food recommendation is of paramount importance to alleviate this problem. Unfortunately, modern multimedia research has enhanced the performance and experience of multimedia recommendation in many fields such as movies and POI, yet largely lags in the food domain. This article proposes a unified framework for food recommendation, and identifies main issues affecting food recommendation including incorporating various context and domain knowledge, building the personal model, and analyzing unique food characteristics. We then review existing solutions for these issues, and finally elaborate research challenges and future directions in this field. To our knowledge, this is the first survey that targets the study of food recommendation in the multimedia field and offers a collection of research studies and technologies to benefit researchers in this field. Weiqing Min, Shuqiang Jiang, Ramesh Jain 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | Opportunities Created by Digitalization of Life and HealthabstractSummary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. We all are in the midst of every aspect of life and health being digitalized. This digitalization may be more disruptive and transformative to human health than any other aspects of life. Moreover, this may result in benefitting all sectors of society, even the poorest and remote groups of people. This is enabled by digitalizing and creating event streams of lifestyle and health and utilizing them to build models of each individual as well as determining their health state. Personal models and health states enable cybernetic principles to provide help and guidance perpetually rather than current episodic approaches used in health care. The data thus collected for individuals enables formation of living laboratories that may result in more precise and actionable disease models. We will present our approach and early results in this talk. Ramesh Jain 0001 |
SERVICES | 1 |
| 2018 | HealthMedia 2018: Third International Workshop on Multimedia for Personal Health and Health CareabstractResearch in multimedia and health is driven by the current technological advancements in sensors and personalized healthcare. There is an increasing amount of work that shows how core multimedia research is becoming an important enabler for solutions with applications and relevance for the societal questions of health. This workshop brings together researchers from diverse topics such as multimedia, pervasive health, lifelogging, accessibility, HCI, but also health, medicine, and psychology to address challenges and opportunities of multimedia in and for health. Jochen Meyer 0001, Susanne Boll, Noel E. O'Connor, Ramesh Jain 0001, Troy McDaniel |
ACM Multimedia | 4 |
| 2018 | Cross-Modal Health State EstimationabstractIndividuals create and consume more diverse data about themselves today than any time in history. Sources of this data include wearable devices, images, social media, geo-spatial information and more. A tremendous opportunity rests within cross-modal data analysis that leverages existing domain knowledge methods to understand and guide human health. Especially in chronic diseases, current medical practice uses a combination of sparse hospital based biological metrics (blood tests, expensive imaging, etc.) to understand the evolving health status of an individual. Future health systems must integrate data created at the individual level to better understand health status perpetually, especially in a cybernetic framework. In this work we fuse multiple user created and open source data streams along with established biomedical domain knowledge to give two types of quantitative state estimates of cardiovascular health. First, we use wearable devices to calculate cardiorespiratory fitness (CRF), a known quantitative leading predictor of heart disease which is not routinely collected in clinical settings. Second, we estimate inherent genetic traits, living environmental risks, circadian rhythm, and biological metrics from a diverse dataset. Our experimental results on 24 subjects demonstrate how multi-modal data can provide personalized health insight. Understanding the dynamic nature of health status will pave the way for better health based recommendation engines, better clinical decision making and positive lifestyle changes. Nitish Nag, Vaibhav Pandey, Preston J. Putzel, Hari Bhimaraju, Srikanth Krishnan, Ramesh Jain 0001 |
ACM Multimedia | 6 |
| 2018 | Deep Learning for Multimedia: Science or Technology?abstractDeep learning has been successfully explored in addressing different multimedia topics recent years, ranging from object detection, semantic classification, entity annotation, to multimedia captioning, multimedia question answering and storytelling. Open source libraries and platforms such as Tensorflow, Caffe, MXnet significantly help promote the wide deployment of deep learning in solving real-world applications. On one hand, deep learning practitioners, while not necessary to understand the involved math behind, are able to set up and make use of a complex deep network. One recent deep learning tool based on Keras even provides the graphical interface to enable straightforward 'drag and drop' operation for deep learning programming. On the other hand, however, some general theoretical problems of learning such as the interpretation and generalization, have only achieved limited progress. Most deep learning papers published these days follow the pipeline of designing/modifying network structures - tuning parameters - reporting performance improvement in specific applications. We have even seen many deep learning application papers without one single equation. Theoretical interpretation and the science behind the study are largely ignored. While excited about the successful application of deep learning in classical and novel problems, we multimedia researchers are responsible to think and solve the fundamental topics in deep learning science. Prof. Guanrong Chen recently wrote an editorial note titled 'Science and Technology, not SciTech' [1]. This panel falls into similar discussion and aims to invite prestigious multimedia researchers and active deep learning practitioners to discuss the positioning of deep learning research now and in the future. Specifically, each panelist is asked to present their opinions on the following five questions: 1)How do you think the current phenomenon that deep learning applications are explosively growing, while the general theoretical problems remain slow progress? 2)Do you agree that deployment of deep learning techniques is getting easy (with a low barrier), while deep learning research is difficult (with a high barrier) 3)What do you think are the core problems for deep learning techniques? 4)What do you think are the core problems for deep learning science? 5)What's your suggestion on the multimedia research in the post-deep learning era? Jun Yu 0002, Ramesh Jain 0001, Rainer Lienhart, Peng Cui 0001, Jiashi Feng |
ACM Multimedia | 3 |
| 2018 | Hookworm Detection in Wireless Capsule Endoscopy Images With Deep LearningabstractAs one of the most common human helminths, hookworm is a leading cause of maternal and child morbidity, which seriously threatens human health. Recently, wireless capsule endoscopy (WCE) has been applied to automatic hookworm detection. Unfortunately, it remains a challenging task. In recent years, deep convolutional neural network (CNN) has demonstrated impressive performance in various image and video analysis tasks. In this paper, a novel deep hookworm detection framework is proposed for WCE images, which simultaneously models visual appearances and tubular patterns of hookworms. This is the first deep learning framework specifically designed for hookworm detection in WCE images. Two CNN networks, namely edge extraction network and hookworm classification network, are seamlessly integrated in the proposed framework, which avoid the edge feature caching and speed up the classification. Two edge pooling layers are introduced to integrate the tubular regions induced from edge extraction network and the feature maps from hookworm classification network, leading to enhanced feature maps emphasizing the tubular regions. Experiments have been conducted on one of the largest WCE datasets with WCE images, which demonstrate the effectiveness of the proposed hookworm detection framework. It significantly outperforms the state-of-the-art approaches. The high sensitivity and accuracy of the proposed method in detecting hookworms shows its potential for clinical application. Jun-Yan He, Xiao Wu 0001, Yu-Gang Jiang 0001, Qiang Peng, Ramesh Jain 0001 |
IEEE Trans. Image Process. | 5 |
| 2017 | Multimedia Rescue Systems for FloodsabstractIn the recent years, digitalization of disaster management has gained much importance in aspects of response, recovery and mitigation. Often, issues identified in past disaster response and relief efforts include lack of communication, delayed ordering of actions (e.g., evacuations), and low levels of preparedness by authorities during disasters. A rapid response in the event of a disaster requires good situational awareness and fast sharing of that knowledge across the community. We believe that a digital ecosystem based on multimedia computing can address this major challenge in the society. The role of technology in disaster management is to connect, inform and ultimately save the lives of those impacted by disasters. Multimedia analysis must expand to take advantage of this fact to benefit society, in particular, help during disasters. Data driven tools as part of new ecosystem can link resources and need to build an effective platform for use during disasters. To accomplish ecosystem coevolution, creating a collaborative system with crowdsourced data, spatial data, and historical data is essential. EventShop is one such tool used to detect and observe changing real-world situations in real-time. Multimedia micro-reports are global, powerful, impactful and spontaneous data streams which are obtained from applications like Krumbs, Twitter, Flickr, etc. These micro-reports assist EventShop in building a rescue system that provides rescue responses to victims during disasters. In this paper, we present a framework that uses EventShop and multimedia micro-reports to correspond with victims during one such disaster, the flood. Animesh Sahay, Anjana Anil Kumar, Siripen Pongpaichet, Ramesh Jain 0001 |
MEDES | 4 |
| 2017 | Health Multimedia: Lifestyle Recommendations Based on Diverse ObservationsabstractManaging health lays the core foundation to enabling quality life experiences. Modern multimedia research has enhanced the quality of experiences in fields such as entertainment, social media, and advertising; yet lags in the health domain. We are developing an approach to leverage multimedia systems for human health. Health is primarily a product of our everyday lifestyle actions, yet we have minimal health guidance on making everyday choices. Recommendations are the key to modern content consumption and decisions. Cybernetic navigation principles that integrate health media sources can power dynamic recommendations to dramatically improve our health decisions. Cybernetic components give real-time feedback on health status, while the navigational approach plots health trajectory. These two principles coalesce data to enable personalized, predictive, and precise health knowledge that can contextually disseminate the right actions to keep individuals on a path to wellness. Nitish Nag, Vaibhav Pandey, Ramesh Jain 0001 |
ICMR | 3 |
| 2017 | MMHealth 2017: Workshop on Multimedia for Personal Health and Health CareabstractEver since the emergence of digitization, we've used the term multimedia to represent a combination of different kinds of media types, such as images, audio, and videos. As new sensing technologies emerge and are now becoming omnipresent in daily lives, the definition, role and significance of multimedia is changing. Multimedia now represents the means for communicating, cooperating, and also for monitoring numerous aspects of daily life, at various levels of granularity and application, ranging from personal to societal. With this shift, we have since moved from comprehending single media and its state toward comprehending media in terms of its use context. Susanne Boll, Touradj Ebrahimi, Cathal Gurrin, Laleh Jalali, Ramesh Jain 0001, Jochen Meyer 0001, Noel E. O'Connor |
ACM Multimedia | 5 |
| 2017 | From Multimedia Logs to Personal ChroniclesabstractMultimodal data streams are essential for analyzing personal life, environmental conditions, and social situations. Since these data streams have different granularities and semantics, the semantic gap becomes even more formidable. To make sense of all the multimodal correlated streams we must first synchronize them in the context of the application, and then analyze them to extract meaningful information. In this paper, we consider the problem of modeling an individual by using daily activity in order to understand their health and behavior. The first step is to correlate diverse data streams with atomic-interval, and segment a person's day into her daily activities. We collect the diverse data streams from the person's smartphone to classify every atomic-interval into a daily activity. Next, we use an interval growing technique for determining daily-activity-intervals and their attributes. Then, these daily-activity-intervals are labeled as the daily activities by using Bagging Formal Concept Analysis (BFCA). Finally, we build a personal chronicle, which is a person's time-ordered list of daily activities. This personal chronicle can then be used to model the person using learning techniques applied to daily activities in the chronicle and relating them to biomedical or behavioral signals. We present the results for this daily activity segmentation and recognition by using lifelogs of 23 participants. Hyungik Oh, Ramesh Jain 0001 |
ACM Multimedia | 2 |
| 2017 | Panel: Cross-media IntelligenceabstractIn this panel, we attempt to review and discuss the recent emerging theoretical and technological advances and trends of cross-media. Integrating data-driven machine learning with human knowledge can effectively lead to explainable, robust, and general models. Thus, the effective employment of the interaction between cross-media data during inference and reasoning becomes a challenge to populate the cross-media knowledge graph. Some other fundamental and controversial issues such as leveraging the auxiliary information to boost the cross-media understanding, the existence of unified framework to bridge the gap between multi-modality will also be discussed in this panel. Yueting Zhuang, Ramesh Jain 0001, Wen Gao 0001, Kiyoharu Aizawa |
ACM Multimedia | 2 |
| 2017 | A graph regularized dimension reduction method for out-of-sample data
Mengfan Tang, Feiping Nie 0001, Ramesh Jain 0001 |
Neurocomputing | 3 |
| 2017 | Semi-supervised learning on large-scale geotagged photos for situation recognition
Mengfan Tang, Feiping Nie 0001, Siripen Pongpaichet, Ramesh Jain 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | Guest editorial: mobile visual tagging with mobile context
Shuqiang Jiang, Liangliang Cao, Jiebo Luo 0001, Ramesh Jain 0001 |
Multim. Syst. | 5 |
| 2017 | Tri-Clustered Tensor Completion for Social-Aware Image Tag RefinementabstractSocial image tag refinement, which aims to improve tag quality by automatically completing the missing tags and rectifying the noise-corrupted ones, is an essential component for social image search. Conventional approaches mainly focus on exploring the visual and tag information, without considering the user information, which often reveals important hints on the (in)correct tags of social images. Towards this end, we propose a novel tri-clustered tensor completion framework to collaboratively explore these three kinds of information to improve the performance of social image tag refinement. Specifically, the inter-relations among users, images and tags are modeled by a tensor, and the intra-relations between users, images and tags are explored by three regularizations respectively. To address the challenges of the super-sparse and large-scale tensor factorization that demands expensive computing and memory cost, we propose a novel tri-clustering method to divide the tensor into a certain number of sub-tensors by simultaneously clustering users, images and tags into a bunch of tri-clusters. And then we investigate two strategies to complete these sub-tensors by considering (in)dependence between the sub-tensors. Experimental results on a real-world social image database demonstrate the superiority of the proposed method compared with the state-of-the-art methods. Jinhui Tang 0001, Xiangbo Shu, Guo-Jun Qi, Zechao Li, Meng Wang 0001, Shuicheng Yan, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2017 | Integration of Diverse Data Sources for Spatial PM2.5 Data InterpolationabstractHeterogeneous data fusion from disparate geospatial sensors has drawn increasing attention in multimedia. Unfortunately, environmental sensors are usually sparsely and preferentially located, which restricts situation recognition of geographical regions and results in uncertainty in derived inferences. Spatial interpolation is an effective way to solve the problem of data sparsity, which demands the availability of related data sources. However, these data sources are usually in different resolutions, distributions, scales, and densities, which poses a major challenge in data integration. To address this problem, we present a novel spatial interpolation framework to incorporate diverse data sources and model the spatial processes explicitly at multiple resolutions. Spectral analysis is deployed to generate features at multiple spatial resolutions and to improve the interpolation accuracy at unobserved locations. A statistical operator based on the spatial Gaussian process is implemented and integrated into a geospatial situation recognition system, which can analyze heterogeneous spatio-temporal data streams derived from sensors. To verify the effectiveness and efficiency of the proposed framework, this framework is applied to the PM2.5 air pollution application. Experiments conducted in California, USA, demonstrate that the proposed method outperforms state-of-the-art approaches. Mengfan Tang, Xiao Wu 0001, Pranav Agrawal 0002, Siripen Pongpaichet, Ramesh Jain 0001 |
IEEE Trans. Multim. | 5 |
| 2016 | Digital Knowledge Ecosystem for Achieving Sustainable Agriculture Production: A Case Study from Sri LankaabstractCrop production problems are common in Sri Lanka which severely effect rural farmers, agriculture sector and the country's economy as a whole. A deeper analysis revealed that the root cause was farmers and other stakeholders in the domain not receiving right information at the right time in the right format. Inspired by the rapid growth of mobile phone usage among farmers a mobile-based solution is sought to overcome this information gap. Farmers needed published information (quasi static) about crops, pests, diseases, land preparation, growing and harvesting methods and real-time situational information (dynamic) such as current crop production and market prices. This situational information is also needed by agriculture department, agro-chemical companies, buyers and various government agencies to ensure food security through effective supply chain planning whilst minimising waste. We developed a notion of context specific actionable information which enables user to act with least amount of further processing. User centered agriculture ontology was developed to convert published quasi static information to actionable information. We adopted empowerment theory to create empowerment-oriented farming processes to motivate farmers to act on this information and aggregated the transaction data to generate situational information. This created a holistic information flow model for agriculture domain similar to energy flow in biological ecosystems. Consequently, the initial Mobile-based Information System evolved into a Digital Knowledge Ecosystem that can predict current production situation in near real enabling government agencies to dynamically adjust the incentives offered to farmers for growing different types of crops to achieve sustainable agriculture production through crop diversification. Athula Ginige, Anusha Indika Walisadeera, Tamara Ginige, Lasanthi N. C. De Silva, Pasquale Di Giovanni, Maneesh Mathai, Jeevani S. Goonetillake, Gihan N. Wikramanayake, Giuliana Vitiello, Monica Sebillo, Genny Tortora, Debbie Richards 0001, Ramesh Jain 0001 |
DSAA | 13 |
| 2016 | A graph based multimodal geospatial interpolation frameworkabstractRecent multimedia research has increasingly focused on large scale multimodal data from disparate geospatial sensors. In addition to the volume of the data, the diversity and granularity of the data poses a major challenge in extracting meaningful and actionable information. To address this, we present a novel spatial interpolation framework, capable of incorporating multimodal data sources and modeling the spatial processes comprehensively at multiple resolutions. The framework transforms the spatial interpolation problem into a graph structure learning problem, based on the latent structure of the data. This enables more efficient and accurate predictions at unobserved locations. We demonstrate the effectiveness of our approach by testing it on air pollution interpolation. Mengfan Tang, Pranav Agrawal 0002, Feiping Nie 0001, Siripen Pongpaichet, Ramesh Jain 0001 |
ICME | 5 |
| 2016 | Social Multimedia Ming: From Special to GeneralabstractSocial multimedia is the hybrid of social media and multimedia. Social multimedia mining employs data mining techniques to understand the association among content, user and interaction in social multimedia. It has three basic problems at basis, micro and macro association levels. This paper extends the current interpretation of social multimedia mining as processing multi-modal individual objects, to a generalized definition beyond modality and beyond object. Cross-OSN (Online Social Networking) data mining, an instantiation to general social multimedia mining, is then positioned with main challenges, review of current studies and future directions. Changsheng Xu, Ramesh Jain 0001 |
ISM | 3 |
| 2016 | Using Photos as Micro-Reports of EventsabstractPhotos serve dual role. Photos are important for capturing, saving, sharing, and reminiscing memories of events and people. Modern photos, however, are becoming more spontaneous, objective, compelling, and universal reports of a moment in an event also. In this paper our focus is on millions of photos being captured as informative reports and using them for emerging applications including situation recognition, trend analysis, and cultural dynamics. EventShop is an open source platform for situation recognition. Utilizing this platform and using a stream of photo reports from various sources as one of the data streams in this platform, we build a visual analytics system to understand the information that could be gleaned from such photo report streams. Our early experiments are based on the Yahoo Flickr Creative Commons 100 Million photos set released recently. We are also using other sources to import and understand the efficacy of these reports for various important applications. Siripen Pongpaichet, Mengfan Tang, Laleh Jalali, Ramesh Jain 0001 |
ICMR | 4 |
| 2016 | Situation Recognition from Multimodal DataabstractSituation recognition is the problem of deriving actionable insights from heterogeneous, real-time, big multimedia data to benefit human lives and resources in different applications. This tutorial will discuss the recent developments towards converting multitudes of data streams including weather patterns, stock prices, social media, traffic information, and disease incidents into actionable insights. Vivek K. Singh 0001, Siripen Pongpaichet, Ramesh Jain 0001 |
ICMR | 3 |
| 2016 | Exploration of Large Image Corpuses in Virtual RealityabstractWith the increasing capture of photos and their proliferation on social media, there is a pressing need for a more intuitive and versatile image search and exploration system. Image search systems have long been confined to the binds of the 2D legacy screens and the keyword text-box. With the recent advances in Virtual Reality (VR) technology, a move towards an immersive VR environment will redefine the image navigation experience. To this end, we propose a VR platform that gathers images from various sources, and addresses the 5 Ws of image search - what, where, when, who and why. We achieve this by providing the user with two modes of interactive exploration - (i) A mode that allows for a graph based navigation of an image dataset, using a steering wheel visualization, along multiple dimensions of time, location, visual concept, people, etc. and (ii) Another mode that provides an intuitive exploration of the image dataset using a logical hierarchy of visual concepts. Our contributions include creating a VR image exploration experience that is intuitive and allows image navigation along multiple dimensions. Sanket Khanwalkar, Shonali Balakrishna, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2016 | Situation Recognition from Multimodal DataabstractNo abstract available. Vivek K. Singh 0001, Siripen Pongpaichet, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2016 | Capped Lp-Norm Graph Embedding for Photo ClusteringabstractPhotos are a predominant source of information on a global scale. Cluster analysis of photos can be applied to situation recognition and understanding cultural dynamics. Graph-based learning provides a current approach for modeling data in clustering problems. However, the performance of this framework depends heavily on initial graph construction by input data. Data outliers degrade graph quality, leading to poor clustering results. We designed a new capped lp-norm graph-based model to reduce the impact of outliers. This is accomplished by allowing the data graph to self adjust as part of the graph embedding. Furthermore, we derive an iterative algorithm to solve the objective function optimization problem. Experiments on four real-world benchmark data sets and Yahoo Flickr Creative Commons data set show the effectiveness of this new graph-based capped lp-norm clustering method. Mengfan Tang, Feiping Nie 0001, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2016 | Research Challenges in Developing Multimedia Systems for Managing Emergency SituationsabstractWith an increasing amount of diverse heterogeneous data and information, the methodology of multimedia analysis has become increasingly relevant in solving challenging societal problems such as managing emergency situations during disasters. Using cybernetic principles combined with multimedia technology, researchers can develop effective frameworks for using diverse multimedia (including traditional multimedia as well as diverse multimodal) data for situation recognition, and determining and communicating appropriate actions to people stranded during disasters. We present known issues in disaster management and then focus on emergency situations. We show that an emergency management problem is fundamentally a multimedia information assimilation problem for situation recognition and for connecting people's needs to available resources effectively, efficiently, and promptly. Major research challenges for managing emergency situations are identified and discussed. We also present a intelligently detecting evolving environmental situations, and discuss the role of multimedia micro-reports as spontaneous participatory sensing data streams in emergency responses. Given enormous progress in concept recognition using machine learning in the last few years, situation recognition may be the next major challenge for learning approaches in multimedia contextual big data. The data needed for developing such approaches is now easily available on the Web and many challenging research problems in this area are ripe for exploration in order to positively impact our society during its most difficult times. Mengfan Tang, Siripen Pongpaichet, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2016 | Introduction to Special Issue on Multimedia Big Data: Networkingabstracteditorial Free Access Share on Introduction to Special Issue on Multimedia Big Data: Networking Editors: Mianxiong Dong Muroran Institute of Technology, Japan Muroran Institute of Technology, JapanView Profile , Vincenzo Piuri Università degli Studi di Milano, Italy Università degli Studi di Milano, ItalyView Profile , Shueng-Han Gary Chan The Hong Kong University of Science and Technology, Hong Kong The Hong Kong University of Science and Technology, Hong KongView Profile , Ramesh Jain University of California, Irvine, CA University of California, Irvine, CAView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 12Issue 5sDecember 2016 Article No.: 70pp 1–3https://doi.org/10.1145/2989214Published:21 September 2016Publication History 1citation336DownloadsMetricsTotal Citations1Total Downloads336Last 12 Months17Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Mianxiong Dong, Vincenzo Piuri, Shueng-Han Gary Chan, Ramesh Jain 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2015 | Exploring spatio-temporal-theme correlation between physical and social streaming data for event detection and pattern interpretation from heterogeneous sensorsabstractIn this paper, we introduce a new method that explores spatio-temporal-theme correlations between physical and social streaming data for event detection and pattern interpretation from heterogeneous sensors. Particularly, we employ a basic two-phase framework in pattern recognition (i.e. feature extraction and detection) with the novel improvement that concerns the use of semantic information acquired from social sensors to automatically label the low-level features extracted from physical sensors. Moreover, by symbolizing the trend component of time-series data, the proposed method has an ability to interpret event's patterns to help users get insights of how events happen. Differentiating from conventional supervised learning methods whose training data are labeled manually and in an off-line mode, the proposed method can collect labels for training data automatically and in an on-line mode. Moreover, after running for a certain time, a training stage can run parallel with the detecting stage when an event model is totally built. After that, the training stage continues learning to increase the accuracy of the event model by nonstop collecting new samples with labels from streaming data. The problem of environmental factors and particularly air pollution impacts on asthma exacerbation is considered for evaluating the proposed method. The experimental results show that the proposed method can probably detect the prevalence of asthma risks in a specific spatio-temporal context as well as help users understand how a change in the surrounding environment (e.g. weather condition and air pollution) can influence their health (e.g. asthma attack) by interpreting detected event's patterns. Minh-Son Dao, Koji Zettsu, Siripen Pongpaichet, Laleh Jalali, Ramesh Jain 0001 |
IEEE BigData | 5 |
| 2015 | An intelligent notification system using context from real-time personal activity monitoringabstractPeople can now receive custom-made information through smartphones, tablets or wearable devices. However, people often tend to miss vital information, even reminders, in the flood of notifications. The problem of finding convenient moments for need-to-know information should be investigated. Because each person's message awareness pattern on a smart medium might be different, the necessity of personalized notification time should be emphasized. We believe that tracking changes in a user's physical activity and other contextual factors will reveal the most convenient moments. We propose a mobile framework, smartNoti, to carefully examine the user environment. The main contributions of our framework are: 1) developing an architecture to provide crucial information in a timely manner at a recognizable moment; 2) integrating, processing, training, and storing personalized latent features from heterogeneous data streams; 3) detecting user context transitions that might provide recognizable and available moments; and 4) predicting these moment and providing a notification message. The experimental validation on Intelligent Callback Reminder, which we implemented on an android application to notify a user missed or rejected call, demonstrates that our approach is effective. We believe that our findings can lead to intelligent strategies to issue unobtrusive notifications on today's smart phones at no extra cost, by using sensors and contextual factors. Hyungik Oh, Laleh Jalali, Ramesh Jain 0001 |
ICME | 3 |
| 2015 | Geospatial interpolation analytics for data streams in eventshopabstractEventShop is an open-source software which provides a generic infrastructure for the analysis of heterogeneous spatio-temporal data streams. Efficient interpolation of data from spatially sparse sources is critical but currently missing in EventShop. To address this challenge, we implement a Spatial Gaussian Process based statistical operator into the EventShop framework. Spectral analysis is employed to generate features at higher spatial resolution and to improve interpolation accuracy at unsampled locations. Further, we test this operator by interpolating air pollution levels in California. The evaluations of multiple metrics demonstrate that our operators outperform earlier EventShop operators, chemical transportation models, and state-of-the-art methods. Mengfan Tang, Pranav Agrawal 0002, Siripen Pongpaichet, Ramesh Jain 0001 |
ICME | 4 |
| 2015 | Bringing Deep Causality to Multimedia Data StreamsabstractWe live in a data abundance era. Availability of large volume of diverse multimedia data streams (ranging from video, to tweets, to activity, and to PM2.5) can now be used to solve many critical societal problems. Causal modeling across multimedia data streams is essential to reap the potential of this data. However, effective frameworks combining formal abstract approaches with practical computational algorithms for causal inference from such data are needed to utilize available data from diverse sensors. We propose a causal modeling framework that builds on data-driven techniques while emphasizing and including the appropriate human knowledge in causal inference. We show that this formal framework can help in designing a causal model with a systematic approach that facilitates framing sharper scientific questions, incorporating expert's knowledge as causal assumptions, and evaluating the plausibility of these assumptions. We show the applicability of the framework in a an important Asthma management application using meteorological and pollution data streams. Laleh Jalali, Ramesh Jain 0001 |
ACM Multimedia | 2 |
| 2015 | Guest Editorial Multimedia: The Biggest Big DataabstractThe goal of this special issue is to provide a premier forum for researchers to present their recent research results on multimedia big data. It follows the recent success event—the First IEEE International Conference on Multimedia Big Data (BigMM 2015) that took place at the Chinese National Convention Center in Beijing, China, from April 20–22, 2015 It also provides an important opportunity for multidisciplinary work connecting big data to multimedia computing. Shu-Ching Chen, Ramesh Jain 0001, Yonghong Tian 0001, Haohong Wang |
IEEE Trans. Multim. | 2 |
| 2015 | Geolocalized Modeling for Dish RecognitionabstractFood-related photos have become increasingly popular , due to social networks, food recommendations, and dietary assessment systems. Reliable annotation is essential in those systems, but unconstrained automatic food recognition is still not accurate enough. Most works focus on exploiting only the visual content while ignoring the context. To address this limitation, in this paper we explore leveraging geolocation and external information about restaurants to simplify the classification problem. We propose a framework incorporating discriminative classification in geolocalized settings and introduce the concept of geolocalized models, which, in our scenario, are trained locally at each restaurant location. In particular, we propose two strategies to implement this framework: geolocalized voting and combinations of bundled classifiers. Both models show promising performance, and the later is particularly efficient and scalable. We collected a restaurant-oriented food dataset with food images, dish tags, and restaurant-level information, such as the menu and geolocation. Experiments on this dataset show that exploiting geolocation improves around 30% the recognition performance, and geolocalized models contribute with an additional 3-8% absolute gain, while they can be trained up to five times faster. Ruihan Xu 0001, Luis Herranz, Shuqiang Jiang, Xinhang Song, Ramesh Jain 0001 |
IEEE Trans. Multim. | 6 |
| 2014 | A Real-time Complex Event Discovery Platform for Cyber-Physical-Social SystemsabstractWe are living in the Internet of Things (IoT) era where all the (smart) objects around us are connected and communicated with each other to serve our life better without the need of explicit instruction. Soon we have to cope with trillions of heterogeneous data streams coming from IoT. Since data is not information, methods for discovering useful and correlative information from data and utilising them for the better life, in real-time mode, are the utmost requirements. Minh-Son Dao, Siripen Pongpaichet, Laleh Jalali, Kyoung-Sook Kim 0001, Ramesh Jain 0001, Koji Zettsu |
ICMR | 5 |
| 2014 | Real-life events in multimedia: detection, representation, retrieval, and applications
Vasileios Mezaris, Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli |
Multim. Tools Appl. | 3 |
| 2014 | Collaborative Online Multitask LearningabstractWe study the problem of online multitask learning for solving multiple related classification tasks in parallel, aiming at classifying every sequence of data received by each task accurately and efficiently. One practical example of online multitask learning is the micro-blog sentiment detection on a group of users, which classifies micro-blog posts generated by each user into emotional or non-emotional categories. This particular online learning task is challenging for a number of reasons. First of all, to meet the critical requirements of online applications, a highly efficient and scalable classification solution that can make immediate predictions with low learning cost is needed. This requirement leaves conventional batch learning algorithms out of consideration. Second, classical classification methods, be it batch or online, often encounter a dilemma when applied to a group of tasks, i.e., on one hand, a single classification model trained on the entire collection of data from all tasks may fail to capture characteristics of individual task; on the other hand, a model trained independently on individual tasks may suffer from insufficient training data. To overcome these challenges, in this paper, we propose a collaborative online multitask learning method, which learns a global model over the entire data of all tasks. At the same time, individual models for multiple related tasks are jointly inferred by leveraging the global model through a collaborative online learning approach. We illustrate the efficacy of the proposed technique on a synthetic dataset. We also evaluate it on three real-life problems-spam email filtering, bioinformatics data classification, and micro-blog sentiment detection. Experimental results show that our method is effective and scalable at the online classification of multiple related tasks. Guangxia Li, Steven C. H. Hoi, Kuiyu Chang, Ramesh Jain 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2013 | Visualizing progressive discoveryabstractComputational problems are increasingly relying on context-aware approaches for tractable solutions. Usually, these approaches statically link additional sources of information to those already present in the problem space. We have been building CueNet, a context discovery framework, which will dynamically discover the most relevant context for a given application problem. In this demonstration, we will show how the identities of people in personal photos can be discovered through contextual information. In this demonstration, we present Picatrix: an event based photo browsing web interface. Users can select a photo, and see a live visualization of how our context discovery algorithm, seeded with the initial information, discovers context from different data sources, and uses it to tag the faces in the given photo. Arjun Satish, Ramesh Jain 0001, Amarnath Gupta |
ICMR | 2 |
| 2013 | Social life networks: a multimedia problem?abstractConnecting people to the resources they need is a fundamental task for any society. We present the idea of a technology that can be used by the middle tier of a society so that it uses people's mobile devices and social networks to connect the needy with providers. We conceive of a world observatory called the Social Life Network (SLN) that connects together people and things and monitors for people's needs as their life situations evolve. To enable such a system we need SLN to register and recognize situations by combining people's activities and data streaming from personal devices and environment sensors, and based on the situations make the connections when possible. But is this a multimedia problem? We show that many pattern recognition, machine learning, sensor fusion and information retrieval techniques used in multimedia-related research are deeply connected to the SLN problem. We sketch the functional architecture of such a system and show the place for these techniques. Amarnath Gupta, Ramesh Jain 0001 |
ACM Multimedia | 2 |
| 2013 | Summary abstract for the 1st ACM international workshop on personal data meets distributed multimediaabstractMultimedia data are now created at a macro, public scale as well as individual personal scale. While distributed multimedia streams (e.g. images, microblogs, and sensor readings) have recently been combined to understand multiple spatio-temporal phenomena like epidemic spreads, seasonal patterns, and political situations; personal data (via mobile sensors, quantified-self technologies) are now being used to identify user behavior, intent, affect, social connections, health, gaze, and interest level in real time. An effective combination of the two types of data can revolutionize multiple applications ranging from healthcare, to mobility, to product recommendation, to content delivery. Building systems at this intersection can lead to better orchestrated media systems that may also improve users' social, emotional and physical well-being. For example, users trapped in risky hurricane situations can receive personalized evacuation instructions based on their health, mobility parameters, and distance to nearest shelter. This workshop bring together researchers interested in exploring novel techniques that combine multiple streams at different scales (macro and micro) to understand and react to each user's needs. Vivek K. Singh 0001, Tat-Seng Chua, Ramesh Jain 0001, Alex Pentland |
ACM Multimedia | 3 |
| 2013 | Creation of individual photo selections: read preferences from the users' eyesabstractThe automated selection of satisfying subsets from large collections of photos is a central challenge in multimedia research. Objective criteria like the depiction of persons or the photo quality are met by existing approaches. But it is difficult to know the users' personal interest, which plays an important role in the selection process. The expected spread of devices with eye tracking support in the near future allows us to measure this interest in a new way. In an experiment with 12 participants, we derive the most interesting photos of a collection for every person from gaze information recorded during the free viewing of the photos. We can show that the eye tracking information delivers valuable information about the users' preferences by comparing the results to a manual selection. The selection based on gaze information significantly outperforms baseline approaches and improves the results by up to 17%. For photo sets of personal interest this improvement is even up to 23%. Tina Walber, Chantal Neuhaus, Steffen Staab, Ansgar Scherp, Ramesh Jain 0001 |
ACM Multimedia | 5 |
| 2013 | Label-specific training set construction from web resource for image annotation
Jinhui Tang 0001, Shuicheng Yan, Chunxia Zhao, Tat-Seng Chua, Ramesh Jain 0001 |
Signal Process. | 5 |
| 2013 | Social image tagging using graph-based reinforcement on multi-type interrelated objects
Xiaoming Zhang 0001, Xiaojian Zhao, Zhoujun Li 0001, Jiali Xia, Ramesh Jain 0001, Wen-Han Chao |
Signal Process. | 5 |
| 2013 | Identification of scene locations from geotagged imagesabstractDue to geotagging capabilities of consumer cameras, it has become easy to capture the exact geometric location where a picture is taken. However, the location is not the whereabouts of the scene taken by the photographer but the whereabouts of the photographer himself. To determine the actual location of an object seen in a photo some sophisticated and tiresome steps are required on a special camera rig, which are generally not available in common digital cameras. This article proposes a novel method to determine the geometric location corresponding to a specific image pixel. A new technique of stereo triangulation is introduced to compute the relative depth of a pixel position. Geographical metadata embedded in images are utilized to convert relative depths to absolute coordinates. When a geographic database is available we can also infer the semantically meaningful description of a scene object from where the specified pixel is projected onto the photo. Experimental results demonstrate the effectiveness of the proposed approach in accurately identifying actual locations. Jong-Seung Park, Ramesh Jain 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Towards optimizing human labeling for interactive image taggingabstractInteractive tagging is an approach that combines human and computer to assign descriptive keywords to image contents in a semi-automatic way. It can avoid the problems in automatic tagging and pure manual tagging by achieving a compromise between tagging performance and manual cost. However, conventional research efforts on interactive tagging mainly focus on sample selection and models for tag prediction. In this work, we investigate interactive tagging from a different aspect. We introduce an interactive image tagging framework that can more fully make use of human's labeling efforts. That means, it can achieve a specified tagging performance by taking less manual labeling effort or achieve better tagging performance with a specified labeling cost. In the framework, hashing is used to enable a quick clustering of image regions and a dynamic multiscale clustering labeling strategy is proposed such that users can label a large group of similar regions each time. We also employ a tag refinement method such that several inappropriate tags can be automatically corrected. Experiments on a large dataset demonstrate the effectiveness of our approach Jinhui Tang 0001, Qiang Chen 0007, Meng Wang 0001, Shuicheng Yan, Tat-Seng Chua, Ramesh Jain 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2013 | GPSView: A scenic driving route plannerabstractGPS devices have been widely used in automobiles to compute navigation routes to destinations. The generated driving route targets the minimal traveling distance, but neglects the sightseeing experience of the route. In this study, we propose an augmented GPS navigation system, GPSView , to incorporate a scenic factor into the routing. The goal of GPSView is to plan a driving route with scenery and sightseeing qualities, and therefore allow travelers to enjoy sightseeing on the drive. To do so, we first build a database of scenic roadways with vistas of landscapes and sights along the roadside. Specifically, we adapt an attention-based approach to exploit community-contributed GPS-tagged photos on the Internet to discover scenic roadways. The premise is: a multitude of photos taken along a roadway imply that this roadway is probably appealing and catches the public's attention. By analyzing the geospatial distribution of photos, the proposed approach discovers the roadside sight spots, or Points-Of-Interest (POIs), which have good scenic qualities and visibility to travelers on the roadway. Finally, we formulate scenic driving route planning as an optimization task towards the best trade-off between sightseeing experience and traveling distance. Testing in the northern California area shows that the proposed system can deliver promising results. Yantao Zheng, Shuicheng Yan, Zhengjun Zha, Tat-Seng Chua, Ramesh Jain 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2012 | Situation recognition: an evolving problem for heterogeneous dynamic big multimedia dataabstractWith the growth in social media, internet of things, and planetary-scale sensing there is an unprecedented need to assimilate spatio-temporally distributed multimedia streams into actionable information. Consequently the concepts like objects, scenes, and events, need to be extended to recognize situations (e.g. epidemics, traffic jams, seasons, flash mobs). This paper motivates and computationally grounds the problem of situation recognition. It describes a systematic approach for combining multimodal real-time big data into actionable situations. Specifically it presents a generic approach for modeling and recognizing situations. A set of generic building blocks and guidelines help the domain experts model their situations of interest. The created models can be tested, refined, and deployed into practice using a developed system (EventShop). Results of applying this approach to create multiple situation-aware applications by combining heterogeneous streams (e.g. Twitter, Google Insights, Satellite imagery, Census) are presented. Vivek K. Singh 0001, Mingyan Gao, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2012 | Societally connected multimedia across culturesabstractThe advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu |
J. Zhejiang Univ. Sci. C | 3 |
| 2012 | Introduction to the special issue of the multimedia tools and applications journal on events in multimedia
Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli |
Multim. Tools Appl. | 2 |
| 2012 | Large-Scale Situation Awareness With Camera Networks and Multimodal SensingabstractSensors of various modalities and capabilities, especially cameras, have become ubiquitous in our environment. Their intended use is wide ranging and encompasses surveillance, transportation, entertainment, education, healthcare, emergency response, disaster recovery, and the like. Technological advances and the low cost of such sensors enable deployment of large-scale camera networks in large metropolises such as London and New York. Multimedia algorithms for analyzing and drawing inferences from video and audio have also matured tremendously in recent times. Despite all these advances, large-scale reliable systems for media-rich sensor-based applications, often classified as situation-awareness applications, are yet to become commonplace. Why is that? There are several forces at work here. First, the system abstractions are just not at the right level for quickly prototyping such applications on a large scale. Second, while Moore's law has held true for predicting the growth of processing power, the volume of data that applications are called upon to handle is growing similarly, if not faster. Enormous amount of sensing data is continually generated for real-time analysis in such applications. Further, due to the very nature of the application domain, there are dynamic and demanding resource requirements for such analyses. The lack of right set of abstractions for programing such applications coupled with their data-intensive nature have hitherto made realizing reliable large-scale situation-awareness applications difficult. Incidentally, situation awareness is a very popular but ill-defined research area that has attracted researchers from many different fields. In this paper, we adopt a strong systems perspective and consider the components that are essential in realizing a fully functional situation-awareness system. Umakishore Ramachandran, Kirak Hong, Liviu Iftode, Ramesh Jain 0001, Kurt Rothermel, JunSuk Shin, Raghupathy Sivakumar |
Proc. IEEE | 4 |
| 2012 | Introduction to the Special Section on Intelligent Multimedia Systems and Technology Part IIabstractNo abstract available. Xian-Sheng Hua 0001, Qi Tian 0001, Alberto Del Bimbo, Ramesh Jain 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2012 | Quantitative Characterization of Semantic Gaps for Learning Complexity Estimation and Inference Model SelectionabstractIn this paper, a novel data-driven algorithm is developed for achieving quantitative characterization of the semantic gaps directly in the visual feature space, where the visual feature space is the common space for concept classifier training and automatic concept detection. By supporting quantitative characterization of the semantic gaps, more effective inference models can automatically be selected for concept classifier training by: (1) identifying the image concepts with small semantic gaps (i.e., the isolated image concepts with high inner-concept visual consistency) and training their one-against-all SVM concept classifiers independently; (2) determining the image concepts with large semantic gaps (i.e., the visually-related image concepts with low inner-concept visual consistency) and training their inter-related SVM concept classifiers jointly; and (3) using more image instances to achieve more reliable training of the concept classifiers for the image concepts with large semantic gaps. Our experimental results on NUS-WIDE and ImageNet image sets have obtained very promising results. Jianping Fan 0001, Xiaofei He 0001, Jinye Peng 0001, Ramesh Jain 0001 |
IEEE Trans. Multim. | 5 |
| 2011 | Collaborative online learning of user generated contentabstractWe study the problem of online classification of user generated content, with the goal of efficiently learning to categorize content generated by individual user. This problem is challenging due to several reasons. First, the huge amount of user generated content demands a highly efficient and scalable classification solution. Second, the categories are typically highly imbalanced, i.e., the number of samples from a particular useful class could be far and few between compared to some others (majority class). In some applications like spam detection, identification of the minority class often has significantly greater value than that of the majority class. Last but not least, when learning a classification model from a group of users, there is a dilemma: A single classification model trained on the entire corpus may fail to capture personalized characteristics such as language and writing styles unique to each user. On the other hand, a personalized model dedicated to each user may be inaccurate due to the scarcity of training data, especially at the very beginning; when users have written just a few articles. To overcome these challenges, we propose learning a global model over all users' data, which is then leveraged to continuously refine the individual models through a collaborative online learning approach. The class imbalance problem is addressed via a cost-sensitive learning approach. Experimental results show that our method is effective and scalable for timely classification of user generated content. Guangxia Li, Kuiyu Chang, Steven C. H. Hoi, Ramesh Jain 0001 |
CIKM | 5 |
| 2011 | Extractive summarization of personal photos from life eventsabstractManually sifting through large collections of personal photos shot at various life events is both tedious and inefficient. In this paper, we propose a photo summarization system which creates a representative subset summary by extracting photos from a larger set shot at an event (e.g., in a trip, birthday, etc). We define three properties that are necessary to generate an effective summary: relevance, diversity and coverage. We propose methods to compute them using multimodal content and context data. The objective for automatic photo summarization is formulated as an optimization of these properties. We discuss algorithms that solve the problem efficiently. A dataset of 7,700 photos from personal life events is created with user-generated ground truth summary. We also propose objective metrics to evaluate summaries automatically. Evaluations using both objective metrics and user feedback show our models can generate summaries which are much better than baselines. Pinaki Sinha, Ramesh Jain 0001 |
ICME | 2 |
| 2011 | Summarization of personal photologs using multidimensional content and contextabstractIn this paper, we propose a framework for generation of representative subset summaries from large personal photo collections. These summaries will help in effective sharing and browsing of the personal photos. We define three salient properties: quality, diversity and coverage that an informative summary should satisfy. We propose methods to compute these properties using multidimensional content and context data. The objective of summarization is modeled as an optimization of these properties, given the size constraints. We also propose metrics which will evaluate the photo summaries based on their representation of the larger corpus and the ability to satisfy user's information needs. We use a dataset of 40K personal photos collected by crawling photo sharing and storage sites of sixteen users. Our experiments show that the summarization algorithm works better than the baseline algorithms. Pinaki Sinha, Sharad Mehrotra, Ramesh Jain 0001 |
ICMR | 3 |
| 2011 | Modeling and representing events in multimediaabstractThis paper presents an overview of the Joint Workshop on Modeling and Representing Events (JMRE), which is held as part of ACM Multimedia 2011. JMRE is concerned with the understanding of events from multimedia, and with using events in order to better organize and consume multimedia. Vasileios Mezaris, Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli, Huiyu Zhou 0001, Jianguo Zhang 0001, Liang Wang 0001, Zhengyou Zhang |
ACM Multimedia | 3 |
| 2011 | Event based experiential computing in agro-advisory system for rural farmersabstractIn our earlier proposed mKRISHI framework, which is a mobile phone based agriculture advisory system, a farmer can record a query with the help of audio-visual interface on the mobile phone application. The multi-modal sensory devices have been deployed to capture the context of the farm so that the query along with the context are presented to the agriculture expert at the web console for the investigation. In this paper, we propose an event based experiential computing approach for agro-advisory services. The classification and modeling of agricultural events, modeling of the agricultural experiences, and a method to browse through the history of agriculture experiences are the key contributions in the present work. Bhushan G. Jagyasi, Arun Pande, Ramesh Jain 0001 |
WiMob | 3 |
| 2011 | Survey papers in multimedia - guest editorial
Ramesh Jain 0001, Alberto Del Bimbo, Tat-Seng Chua, Borko Furht |
Multim. Tools Appl. | 1 |
| 2011 | Hot research topics - guest editorial
Ramesh Jain 0001, Alberto Del Bimbo, Tat-Seng Chua, Borko Furht |
Multim. Tools Appl. | 1 |
| 2011 | A comprehensive study of visual event computing
Wei Qi Yan 0001, Declan F. Kieran, Setareh Rafatirad, Ramesh Jain 0001 |
Multim. Tools Appl. | 4 |
| 2011 | Introduction to the special issue on intelligent multimedia systems and technologyabstractNo abstract available. Xian-Sheng Hua 0001, Qi Tian 0001, Alberto Del Bimbo, Ramesh Jain 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2011 | Image annotation by kNN-sparse graph-based label propagation over noisily tagged web imagesabstractIn this article, we exploit the problem of annotating a large-scale image corpus by label propagation over noisily tagged web images. To annotate the images more accurately, we propose a novel k NN-sparse graph-based semi-supervised learning approach for harnessing the labeled and unlabeled data simultaneously. The sparse graph constructed by datum-wise one-vs- k NN sparse reconstructions of all samples can remove most of the semantically unrelated links among the data, and thus it is more robust and discriminative than the conventional graphs. Meanwhile, we apply the approximate k nearest neighbors to accelerate the sparse graph construction without loosing its effectiveness. More importantly, we propose an effective training label refinement strategy within this graph-based learning framework to handle the noise in the training labels, by bringing in a dual regularization for both the quantity and sparsity of the noise. We conduct extensive experiments on a real-world image database consisting of 55,615 Flickr images and noisily tagged training labels. The results demonstrate both the effectiveness and efficiency of the proposed approach and its capability to deal with the noise in the training labels. Jinhui Tang 0001, Richang Hong, Shuicheng Yan, Tat-Seng Chua, Guo-Jun Qi, Ramesh Jain 0001 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2011 | Introduction to special issue on social mediaabstractNo abstract available. Susanne Boll, Ramesh Jain 0001, Jiebo Luo 0001, Dong Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2010 | Efficient Approximate Visibility Query in Large Dynamic Environments
Leyla Kazemi, Farnoush Banaei Kashani, Cyrus Shahabi, Ramesh Jain 0001 |
DASFAA (1) | 4 |
| 2010 | Micro-blogging Sentiment Detection by Collaborative Online LearningabstractWe study the online micro-blog sentiment detection problem, which aims to determine whether a micro-blog post expresses emotions. This problem is challenging because a micro-blog post is very short and individuals have distinct ways of expressing emotions. A single classification model trained on the entire corpus may fail to capture characteristics unique to each user. On the other hand, a personalized model for each user may be inaccurate due to the scarcity of training data, especially at the very beginning where users have just posted a few entries. To overcome these challenges, we propose learning a global model over all micro-bloggers, which is then leveraged to continuously refine the individual models through a collaborative online learning way. We evaluate our algorithm on a real-life micro-blog dataset collected from the popular micro-blog site - Twitter. Results show that our algorithm is effective and efficient for timely sentiment detection in real micro-blogging applications. Guangxia Li, Steven C. H. Hoi, Kuiyu Chang, Ramesh Jain 0001 |
ICDM | 4 |
| 2010 | W2Go: a travel guidance system by automatic landmark rankingabstractIn this paper, we present a travel guidance system W2Go (Where to Go), which can automatically recognize and rank the landmarks for travellers. In this system, a novel Automatic Landmark Ranking (ALR) method is proposed by utilizing the tag and geo-tag information of photos in Flickr and user knowledge from Yahoo Travel Guide. ALR selects the popular tourist attractions (landmarks) based on not only the subjective opinion of the travel editors as is currently done on sites like WikiTravel and Yahoo Travel Guide, but also the ranking derived from popularity among tourists. Our approach utilizes geo-tag information to locate the positions of the tag-indicated places, and computes the probability of a tag being a landmark/site name. For potential landmarks, impact factors are calculated from the frequency of tags, user numbers in Flickr, and user knowledge in Yahoo Travel Guide. These tags are then ranked based on the impact factors. Several representative views for popular landmarks are generated from the crawled images with geo-tags to describe and present them in context of information derived from several relevant reference sources. The experimental comparisons to the other systems are conducted on eight famous cities over the world. User-based evaluation demonstrates the effectiveness of the proposed ALR method and the W2Go system. Yue Gao 0002, Jinhui Tang 0001, Richang Hong, Qionghai Dai, Tat-Seng Chua, Ramesh Jain 0001 |
ACM Multimedia | 6 |
| 2010 | Content without context is meaninglessabstractWe revisit one of the most fundamental problems in multimedia that is receiving enormous attention from researchers without making much progress in solving it: the problem of bridging the semantic gap. Research in this area has focused on developing increasingly rigorous techniques using the content. Researchers consider that Content is King and ignore everything else. In this paper, first we will discuss how this infatuation with content continues to be the biggest hurdle in the success of, ironically, content based approaches for multimedia search. Lately, many commercial systems have ignored content in favor of context and demonstrated better success. Given that the mobile phones are the major platform for the next generation of computing, context becomes easily available and more relevant. We show that it is not Content Versus Context; rather it is Content and Context that is required to bridge the semantic gap. In this paper, first we will discuss reasons for our approach and then present approaches that appropriately combine context with content to help bridge the semantic gap and solve important problems in multimedia computing. Ramesh Jain 0001, Pinaki Sinha |
ACM Multimedia | 1 |
| 2010 | Modeling, detecting, and processing events in multimediaabstractNo abstract available. Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli, Vasileios Mezaris |
ACM Multimedia | 2 |
| 2010 | Social pixels: genesis and evaluationabstractHuge amounts of social multimedia is being created daily by a combination of globally distributed disparate sensors, including human-sensors (e.g. tweets) and video cameras. Taken together, this represents information about multiple aspects of the evolving world. Understanding the various events, patterns and situations emerging in such data has applications in multiple domains. We develop abstractions and tools to decipher various spatio-temporal phenomena which manifest themselves across such social media data. We describe an approach for aggregating social interest of users about any particular theme from any particular location into 'social pixels'. Aggregating such pixels spatio-temporally allows creation of social versions of images and videos, which then become amenable to various media processing techniques (like segmentation, convolution) to derive semantic situation information. We define a declarative set of operators upon such data to allow for users to formulate queries to visualize, characterize, and analyze such data. Results of applying these operations over an evolving corpus of millions of Twitter and Flickr posts, to answer situation-based queries in multiple application domains are promising. Vivek K. Singh 0001, Mingyan Gao, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2010 | One person labels one million imagesabstractTargeting the same objective of alleviating the manual work as automatic annotation, in this paper, we propose a novel framework with minimal human effort to manually annotate a large-scale image corpus. In this framework, a dynamic multi-scale cluster labeling strategy is proposed to manually label the clusters of similar image regions. The users label the multi-scale clusters of regions instead of individual images, thus each labeling operation can annotate hundreds or even thousands of images simultaneously with much reduced manual work. Meanwhile the manual labeling guarantees the accuracy of the labels. Compared to automatic annotation, the proposed framework is more flexible, general and effective, especially for annotating those labels with large semantic gaps. Experiments on NUS-WIDE dataset demonstrate that the proposed fast manual annotation framework is much more effective than automatic annotation and comparatively efficient. Jinhui Tang 0001, Qiang Chen 0007, Shuicheng Yan, Tat-Seng Chua, Ramesh Jain 0001 |
ACM Multimedia | 5 |
| 2010 | Overview of ACM international workshop on connected multimediaabstractFollowing the very first international workshop on connected multimedia held in Hangzhou, China, in October of 2009 jointly sponsored by US National Science Foundation and Zhejiang University of China, this is the very first ACM International Workshop on Connected Multimedia in conjunction with ACM International Conference on Multimedia held in Florence, Italy, in October of 2010. In this workshop overview, we first define what we mean by connected multimedia, and then briefly overview the program of this workshop. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang |
ACM Multimedia | 3 |
| 2010 | Spatio-temporal Event Stream Processing in Multimedia Communication Systems
Mingyan Gao, Ramesh Jain 0001, Beng Chin Ooi |
SSDBM | 3 |
| 2010 | What the web can't doabstractThis panel discusses how polling in the HTTPd protocol affects how we are building the next generation of the web and its applications. As other technologies (HTML, Javascript, etc.) move forward, we ask should the web's protocol also evolve or is it sufficient for the web to continue through just GET and POST? David A. Shamma, Seth Fitzsimmonds, Joe Gregorio, Adam Hupp, Ramesh Jain 0001, Kevin Marks |
WWW | 5 |
| 2010 | Situation detection and control using spatio-temporal analysis of microblogsabstractLarge volumes of spatio-temporal-thematic data being created using sites like Twitter and Jaiku, can potentially be combined to detect events, and understand various 'situations' as they are evolving at different spatio-temporal granularity across the world. Taking inspiration from traditional image pixels which represent aggregation of photon energies at a location, we consider aggregation of user interest levels at different geo-locations as social pixels. Combining such pixels spatio-temporally allows for creation of social images and video. Here, we describe how the use of relevant (media processing inspired) situation detection operators upon such 'images', and domain based rules can be used to decide relevant control actions. The ideas are showcased using a Swine flu monitoring application which uses Twitter data. Vivek K. Singh 0001, Mingyan Gao, Ramesh Jain 0001 |
WWW | 3 |
| 2010 | Structural analysis of the emerging event-webabstractEvents are the fundamental abstractions to study the dynamic world. We believe that the next generation of web (i.e. event-web), will focus on interconnections between events as they occur across space and time [3]. In fact we argue that the real value of large volumes of microblog data being created daily lies in its inherent spatio-temporality, and its correlation with the real-world events. In this context, we studied the structural properties of a corpus of 5,835,237 Twitter microblogs, and found it to exhibit Power laws across space and time, much like those exhibited by events in multiple domains. The properties studied over microblogs on different topics can be applied to study relationships between related events, as well as data organization for event-based, real-time, and location-aware applications. Vivek K. Singh 0001, Ramesh Jain 0001 |
WWW | 2 |
| 2009 | A Flexible Surveillance System ArchitectureabstractTraditional multimedia surveillance systems are task specific and tightly coupled to the environment. Moreover, system designs generally start with the assumption that the environment, context, and sensors always remain static. With such a tight coupling, it becomes very difficult to port the system to new environments. Furthermore, for most of the systems, there is no straightforward way to upgrade the existing system to incorporate technological advancements such as new sensors or novel feature extraction techniques. We propose a flexible surveillance system architecture which can be easily ported in different environments, is dynamic without any significant compromise in system performance, and can be extended to integrate newer technological developments. We also introduce the notion of environment model (EM), which completely defines the coupling between system and the physical environment. The isolation of environment specific variables in EM makes the system easily portable in different environments. We present results of a prototype implementation of the system that highlights our design goals. Mukesh Saini, Mohan Kankanhalli, Ramesh Jain 0001 |
AVSS | 3 |
| 2009 | SMPL, a Specification Based Framework for the Semantic Structure, Annotation and Control of SMIL DocumentsabstractIn this paper we describe the design and implementation of a framework for declarative XML languages-SMPL. The framework is intended to support the semantic structure, annotation and control of SMIL documents. SMPL is supported by all browsers de-facto. Unlike perhaps other Web standards, where adding functionality has to be included in a release of a new version for a plug-in or a Web browser, SMPL is extendable and yet requires no specific changes to the implementation of a plug-in or the browser, beyond it's support for SMIL. We envision SMPL as the framework for the next generation semantic Multimedia document management languages. It supports the creation of different languages that facilitate different requirements for the semantic structuring, annotation or control of SMIL synchronized multimedia documents. We evaluate SMPL in the context of a distant learning environment and a drill replay and management application. Ronen Vaisenberg, Ramesh Jain 0001, Sharad Mehrotra |
ISM | 2 |
| 2009 | Events in multimediaabstractNo abstract available. Ansgar Scherp, Ramesh Jain 0001, Mohan Kankanhalli |
ACM Multimedia | 2 |
| 2009 | Personal photo album summarizationabstractPhoto album summarization is the process of selecting a subset of photos from a larger collection which best preserves the information in the entire set and is semantically coherent. In this paper we propose a system which uses heterogeneous information sources associated with digital photos and generates a summary. Our algorithm adapts itself based on the type of event it is summarizing (Yearbook, Week or Single Day Event) We model the summarization problem as a retrieval problem based on different types of queries. We propose some evaluation metrics for the summary. We use an intuitive web based interface to present the results so that users can further explore the summary in an interactive way. This system is our submission to the CeWe Challenge for the Next Generation of Tangible Multimedia Products. Pinaki Sinha, Hamed Pirsiavash, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2009 | MEDIALIFE: from images to a life chronicleabstractdemonstration Share on MEDIALIFE: from images to a life chronicle Authors: Amarnath Gupta University of California San Diego, La Jolla, CA, USA University of California San Diego, La Jolla, CA, USAView Profile , Setareh Rafatirad University of California Irvine, Irvine, CA, USA University of California Irvine, Irvine, CA, USAView Profile , Mingyan Gao University of California Irvine, Irvine, CA, USA University of California Irvine, Irvine, CA, USAView Profile , Ramesh Jain University of California Irvine, Irvine, CA, USA University of California Irvine, Irvine, CA, USAView Profile Authors Info & Claims SIGMOD '09: Proceedings of the 2009 ACM SIGMOD International Conference on Management of dataJune 2009 Pages 1119–1122https://doi.org/10.1145/1559845.1559998Published:29 June 2009Publication History 4citation280DownloadsMetricsTotal Citations4Total Downloads280Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Amarnath Gupta, Setareh Rafatirad, Mingyan Gao, Ramesh Jain 0001 |
SIGMOD Conference | 4 |
| 2009 | Towards Environment-to-Environment (E2E) multimedia communication systems
Vivek K. Singh 0001, Hamed Pirsiavash, Ish Rishabh, Ramesh Jain 0001 |
Multim. Tools Appl. | 4 |
| 2009 | Tolkien: An Event Based Storytelling SystemabstractSince the dawn of human civilization, stories have been a popular medium of communication, both synchronously and asynchronously. Technically, a story is a time-ordered coherent sequence of events. In many applications, heterogeneous data is collected and organized so appropriate stories could be told. In this paper, we present a system that helps in generation of stories using a large database of events with associated multimodal data, called eventbase. We define storytelling as a two step process in which a storyteller can retrieve appropriate events and associated data, and then those are further filtered using preferences of the viewer. We develop this model using a measure of interestingness based on attributes of selected events and the preferences. Using an event system developed in our laboratory, we demonstrate the story telling process in Tolkien as one that generates multiple queries to select coherent interesting events to form a story. Arjun Satish, Ramesh Jain 0001, Amarnath Gupta |
Proc. VLDB Endow. | 2 |
| 2008 | Multimodal observation systemsabstractIn recent years, we have seen a significant research interest in a number of multimodal sensing applications like surveillance, video ethnography, tele-presence, assisted living, life blogging etc. However, these applications are currently evolving as separate silos with no interconnection. Further, the individual application-centric architectures typically tend to focus on specific sensors, specific (hardwired) queries and deal with specific environments. We present a generic sensing architecture 'Observation System', which allows multiple users to undertake different applications through abstracted interaction with a common set of sensors. The observation system observes behavior of various objects in an environment and keeps a record of important events and activities in an eventbase. In this system, multifarious data collected from disparate sensors and other sources are correlated to understand and gain insights in the environment. The observation system has applications in many areas including but not limited to surveillance, traffic monitoring, ethnography, marketing, and healthcare. In this paper, we present the architecture and functionality of such a system and present details of activity detection using multiple sensor streams in a distributed sensing environment. We also present results of such an approach and potential extensions to the analysis of more complex activities and events. Mukesh Saini, Vivek K. Singh 0001, Ramesh Jain 0001, Mohan Kankanhalli |
ACM Multimedia | 3 |
| 2008 | Message from the general chairsabstractPresents the introductory welcome message from the conference proceedings. Ramesh Jain 0001, Mohan Kumar |
WOWMOM | 1 |
| 2008 | MedSMan: a live multimedia stream querying system
Bin Liu 0010, Amarnath Gupta, Ramesh Jain 0001 |
Multim. Tools Appl. | 3 |
| 2008 | Mining Multilevel Image Semantics via Hierarchical ClassificationabstractIn this paper, we have proposed a novel framework for mining multilevel image semantics via hierarchical classification. To bridge the semantic gap more successfully, salient objects are used to characterize the intermediate image semantics effectively. The salient objects are defined as the connected image regions that capture the dominant visual properties linked to the corresponding physical objects in an image. To achieve a more reliable and tractable concept learning in high-dimensional feature space, a novel algorithm calledproduct of mixture-experts(PoM) is proposed to reduce the size of training images and speed up concept learning. A novel hierarchical concept learning algorithm is proposed by incorporating concept ontology and multitask learning to enhance the discrimination power of the concept models and reduce the computational complexity for learning the concept models for large amount of image concepts, which may have huge intra-concept variations and inter-concept similarities on their visual properties. A hyperbolic image visualization algorithm has been developed for allowing users to specify their queries easily and assess the query results interactively. Our experiments on large-scale image collections have also obtained very positive results. Jianping Fan 0001, Yuli Gao, Hangzai Luo, Ramesh Jain 0001 |
IEEE Trans. Multim. | 4 |
| 2008 | Editorial: Introduction to the Special Issue on Multimedia Data MiningabstractThe twelve papers in this special issue focus on multimedia data mining. The special issue evolved from a successful workshop organized in conjunction with the 2006 ACM KDD conference, but the special issue was open to the whole community. Zhongfei Zhang, Florent Masseglia, Ramesh Jain 0001, Alberto Del Bimbo |
IEEE Trans. Multim. | 3 |
| 2007 | Design and Implementation of a Middleware for Sentient SpacesabstractSurveillance is an important task for guaranteeing the security of individuals. Being able to intelligently monitor the activity in given spaces is essential to achieve such surveillance. Sentient spaces based on a large set of sensors provide the potential for such intelligent monitoring. However, heavily instrumenting a space with sensors it is not enough to build a sentient space. One needs a software architecture that allows programming all these sensors in a transparent and efficient manner. In this paper, we present SATware, a stream acquisition and transformation middleware we are developing to analyze, query, and transform multimodal sensor data streams to facilitate flexible development of sentient environments. SATware provides a powerful application development environment in which users (i.e., application builders) can focus on the specifics of the application without having to deal with the technical peculiarities of accessing a large number of diverse sensors via different protocols. Bijit Hore, Hojjat Jafarpour, Ramesh Jain 0001, Shengyue Ji, Daniel Massaguer, Sharad Mehrotra, Nalini Venkatasubramanian, Utz Westermann |
ISI | 3 |
| 2007 | Annotation of paintings with high-level semantic concepts using transductive inference and ontology-based concept disambiguationabstract10.1145/1291233.1291335 Liza Leslie, Tat-Seng Chua, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2007 | Ontology-Based Annotation of Paintings Using Transductive Inference Framework
Liza Leslie, Tat-Seng Chua, Ramesh Jain 0001 |
MMM (1) | 3 |
| 2007 | Incorporating Concept Ontology for Hierarchical Video Classification, Annotation, and VisualizationabstractMost existing content-based video retrieval (CBVR) systems are now amenable to support automatic low-level feature extraction, but they still have limited effectiveness from a user's perspective because of the semantic gap. Automatic video concept detection via semantic classification is one promising solution to bridge the semantic gap. To speed up SVM video classifier training in high-dimensional heterogeneous feature space, a novel multimodal boosting algorithm is proposed by incorporating feature hierarchy and boosting to reduce both the training cost and the size of training samples significantly. To avoid the inter-level error transmission problem, a novel hierarchical boosting scheme is proposed by incorporating concept ontology and multitask learning to boost hierarchical video classifier training through exploiting the strong correlations between the video concepts. To bridge the semantic gap between the available video concepts and the users' real needs, a novel hyperbolic visualization framework is seamlessly incorporated to enable intuitive query specification and evaluation by acquainting the users with a good global view of large-scale video collections. Our experiments in one specific domain of surgery education videos have also provided very convincing results. Jianping Fan 0001, Hangzai Luo, Yuli Gao, Ramesh Jain 0001 |
IEEE Trans. Multim. | 4 |
| 2006 | Automatic image annotation by incorporating feature hierarchy and boosting to scale up SVM classifiersabstractThe performance of image classifiers largely depends on two inter-related issues:(1)suitable frameworks for image content representation and automatic feature extraction;(2) effective algorithms for image classifier training and feature subset selection. To address the first issue, a multiresolution grid-based framework is proposed for image content representation and feature extraction to bypass the time-consuming and erroneous process for image segmentation. To address the second issue, a hierarchical boosting algorithm is proposed by incorporating feature hierarchy and boosting to scale up SVM image classifier training in high-dimensional feature space. The high-dimensional multi-modal heterogeneous visual features are partitioned into multiple low-dimensional single-modal homogeneous feature subsets and each of them characterizes certain visual property of images. For each homogeneous feature subset, principal component analysis (PCA)is performed to exploit the feature correlations and a weak classifier is learned simultaneously. After the weak classifiers for different feature subsets and grid sizes are available, they are combined to boost an optimal classifier for the given object class or image concept, and the most representative feature subsets and grid sizes are selected. Our experiments on a specific domain of natural images have obtained very positive results. Yuli Gao, Jianping Fan 0001, Xiangyang Xue 0001, Ramesh Jain 0001 |
ACM Multimedia | 4 |
| 2006 | Event-centric multimedia data management for reconnaissance mission analysis and reportingabstractWe demonstrate the concept of event-centric multimedia data management in the context of a multimedia eChronicle for the analysis, exploration, and reporting of events in military reconnaissance missions. Unlike the traditional media-centric approach, event-centricmultimedia data management focuses on the management of real-world events; documenting media are regarded as event metadata. For the detection of mission events, we apply simple but robust spatio-temporal clustering of basic soldier state and media events with good results considering the uncontrolled environment a military patrol constitutes. The core of the architecture is generic and applicable for the event-centric management of multimedia data in other domains as well. Utz Westermann, Srikanth Agaram, Ramesh Jain 0001 |
ACM Multimedia | 4 |
| 2006 | Transductive inference using multiple experts for brushwork annotation in paintings domainabstractMany recent studies perform annotation of paintings based on brushwork. In these studies the brushwork is modeled indirectly as part of the annotation of high-level artistic concepts such as the artist name using low-level texture. In this paper, we develop a serial multi-expert framework for explicit annotation of paintings with brushwork classes. In the proposed framework, each individual expert implements transductive inference by exploiting both labeled and unlabelled data. To minimize the problem of noise in the feature space, the experts select appropriate features based on their relevance to the brushwork classes. The selected features are utilized to generate several models to annotate the unlabelled patterns. The experts select the best performing model based on Vapnik combined bound. The transductive annotation using multiple experts out-performs the conventional baseline method in annotating patterns with brushwork classes. Liza Leslie, Tat-Seng Chua, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2006 | Semi-supervised annotation of brushwork in paintings domain using serial combinations of multiple expertsabstract10.1145/1180639.1180752 Liza Leslie, Tat-Seng Chua, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2006 | Events in Multimedia Electronic Chronicles (E-Chronicles)abstractThe nature of information has changed significantly in the last two decades. Now information is multimedia, sensitive to its spatio-temporal roots, live, and dynamic. Current database and search technology is very limited in addressing organization, management, and access of emerging information systems. In this paper, we address some of the fundamental issues that must be addressed in general multimedia information management systems. We present multimedia electronic chronicles (e-chronicles) as an example of emerging systems that need such technology. We believe that to deal with dynamic information, events should be used as the fundamental basis in organizing and accessing information. Event models used in our approach capture the semantics involved in supporting e-chronicles. Our ideas are demonstrated in the context of an e-chronicle that is being implemented for a reconnaissance application using multiple disparate sensors. Utz Westermann, Ramesh Jain 0001 |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2006 | Information assimilation framework for event detection in multimedia surveillance systems
Pradeep K. Atrey, Mohan Kankanhalli, Ramesh Jain 0001 |
Multim. Syst. | 3 |
| 2006 | Experiential Sampling in Multimedia SystemsabstractMultimedia systems must deal with multiple data streams. Each data stream usually contains significant volume of redundant noisy data. In many real-time applications, it is essential to focus the computing resources on a relevant subset of data streams at any given time instant and use it to build the model of the environment. We formulate this problem as an experiential sampling problem and propose an approach to utilize computing resources efficiently on the most informative subset of data streams. First, in this paper, we focus on theoretical background and develop a theoretical framework for a single data stream. We generalize the notion of static visual attention in a dynamical systems setting and propose a dynamical attention-orientated analysis method. This is achieved by a sampling representation that utilizes the current context and past experience for attention evolution. Hence, the multimedia analysis task at hand can select its data of interest while immediately discarding the irrelevant data to achieve efficiency and adaptability. Mohan Kankanhalli, Jun Wang 0012, Ramesh Jain 0001 |
IEEE Trans. Multim. | 3 |
| 2006 | Experiential Sampling on Multiple Data StreamsabstractMultimedia systems must deal with multiple data streams. Each data stream usually contains significant volume of redundant noisy data. In many real-time applications, it is essential to focus the computing resources on a relevant subset of data streams at any given time instant and use it to build the model of the environment. We formulate this problem as an experiential sampling problem and propose an approach to utilize computing resources efficiently on the most informative subset of data streams. In this paper, we generalize our experiential sampling framework to multiple data streams and provide an evaluation measure for this technique. We have successfully applied this framework to the problems of traffic monitoring, face detection and monologue detection. Mohan Kankanhalli, Jun Wang 0012, Ramesh Jain 0001 |
IEEE Trans. Multim. | 3 |
| 2006 | Content-based multimedia information retrieval: State of the art and challengesabstractExtending beyond the boundaries of science, art, and culture, content-based multimedia information retrieval provides new paradigms and methods for searching through the myriad variety of media all over the world. This survey reviews 100+ recent articles on content-based multimedia information retrieval and discusses their role in current research directions which include browsing and search paradigms, user studies, affective computing, learning, semantic queries, new features and media types, high performance indexing, and evaluation techniques. Based on the current state of the art, we discuss the major challenges for the future. Michael S. Lew, Nicu Sebe, Chaabane Djeraba, Ramesh Jain 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2005 | Electronic Chronicles: Empowering Individuals, Groups, and OrganizationsabstractContinuing strides in processing, storage, sensing, and networking technologies are enabling people to capture their activities and experiences as greater volumes of ever-richer media. A big emerging challenge today is the organization, retrieval, and exploitation of such multimedia data surrounding the activities of individuals or enterprises. The field of multimedia electronic chronicles deals with the unified contextual organization, presentation, and analysis of temporal streams of multimedia data captured by individuals, groups, or organizations. The value of electronic chronicles is in converting activity and experience from the past into actionable intelligence in the present. Such multimedia electronic chronicles, with their associated techniques for search and navigation, analysis and reasoning, and prediction and alerting, will have enormous impact on various spheres of life spanning enhancement of personal life, business productivity, entertainment, and government operations. This paper, which serves as a companion to a keynote talk by the first author at ICME 2005, explains the notion of electronic chronicles and outlines the research challenges in this area Gopal Sarma Pingali, Ramesh Jain 0001 |
ICME | 2 |
| 2005 | MedSMan: a streaming data management system over live multimediaabstractQuerying live media streams is a challenging problem that is becoming an essential requirement in a growing number of applications. Research in multimedia information systems has addressed and made good progress in dealing with archived data. Meanwhile, research in stream databases has received significant attention for querying alphanumeric symbolic streams. The lack of a unifying data model capable of representing multimedia data and providing reasonable abstractions for querying live multimedia streams poses the challenge of how to make the best use of data in video and other sensor networks for various applications including video surveillance, live conferencing and Eventweb. This paper presents a system that enables direct capture of media streams from sensors and automatically generates meaningful feature streams that can be queried by a data stream processor. The system provides an effective combination of extensible digital processing techniques and general data stream management research. Bin Liu 0010, Amarnath Gupta, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2005 | What is the state of our community?abstract10.1145/1101149.1101297 Yong Rui, Ramesh Jain 0001, Nicolas D. Georganas, HongJiang Zhang, Klara Nahrstedt, John R. Smith, Mohan Kankanhalli |
ACM Multimedia | 2 |
| 2005 | MediaBroker: A pervasive computing infrastructure for adaptive transformation and sharing of stream data
Umakishore Ramachandran, Martin Modahl, Ilya Bagrak, Matthew Wolenetz, David J. Lillethun, Bin Liu 0010, James Kim, Phillip W. Hutto, Ramesh Jain 0001 |
Pervasive Mob. Comput. | 9 |
| 2005 | IMCE: Integrated media creation environmentabstractWe discuss the design goals for an integrated media creation environment (IMCE) aimed at enabling the average user to create media artifacts with professional qualities. The resulting requirements are implemented and we demonstrate the efficacy of the resulting system with the generation of two simple home movies. The significance for the average user seeking to create home movies lies in the flexible and automatic application of film principles to the task, removal of tedious low-level editing by means of well-formed media transformations in terms of high-level film constructs (e.g., tempo), and content repurposing powered by those same transformations added to the rich semantic information maintained at each phase of the process. Brett Adams, Svetha Venkatesh, Ramesh Jain 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2005 | Guest editorial: the international ACM Multimedia conference 1993 - ten years afterabstractNo abstract available. Ramesh Jain 0001, Thomas Plagemann, Ralf Steinmetz |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2005 | ACM SIGMM retreat report on future directions in multimedia researchabstractThe ACM Multimedia Special Interest Group was created ten years ago. Since that time, researchers have solved a number of important problems related to media processing, multimedia databases, and distributed multimedia applications. A strategic retreat was organized as part of ACM Multimedia 2003 to assess the current state of multimedia research and suggest directions for future research. This report presents the recommendations developed during the retreat. The major observation is that research in the past decade has significantly advanced hardware and software support for distributed multimedia applications and that future research should focus on identifying and delivering applications that impact users in the real-world.The retreat suggested that the community focus on solving three grand challenges: (1) make authoring complex multimedia titles as easy as using a word processor or drawing program, (2) make interactions with remote people and environments nearly the same as interactions with local people and environments, and (3) make capturing, storing, finding, and using digital media an everyday occurrence in our computing environment. The focus of multimedia researchers should be on applications that incorporate correlated media, fuse data from different sources, and use context to improve application performance. Lawrence A. Rowe, Ramesh Jain 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2004 | eVitae: An Event-Based Electronic Chronicle
Punit Gupta, Ramesh Jain 0001 |
EDBT | 4 |
| 2004 | Using Stream Semantics for Continuous Queries in Media Stream ProcessorsabstractIn the case of media and feature streams, explicit inter-stream constraints exist and can be exploited in the evaluation of continuous queries in the spirit of semantic query optimization. We express these properties using a media stream declaration language MSDL. In the demonstration, we present IMMERSI-MEET, an application built around an immersive environment. The IMMERSI-MEET system distinguishes between continuous streams, where values of different types come at a specified data rate, and discrete streams where sources push values intermittently. In MSDL, any dependence declaration must have at least one dependency specifying predicate in the body. As stream declarations are registered, the stream constraints are interpreted to construct a set of evaluation directives. Amarnath Gupta, Bin Liu 0010, Pilho Kim, Ramesh Jain 0001 |
ICDE | 4 |
| 2004 | IMCE: integrated media creation environmentabstractWe discuss the design and implementation of an integrated media creation environment, and demonstrate its efficacy in the generation of two simple home movies. The significance for the average user seeking to create home movies lies in the flexible and automatic application of film principles to the task, removal of tedious low-level editing by means of well-formed media transformations in terms of high-level film constructs (e.g. tempo), and content repurposing powered by those same transformations added to the rich semantic information maintained at each phase of the process. Brett Adams, Svetha Venkatesh, Ramesh Jain 0001 |
ICME | 3 |
| 2004 | An event model and its implementation for multimedia information representation and retrievalabstractThis work presents a novel framework, built around the notion of an event, for modeling, storage, analysis, and querying of multimedia data. We present an event model using which multimedia data, its spatio-temporal characteristics, and a variety of other potentially flexibly defined attributes can be represented. We also describe the design of a distributed event-based information management system that allows interpretation of multimedia data based on event definitions, storage of such event-based media data, and definition of various types of queries on such information. Examples involving event-based modeling and querying of multimedia data from different settings are presented to illustrate the approach. Derik Pack, Ramesh Jain 0001 |
ICME | 4 |
| 2004 | Representation and retrieval of paintings based on art history conceptsabstractThis work presents a framework for the retrieval of paintings using concepts defined in the field of art and design. These concepts are organized into meta and application-specific layers in the system. Concepts of the meta-layer take into account the visual attributes of the paintings, and their manipulations, which infer abstract attributes that influence the interpretation of the paintings by the observers. The concepts of the application-specific layer are built above the meta-layer. They represent art categories that are widely used by general users for navigating the collections of paintings. Potentially, the usage of artistic concepts for annotation supports a wider query base making the retrieval much more powerful. Liza Leslie, Tat-Seng Chua, A. Irina, Ramesh Jain 0001 |
ICME | 4 |
| 2004 | Between context-aware media capture and multimedia content analysis: where do we find the promised land?abstractNo abstract available. Susanne Boll, Dick C. A. Bulterman, Ramesh Jain 0001, Tat-Seng Chua, Rainer Lienhart, Lynn Wilcox, Marc Davis, Svetha Venkatesh |
ACM Multimedia | 3 |
| 2004 | Designing experiential environments for management of personal multimediaabstractWith the increasing ubiquity of sensors and computational resources, it is becoming easier and increasingly common for people to electronically record, photographs, text, audio, and video gathered over their lifetime. Assimilating and taking advantage of such data requires recognition of its multimedia nature, development of data models that can represent semantics across different media, representation of complex relationships in the data (such as spatio-temporal, causal, or evolutionary), and finally, development of paradigms to mediate user-media interactions. There is currently a paucity of theoretical frameworks and implementations that allow management of diverse and rich multimedia data collections in context of the aforementioned requirements. This paper presents our research in designing an experiential Multimedia Electronic Chronicle system that addresses many of these issues in the concrete context of personal multimedia information. Central to our approach is the characterization and organization of media using the concept of an "event" for unified modeling and indexing. The event-based unified multimedia model underlies the experiential user interface, which supports direct interactions with the data within a unified presentation-exploration-query environment. In this environment, explicit facilities to model space and time aid in exploration and querying as well as in representation and reasoning with dynamic relationships in the data. Experimental and comparative studies demonstrate the promise of this research. Rachel Lee Knickmeyer, Punit Gupta, Ramesh Jain 0001 |
ACM Multimedia | 4 |
| 2003 | Out-of-the Box Data Engineering - Events in Heterogeneous EnvironmentsabstractData has changed significantly over the last few decades. Computing systems that initially dealt with data and computation rapidly moved to information and communication. The next step on the evolutionary scale is insight and experience. Applications are demanding the use of live, spatio-temporal, heterogeneous data. Data engineering must keep pace by designing experiential environments that let users apply their senses to observe data and information about an event and to interact with aspects of the event that are of particular interest. We call this out-of-the-box data engineering because it means we must think beyond many of our time-worn perspectives and technical approaches. Ramesh Jain 0001 |
ICDE | 1 |
| 2003 | Semantics in Multimedia Systems
Ramesh Jain 0001 |
MMM | 1 |
| 2002 | A survey on the use of pattern recognition methods for abstraction, indexing and retrieval of images and video
Sameer K. Antani, Rangachar Kasturi, Ramesh Jain 0001 |
Pattern Recognit. | 3 |
| 2001 | TeleExperience: communicating compelling experienceabstractWe experience our environment using our natural senses: sight, sound, touch, taste, and smell. These senses combined with the knowledge of the world allow us to experience and function in the world. Data are observed facts or measurements. Information is derived from data in a specific context. Experience is direct observation or participation in an event. The development of civilization is the story of the development of understanding of 'experience' and how to share it with fellow human beings of current and future generations. The desire to share experiences and desire to experience various events where one can not be present, will continue to be the motivating factor in the development of exciting technology in the future. A look at history shows how our society evolved to become an 'information society' and is on its way to becoming an 'experience society'.Compelling and engaging experiences require immersion in a rich set of multimedia data and information so that one can directly observe a subset of the data and information. TeleExperience is a natural major step in technology evolution. It will impact every aspect of our society including education, business, sexual behavior, and health care. TeleExperience will give rise to an experience society. In this presentation, we will examine the role of multimedia information systems, multimedia presentations, situated computing, perception systems, and personalization approaches to realize TeleExperience. Then we will discuss aspects of taking a promising and exciting technology from 'laboratory to popular practice' based on practical experience. Ramesh Jain 0001 |
ACM Multimedia | 1 |
| 2001 | Emergent Semantics through Interaction in Image DatabasesabstractIn this paper, we briefly discuss some aspects of image semantics and the role that it plays for the design of image databases. We argue that images don't have an intrinsic meaning, but that they are endowed with a meaning by placing them in the context of other images and by the user interaction. From this observation, we conclude that, in an image, database users should be allowed to manipulate not only the individual images, but also the relation between them. We present an interface model based on the manipulation of configurations of images. Simone Santini, Amarnath Gupta, Ramesh Jain 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2000 | Semiorder Database for Complex Activity Recognition in Multi-Sensory EnvironmentsabstractA prototype semiorder database used for activity recognition in multi-sensory monitoring environments is described. Activities are spatio-temporal compositions of events, which are a type of atomic semantic units for such compositions. The focus is on the temporal composition of activities from events in the presence of bounded duration of temporal uncertainty in an event occurrence. Such temporal uncertainty forces the concurrency between event occurrences to be intransitive. Under certain assumptions, a subclass of partial orders, known as semiorders, models such intransitive concurrency appropriately. The semiorder database stores events and their semiorder temporal order of occurrences. A semiorder data model and the corresponding query language that embeds a semiorder pattern language are the main constituents of the semiorder database. We demonstrate this database and queries for activity recognition in a real time environment. The demonstration also includes a transducer subsystem for detection of events. Shailendra K. Bhonsle, Amarnath Gupta, Simone Santini, Ramesh Jain 0001 |
ICDE | 4 |
| 2000 | Content-Based Image Retrieval at the End of the Early YearsabstractPresents a review of 200 references in content-based image retrieval. The paper starts with discussing the working conditions of content-based retrieval: patterns of use, types of pictures, the role of semantics, and the sensory gap. Subsequent sections discuss computational steps for image retrieval systems. Step one of the review is image processing for retrieval sorted by color, texture, and local geometry. Features for retrieval are discussed next, sorted by: accumulative and global features, salient points, object and shape features, signs, and structural combinations thereof. Similarity of pictures and objects in pictures is reviewed for each of the feature types, in close connection to the types and means of feedback the user of the systems is capable of giving by interaction. We briefly discuss aspects of system engineering: databases, system architecture, and evaluation. In the concluding section, we present our view on: the driving force of the field, the heritage from computer vision, the influence on computer vision, the role of similarity and of interaction, the need for databases, the problem of evaluation, and the role of the semantic gap. Arnold W. M. Smeulders, Marcel Worring, Simone Santini, Amarnath Gupta, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 1999 | VisualGREP: A Systematic Method to Compare and Retrieve Video Sequences
Rainer Lienhart, Wolfgang Effelsberg, Ramesh Jain 0001 |
Multim. Tools Appl. | 3 |
| 1999 | Similarity MeasuresabstractWith complex multimedia data, we see the emergence of database systems in which the fundamental operation is similarity assessment. Before database issues can be addressed, it is necessary to give a definition of similarity as an operation. We develop a similarity measure, based on fuzzy logic, that exhibits several features that match experimental findings in humans. The model is dubbed fuzzy feature contrast (FFC) and is an extension to a more general domain of the feature contrast model due to Tversky (1977). We show how the FFC model can be used to model similarity assessment from fuzzy judgment of properties, and we address the use of fuzzy measures to deal with dependencies among the properties. Simone Santini, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Lip Posture Estimation using Kinematically Constrained Mixture ModelsabstractA novel approach for estimating 3D lip posture from monocular video se-quences is presented. The lips are modeled as a four body closed kinematic chain with each body possessing translational, rotational and prismatic (to ac-count for deformations) degrees of freedom. Geometric constraints relating these bodies to each other, and to the face as a whole, are used to constrain the space of possible lip postures recovered from each image. These constraints are used with the recently proposed Expectation Constrained Maximization algorithm to estimate the lip posture from video frames that have been pro-cessed (using a color segmentation algorithm described here) to identify lip regions. 1 Patrick H. Kelly, Edward Hunter, Kenneth Kreutz-Delgado, Ramesh Jain 0001 |
BMVC | 4 |
| 1998 | Content-based Multimedia Information ManagementabstractSummary form only given. All image search engines provide mechanisms to search based on keywords and provide the ability to do content-based searching using querying by pictorial example. In this paper, first we present some results from current image databases. Then we present a new approach to exploring image databases. Most of our current results are drawn from image and video asset management systems designed at Virage. The new approach is based on a navigational paradigm being developed at the University of California, San Diego. Ramesh Jain 0001 |
ICDE | 1 |
| 1998 | Event detection from continuous mediaabstractIt is difficult to extract semantical contents like events that occurred in the real world from continuous media consisting of visual, auditory, and linguistic information, because machines' ability of understanding media is still incomplete. To avoid this, we propose a strategy of collaborating multimodal information processing for continuous media, called intermodal collaboration and making full use of domain knowledge. The basic experimental results about event detection indicate that our strategy is effective. Noboru Babaguchi, Ramesh Jain 0001 |
ICPR | 2 |
| 1998 | Beyond Query by ExampleabstractThis paper considers some of the problems we found trying to atract mtig from image in database applications, and propos= some wa~s to solve th-~k argue that the m-g of an image xs an N-d@ed enti@, and it is not in general possible to derive from m image the mwning that the user of the database ~ts.&tier, we shodd be contrmt tith a wrreKation between the intended meaning ad simple perceptual du~that databases -~.&ther than ~orking on the impossible task of ~g unaxnblguous meaning from ixnag~, we shoxdd provide the user with the took he needs to tire the database in the arm of the f=ture space where %ter~~dges are -.-.. .-=. Simone Santini, Ramesh Jain 0001 |
ACM Multimedia | 2 |
| 1998 | Beyond query by exampleabstractThis paper considers some of the problems we found trying to extract meaning from images in database applications, and proposes some ways to solve them. We argue that the meaning of an image is an ill-defined entity, and it is not in general possible to derive from an image the meaning that the user of the database wants. Rather, we should be content with a correlation between the intended meaning and simple perceptual clues that databases can extract. Rather than working on the impossible task of extracting unambiguous meaning from images, we should provide the user with the tools he needs to drive the database in the areas of the feature space where "interesting" images are. Simone Santini, Ramesh Jain 0001 |
MMSP | 2 |
| 1998 | The virtual museum: an integrated text and image databaseabstractWe describe our "virtual museum" project: an union of image and text retrieval technologies that allows users to visit art collections on the Web. The virtual museum is composed of a series of rooms that the user can visit. The connections between the rooms are variable: passing from a room to another yields a result which is the outcome of a query, and depends on the query criterion that the user has selected. This "variable topology" adds interest to the museum visit since it allows the visitor, within certain limits, to customize his/her museum experience and to make it different every time. Simone Santini, Ramesh Jain 0001, Marco Corvi |
MMSP | 2 |
| 1997 | Do Images Mean Anything?abstractWe analyze the operations that form the conceptual foundation of a traditional database: the determination and matching of the meaning of a record of data. With a careful definition of the "meaning" of a piece of data, we show that matching can be naturally defined. We then analyze whether a similar concept can be applied to image databases. We argue that several concepts, that led to image database models based on matching, cannot be reasonably extended to the visual world of a perceptual search engine. Simone Santini, Ramesh Jain 0001 |
ICIP (1) | 2 |
| 1997 | 3D Video Generation with Multiple Perspective Camera ViewsabstractConventional video limits the viewers to a fixed predetermined viewpoint. With the advancement of image generation and analysis techniques, it becomes possible to overcome this limitation. We present a framework for the generation of 3D video, based on the uses of multiple cameras from different perspectives. Exploiting a priori knowledge of the scene geometry and dynamic objects, we can reconstruct 3D shapes of objects and generate virtual views of spatio-temporal events. An overview of the framework is first presented. Then we report our ongoing research to overcome issues in segmentation and playback. Li-Cheng Tai, Ramesh Jain 0001 |
ICIP (1) | 2 |
| 1997 | Content-Centric Interactive Video on the World Wide Web
Arun Katkere, Jennifer Schlenzig, Ramesh Jain 0001 |
Comput. Networks | 3 |
| 1997 | Towards Video-Based Immersive Environments
Arun Katkere, Saied Moezzi, Don Y. Kuramura, Patrick H. Kelly, Ramesh Jain 0001 |
Multim. Syst. | 5 |
| 1997 | Similarity is a Geometer
Simone Santini, Ramesh Jain 0001 |
Multim. Tools Appl. | 2 |
| 1996 | Similarity Queries in Image DatabaseabstractQuery-by-content image database will be based on similarity, rather than on matching, where similarity is a measure that is defined and meaningful for every pair of images in the image space. Since it is the human user that, in the end, has to be satisfied with the results of the query, it is natural to base the similarity measure that we will use on the characteristics of human similarity assessment. In the first part of this paper, we review some of these characteristics and define a similarity measure based on them. Another problem that similarity-based databases will have to face is how to combine different queries into a single complex query. We present a solution based on three operators that are the analogous of the and, or, and not operators one uses in traditional databases. These operators are powerful enough to express queries of unlimited complexity, yet have a very intuitive behavior, making easy for the user to specify a query tailored to a particular need. Simone Santini, Ramesh Jain 0001 |
CVPR | 2 |
| 1996 | Similarity Indexing with the SS-treeabstractEfficient indexing of high dimensional feature vectors is important to allow visual information systems and a number other applications to scale up to large databases. We define this problem as "similarity indexing" and describe the fundamental types of "similarity queries" that we believe should be supported. We also propose a new dynamic structure for similarity indexing called the similarity search tree or SS-tree. In nearly every test we performed on high dimensional data, we found that this structure performed better than the R*-tree. Our tests also show that the SS-tree is much better suited for approximate queries than the R*-tree. David A. White 0002, Ramesh Jain 0001 |
ICDE | 2 |
| 1996 | Automated diagnosis and image understanding with object extraction, object classification, and inferencing in retinal imagesabstractMedical imaging is shifting from film to electronic images. The STARE (structured analysis of the retina) system is a sophisticated image management system that will automatically diagnose images, compare images, measure key features in images, annotate image contents, and search for images similar in content. We concentrate on automated diagnosis. The images are annotated by segmentation of objects of interest, classification of the extracted objects, and reasoning about the image contents. The inferencing is accomplished with Bayesian networks that learn from image examples of each disease. This effort at image understanding in fundus images anticipates the future use of medical images. As these capabilities mature, we expect that ophthalmologists and physicians in other fields that rely in images will use a system like STARE to reduce repetitive work, to provide assistance to physicians in difficult diagnoses or with unfamiliar diseases, and to manage images in large image databases. 1. Michael H. Goldbaum, Saied Moezzi, Adam L. Taylor, Shankar Chatterjee, Jeffrey E. Boyd, Edward Hunter, Ramesh Jain 0001 |
ICIP (3) | 7 |
| 1996 | Content-based retrieval of ophthalmological imagesabstractThis paper describes steps towards an information system for the storage and content-based retrieval of ocular fundus images. Based on the Virage Incorporated framework for defining similarity metrics, the authors have developed a number of primitives for the representation of ocular fundus images. A prototype Query By Pictorial Example (QBPE) system yields similarity rankings in approximate agreement with those of a human expert. Amarnath Gupta, Saied Moezzi, Adam L. Taylor, Shankar Chatterjee, Ramesh Jain 0001, Michael H. Goldbaum, S. Burgess |
ICIP (3) | 5 |
| 1996 | Visual computing education at UCSDabstractOur goal at the University of California, San Diego, is to emphasize the increased importance of visual computing and image engineering by expanding and redesigning our curriculum. The planned curriculum development includes a 1-year undergraduate sequence in image processing and machine perception, and a redesign of the current core graduate sequence to put increased emphasis on synthesis of ideas from separate areas into a core sequence of visual computing. We are also introducing many advanced courses at both the undergraduate and graduate levels. Ramesh Jain 0001, Pamela C. Cosman |
ICIP (1) | 1 |
| 1996 | Gabor space and the development of preattentive similarityabstractWe show that a certain class of similarity measures, which is based on set-theoretic concepts, and explains many of the characteristics of human similarity assessment, can be interpreted as a distance in a suitable psychological space. This view unifies a number of different measures of similarity that psychological experiments have determined to be active in humans for different classes of stimuli. The study arises out of a consideration of similarity in retrieval from a multimedia database. Simone Santini, Ramesh Jain 0001 |
ICPR | 2 |
| 1996 | Interactive Video on WWW: Beyond VCR-Like Interfaces
Arun Katkere, Jennifer Schlenzig, Amarnath Gupta, Ramesh Jain 0001 |
Comput. Networks | 4 |
| 1995 | Similarity Matching
Simone Santini, Ramesh Jain 0001 |
ACCV | 2 |
| 1995 | Multiple Perspective Interactive Video
Arun Katkere, Don Y. Kuramura, Saied Moezzi, Patrick H. Kelly, Deborah Swanberg, Koji Wakimoto, Edward Hunter, Li-Cheng Tai, Shankar Chatterjee, Ramesh Jain 0001 |
IJCAI | 10 |
| 1995 | An Architecture for Multiple Perspective Interactive VideoabstractNo abstract available. Patrick H. Kelly, Arun Katkere, Don Y. Kuramura, Saied Moezzi, Shankar Chatterjee, Ramesh Jain 0001 |
ACM Multimedia | 6 |
| 1995 | Production Model Based Digital Video Segmentation
Arun Hampapur, Ramesh Jain 0001, Terry E. Weymouth |
Multim. Tools Appl. | 2 |
| 1995 | Block-structured recurrent neural networks
Simone Santini, Alberto Del Bimbo, Ramesh Jain 0001 |
Neural Networks | 3 |
| 1995 | VLSI Architectures for High-Speed Range EstimationabstractDepth recovery from gray-scale images is an important topic in the field of computer and robot vision. Intensity gradient analysis (IGA) is a robust technique for inferring depth information from a sequence of images acquired by a sensor undergoing translational motion. IGA obviates the need for explicitly solving the correspondence problem and hence is an efficient technique for range estimation. Many applications require real time processing at very high frame rates. The design of special purpose hardware could significantly speed up the computations in IGA. In this paper, we propose two VLSI architectures for high-speed range estimation based on IGA. The architectures fully utilize the principles of pipelining and parallelism in order to obtain high speed and throughput. The designs are conceptually simple and suitable for implementation in VLSI.> Raghu Sastry, N. Ranganathan, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1995 | Evidential reasoning for building environment mapsabstractWe address the problem of building a map of the environment utilizing sensory depth information obtained from multiple viewpoints. The desired representation of the environment is in the form of a finite-resolution three-dimensional grid of voxels. Each voxel within the grid is assigned a binary value corresponding to its occupancy state. We present an approach for multi-sensory depth information assimilation based on the Dempster-Shafer theory for evidential reasoning. This approach provides a mechanism to explicitly model ignorance which is desirable when dealing with an unknown environment. We present results obtained from this approach on a laboratory stereo sequence and contrast these with results obtained from the traditional Bayesian approach.> Arun P. Tirumalai, Brian G. Schunck, Ramesh Jain 0001 |
IEEE Trans. Syst. Man Cybern. | 3 |
| 1994 | Digital Video SegmentationabstractThe data driven, bottom up approach to video segmentation has ignored the inherent structure that exists in video. This work uses the model driven approach to digital video segmentation. Mathematical models of video based on video production techniques are formulated. These models are used to classify the edit effects used in video and film production. The classes and models are used to systematically design the feature detectors for detecting edit effects in digital video. Digital video segmentation is formulated as a feature based classification problem. Experimental results from segmenting cable television programming with cuts, fades, dissolves and page translate edits are presented. Arun Hampapur, Terry E. Weymouth, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 1994 | Recursive identification of gesture inputs using hidden Markov modelsabstractHuman-machine interfaces play a role of growing importance as computer technology continues to evolve. Motivated by the desire to provide users with an intuitive gesture input system, we describe the design of a recursive filter applied to the vision-based gesture interpretation problem. The gestures are modeled as a hidden Markov model with the state representing the gesture sequences, and the observations being the current static hand pose. At each time step the recursive filter updates its estimate of what gesture is occurring based on the current extracted pose information. The result is a robust system which provides the user with continual feedback during compound gestures.> Jennifer Schlenzig, Edward Hunter, Ramesh Jain 0001 |
WACV | 3 |
| 1994 | Knowledge Caching for Sensor-Based Systems
Yuval Roth, Ramesh Jain 0001 |
Artif. Intell. | 2 |
| 1994 | Vector Field Analysis for Oriented PatternsabstractPresents a method, based on the properties of vector fields, for the estimation of a set of symbolic descriptors (node, saddle, star-node, improper-node, center, and spiral) from linear orientation fields. Planar first-order phase portraits are used to model the linear orientation fields. A weighted linear estimator is developed to estimate linear phase portraits, using only the flow orientation. A classification scheme for planar first-order phase portraits, based on their local properties: curl, divergence, and deformation is developed. The authors present results of experiments on noise-added synthetic flow patterns and real oriented textures.> Chiao-Fe Shu, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | A New Generalized Computational Framework for Finding Object Orientation Using Perspective Trihedral Angle ConstraintabstractThis paper investigates a fundamental problem of determining the position and orientation of a three-dimensional (3-D) object using a single perspective image view. The technique is focused on the interpretation of trihedral angle constraint information. A new closed form solution based on Kanatani's formulation is proposed. The main distinguishing feature of the authors' method over the original Kanatani formulation is that their approach gives an effective closed form solution for a general trihedral angle constraint. The method also provides a general analytic technique for dealing with a class of problem of shape from inverse perspective projection by using "angle to angle correspondence information." A detailed implementation of the authors' technique is presented. Different trihedral angle configurations were generated using synthetic data for testing the authors' approach of finding object orientation by angle to angle constraint. The authors performed simulation experiments by adding some noise to the synthetic data for evaluating the effectiveness of their method in a real situation. It has been found that the authors' method worked effectively in a noisy environment which confirms that the method is robust in practical application.> Yuyan Wu, S. Sitharama Iyengar, Ramesh Jain 0001, Santanu Bose |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1994 | A robust backpropagation learning algorithm for function approximationabstractThe backpropagation (BP) algorithm allows multilayer feedforward neural networks to learn input-output mappings from training samples. Due to the nonlinear modeling power of such networks, the learned mapping may interpolate all the training points. When erroneous training data are employed, the learned mapping can oscillate badly between data points. In this paper we derive a robust BP learning algorithm that is resistant to the noise effects and is capable of rejecting gross errors during the approximation process. The spirit of this algorithm comes from the pioneering work in robust statistics by Huber and Hampel. Our work is different from that of M-estimators in two aspects: 1) the shape of the objective function changes with the iteration time; and 2) the parametric form of the functional approximator is a nonlinear cascade of affine transformations. In contrast to the conventional BP algorithm, three advantages of the robust BP algorithm are: 1) it approximates an underlying mapping rather than interpolating training samples; 2) it is robust against gross errors; and 3) its rate of convergence is improved since the influence of incorrect samples is gracefully suppressed. David S. Chen, Ramesh Jain 0001 |
IEEE Trans. Neural Networks | 2 |
| 1993 | Shape from perspective trihedral angle constraintabstractA fundamental problem of determining the position and orientation of a 3-D object using a single perspective image view is defined and investigated. The technique is based on the interpretation of trihedral angle constraint information. A new closed-form solution to the problem is proposed. The method also provides a general analytic technique for dealing with a class of problem of shape from inverse perspective projection by using angle to angle correspondence information. Simulation experiments show that the authors' method is effective and robust for real application.> Yuyan Wu, S. Sitharama Iyengar, Ramesh Jain 0001, Santanu Bose |
CVPR | 3 |
| 1993 | Concept Clustering in a Query Interface to an Image DatabaseabstractArticle Concept clustering ina query interface to an image database Share on Authors: Jill Kliger The University of Michigan, Advanced Technology Laboratory, 1101 Beal Avenue, Ann Arbor, Michigan The University of Michigan, Advanced Technology Laboratory, 1101 Beal Avenue, Ann Arbor, MichiganView Profile , Deborah Swanberg The University of Michigan, Advanced Technology Laboratory, 1101 Beal Avenue, Ann Arbor, Michigan The University of Michigan, Advanced Technology Laboratory, 1101 Beal Avenue, Ann Arbor, MichiganView Profile , Ramesh Jain The University of Michigan, Advanced Technology Laboratory, 1101 Beal Avenue, Ann Arbor, Michigan The University of Michigan, Advanced Technology Laboratory, 1101 Beal Avenue, Ann Arbor, MichiganView Profile Authors Info & Claims UIST '93: Proceedings of the 6th annual ACM symposium on User interface software and technologyDecember 1993 Pages 11–21https://doi.org/10.1145/168642.168644Online:01 December 1993Publication History 1citation411DownloadsMetricsTotal Citations1Total Downloads411Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Jill Kliger, Deborah Swanberg, Ramesh Jain 0001 |
ACM Symposium on User Interface Software and Technology | 3 |
| 1993 | Simulation and Expectation in Sensor-Based SystemsabstractSimulations have traditionally been used as off-line tools for examining process models and experimenting with system models for which it would have been either impossible or too dangerous, expensive, or time-consuming, to perform with physical systems. We propose a novel way of regarding simulations as part of both the development and the working phases of systems. In our approach simulation is used within the processing and control loop of the system to provide sensor and state expectations. This minimizes the inverse sensory data analysis and model maintenance problems. We refer to this mode of operation as the verification mode, in contrast to the traditional discovery mode. In order to provide simulations and planning that are intertwined with the control of a physical system, temporal issues have to be considered. By limiting the focus of the system to small portions of complex models which are temporarily relevant to the system’s operation, the system is able to maintain its models and respond faster. For this we employ the Context-based Caching (CbC) mechanism within our Mobile Platform Control and Simulation Program (MOSIM). CbC is a knowledge management technique which maintains large knowledge bases by making the necessary information available at the right time. Yuval Roth, Ramesh Jain 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 1993 | An Abstraction-Based Approach to 3-D Pose Determination from Range ImagesabstractAn abstraction-based paradigm that makes explicit the process of imposing assumptions on data is discussed. The units of abstraction are models in which levels of abstraction are determined by the degree of assumption necessary for their application. A general-to-specific refinement process provides a mechanism to proceed gracefully through the abstraction hierarchy. This strategy was applied to the recognition and pose determination of objects comprising simple and compound cylindrical and planar surfaces in dense range data. A method of computing reliable Gaussian and mean curvature sign-map descriptors from the polynomial approximations of surfaces is demonstrated. A means for determining the pose of constructed geometric forms whose algebraic surface descriptions are nonlinear in terms of their orienting parameters is developed. It is shown that biquadratic surfaces are suitable companion-linear forms for cylinder approximation and parameter estimation. The estimates provide the initial parametric approximations necessary for a nonlinear regression stage to fine tune the estimates by fitting the actual nonlinear form to the data.> Francis K. H. Quek, Ramesh Jain 0001, Terry E. Weymouth |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1993 | A Visual Information Management System for the Interactive Retrieval of FacesabstractThe complex nature of two-dimensional image data has presented problems for traditional information systems designed strictly for alphanumeric data. Systems aimed at effectively managing image data have generally approached the problem from two different views: They either possess a strong database component with little image understanding, or they serve as an image repository for computer vision applications, with little emphasis on the image retrieval process. A general architecture for visual information-management systems (VIMS), which combine the strengths of both approaches, is presented. The system utilizes computer vision routines for both insertion and retrieval and allows easy query-by-example specifications. The vision routines are used to segment and evaluate objects based on domain-knowledge describing the objects and their attributes. The vision system can then assign feature values to be used for similarity-measures and image retrieval. A VIMS developed for face-image retrieval is presented to demonstrate these ideas.> Jeffrey R. Bach, Santanu Paul, Ramesh Jain 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1992 | An Interactive Image Management System for Face Information Retrieval
Jeffrey R. Bach, Santanu Paul, Ramesh Jain 0001 |
CIKM | 3 |
| 1992 | Surface reconstruction using neural networksabstractA surface reconstruction method using multilayer feedforward neural networks is proposed. The parametric form represented by multilayer neural networks can model piecewise smooth surfaces in a way that is more general and flexible than many of the classical methods. The approximation method is based on a robust backpropagation (BP) algorithm, which extends the basic BP algorithm to handle errors, especially others, in the training data.> David S. Chen, Ramesh Jain 0001, Brian G. Schunck |
CVPR | 2 |
| 1992 | Vector field analysis for oriented patternsabstractA method based on the properties of orientation fields is presented for the estimation of a set of symbolic descriptors from nondegenerate linear orientation fields, modeled by two-dimensional first-order phase portraits. It was previously shown that flow orientation is sufficient to characterize a flow pattern and locate its critical point position. A linear estimator is designed to estimate nondegenerate phase portraits. A classification scheme based on their local properties-curl, divergence, and deformation-is developed. Some experimental results are reported.> Chiao-Fe Shu, Ramesh Jain 0001 |
CVPR | 2 |
| 1992 | Model-driven pose correctionabstractPose determination for robot navigation is discussed. The problem is to maintain the system's instantaneous precept of its position and orientation in space for performing various tasks. The authors describe a system in which models were used to guide the sensory interpretation and to correct expectations. In this system, simulated images were used to analyze the real images and to correct the pose parameters. The reported techniques have been implemented and experiments with real images in a real environment have been performed.> Yuval Roth, Annie S. Win, Remzi H. Arpaci-Dusseau, Terry E. Weymouth, Ramesh Jain 0001 |
ICRA | 5 |
| 1992 | Architecture of a Multimedia Information System for Content-Based Retrieval
Deborah Swanberg, Chiao-Fe Shu, Ramesh Jain 0001 |
NOSSDAV | 3 |
| 1992 | Restoration of scanning probe microscope imagesabstractScanning probe microscopy (SXM), which includes techniques such as scanning tunneling microscopy (STM) and scanning force microscopy (SFM), is becoming popular for 3D metrology in the semiconductor industry and for high resolution 3D imaging of surfaces in Materials Science and Biology. The authors present imaging models for SXM that take into account the effect of probe geometry on topographic images produced by SXM in 'contact' and 'non-contact' modes. The authors formulate methods for restoring an SXM image to obtain the original surface. Criteria for determining certainty of restoration are developed. It is shown that the methods developed can be expressed in terms of gray scale morphological operators. The efficacy of the approach is demonstrated by applying it to synthetic and real data.> Gopal Sarma Pingali, Ramesh Jain 0001 |
WACV | 2 |
| 1992 | Why aspect graphs are not (yet) practical for computer vision
Olivier D. Faugeras, Joseph L. Mundy, Narendra Ahuja, Charles R. Dyer, Alex Pentland, Ramesh Jain 0001, Katsushi Ikeuchi, Kevin W. Bowyer |
CVGIP Image Underst. | 6 |
| 1992 | Recovering a boundary-level structural description from dynamic stereo
Arun P. Tirumalai, Brian G. Schunck, Ramesh Jain 0001 |
Image Vis. Comput. | 3 |
| 1992 | Reasoning About Edges in Scale SpaceabstractExplores the role of reasoning in early vision processing. In particular, the problem of detecting edges is addressed. The authors do not try to develop another edge detector, but rather, they study an edge detector rigorously to understand its behavior well enough to formulate a reasoning process that allow appliance of the detector judiciously to recover useful information. They present a multiscale reasoning algorithm for edge recovery: reasoning about edges in scale space (RESS). The knowledge in RESS is acquired from the theory of edge behavior in scale space and represented by a number of procedures. RESS recovers desired edge curves through a number of reasoning processes on zero crossing images at various scales. The knowledge of edge behavior in scale space enables RESS to select proper scale parameters, recover missing edges, eliminate noise or false edges, and correct the locations of edges. A brief evaluation of RESS is performed by comparing it with two well-known multistage edge detection algorithms.> Yi Lu 0014, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | Computerized Flow Field Analysis: Oriented Texture FieldsabstractAn approach to the solution of signal-to-symbol transformation in the domain of flow fields, such as oriented texture fields and velocity vector fields, is discussed. The authors use the geometric theory of differential equations to derive a symbol set based on the visual appearance of phase portraits which are a geometric representation of the solution curves of a system of differential equations. They also provide the computational framework to start with a given flow field and derive its symbolic representation. Specifically, they segment the given texture, derive its symbolic representation, and perform a quantitative reconstruction of the salient features of the original texture based on the symbolic descriptors. Results of applying this technique to several real texture images are presented. This technique is useful in describing complex flow visualization pictures, defects in lumber processing, defects in semiconductor wafer inspection, and optical flow fields.> A. Ravishankar Rao, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | Dynamic Stereo with Self-CalibrationabstractAn approach for incremental refinement of disparity maps obtained from a dynamic stereo sequence of a static scene is presented. The approach has been implemented using a binocular stereo vision system mounted on a mobile robot. A robust least median of squares based algorithm is given for recovering the camera motion between successive viewpoints, which provides a self-calibration mechanism. The recovered motion is utilized for recursive disparity prediction and refinement using a robust Kalman filter model.> Arun P. Tirumalai, Brian G. Schunck, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1991 | Trajectories and eventsabstractA solution to the correspondence problem using constraint satisfaction is described. It uses a real-time line-fitting algorithm to detect changes in a point's motion parameters as they happen. A trajectory is hypothesized for a single point's motion. Since the event detected may be wrong, multiple trajectories are hypothesized for each point. A correspondence is drawn from the set of hypothesized trajectories. The algorithm is robust, working on noisy and non-rigid motion data, and does not use false precision.> Susan M. Haynes, Ramesh Jain 0001 |
CVPR | 2 |
| 1991 | Qualitative motion analysis using a spatio-temporal approachabstractA spatio-temporal approach is presented which can extract motion information without requiring the computation of image flow. In this approach, an image sequence is treated as a single 3-D image. The sequence is first segmented through the grouping of voxels in the spatio-temporal space. After the image sequence is segmented, the area in the image which corresponds to the moving objects in the scene is identified. The direction of motion in space for each identified moving object is reported.> Shih-Ping Liou, Ramesh Jain 0001 |
CVPR | 2 |
| 1991 | A linear algorithm for computing the phase portraits of oriented texturesabstractPhase portraits are a powerful mathematical model for describing oriented textures. An isotangent-based approach is presented which is a linear formulation to the problem, to locate the critical points and compute the parameter sets of this model for the nonsingular two-dimensional first-order phase portraits. The authors classify flow patterns by Jordan canonical forms of the characteristic matrix made up of the estimated parameters. For these systems, they prove that all the isotangent curves are straight lines which intersect at a critical point. They also apply least median of squares (LMS) estimators to find the isotangent lines and locate the critical point. A linear regression technique is used to estimate the parameters of the two-dimensional first-order phase portrait of a given flow pattern. Results of applying the algorithm to synthetic and real images are presented.> Chiao-Fe Shu, Ramesh Jain 0001, Francis K. H. Quek |
CVPR | 2 |
| 1991 | Semantic Queries with Pictures: The VIMSYS Model
Amarnath Gupta, Terry E. Weymouth, Ramesh Jain 0001 |
VLDB | 3 |
| 1991 | Ignorance, myopia, and naiveté in computer vision systems
Ramesh Jain 0001, Thomas O. Binford |
CVGIP Image Underst. | 1 |
| 1991 | An approach to three-dimensional image segmentation
Shih-Ping Liou, Ramesh Jain 0001 |
CVGIP Image Underst. | 2 |
| 1991 | Machine vision in the 1990s: Applications and how to get there
Dragutin Petkovic, J. Wilder, Masakazu Ejiri, Robert M. Haralick, Ramesh Jain 0001, Peter A. Ruetz, Jack Sklansky, C. W. Swonger |
Mach. Vis. Appl. | 5 |
| 1991 | A Parallel Technique for Signal-Level Perceptual OrganizationabstractDue to the potential for essentially unbounded scene complexity, it is often necessary to translate the sensor-derived signals into richer symbolic representations. A key initial stage in this abstraction process is signal-level perceptual organization (SLPO) involving the processes of partitioning and identification. A parallel SLPO algorithm that follows the global hypothesis testing paradigm, but breaks the iterative structure of conventional region growing through the use of alpha -partitioning and region filtering is presented. These two techniques segment an image such that the gray-level variation within each region can be described by a regression model. Experimental results demonstrate the effectiveness of this algorithm.> Shih-Ping Liou, Arnold H. Chiu, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1991 | Structure from motion-a critical analysis of methodsabstractExisting methods for structure from motion are classified and analyzed and mathematical relationships among them are shown. Their behavior on the same synthetic data is evaluated, examining the effect of noise, initial estimates, and the amount of motion. The missing necessary polynomial constraint on the singular values of R.Y. Tsai and T.S. Huang's (1981) essential parameter matrix is shown, and related linear methods for less general SFM problems and their necessary polynomial constraints are derived. The experiments show that for data which the constraints are satisfied, linear methods work correctly and quickly, and where they are not applicable, polynomial methods may work but are slower.> Charles Jerian, Ramesh Jain 0001 |
IEEE Trans. Syst. Man Cybern. | 2 |
| 1990 | Dynamic stereo with self-calibrationabstractDynamic stereo is useful for constructing a complete map of the environment as only a portion of the actual environment is visible from each viewpoint. In addition, there is usually an overlap between the portions of the environment visible from two successive viewpoints. It is then feasible to utilize a prediction-verification approach to combine the individual depth estimates of features visible from both viewpoints to obtain a more accurate estimate. A fundamental requirement for such an approach to be used is accurate knowledge of the camera motion between the two viewpoints. A robust least median of squares (LMS)-based algorithm to recover this motion which provides a self-calibration mechanism is presented. The recovered motion is utilized for recursive disparity prediction and refinement using a robustified Kalman filter formulation. Results are presented for a laboratory stereo sequence.> Arun P. Tirumalai, Brian G. Schunck, Ramesh Jain 0001 |
ICCV | 3 |
| 1990 | A parallel technique for three-dimensional image segmentationabstractThe authors present a parallel 3-D image segmentation algorithm which, through the use of alpha -partitioning and volume filtering, segments 3-D images such that the gray-level variation within each volume can be described by a regression model. Experimental results indicate that the use of discontinuity locations in alpha -partitioning can greatly assist the segmentation of 3-D images. It is also shown that the use of volume filters in removing incorrect partitions permits quick segmentation of 3-D images into functional descriptions for each volume. These two characteristics enabled the parallel 3-D image segmentation algorithm to produce fair segmentation results with large savings in computation time.> Shih-Ping Liou, Ramesh Jain 0001 |
ICPR (1) | 2 |
| 1990 | Analyzing oriented textures through phase portraitsabstractAn attempt is made to develop a solution for signal-to-symbol transformation in the domain of flowlike or oriented texture. The geometric theory of differential equations is used to derive a symbol set based on the visual appearance of phase portraits. This theory provides a technique for describing textures both qualitatively and quantitatively. An attractive feature of this symbol set is that it is domain independent and makes no assumptions about the kind of texture that may be present. The computational framework for starting with a given oriented texture is provided, and its symbolic representation is derived. This is based on computing the orientation field for the texture and then using a nonlinear least-squares technique over successive windows to determine the changing spatial behavior of the texture. Results of the application of this technique to real texture images are presented.> A. Ravishankar Rao, Ramesh Jain 0001 |
ICPR (1) | 2 |
| 1990 | Scanning electron microscope-based stereo analysis
Ali E. Kayaalp, A. Ravishankar Rao, Ramesh Jain 0001 |
Mach. Vis. Appl. | 3 |
| 1990 | Using Dynamic Programming for Solving Variational Problems in VisionabstractDynamic programming is discussed as an approach to solving variational problems in vision. Dynamic programming ensures global optimality of the solution, is numerically stable, and allows for hard constraints to be enforced on the behavior of the solution within a natural and straightforward structure. As a specific example of the approach's efficacy, applying dynamic programming to the energy-minimizing active contours is described. The optimization problem is set up as a discrete multistage decision process and is solved by a time-delayed discrete dynamic programming algorithm. A parallel procedure for decreasing computational costs is discussed.> Amir A. Amini, Terry E. Weymouth, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1990 | Polynomial Methods for Structure from MotionabstractThe authors analyze the limitations of structure from motion (SFM) methods presented in the literature and propose the use of a polynomial system of equations, with the unit quaternions representing rotation, to recover SFM under perspective projection. The authors combine the equations by the method of resultants with the MAXIMA symbolic algebra system, reducing the system to a single polynomial. Its real roots are then found with Sturm sequences. Since this system has multiple solutions, a hypothesize-and-verify scheme is used to eliminate incorrect ones. The scheme diminishes the sensitivity of using polynomial equations. The authors examine the effect of different rotation axes and angles on SFM accuracy and compare the performance of the algorithm to a few earlier approaches. Generally, it is found that a large amount of motion is the most important factor in getting good SFM accuracy.> Charles Jerian, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1989 | Scanning electron microscope based stereo analysis [for semiconductor IC inspection]abstractA novel technique to analyze stereo images generated from a scanning electron microscope (SEM) for inspection of wafers for process control is presented. The two main features of this technique are that it uses a binary linear programming approach to set up and solve the correspondence problem, and that it uses constraints based on the physics of SEM image formation. Binary linear programming is a powerful tool with which to tackle constrained optimization problems, especially in the cases that involve matching between one data set and another. The authors also analyze the process of SEM image formation, and present constraints that are useful in solving the stereo correspondence problem. This technique has been solved on many images. Results for a few wafers are included.> Ali E. Kayaalp, A. Ravishankar Rao, Ramesh Jain 0001 |
CVPR | 3 |
| 1989 | A nonlinear optimization algorithm for the estimation of structure and motion parametersabstractA nonlinear least-squares optimization technique is proposed which uses the Levenberg-Marquardt method and estimates the motion and structure parameters to a global scale factor by minimizing an objective function. This objective function is the mean-square difference between the measured coordinates of feature points in the image plane and the coordinates predicted from the current state estimate. In comparison to existing approaches, this technique converges faster and yields better estimates. A recursive version of this algorithm is developed using the block approach. This algorithm is shown to also track eventful motion effectively. The performance of the proposed technique on real image sequences is also presented. Some performance results are indicated to illustrate the efficacy of this approach.> R. V. Raja Kumar 0002, Arun P. Tirumalai, Ramesh Jain 0001 |
CVPR | 3 |
| 1989 | Range estimation from intensity gradient analysisabstractThe authors have developed a depth-recovery technique that completely avoids the computationally intensive steps of feature selection and correspondence required by conventional approaches. The intensity gradient analysis (IGA) technique is a depth-recovery algorithm that utilizes the properties of the MCSO (moving camera, stationary objects) scenario. Depth values are obtained by analyzing temporal intensity gradients arising from the optic flow field induced by known camera motion. In doing so, IGA avoids the feature extraction and correspondence steps of conventional approaches and is therefore very fast. A detailed description of the algorithm is provided along with experimental results from complex laboratory scenes. It is suggested that the most appealing property of this approach is that IGA places little burden on computational resources, and therefore seems ideally suited for real-world robotic applications.> Kurt D. Skifstad, Ramesh Jain 0001 |
ICRA | 2 |
| 1989 | Motion detection in spatio-temporal space
Shih-Ping Liou, Ramesh Jain 0001 |
Comput. Vis. Graph. Image Process. | 2 |
| 1989 | Illumination independent change detection for real world image sequences
Kurt D. Skifstad, Ramesh Jain 0001 |
Comput. Vis. Graph. Image Process. | 2 |
| 1989 | Range estimation from Intensity Gradient Analysis
Kurt D. Skifstad, Ramesh Jain 0001 |
Mach. Vis. Appl. | 2 |
| 1989 | Behavior of Edges in Scale SpaceabstractAn analysis is presented of the behavior of edges in scale space for deriving rules useful in reasoning. This analysis of liner edges at different scales in images includes the mutual influence of edges and identifies at what scale neighboring edges start influencing the response of a Laplacian or Gaussian operator. Dislocation of edges, false edges, and merging of edges in the scale space are examined to formulate rules for reasoning in the scale space. The theorems, corollaries, and assertions presented can be used to recover edges, and related features, in complex images. The results reported include one lemma, three theorems, a number of corollaries and six assertions. The rigorous mathematical proofs for the theorems and corollaries are presented. These theorems and corollaries are further applied to more general situations, and the results are summarized in six assertions. A qualitative description as well as some experimental results are presented for each assertion.> Yi Lu 0014, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1988 | Spline-based surface fitting on range images for CAD applicationsabstractThe authors present an approach for integration of manually designed parts in CAD databases. They propose a system to derive spline-based descriptions of various component surfaces of objects given only the range image of the object under consideration. The system consists of two modules. The first module, which is centred on a robust segmentation algorithm, generates a segmentation plot of the given range image. The second module then derives spline-based descriptions for each of the different segmented surfaces.> Sanjeev M. Naik, Ramesh Jain 0001 |
CVPR | 2 |
| 1988 | Polynomial Methods For Structure From MotionabstractThere have been many attempts to solve the SFM problem, yet few practical solutions. Previous work has used either linearizations that ignore vital constraints or non-linear equations with multiple solutions, where the existence of multiple solutions has been ignored.
First, we classify and analyze many existing methods and find that they do not work well in the presence of noise, or without a good initial estimate. Then, we propose new polynomial systems solutions for both orthographic and perspective projection that work better than existing methods in the presence of noise. Finally, we examine the effect of such factors as the number of frames and the axis and angle of rotation on the ability to recover structure. We found that additional frames are of no value and that large amounts of rotation resulting in disparate views are very helpful for accurate structure recovery. Charles Jerian, Ramesh Jain 0001 |
ICCV | 2 |
| 1988 | Dynamic visionabstractThe architecture of a dynamic vision system is discussed and issues related to motion detection and segmentation, incremental recovery of structure from motion, the environment and world models, and control structures based on qualitative approaches are examined. The ideas presented here are dynamic, and the approach is evolving. The approach does not require image flow field. In some cases, some properties of image flow fields are used, but computation of the field is not necessary. The emphasis here is on recovering robust information using the principles of least commitment and graceful degradation.> Ramesh Jain 0001 |
ICPR | 1 |
| 1988 | Dissertation abstracts
Ramesh Jain 0001 |
Mach. Vis. Appl. | 1 |
| 1988 | Perception engineering
Ramesh Jain 0001 |
Mach. Vis. Appl. | 1 |
| 1988 | Report on Range Image Understanding Workshop, East Lansing, Michigan, March 21-23, 1988
Ramesh Jain 0001, Anil K. Jain 0001 |
Mach. Vis. Appl. | 1 |
| 1988 | An editorial welcome
Ramesh Jain 0001, André Oosterlinck, Jorge Sanz, Jack Sklansky, Masahiko Yachida |
Mach. Vis. Appl. | 1 |
| 1988 | Automatic Solder Joint InspectionabstractThe task of automating the visual inspection of pin-in-hole solder joints is addressed. Two approaches are explored: statistical pattern recognition and expert systems. An objective dimensionality-reduction method is used to enhance the performance of traditional statistical pattern recognition approaches by decorrelating feature data, generating feature weights, and reducing run-time computations. The expert system uses features in a manner more analogous to the visual clues that a human inspector would rely on for classification. Rules using these cues are developed, and a voting scheme is implemented to accumulate classification evidence incrementally. Both methods compared favorably with human inspector performance.> Sandra L. Bartlett, Paul J. Besl, Charles L. Cole, Ramesh Jain 0001, Debashish Mukherjee, Kurt D. Skifstad |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 1988 | Segmentation through Variable-Order Surface FittingabstractThe solution of the segmentation problem requires a mechanism for partitioning the image array into low-level entities based on a model of the underlying image structure. A piecewise-smooth surface model for image data that possesses surface coherence properties is used to develop an algorithm that simultaneously segments a large class of images into regions of arbitrary shape and approximates image data with bivariate functions so that it is possible to compute a complete, noiseless image reconstruction based on the extracted functions and regions. Surface curvature sign labeling provides an initial coarse image segmentation, which is refined by an iterative region-growing method based on variable-order surface fitting. Experimental results show the algorithm's performance on six range images and three intensity images.> Paul J. Besl, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1987 | Motion Stereo Using Ego-Motion Complex Logarithmic MappingabstractStereo information can be obtained using a moving camera. If a dynamic scene is acquired using a translating camera and the camera motion parameters are known, then the analysis of the scene may be facilitated by ego-motion complex logarithmic mapping (ECLM). It is shown in this paper that by using the complex logarithmic mapping (CLM) with respect to the focus of expansion, the depth of stationary components can be determined easily in the transformed image sequence. The proposed approach for depth recovery avoids the difficult problems of establishing correspondence and computation of optical flow, by using the ego-motion information. An added advantage of the CLM will be the invariances it offers. We report our experiments with synthetic data to show the sensitivity of the depth recovery, and show results of real scenes to demonstrate the efficacy of the proposed motion stereo in applications such as autonomous navigation. Ramesh Jain 0001, Sandra L. Bartlett, Nancy O'Brien |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1987 | Finding Trajectories of Feature Points in a Monocular Image SequenceabstractIdentifying the same physical point in more than one image, the correspondence problem, is vital in motion analysis. Most research for establishing correspondence uses only two frames of a sequence to solve this problem. By using a sequence of frames, it is possible to exploit the fact that due to inertia the motion of an object cannot change instantaneously. By using smoothness of motion, it is possible to solve the correspondence problem for arbitrary motion of several nonrigid objects in a scene. We formulate the correspondence problem as an optimization problem and propose an iterative algorithm to find trajectories of points in a monocular image sequence. A modified form of this algorithm is useful in case of occlusion also. We demonstrate the efficacy of this approach considering synthetic, laboratory, and real scenes. Ishwar K. Sethi, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1986 | Recognizing partially visible objects using feature indexed hypothesesabstractA common task in computer vision is to recognize the objects in an image. Most computer vision systems do this by matching models for each possible object type in turn, recognizing objects by the best matches. This is not ideal, as it does not take advantage of the similarities and differences between the possible object types. The computation time also increases linearly with the number of possible objects, which can become a problem if the number is large. This paper describes a new recognition method, the feature indexed hypotheses method, which takes advantage of the similarities and differences between object types, and is able to handle cases, where there are a large number of possible object types, in sub-linear computation time. A two-dimensional occluded parts recognition system using this method is described. Thomas F. Knoll, Ramesh Jain 0001 |
ICRA | 2 |
| 1986 | Invariant surface characteristics for 3D object recognition in range images
Paul J. Besl, Ramesh Jain 0001 |
Comput. Vis. Graph. Image Process. | 2 |
| 1986 | Pulse and staircase edge models
Mubarak Shah, Arun K. Sood, Ramesh Jain 0001 |
Comput. Vis. Graph. Image Process. | 3 |
| 1986 | A Pyramid-Based Approach to Segmentation Applied to Region MatchingabstractIn this paper, we attempt to place segmentation schemes utilizing the pyramid architecture on a firm footing. We show that there are some images which cannot be segmented in principle. An efficient segmentation scheme is also developed using pyramid relinking. This scheme will normally have a time complexity which is a sublinear function of the image diameter, which compares favorably to other schemes. The efficacy of our approach to segmentation using pyramid schemes is demonstrated in the context of region matching. The global features we use are compared to those used in previous approaches and this comparison will indicate that our approach is more robust than the standard moment-based techniques. William I. Grosky, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1986 | Recognizing partially visible objects using feature indexed hypothesesabstractA common task in computer vision is to recognize the objects in an image. Most computer vision systems do this by matching models for each possible object type in turn, recognizing objects by the best matches. This is not ideal, as it does not take advantage of the similarities and differences between the possible object types. The computation time also increases linearly with the number of possible objects, which can become a problem if the number is large. A new recognition method is described: feature indexed hypotheses, which takes advantage of the similarities and differences between object types and is able to handle cases, where there are a large number of possible object types, in sublinear computation time. A two-dimensional occluded parts recognition system using this method is described. Thomas F. Knoll, Ramesh Jain 0001 |
IEEE J. Robotics Autom. | 2 |
| 1985 | Automatic visual inspection of solder jointsabstractThis paper describes an approach for automatic inspection of solder joints on printed circuit boards using gray-scale images. Common defects in solder joints are recognized using features computed from segmented solder joint subimages. Unacceptable joints are assigned to one of several defective classes. Defect classification, rather than just detection of defective joints, is motivated by the desire to automatically take corrective action on the assembly line. The features used for classification are based on characteristics of intensity surfaces. It is shown that features derived from surface facets are effective in the classification of solder joints using a minimum-distance classification algorithm. Paul J. Besl, Edward J. Delp, Ramesh Jain 0001 |
ICRA | 3 |
| 1985 | Object recognition using multiple viewsabstractA new approach to model based object recognition employing multiple views is described. The emphasis is given on the determination of camera viewpoints for succesive views looking for distinguishing features of objects. The distance and direction of the camera are determined separately. The distance is determined by the size of the object and the feature, while the direction is determined by the shape of the feature and the presence of the occluding objects. Hwang-Soo Kim, Ramesh Jain 0001, Richard A. Volz |
ICRA | 2 |
| 1985 | Uncertainty Management in a Distributed Knowledge Based System
Naseem A. Khan, Ramesh Jain 0001 |
IJCAI | 2 |
| 1985 | Automatic visual solder joint inspectionabstractAn approach is described for the automatic inspection of solder joints on printed circuit boards. Common defects are identified in solder joints and a joint is classified as being good or belonging to one of the defective classes. The motivation for this classification is not just the detection of defective joints, but the desire to automatically take corrective action on the assembly line. The features used for classification are based on characteristics of intensity surfaces. It is shown that features derived from facets and Gaussian curvature are effective in the classification of solder joints using a minimum-distance classification algorithm. Class separation plots are shown to be useful for quickly studying individual effectiveness of a feature or pair of features in classification. Results show the efficacy of the described approach. Paul J. Besl, Edward J. Delp, Ramesh Jain 0001 |
IEEE J. Robotics Autom. | 3 |
| 1984 | Detecting time-varying corners
Mubarak Shah, Ramesh Jain 0001 |
Comput. Vis. Graph. Image Process. | 2 |
| 1984 | Difference and accumulative difference pictures in dynamic scene analysis
Ramesh Jain 0001 |
Image Vis. Comput. | 1 |
| 1984 | Segmentation of Frame Sequences Obtained by a Moving ObserverabstractIn many applications of computer vision, a frame sequence may be acquired using a moving camera. We propose ego-motion polar transformation for segmentation of such sequences. It is shown that segmentation and extraction of motion information become easier in the transformed domain. Our experience with a translating camera indicates that this technique can play a very important role in the analysis of moving observer dynamic scenes. Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1984 | Determining Motion Parameters for Scenes with Translation and RotationabstractA study of methods that determine the rotation parameters of a camera moving through synthetic and real scenes is conducted. Algorithms that combine ideas of Jain and Prazdny, using hypothesizeand-verify paradigm, are developed to find translational and rotational parameters. An argument is made for using hypothesized motion parameters rather than relaxation labeling to find correspondence. Some work with real scenes shows the difficulties introduced by noise, the lack of resolution, and the need for better low-level techniques. Charles Jerian, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1983 | Detection of moving edges
Susan M. Haynes, Ramesh Jain 0001 |
Comput. Vis. Graph. Image Process. | 2 |
| 1983 | An approach to the segmentation of textured dynamic scenes
S. N. Jayaramamurthy, Ramesh Jain 0001 |
Comput. Vis. Graph. Image Process. | 2 |
| 1983 | Optimal Quadtrees for Image SegmentsabstractQuadtrees are compact hierarchical representations of images. In this paper, we define the efficiency of quadtrees in representing image segments and derive the relationship between the size of the enclosing rectangle of an image segment and its optimal quadtree. We show that if an image segment has an enclosing rectangle having sides of lengths x and y, such that 2N-1 × max (x, y) ¿ 2N, then the optimal quadtree may be the one representing an image of size 2N × 2N or 2N+1 × 2N+1. It is shown that in some situations the quadtree corresponding to the larger image has fewer nodes. Also, some necessary conditions are derived to identify segments for which the larger image size results in a quadtree which is no more expensive than the quadtree for the smaller image size. William I. Grosky, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1983 | Direct Computation of the Focus of ExpansionabstractOptical flow carries valuable information about the nature and depth of surfaces and the relative motion between observer and objects. In the extraction of this information, the focus of expansion plays a vital role. In contrast to the current approaches, this paper presents a method for the direct computation of the focus of expansion using an optimization approach. The optical flow can then be computed using the focus of expansion. Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1982 | Obtaining 3-dimensional shape of textured and specular surfaces using four-source photometry
E. North Coleman Jr., Ramesh Jain 0001 |
Comput. Graph. Image Process. | 2 |
| 1982 | Detection of moving edges
Susan M. Haynes, Ramesh Jain 0001 |
Comput. Graph. Image Process. | 2 |
| 1982 | An approach to the segmentation of textured dynamic scenes
S. N. Jayaramamurthy, Ramesh Jain 0001 |
Comput. Graph. Image Process. | 2 |
| 1982 | Normalized quadtrees with respect to translations
William I. Grosky, Ramesh Jain 0001 |
Comput. Graph. Image Process. | 3 |
| 1982 | Normalized quadtrees with respect to translations
William I. Grosky, Ramesh Jain 0001 |
Comput. Graph. Image Process. | 3 |
| 1982 | Segmentation of moving observer frame sequences
Ramesh Jain 0001 |
Pattern Recognit. Lett. | 1 |
| 1982 | A Pipelined Pseudoparallel System Architecture for Real-Time Dynamic Scene AnalysisabstractIn this paper we introduce the concept of pseudoparallelism, in which the serial algorithm is partitioned into several noninteractive independent subtasks so that parallelism can be used within each subtask level. This approach is illustrated by applying it to a real-time dynamic scene analysis. Complete details of such a pseudoparallel architecture with an emphasis to avoid interprocessor communications have been worked out. Problems encountered in the course of designing such a system with a distributed operating system (no master control) have been outlined and necessary justifications have been provided. A scheme indicating various memory modules, processing elements, and their data-path requirements is included and ways to provide continuous flow of partitioned information in the form of a synchronized pipeline are described. Dharma P. Agrawal, Ramesh Jain 0001 |
IEEE Trans. Computers | 2 |
| 1981 | Shape from Shading for Surfaces with Texture and Specularity
E. North Coleman Jr., Ramesh Jain 0001 |
IJCAI | 2 |
| 1981 | Extraction of Motion Information from Peripheral ProcessesabstractThis paper is mainly concerned with low-level processes in machine perception of motion. A motion analysis system should exploit information contained in ``early warning signals'' during the intensity based peripheral phase of motion perception. We show that intensity based difference pictures contain motion information about objects in a dynamic scene, and present methods for the extraction of motion information in the peripheral phase. Some experiments with laboratory generated and real world scenes demonstrate the potential of the technique. Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1979 | Extraction of Moving Object Images Through Change Detection
Ramesh Jain 0001, Worthy N. Martin, Jake K. Aggarwal |
IJCAI | 1 |
| 1979 | On the Analysis of Accumulative Difference Pictures from Image Sequences of Real World ScenesabstractThe count of events where sample areas from the second and subsequent frames of a TV-image sequence are incompatible with the corresponding sample area of the first frame are accumulated in a first-order difference picture (FODP). Analysis of this FODP provides a separate estimate for images of moving objects and of stationary scene components. We start from the hypothesis that the first frame represents the stationary scene component. Once it has been recognized that a subarea of this initial estimate corresponds to the image of a moving object, the grey values in this subarea are replaced by later estimates of the stationary background at this position. No knowledge specific to a particular scene is utilized in the algorithm. The results for two scene sequences are presented. Ramesh Jain 0001, Hans-Hellmut Nagel |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1978 | Comments on "Fuzzy Set Theory versus Bayesian Statistics"abstractRecently, Stallings compared fuzzy set theory with Bayesian statistics. This comparison is improper as it compares Bayesian statistics with a particular case of fuzzy set theory. It is shown here that other interpretation of fuzzy connectives may result in entirely different results. Ramesh Jain 0001, William Stallings |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1977 | Separating Non-Stationary from Stationary Scene Components in a Sequence of Real World TV Images
Ramesh Jain 0001, D. Militzer, Hans-Hellmut Nagel |
IJCAI | 1 |
| 1971 | Review of "Mathematical Programminig" by Claude McMillan Jr
Ramesh Jain 0001 |
IEEE Trans. Syst. Man Cybern. | 1 |