VLDB 2026 Research / reviewers in the wild / expert
Hema Swetha Koppula
dblp:41/7536 · also Hema Koppula
· DBLP profile ↗
17ranked-venue papers
6as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
3D vision · 16% Generative modeling · 15% Video understanding and tracking · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 64% Web and social media mining · 28% Data integration and cleaning · 8% |
Topics — the 28 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis |
0.6 | 1 | 2022 | Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models · ICML 2022 |
Computer vision › 3D vision
3d scene understanding |
0.5 | 3 | 2016 | Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016 Hallucinated Humans as the Hidden Context for Labeling 3D Scenes · CVPR 2013 Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011 |
Robotics › Autonomous driving
maneuver anticipation |
0.5 | 2 | 2016 | Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016 Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models · ICCV 2015 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.4 | 2 | 2016 | Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016 Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011 |
Computer vision › Video understanding and tracking › activity recognition
activity detection |
0.2 | 1 | 2016 | Anticipating Human Activities Using Object Affordances for Reactive Robotic Response · IEEE Trans. Pattern Anal. Mach. Intell. 2016 |
Robotics › Robot manipulation › service robot
assistive robotics |
0.2 | 1 | 2016 | Anticipating Human Activities Using Object Affordances for Reactive Robotic Response · IEEE Trans. Pattern Anal. Mach. Intell. 2016 |
Computer vision › Video understanding and tracking › action anticipation
driver action prediction |
0.2 | 1 | 2016 | Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016 |
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM |
0.2 | 1 | 2016 | Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.2 | 1 | 2016 | Recurrent Neural Networks for driver activity anticipation via sensory-fusion architecture · ICRA 2016 |
Robotics › Autonomous driving
driver behavior modeling |
0.2 | 1 | 2015 | Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models · ICCV 2015 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.2 | 1 | 2014 | Physically Grounded Spatio-temporal Object Affordances · ECCV (3) 2014 |
Computer vision › 3D vision › embodied vision
object affordance |
0.2 | 1 | 2014 | Physically Grounded Spatio-temporal Object Affordances · ECCV (3) 2014 |
Machine learning › Generative modeling
autoregressive model |
0.2 | 1 | 2022 | Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models · ICML 2022 |
Machine learning › Generative modeling › generative model
generative sequence model |
0.2 | 1 | 2022 | Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models · ICML 2022 |
Computer vision › 3D vision › 3d scene understanding
3d scene labeling |
0.2 | 1 | 2013 | Hallucinated Humans as the Hidden Context for Labeling 3D Scenes · CVPR 2013 |
Computer vision › Video understanding and tracking
spatio-temporal modeling |
0.2 | 1 | 2013 | Learning Spatio-Temporal Structure from RGB-D Videos for Human Activity Detection and Anticipation · ICML (3) 2013 |
Computer vision › Image recognition and object detection
object detection |
0.1 | 1 | 2011 | Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011 |
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation |
0.1 | 1 | 2011 | Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011 |
Information retrieval › search engines
web crawling |
0.1 | 1 | 2010 | Learning URL patterns for webpage de-duplication · WSDM 2010 |
Web and social media mining › web mining
web page deduplication |
0.1 | 1 | 2010 | Learning URL patterns for webpage de-duplication · WSDM 2010 |
Information retrieval
web search |
0.1 | 1 | 2010 | Learning URL patterns for webpage de-duplication · WSDM 2010 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field |
0.1 | 1 | 2016 | Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.1 | 1 | 2016 | Modeling 3D Environments through Hidden Human Context · IEEE Trans. Pattern Anal. Mach. Intell. 2016 |
Computer vision › Video understanding and tracking
temporal modeling |
0.1 | 1 | 2015 | Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving Models · ICCV 2015 |
Natural language and speech › Information extraction and text analysis
topic model |
0.0 | 1 | 2013 | Hallucinated Humans as the Hidden Context for Labeling 3D Scenes · CVPR 2013 |
Robotics › Robot navigation and mapping
mobile robot perception |
0.0 | 1 | 2011 | Semantic Labeling of 3D Point Clouds for Indoor Scenes · NIPS 2011 |
Data integration and cleaning › entity resolution
duplicate detection |
0.0 | 1 | 2010 | Learning URL patterns for webpage de-duplication · WSDM 2010 |
Information retrieval
indexing |
0.0 | 1 | 2010 | Learning URL patterns for webpage de-duplication · WSDM 2010 |
Methods — techniques the papers use, named apart from their topics
style transformation module · 0.6style equalization · 0.6sequence-to-sequence prediction · 0.2sensory fusion · 0.2particle filtering · 0.2infinite latent conditional random field · 0.2human-object interaction modeling · 0.2dirichlet process · 0.2anticipatory temporal conditional random field · 0.2autoregressive input-output HMM · 0.2rule mining · 0.1mapreduce · 0.1machine learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Corpus Synthesis for Zero-Shot ASR Domain Adaptation Using Large Language ModelsabstractWhile Automatic Speech Recognition (ASR) systems are widely used in many real-world applications, they often do not generalize well to new domains and need to be fine-tuned on data from these domains. However, target-domain data usually are not readily available in many scenarios. In this paper, we propose a new strategy for adapting ASR models to new target domains without any text or speech from those domains. To accomplish this, we propose a novel data synthesis pipeline that uses a Large Language Model (LLM) to generate a target domain text corpus, and a state-of-the-art controllable speech synthesis model to generate the corresponding speech. We propose a simple yet effective in-context instruction fine-tuning strategy to increase the effectiveness of LLM in generating text corpora for new domains. Experiments on the SLURP dataset show that the proposed method achieves an average relative word error rate improvement of 28% on unseen target domains without any performance drop in source domains. Hsuan Su, Ting-Yao Hu, Hema Swetha Koppula, Raviteja Vemulapalli, Jen-Hao Rick Chang, Karren D. Yang, Gautam Varma Mantena, Oncel Tuzel |
ICASSP | 3 |
| 2023 | Text is all You Need: Personalizing ASR Models Using Controllable Speech SynthesisabstractAdapting generic speech recognition models to specific individuals is a challenging problem due to the scarcity of personalized data. Recent works have proposed boosting the amount of training data using personalized text-to-speech synthesis. Here, we ask two fundamental questions about this strategy: when is synthetic data effective for personalization, and why is it effective in those cases? To address the first question, we adapt a state-of-the-art automatic speech recognition (ASR) model to target speakers from four benchmark datasets representative of different speaker types. We show that ASR personalization with synthetic data is effective in all cases, but particularly when (i) the target speaker is underrepresented in the global data, and (ii) the capacity of the global model is limited. To address the second question of why personalized synthetic data is effective, we use controllable speech synthesis (CSS) to generate speech with varied styles and content. Surprisingly, we find that the text content of the synthetic data, rather than style, is important for speaker adaptation. These results lead us to propose a data selection strategy for ASR personalization based on speech content. Karren D. Yang, Ting-Yao Hu, Jen-Hao Rick Chang, Hema Swetha Koppula, Oncel Tuzel |
ICASSP | 4 |
| 2022 | SYNT++: Utilizing Imperfect Synthetic Data to Improve Speech RecognitionabstractWith recent advances in speech synthesis, synthetic data is becoming a viable alternative to real data for training speech recognition models. However, machine learning with synthetic data is not trivial due to the gap between the synthetic and the real data distributions. Synthetic datasets may contain artifacts that do not exist in real data such as structured noise, content errors, or unrealistic speaking styles. Moreover, the synthesis process may introduce a bias due to uneven sampling of the data manifold. We propose two novel techniques during training to mitigate the problems due to the distribution gap: (i) a rejection sampling algorithm and (ii) using separate batch normalization statistics for the real and the synthetic samples. We show that these methods significantly improve the training of speech recognition models using synthetic data. We evaluate the proposed approach on keyword detection and Automatic Speech Recognition (ASR) tasks, and observe up to 18% and 13% relative error reduction, respectively, compared to naively using the synthetic data. Ting-Yao Hu, Mohammadreza Armandpour, Ashish Shrivastava 0001, Jen-Hao Rick Chang, Hema Swetha Koppula, Oncel Tuzel |
ICASSP | 5 |
| 2022 | Style Equalization: Unsupervised Learning of Controllable Generative Sequence ModelsabstractControllable generative sequence models with the capability to extract and replicate the style of specific examples enable many applications, including narrating audiobooks in different voices, auto-completing and auto-correcting written handwriting, and generating missing training samples for downstream recognition tasks. However, under an unsupervised-style setting, typical training algorithms for controllable sequence generative models suffer from the training-inference mismatch, where the same sample is used as content and style input during training but unpaired samples are given during inference. In this paper, we tackle the training-inference mismatch encountered during unsupervised learning of controllable generative sequence models. The proposed method is simple yet effective, where we use a style transformation module to transfer target style information into an unrelated style input. This method enables training using unpaired content and style samples and thereby mitigate the training-inference mismatch. We apply style equalization to text-to-speech and text-to-handwriting synthesis on three datasets. We conduct thorough evaluation, including both quantitative and qualitative user studies. Our results show that by mitigating the training-inference mismatch with the proposed style equalization, we achieve style replication scores comparable to real data in our user studies. Jen-Hao Rick Chang, Ashish Shrivastava 0001, Hema Swetha Koppula, Xiaoshuai Zhang, Oncel Tuzel |
ICML | 3 |
| 2021 | SapAugment: Learning A Sample Adaptive Policy for Data AugmentationabstractData augmentation methods usually apply the same augmentation (or a mix of them) to all the training samples. For example, to perturb data with noise, the noise is sampled from a Normal distribution with a fixed standard deviation, for all samples. We hypothesize that a hard sample with high training loss already provides strong training signal to update the model parameters and should be perturbed with mild or no augmentation. Perturbing a hard sample with a strong augmentation may also make it too hard to learn from. Furthermore, a sample with low training loss should be perturbed by a stronger augmentation to provide more robustness to a variety of conditions. To formalize these intuitions, we propose a novel method to learn a Sample-Adaptive Policy for Augmentation – SapAugment. Our policy adapts the augmentation parameters based on the training loss of the data samples. In the example of Gaussian noise, a hard sample will be perturbed with a low variance noise and an easy sample with a high variance noise. Furthermore, the proposed method combines multiple augmentation methods into a methodical policy learning framework and obviates hand-crafting augmentation parameters by trial-and-error. We apply our method on an automatic speech recognition (ASR) task, and combine existing and novel augmentations using the proposed framework. We show substantial improvement, up to 21% relative reduction in word error rate on LibriSpeech dataset, over the state-of-the-art speech augmentation method. Ting-Yao Hu, Ashish Shrivastava 0001, Jen-Hao Rick Chang, Hema Swetha Koppula, Kyuyeon Hwang, Ozlem Kalinli, Oncel Tuzel |
ICASSP | 4 |
| 2016 | Recurrent Neural Networks for driver activity anticipation via sensory-fusion architectureabstractAnticipating the future actions of a human is a widely studied problem in robotics that requires spatio-temporal reasoning. In this work we propose a deep learning approach for anticipation in sensory-rich robotics applications. We introduce a sensory-fusion architecture which jointly learns to anticipate and fuse information from multiple sensory streams. Our architecture consists of Recurrent Neural Networks (RNNs) that use Long Short-Term Memory (LSTM) units to capture long temporal dependencies. We train our architecture in a sequence-to-sequence prediction manner, and it explicitly learns to predict the future given only a partial temporal context. We further introduce a novel loss layer for anticipation which prevents over-fitting and encourages early anticipation. We use our architecture to anticipate driving maneuvers several seconds before they happen on a natural driving data set of 1180 miles. The context for maneuver anticipation comes from multiple sensors installed on the vehicle. Our approach shows significant improvement over the state-of-the-art in maneuver anticipation by increasing the precision from 77.4% to 90.5% and recall from 71.2% to 87.4%. Ashesh Jain, Avi Singh, Hema Swetha Koppula, Shane Soh, Ashutosh Saxena |
ICRA | 3 |
| 2016 | Modeling 3D Environments through Hidden Human ContextabstractThe idea of modeling object-object relations has been widely leveraged in many scene understanding applications. However, as the objects are designed by humans and for human usage, when we reason about a human environment, we reason about it through an interplay between the environment, objects and humans. In this paper, we model environments not only through objects, but also through latent human poses and human-object interactions. In order to handle the large number of latent human poses and a large variety of their interactions with objects, we present Infinite Latent Conditional Random Field (ILCRF) that models a scene as a mixture of CRFs generated from Dirichlet processes. In each CRF, we model objects and object-object relations as existing nodes and edges, and hidden human poses and human-object relations as latent nodes and edges. ILCRF generatively models the distribution of different CRF structures over these latent nodes and edges. We apply the model to the challenging applications of 3D scene labeling and robotic scene arrangement. In extensive experiments, we show that our model significantly outperforms the state-of-the-art results in both applications. We further use our algorithm on a robot for arranging objects in a new scene using the two applications aforementioned. Hema Swetha Koppula, Ashutosh Saxena |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Anticipating Human Activities Using Object Affordances for Reactive Robotic ResponseabstractAn important aspect of human perception is anticipation, which we use extensively in our day-to-day activities when interacting with other humans as well as with our surroundings. Anticipating which activities will a human do next (and how) can enable an assistive robot to plan ahead for reactive responses. Furthermore, anticipation can even improve the detection accuracy of past activities. The challenge, however, is two-fold: We need to capture the rich context for modeling the activities and object affordances, and we need to anticipate the distribution over a large space of future human activities. In this work, we represent each possible future using an anticipatory temporal conditional random field (ATCRF) that models the rich spatial-temporal relations through object affordances. We then consider each ATCRF as a particle and represent the distribution over the potential futures using a set of particles. In extensive evaluation on CAD-120 human activity RGB-D dataset, we first show that anticipation improves the state-of-the-art detection results. We then show that for new subjects (not seen in the training set), we obtain an activity anticipation accuracy (defined as whether one of top three predictions actually happened) of 84.1, 74.4 and 62.2 percent for an anticipation time of 1, 3 and 10 seconds respectively. Finally, we also show a robot using our algorithm for performing a few reactive responses. Hema Swetha Koppula, Ashutosh Saxena |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Car that Knows Before You Do: Anticipating Maneuvers via Learning Temporal Driving ModelsabstractAdvanced Driver Assistance Systems (ADAS) have made driving safer over the last decade. They prepare vehicles for unsafe road conditions and alert drivers if they perform a dangerous maneuver. However, many accidents are unavoidable because by the time drivers are alerted, it is already too late. Anticipating maneuvers beforehand can alert drivers before they perform the maneuver and also give ADAS more time to avoid or prepare for the danger. In this work we anticipate driving maneuvers a few seconds before they occur. For this purpose we equip a car with cameras and a computing device to capture the driving context from both inside and outside of the car. We propose an Autoregressive Input-Output HMM to model the contextual information alongwith the maneuvers. We evaluate our approach on a diverse data set with 1180 miles of natural freeway and city driving and show that we can anticipate maneuvers 3.5 seconds before they occur with over 80% F1-score in real-time. Ashesh Jain, Hema Swetha Koppula, Bharad Raghavan, Shane Soh, Ashutosh Saxena |
ICCV | 2 |
| 2014 | Physically Grounded Spatio-temporal Object Affordances
Hema Swetha Koppula, Ashutosh Saxena |
ECCV (3) | 1 |
| 2013 | Hallucinated Humans as the Hidden Context for Labeling 3D ScenesabstractFor scene understanding, one popular approach has been to model the object-object relationships. In this paper, we hypothesize that such relationships are only an artifact of certain hidden factors, such as humans. For example, the objects, monitor and keyboard, are strongly spatially correlated only because a human types on the keyboard while watching the monitor. Our goal is to learn this hidden human context (i.e., the human-object relationships), and also use it as a cue for labeling the scenes. We present Infinite Factored Topic Model (IFTM), where we consider a scene as being generated from two types of topics: human configurations and human-object relationships. This enables our algorithm to hallucinate the possible configurations of the humans in the scene parsimoniously. Given only a dataset of scenes containing objects but not humans, we show that our algorithm can recover the human object relationships. We then test our algorithm on the task of attribute and object labeling in 3D scenes and show consistent improvements over the state-of-the-art. Hema Swetha Koppula, Ashutosh Saxena |
CVPR | 2 |
| 2013 | Learning Spatio-Temporal Structure from RGB-D Videos for Human Activity Detection and AnticipationabstractWe consider the problem of detecting past activities as well as anticipating which activity will happen in the future and how. We start by modeling the rich spatio-temporal relations between human poses and objects (called affordances) using a conditional random field (CRF). However, because of the ambiguity in the temporal segmentation of the sub-activities that constitute an activity, in the past as well as in the future, multiple graph structures are possible. In this paper, we reason about these alternate possibilities by reasoning over multiple possible graph structures. We obtain them by approximating the graph with only additive features, which lends to efficient dynamic programming. Starting with this proposal graph structure, we then design moves to obtain several other likely graph structures. We then show that our approach improves the state-of-the-art significantly for detecting past activities as well as for anticipating future activities, on a dataset of 120 activity videos collected from four subjects. Hema Swetha Koppula, Ashutosh Saxena |
ICML (3) | 1 |
| 2013 | Anticipating human activities for reactive robotic responseabstractAn important aspect of human perception is anticipation, which we use extensively in our day-to-day activities when interacting with other humans as well as with our surroundings. Anticipating which activities will a human do next (and how to do them) can enable an assistive robot to plan ahead for reactive responses in the human environments. In this work, our goal is to enable robots to predict the future activities as well as the details of how a human is going to perform them in short-term (e.g., 1-10 seconds). For example, if a robot has seen a person move his hand to a coffee mug, it is possible he would move the coffee mug to a few potential places such as his mouth, to a kitchen sink or just move it to a different location on the table. If a robot can anticipate this, then it would rather not start pouring milk into the coffee when the person is moving his hand towards the mug, thus avoiding a spill. We represent each possible future using an anticipatory temporal conditional random field (ATCRF) that models the rich spatial-temporal relations through object affordances. We then consider each ATCRF as a particle and represent the distribution over the potential futures using a set of particles. We evaluate our anticipation approach extensively on CAD-120 human activity dataset, which contains 120 RGB-D videos of daily human activities, such as microwaving food, taking medicine, etc. For robotic evaluation, we measure how many times the robot anticipates and performs the correct reactive response. The accompanying video shows a PR2 robot performing assistive tasks based on the anticipations generated by our proposed method. Hema Swetha Koppula, Ashutosh Saxena |
IROS | 1 |
| 2011 | Semantic Labeling of 3D Point Clouds for Indoor ScenesabstractInexpensive RGB-D cameras that give an RGB image together with depth data have become widely available. In this paper, we use this data to build 3D point clouds of full indoor scenes such as an office and address the task of semantic labeling of these 3D point clouds. We propose a graphical model that captures various features and contextual relations, including the local visual appearance and shape cues, object co-occurence relationships and geometric relationships. With a large number of object classes and relations, the model’s parsimony becomes important and we address that by using multiple types of edge potentials. The model admits efficient approximate inference, and we train it using a maximum-margin learning approach. In our experiments over a total of 52 3D scenes of homes and offices (composed from about 550 views, having 2495 segments labeled with 27 object classes), we get a performance of 84.06% in labeling 17 object classes for offices, and 73.38% in labeling 17 object classes for home scenes. Finally, we applied these algorithms successfully on a mobile robot for the task of finding objects in large cluttered rooms. Hema Swetha Koppula, Abhishek Anand, Thorsten Joachims, Ashutosh Saxena |
NIPS | 1 |
| 2010 | Alignment of short length parallel corpora with an application to web searchabstractWith evolving Web, short length parallel corpora is becoming very common and some of these include user queries, web snippets etc. This paper concerns situations where short length parallel corpora has to be analyzed in order to find meaningful unit-alignment. This is similar to dealing with parallel corpora where a sentence level alignment of translations is required, but differs in that the alignment is to be inferred at unit (word or phrase) level. A Conditional Random Field (CRF) based approach is proposed to discover this unit alignment. Given pairs of semantically or syntactically similar entities, the problem is formulated as that of mutual segmentation and sequence alignment problem. The mutual segmentation refers to the process of segmenting the first entity based on units (or labels) in the second entity and vice-versa. The process of optimizing this mutual segmentation also results in optimal unit alignment. Since our training data is not segmented and unit-aligned, we modify the CRF objective function to accommodate unsupervised data and iterative learning. We have applied this framework to Web Search domain and specifically for query reformulation task. Finally, our experiments suggest that the proposed approach indeed results in meaningful alternatives of the original query. Jitendra Ajmera, Hema Swetha Koppula, Krishna P. Leela, Shibnath Mukherjee, Mehul Parsana |
CIKM | 2 |
| 2010 | Learning URL patterns for webpage de-duplicationabstractPresence of duplicate documents in the World Wide Web adversely affects crawling, indexing and relevance, which are the core building blocks of web search. In this paper, we present a set of techniques to mine rules from URLs and utilize these rules for de-duplication using just URL strings without fetching the content explicitly. Our technique is composed of mining the crawl logs and utilizing clusters of similar pages to extract transformation rules, which are used to normalize URLs belonging to each cluster. Preserving each mined rule for de-duplication is not efficient due to the large number of such rules. We present a machine learning technique to generalize the set of rules, which reduces the resource footprint to be usable at web-scale. The rule extraction techniques are robust against web-site specific URL conventions. We compare the precision and scalability of our approach with recent efforts in using URLs for de-duplication. Experimental results demonstrate that our approach achieves 2 times more reduction in duplicates with only half the rules compared to the most recent previous approach. Scalability of the framework is demonstrated by performing a large scale evaluation on a set of 3 Billion URLs, implemented using the MapReduce framework. Hema Swetha Koppula, Krishna P. Leela, Krishna Prasad Chitrapura, Sachin Garg, Amit Sasturkar |
WSDM | 1 |
| 2009 | URL normalization for de-duplication of web pagesabstractPresence of duplicate documents in the World Wide Web adversely affects crawling, indexing and relevance, which are the core building blocks of web search. In this paper, we present a set of techniques to mine rules from URLs and utilize these learnt rules for de-duplication using just URL strings without fetching the content explicitly. Our technique is composed of mining the crawl logs and utilizing clusters of similar pages to extract specific rules from URLs belonging to each cluster. Preserving each mined rules for de-duplication is not efficient due to the large number of specific rules. We present a machine learning technique to generalize the set of rules, which reduces the resource footprint to be usable at web-scale. The rule extraction techniques are robust against web-site specific URL conventions. We demonstrate the effectiveness of our techniques through experimental evaluation. Hema Swetha Koppula, Krishna P. Leela, Krishna Prasad Chitrapura, Sachin Garg, Pavan Kumar GM, Chittaranjan Haty, Amit Sasturkar |
CIKM | 2 |