VLDB 2026 Research / reviewers in the wild / expert
Shubhra Kanti Karmaker Santu
dblp:150/3985 · also Santu Karmaker, Shubhra (Santu) K. Karmaker, Shubhra Kanti Karmaker
· DBLP profile ↗
25ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0001-5744-6925ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Path Not Taken: Duality in Reasoning about Program ExecutionabstractLarge language models (LLMs) have shown remarkable capabilities across diverse coding tasks.However, their adoption requires a true understanding of program execution rather than relying on surface-level patterns.Existing benchmarks primarily focus on predicting program properties tied to specific inputs (e.g., code coverage, program outputs).As a result, they provide a narrow view of dynamic code reasoning and are prone to data contamination.We argue that understanding program execution requires evaluating its inherent duality through two complementary reasoning tasks: (i) predicting a program's observed behavior for a given input, and (ii) inferring how the input must be mutated toward a specific behavioral objective.Both tasks jointly probe a model's causal understanding of execution flow.We instantiate this duality in DEXBENCH, a benchmark comprising 445 paired instances, and evaluate 13 LLMs.Our results demonstrate that dual-path reasoning provides a robust and discriminative proxy for dynamic code understanding. Eshgin Hasanov, Md. Mahadi Hassan, Shubhra Kanti Karmaker Santu, Aashish Yadavally |
ACL (1) | 3 |
| 2025 | Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled PreferencesabstractMohammad Saqib Hasan, Saikat Chakraborty, Santu Karmaker, Niranjan Balasubramanian. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Mohammad Saqib Hasan, Shubhra Kanti Karmaker Santu, Niranjan Balasubramanian |
ACL (1) | 3 |
| 2025 | Benchmarking LLMs on Semantic Overlap SummarizationabstractSemantic Overlap Summarization (SOS) is a multi-document summarization task focused on extracting the common information shared across alternative narratives which is a capability that is critical for trustworthy generation in domains such as news, law, and healthcare.We benchmark popular Large Language Models (LLMs) on SOS and introduce PrivacyPolicy-Pairs (3P), a new dataset of 135 high-quality samples from privacy policy documents, which complements existing resources and broadens domain coverage.Using the TELeR prompting taxonomy, we evaluate nearly one million LLM-generated summaries across two SOS datasets and conduct human evaluation on a curated subset.Our analysis reveals strong prompt sensitivity, identifies which automatic metrics align most closely with human judgments, and provides new baselines for future SOS research 1 . John Salvador, Naman Bansal, Mousumi Akter 0001, Souvika Sarkar, Anupam Das 0008, Shubhra Kanti Karmaker Santu |
EMNLP | 6 |
| 2025 | LLMs as Meta-Reviewers' Assistants: A Case StudyabstractEftekhar Hossain, Sanjeev Kumar Sinha, Naman Bansal, R. Alexander Knipper, Souvika Sarkar, John Salvador, Yash Mahajan, Sri Ram Pavan Kumar Guttikonda, Mousumi Akter, Md. Mahadi Hassan, Matthew Freestone, Matthew C. Williams Jr., Dongji Feng, Santu Karmaker. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Eftekhar Hossain, Sanjeev Kumar Sinha, Naman Bansal, R. Alexander Knipper, Souvika Sarkar, John Salvador, Yash Mahajan, Sri Guttikonda, Mousumi Akter 0001, Md. Mahadi Hassan, Matthew Freestone, Matthew C. Williams Jr., Dongji Feng, Shubhra Kanti Karmaker Santu |
NAACL (Long Papers) | 14 |
| 2024 | Challenges and Preferences of Learning Machine Learning: A Student PerspectiveabstractThis research paper systematically identifies the perceptions of learning machine learning (ML) topics. To keep up with the ever-increasing need for professionals with ML expertise, for-profit and non-profit organizations conduct a wide range of ML-related courses at undergraduate and graduate levels. Despite the availability of ML-related education materials, there is lack of understanding how students perceive ML-related topics and the dissemination of ML-related topics. A systematic categorization of students' perceptions of these courses can aid educators in understanding the challenges that students face, and use that understanding for better dissemination of ML-related topics in courses. The goal of this paper is to help educators teach machine learning (ML) topics by providing an experience report of students' perceptions related to learning ML. We accomplish our research goal by conducting an empirical study where we deploy a survey with 83 students across five academic institutions. These students are recruited from a mixture of undergraduate and graduate courses. We apply a qualitative analysis technique called open coding to identify challenges that students encounter while studying ML-related topics. Using the same qualitative analysis technique we identify quality aspects do students prioritize ML-related topics. From our survey, we identify 11 challenges that students face when learning about ML topics, amongst which data quality is the most frequent, followed by hardware-related challenges. We observe the majority of the students prefer hands-on projects over theoretical lectures. Furthermore, we find the surveyed students to consider ethics, security, privacy, correctness, and performance as essential considerations while developing ML-based systems. Based on our findings, we recommend educators who teach ML-related courses to (i) incorporate hands-on projects to teach ML-related topics, (ii) dedicate course materials related to data quality, (iii) use lightweight virtualization tools to showcase computationally intensive topics, such as deep neural networks, and (iv) empirical evaluation of how large language models can be used in ML-related education. Effat Farhana, Fan Wu 0013, Hossain Shahriar, Shubhra Kanti Karmaker Santu, Akond Ashfaque Ur Rahman |
FIE | 4 |
| 2024 | SimPal: Towards a Meta-Conversational Framework to Understand Teacher's Instructional Goals for K-12 PhysicsabstractSimulations are widely used to teach science in grade schools.These simulations are often augmented with a conversational artificial intelligence (AI) agent to provide real-time scaffolding support for students conducting experiments using the simulations.AI agents are highly tailored for each simulation, with a predesigned set of Instructional Goals (IGs), making it difficult for teachers to adjust IGs as the agent may no longer align with the revised IGs.Additionally, teachers are hesitant to adopt new third-party simulations for the same reasons.In this research, we introduce SimPal, a Large Language Model (LLM) based meta-conversational agent, to solve this misalignment issue between a pre-trained conversational AI agent and the constantly evolving pedagogy of instructors.Through natural conversation with SimPal, teachers first explain their desired IGs, based on which SimPal identifies a set of relevant physical variables and their relationships to create symbolic representations of the desired IGs.The symbolic representations can then be leveraged to design prompts for the original AI agent to yield better alignment with the desired IGs.We empirically evaluated SimPal using two LLMs, ChatGPT-3.5 and PaLM 2, on 63 Physics simulations from PhET and Golabz.Additionally, we examined the impact of different prompting techniques on LLM's performance by utilizing the TELeR taxonomy to identify relevant physical variables for the IGs.Our findings showed that SimPal can do this task with a high degree of accuracy when provided with a well-defined prompt. Effat Farhana, Souvika Sarkar, R. Alexander Knipper, Indrani Dey, Hari Narayanan, Sadhana Puntambekar, Shubhra Kanti Karmaker Santu |
L@S | 7 |
| 2024 | Processing Natural Language on Embedded Devices: How Well Do Modern Models Perform?abstractVoice-controlled systems are becoming ubiquitous in many IoT-specific applications such as home/industrial automation, automotive infotainment, and healthcare. While cloud-based voice services (\eg Alexa, Siri) can leverage high-performance computing servers, some use cases (\eg robotics, automotive infotainment) may require to execute the natural language processing (NLP) tasks offline, often on resource-constrained embedded devices. Transformer-based language models such as BERT and its variants are primarily developed with compute-heavy servers in mind. Despite the great performance of BERT models across various NLP tasks, their large size and numerous parameters pose substantial obstacles to offline computation on embedded systems. Lighter replacement of such language models (\eg DistilBERT and TinyBERT) often sacrifice accuracy, particularly for complex NLP tasks. Until now, it is still unclear \ca whether the state-of-the-art language models, \viz BERT and its variants are deployable on embedded systems with a limited processor, memory, and battery power and \cb if they do, what are the "right'' set of configurations and parameters to choose for a given NLP task. This paper presents aperformance study of transformer language models under different hardware configurations and accuracy requirements and derives empirical observations about these resource/accuracy trade-offs. In particular, we study how the most commonly used BERT-based language models (\viz BERT, RoBERTa, DistilBERT, and TinyBERT) perform on embedded systems. We tested them on \textitfour off-the-shelf embedded platforms (\hardware) with 2 GB and 4 GB memory (\ie a total of \textiteight hardware configurations) and \textitfour datasets (\ie HuRIC, GoEmotion, CoNLL, WNUT17) running various NLP tasks. Our study finds that executing complex NLP tasks (such as "sentiment'' classification) on embedded systems isfeasible even without any GPUs (\eg \rpi with 2 GB of RAM). We release our implementations for community use. Our findings can help designers understand the deployability and performance of transformer language models, especially those based on BERT architectures. Souvika Sarkar, Mohammad Fakhruddin Babar, Md. Mahadi Hassan, Monowar Hasan, Shubhra Kanti Karmaker Santu |
ICPE | 5 |
| 2023 | On Evaluation of Bangla Word AnalogiesabstractThis paper presents a benchmark dataset of Bangla word analogies for evaluating the quality of existing Bangla word embeddings.Despite being the 7 th largest spoken language in the world, Bangla is still a low-resource language and popular NLP models often struggle to perform well on Bangla data sets.Therefore, developing a robust evaluation set is crucial for benchmarking and guiding future research on improving Bangla word embeddings, which is currently missing.To address this issue, we introduce a new evaluation set of 16,678 unique word analogies in Bangla as well as a translated and curated version of the original Mikolov dataset (10,594 samples) in Bangla.Our experiments with different state-of-the-art embedding models reveal that current Bangla word embeddings struggle to achieve high accuracy on both data sets, demonstrating a significant gap in multilingual NLP research. Mousumi Akter 0001, Souvika Sarkar, Shubhra Kanti Karmaker Santu |
EMNLP | 3 |
| 2023 | Zero-Shot Multi-Label Topic Inference with Sentence Encoders and LLMsabstractIn this paper, we conducted a comprehensive study with the latest Sentence Encoders and Large Language Models (LLMs) on the challenging task of "definition-wild zero-shot topic inference", where users define or provide the topics of interest in real-time.Through extensive experimentation on seven diverse data sets, we observed that LLMs, such as ChatGPT-3.5 and PaLM, demonstrated superior generality compared to other LLMs, e.g., BLOOM and GPT-NeoX.Furthermore, Sentence-BERT, a BERT-based classical sentence encoder, outperformed PaLM and achieved performance comparable to ChatGPT-3.5. Souvika Sarkar, Dongji Feng, Shubhra Kanti Karmaker Santu |
EMNLP | 3 |
| 2023 | Joint upper & expected value normalization for evaluation of retrieval systems: A case study with Learning-to-Rank methods
Dongji Feng, Shubhra Kanti Karmaker Santu |
Inf. Process. Manag. | 2 |
| 2023 | Ad-Hoc Monitoring of COVID-19 Global Research Trends for Well-Informed Policy MakingabstractThe COVID-19 pandemic has affected millions of people worldwide with severe health, economic, social, and political implications. Healthcare Policy Makers (HPMs) and medical experts are at the core of responding to this continuously evolving pandemic situation and are working hard to contain the spread and severity of this relatively unknown virus. Biomedical researchers are continually discovering new information about this virus and communicating the findings through scientific articles. As such, it is crucial for HPMs and funding agencies to monitor the COVID-19 research trend globally on a regular basis. However, given the influx of biomedical research articles, monitoring COVID-19 research trends has become more challenging than ever, especially when HPMs want on-demand guided search techniques with a set of topics of interest in mind. Unfortunately, existing topic trend modeling techniques are unable to serve this purpose as (1) traditional topic models are unsupervised, and (2) HPMs in different regions may have different topics of interest that they want to track. To address this problem, we introduce a novel computational task in this article calledAd-Hoc Topic Tracking, which is essentially a combination ofzero-shottopic categorization and the spatio-temporal analysis task. We then propose multiplezero-shotclassification methods to solve this task by building on state-of-the-art language understanding techniques. Next, we picked the best-performing method based on its accuracy on a separate validation dataset and then applied it to a corpus of recent biomedical research articles to track COVID-19 research endeavors across the globe using a spatio-temporal analysis. A demo website has also been developed for HPMs to create custom spatio-temporal visualizations of COVID-19 research trends. The research outcomes demonstrate that the proposedzero-shotclassification methods can potentially facilitate further research on this important subject matter. At the same time, the spatio-temporal visualization tool will greatly assist HPMs and funding agencies in making well-informed policy decisions for advancing scientific research efforts. Souvika Sarkar, Biddut Sarker Bijoy, Syeda Jannatus Saba, Dongji Feng, Yash Mahajan, Mohammad Ruhul Amin, Sheikh Rabiul Islam, Shubhra Kanti Karmaker Santu |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2022 | Data-Driven Estimation of Effectiveness of COVID-19 Non-pharmaceutical Intervention PoliciesabstractNon-pharmaceutical Interventions (NPIs), such as Stay-at-Home, and Face-Mask-Mandate, are essential components of the public health response to contain an outbreak like COVID-19. However, it is very challenging to quantify the individual or joint effectiveness of NPIs and their impact on people from different racial and ethnic groups or communities in general. Therefore, in this paper, we study the following two research questions: 1) How can we quantitatively estimate the effectiveness of different NPI policies pertaining to the COVID-19 pandemic?; and 2) Do these policies have considerably different effects on communities from different races and ethnicity? To answer these questions, we model the impact of an NPI as a joint function of stringency and effectiveness over a duration of time. Consequently, we propose a novel stringency function that can provide an estimate of how strictly an NPI was implemented on a particular day. Next, we applied two popular tree-based discriminative classifiers, considering the change in daily COVID cases and death counts as binary target variables, while using stringency values of different policies as independent features. Finally, we interpreted the learned feature weights as the effectiveness of COVID-19 NPIs. Our experimental results suggest that, at the country level, restaurant closures and stay-at-home policies were most effective in restricting the COVID-19 confirmed cases and death cases respectively; and overall, restaurant closing was most effective in hold-down of COVID-19 cases at individual community levels such as Asian, White, Black, AIAN and, NHPI. Additionally, we also performed a comparative analysis between race-specific effectiveness and country-level effectiveness to see whether different communities were impacted differently. Our findings suggest that the different policies impacted communities (race and ethnicity) differently. Yash Mahajan, Sheikh Rabiul Islam, Mohammad Ruhul Amin, Shubhra Kanti Karmaker Santu |
IEEE Big Data | 4 |
| 2022 | Semantic Overlap Summarization among Multiple Alternative Narratives: An Exploratory StudyabstractIn this paper, we introduce an important yet relatively unexplored NLP task called Semantic Overlap Summarization (SOS), which entails generating a single summary from multiple alternative narratives which can convey the common information provided by those narratives. As no benchmark dataset is readily available for this task, we created one by collecting 2,925 alternative narrative pairs from the web and then, went through the tedious process of manually creating 411 different reference summaries by engaging human annotators. As a way to evaluate this novel task, we first conducted a systematic study by borrowing the popular ROUGE metric from text-summarization literature and discovered that ROUGE is not suitable for our task. Subsequently, we conducted further human annotations to create 200 document-level and 1,518 sentence-level ground-truth overlap labels. Our experiments show that the sentence-wise annotation technique with three overlap labels, i.e., Absent (A), Partially-Present (PP), and Present (P), yields a higher correlation with human judgment and higher inter-rater agreement compared to the ROUGE metric. Naman Bansal, Mousumi Akter 0001, Shubhra Kanti Karmaker Santu |
COLING | 3 |
| 2022 | SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at ScaleabstractRecent work has introduced an important yet relatively under-explored NLP task called Semantic Overlap Summarization (SOS) that entails generating a summary from multiple alternative narratives which conveys the common information provided by those narratives.Previous work also published a benchmark dataset for this task by collecting 2, 925 alternative narrative pairs from the web and manually annotating 411 different reference summaries by engaging human annotators.In this paper, we exclusively focus on the automated evaluation of the SOS task using the benchmark dataset.More specifically, we first use the popular ROUGE metric from text-summarization literature and conduct a systematic study to evaluate the SOS task.Our experiments discover that ROUGE is not suitable for this novel task and therefore, we propose a new sentencelevel precision-recall style automated evaluation metric, called SEM-F 1 (Semantic F 1 ).It is inspired by the benefits of the sentence-wise annotation technique using overlap labels reported by the previous work.Our experiments show that the proposed SEM-F 1 metric yields a higher correlation with human judgment and higher inter-rater agreement compared to the ROUGE metric. Naman Bansal, Mousumi Akter 0001, Shubhra Kanti Karmaker Santu |
EMNLP | 3 |
| 2022 | Learning to Generate Overlap Summaries through Noisy Synthetic DataabstractSemantic Overlap Summarization (SOS) is a novel and relatively under-explored seq-to-seq task which entails summarizing common information from multiple alternate narratives.One of the major challenges for solving this task is the lack of existing datasets for supervised training.To address this challenge, we propose a novel data augmentation technique, which allows us to create large amount of synthetic data for training a seq-to-seq model that can perform the SOS task.Through extensive experiments using narratives from the news domain, we show that the models finetuned using the synthetic dataset provide significant performance improvements over the pre-trained vanilla summarization techniques and are close to the models fine-tuned on the golden training data; which essentially demonstrates the effectiveness of out proposed data augmentation technique for training seq-to-seq models on the SOS task. Naman Bansal, Mousumi Akter 0001, Shubhra Kanti Karmaker Santu |
EMNLP | 3 |
| 2020 | Empirical Analysis of Impact of Query-Specific Customization of nDCG: A Case-Study with Learning-to-Rank MethodsabstractIn most existing works, nDCG is computed for a fixed cutoff k, i.e., [email protected] and some fixed discounting coefficient. Such a conventional query-independent way to compute nDCG does not accurately reflect the utility of search results perceived by an individual user and is thus non-optimal. In this paper, we conduct a case study of the impact of using query-specific nDCG on the choice of the optimal Learning-to-Rank (LETOR) methods, particularly to see whether using a query-specific nDCG would lead to a different conclusion about the relative performance of multiple LETOR methods than using the conventional query-independent nDCG would otherwise. Our initial results show that the relative ranking of LETOR methods using query-specific nDCG can be dramatically different from those using the query-independent nDCG at the individual query level, suggesting that query-specific nDCG may be useful in order to obtain more reliable conclusions in retrieval experiments. Shubhra Kanti Karmaker Santu, Parikshit Sondhi, ChengXiang Zhai |
CIKM | 1 |
| 2020 | Towards Automated Sexual Violence Report Tracking
Naeemul Hassan, Amrit Poudel, Jason G. Hale, Claire Hubacek, Khandaker Tasnim Huq, Shubhra Kanti Karmaker Santu, Syed Ishtiaque Ahmed |
ICWSM | 6 |
| 2019 | Analysis of Adaptive Training for Learning to Rank in Information RetrievalabstractLearning to Rank is an important framework used in search engines to optimize the combination of multiple features in a single ranking function. In the existing work on learning to rank, such a ranking function is often trained on a large set of different queries to optimize the overall performance on all of them. However, the optimal parameters to combine those features are generally query-dependent, making such a strategy of "one size fits all" non-optimal. Some previous works have addressed this problem by suggesting a query-level adaptive training for learning to rank with promising results. However, previous work has not analyzed the reasons for the improvement. In this paper, we present a Best-Feature Calibration (BFC) strategy for analyzing learning to rank models and use this strategy to examine the benefit of query-level adaptive training. Our results show that the benefit of adaptive training mainly lies in the improvement of the robustness of learning to rank in cases where it does not perform as well as the best single feature. Saar Kuzi, Sahiti Labhishetty, Shubhra Kanti Karmaker Santu, Prasad Pradip Joshi, ChengXiang Zhai |
CIKM | 3 |
| 2019 | TILM: Neural Language Models with Evolving Topical InfluenceabstractContent of text data are often influenced by contextual factors which often evolve over time (e.g., content of social media are often influenced by topics covered in the major news streams).Existing language models do not consider the influence of such related evolving topics, and thus are not optimal.In this paper, we propose to incorporate such topicalinfluence into a language model to both improve its accuracy and enable cross-stream analysis of topical influences.Specifically, we propose a novel language model called Topical Influence Language Model (TILM), which is a novel extension of a neural language model to capture the influences on the contents in one text stream by the evolving topics in another related (or possibly same) text stream.Experimental results on six different text stream data comprised of conference paper titles show that the incorporation of evolving topical influence into a language model is beneficial and TILM outperforms multiple baselines in a challenging task of text forecasting.In addition to serving as a language model, TILM further enables interesting analysis of topical influence among multiple text streams. Shubhra Kanti Karmaker Santu, Kalyan Veeramachaneni, ChengXiang Zhai |
CoNLL | 1 |
| 2018 | JIM: Joint Influence Modeling for Collective Search BehaviorabstractPrevious work has shown that popular trending events are important external factors which pose significant influence on user search behavior and also provided a way to computationally model this influence. However, their problem formulation was based on the strong assumption that each event poses its influence independently. This assumption is unrealistic as there are many correlated events in the real world which influence each other and thus, would pose a joint influence on the user search behavior rather than posing influence independently. In this paper, we study this novel problem of Modeling the Joint Influences posed by multiple correlated events on user search behavior. We propose a Joint Influence Model based on the Multivariate Hawkes Process which captures the inter-dependency among multiple events in terms of their influence upon user search behavior. We evaluate the proposed Joint Influence Model using two months query-log data from https://search.yahoo.com/. Experimental results show that the model can indeed capture the temporal dynamics of the joint influence over time and also achieves superior performance over different baseline methods when applied to solve various interesting prediction problems as well as real-word application scenarios, e.g., query auto-completion. Shubhra Kanti Karmaker Santu, Liangda Li, Yi Chang 0001, ChengXiang Zhai |
CIKM | 1 |
| 2017 | A Study of Feature Construction for Text-based Forecasting of Time Series VariablesabstractTime series are ubiquitous in the world since they are used to measure various phenomena (e.g., temperature, spread of a virus, sales, etc.). Forecasting of time series is highly beneficial (and necessary) for optimizing decisions, yet is a very challenging problem; using only the historical values of the time series is often insufficient. In this paper, we study how to construct effective additional features based on related text data for time series forecasting. Besides the commonly used n-gram features, we propose a general strategy for constructing multiple topical features based on the topics discovered by a topic model. We evaluate feature effectiveness using a data set for predicting stock price changes where we constructed additional features from news text articles for stock market prediction. We found that: 1) Text-based features outperform time series-based features, suggesting the great promise of leveraging text data for improving time series forecasting. 2) Topic-based features are not very effective stand-alone, but they can further improve performance when added on top of n-gram features. 3) The best topic-based feature appears to be a long-term aggregation of topics over time with high weights on recent topics. Dominic Seyler, Shubhra Kanti Karmaker Santu, ChengXiang Zhai |
CIKM | 3 |
| 2017 | On Application of Learning to Rank for E-Commerce SearchabstractE-Commerce (E-Com) search is an emerging important new application of information retrieval. Learning to Rank (LETOR) is a general effective strategy for optimizing search engines, and is thus also a key technology for E-Com search. While the use of LETOR for web search has been well studied, its use for E-Com search has not yet been well explored. In this paper, we discuss the practical challenges in applying learning to rank methods to E-Com search, including the challenges in feature representation, obtaining reliable relevance judgments, and optimally exploiting multiple user feedback signals such as click rates, add-to-cart ratios, order rates, and revenue. We study these new challenges using experiments on industry data sets and report several interesting findings that can provide guidance on how to optimally apply LETOR to E-Com search: First, popularity-based features defined solely on product items are very useful and LETOR methods were able to effectively optimize their combination with relevance-based features. Second, query attribute sparsity raises challenges for LETOR, and selecting features to reduce/avoid sparsity is beneficial. Third, while crowdsourcing is often useful for obtaining relevance judgments for Web search, it does not work as well for E-Com search due to difficulty in eliciting sufficiently fine grained relevance judgments. Finally, among the multiple feedback signals, the order rate is found to be the most robust training objective, followed by click rate, while add-to-cart ratio seems least robust, suggesting that an effective practical strategy may be to initially use click rates for training and gradually shift to using order rates as they become available. Shubhra Kanti Karmaker Santu, Parikshit Sondhi, ChengXiang Zhai |
SIGIR | 1 |
| 2016 | Generative Feature Language Models for Mining Implicit Features from Customer ReviewsabstractOnline customer reviews are very useful for both helping consumers make buying decisions on products or services and providing business intelligence. However, it is a challenge for people to manually digest all the opinions buried in large amounts of review data, raising the need for automatic opinion summarization and analysis. One fundamental challenge in automatic opinion summarization and analysis is to mine implicit features, i.e., recognizing the features implicitly mentioned (referred to) in a review sentence. Existing approaches require many ad hoc manual parameter tuning, and are thus hard to optimize or generalize; their evaluation has only been done with Chinese review data. In this paper, we propose a new approach based on generative feature language models that can mine the implicit features more effectively through unsupervised statistical learning. The parameters are optimized automatically using an Expectation-Maximization algorithm. We also created eight new data sets to facilitate evaluation of this task in English. Experimental results show that our proposed approach is very effective for assigning features to sentences that do not explicitly mention the features, and outperforms the existing algorithms by a large margin. Shubhra Kanti Karmaker Santu, Parikshit Sondhi, ChengXiang Zhai |
CIKM | 1 |
| 2014 | Towards better generalization in Pittsburgh learning classifier systemsabstractGeneralization ability of a classifier is an important issue for any classification task. This paper proposes a new evolutionary system, i.e., EDARIC, based on the Pittsburgh approach for evolutionary machine learning and classification. The new system uses a destructive approach that starts with large-sized rules and gradually decreases the sizes as evolution progresses. Unlike most previous works, EDARIC adopts an intelligent deletion mechanism, evolves a separate population for each class of a given problem and uses an ensemble system to classify unknown instances. These features help in avoiding over-fitting and class-imbalance problems, which are beneficial for improving generalization ability of a classification system. EDARIC also applies a rule post-processing step to exempt the evolution phase from the burden of tuning a large number of parameters. Experimental results on various benchmark classification problems reveal that EDARIC has better generalization ability in case of both standard and imbalanced datasets compared to many existing algorithms in the literature. Shubhra Kanti Karmaker Santu, Kazuyuki Murase |
IEEE Congress on Evolutionary Computation | 1 |
| 2014 | Forecasting time series - A layered ensemble architectureabstractTime series forecasting (TSF) have been widely used in many application areas such as science, engineering and finance. The characteristics of phenomenon generating a series are usually unknown and information available for forecasting is only limited to the past values of the series. It is, therefore, necessary to use an appropriate number of past values, termed lag, for forecasting. This paper presents a layered ensemble architecture (LEA) for TSF problems. Our architecture is consisted of two layers, each of which uses an ensemble of neural networks. Unlike most previous studies on TSF, LEA puts emphasis on both accuracy and diversity among individual networks in an ensemble. While the ensemble of the first layer tries to find an appropriate lag of a given time series, it of the second layer makes forecasting using the obtained lag. The use of the appropriate lag signifies LEA's effort in producing accurate networks for constructing the ensemble. In order to maintain diversity among networks in the ensemble, LEA trains each network in the ensemble using a different training set. The proposed architecture uses a clustering based selection method that considers both accuracy and diversity in selecting networks to construct the ensemble. Accuracy is maintained here by selecting the best networks from each cluster. On the other hand, diversity is ensured by using the variance information in constructing clusters. LEA has been tested extensively on the time series data sets of NN3 competition. In terms of prediction accuracy, our experimental results have showed clearly that LEA is better than other ensemble and non-ensemble algorithms. Shubhra Kanti Karmaker Santu, Kazuyuki Murase |
IJCNN | 2 |