Sumit Mishra

dblp:133/0965 · DBLP profile ↗
← Back
32ranked-venue papers
26as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 18 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 IMUDistill: Knowledge Transfer for Enhancing Low Precision IMU Performance
abstract
Inertial measurement unit (IMU) is a critical component for autonomous robot navigation. This paper presents a novel training framework that uses knowledge distillation to enhance low end, or low precision, IMU performance by transferring learned representations from high end IMU. In the proposed framework consisting of dual branch, teacher branch trained with high end IMU sensor data is used to train the student branch with low end IMU. For training the low end IMU branch, combined loss function that balances signal matching and knowledge transfer is used. Testing with the MAGF-ID dataset demonstrates substantial reduction of mean absolute error as much as 39-65% with accelerometer measurements and 47-70% for gyroscope readings with best improvement in the (gyroscope) X-axis, as compared to original low end IMU measurements. Qualitative analysis also reveals that proposed framework is efficient in systematic error correction and noise suppression.
Sumit Mishra, Hitesh Kumar, Kuk Won Ko, Dong-Soo Har
IEEE Signal Process. Lett.1
2026 TIME-VAD: Text-Informed Magnitude Enhancement Feature Learning for Vehicle Accident Detection and Anticipation
abstract
Vehicular accidents pose a substantial risk to drivers, underscoring the persistent and vital need for heightening safety measures. Early accident anticipation mechanisms are imperative for proactive measures, while detection accuracy is pivotal for prompt response and effective post-accident mitigation. Accurate and early anticipation of accidents for automated driving assistance systems in vehicles or CCTV in cities remains a complex task due to the intricate spatial-temporal interactions within traffic videos. This study presents text-informed magnitude enhancement in contrastive multiple-instance feature learning for vehicle accident detection and anticipation (TIME-VAD). Text is a better representative of concepts when compared to images in video, thus multi-modal learning is suitable. Also, the traditional assumption about feature magnitude of accidents and normal frames in magnitude based multiple-instance learning using weak supervision may not hold. This has led to the development of a novel weak-supervised learning strategy involving magnitude enhancement from textual concepts. For a better frame-level perception of accident risks in videos, dynamic temporal attentions are refined using the proposed dilated temporal conv-attention (DTCA) block. In-depth component-level analysis is performed to showcase the model’s efficacy while elucidating its operational mechanisms. Evaluation is conducted on three benchmark datasets, considering both earliness and accuracy-related metrics. Extensive experiments demonstrate that our TIME-VAD model outperforms the existing models. Compared to the previous top-performing supervised model achieving 84.7% accuracy, TIME-VAD achieves a 94.44% accuracy (measured by ROC-AUC) on the largest DoTA dataset. Notably, our model also excels in measuring how early it detects accidents compared to previous methods. The code will be released onhttps://github.com/sumitmishra209/TIME-VAD
Sumit Mishra, Medhavi Mishra, Pranjay Shyam, Dong-Soo Har
IEEE Trans. Intell. Transp. Syst.1
2025 Analysis of Merge Non-dominated Sorting Algorithm
Sumit Mishra, Ved Prakash, Carlos A. Coello Coello
EMO (2)1
2024 Know How Much Sensitive Precision and Recall Validity Measures Are?
Sumit Mishra, Srinibas Swain, Ved Prakash
ICPR (2)1
2023 On the Computational Complexity of Efficient Non-dominated Sort Using Binary Search
Ved Prakash, Sumit Mishra, Carlos A. Coello Coello
EMO2
2023 MMCo-Clus - An Evolutionary Co-clustering Algorithm for Gene Selection (Extended abstract)
abstract
Dimensionality reduction through feature selection becomes inevitable to overcome the problem of the Curse of dimensionality. In this article, we propose a feature (gene) selection method for high dimensional gene expression (GE) data through a Multi-objective optimization-based Multi-view Co-Clustering algorithm (named MMCo-Clus). A thorough comparative analysis with existing feature selection algorithms using external/internal evaluation metrics supports our proposed method’s potency.
Laizhong Cui, Sudipta Acharya, Sumit Mishra, Yi Pan 0001, Joshua Zhexue Huang
ICDE3
2023 Sensing Accident-Prone Features in Urban Scenes for Proactive Driving and Accident Prevention
abstract
In urban cities, visual information on and along roadways is likely to distract drivers and lead to missing traffic signs and other accident-prone (AP) features. To avoid accidents due to missing these visual cues, this paper proposes a visual notification of AP-features to drivers based on real-time images obtained via dashcam. For this purpose, Google Street View images around accident hotspots (areas of dense accident occurrence) identified by a real-accident dataset are used to train a novel attention module to classify a given urban scene into an accident hotspot or a non-hotspot (area of sparse accident occurrence). The proposed module leverages channel, point, and spatial-wise attention learning on top of different CNN backbones. This leads to better classification results and more certain AP-features with better contextual knowledge when compared with CNN backbones alone. Our proposed module achieves up to 92% classification accuracy. The capability of detecting AP-features by the proposed model were analyzed by a comparative study of three different class activation map (CAM) methods, which are used to inspect specific AP-features causing the classification decision. The outputs of the CAM methods were processed by an image processing pipeline to extract only the AP-features that are explainable to drivers and notified using a visual notification system. Range of experiments was performed to prove the efficacy and AP-features of the system. Ablation of the AP-features taking 9.61%, on average, of the total area in each image sample increased the chance of a given area to be classified as a non-hotspot by up to 21.8%.
Sumit Mishra, Praveen Kumar Rajendran, Luiz Felipe Vecchietti, Dong-Soo Har
IEEE Trans. Intell. Transp. Syst.1
2022 Infra Sim-to-Real: An efficient baseline and dataset for Infrastructure based Online Object Detection and Tracking using Domain Adaptation
abstract
Increasing usage of traffic cameras provides an opportunity to utilize them for smart city applications. However, the efficacy of such systems is determined by their ability to detect and track objects of interest from diverse viewpoints accurately. This is challenging due to the diverse viewpoints, elevations, and distinct properties of camera sensors. Thus, to ensure robust performance, the training dataset should cover many variations, including viewpoints, illumination changes, and diverse weather conditions. However, constructing such a dataset is expensive in terms of data collection and annotation. This paper proposes an unsupervised domain adaptation approach wherein a synthetic dataset is generated using a simulator and subsequently used to ensure performance consistency of multi-object-tracking (MOT) algorithms across a diverse range of manually annotated natural scenes. Towards this end, we emphasize achieving domain invariant object detection by combining image stylization and class-balancing augmentation. Furthermore, we extend the robust detection algorithm to track detected objects across a large time scale using feature embeddings generated by the detector. Based on qualitative and quantitative results, we demonstrate the viability of such a system that is invariant to illumination, weather, viewpoint, and scene changes while providing a baseline for future research. Codebase and datasets would be made available at https://github.com/pranjay-dev/IS2R.
Pranjay Shyam, Sumit Mishra, Kuk-Jin Yoon, Kyung-Soo Kim 0001
IV2
2022 A multi-objective worker selection scheme in crowdsourced platforms using NSGA-II
Akash Yadav, Sumit Mishra, Ashok Singh Sairam
Expert Syst. Appl.2
2022 RTT-Based Rogue UAV Detection in IoV Networks
abstract
Unmanned aerial vehicles (UAVs) are being used in different emerging domains for accomplishing many critical tasks. However, due to the various constraints, such as battery life, computational resources, etc., a UAV under a mission (M-UAV) often needs assistance from an edge/cloud server that is reachable from the M-UAV’s location. A connection between an M-UAV and edge server can be established via an access point or AP. Therefore, before sharing any sensitive information with the edge server, it is essential for an M-UAV to determine the legitimacy of the selected AP. Recently, some works in this direction indicate that a rogue UAV (R-UAV) can successfully mimic a legitimate AP for intercepting the communication channel. Hence, there should be a robust detection mechanism in place for addressing such a threat scenario. In this article, considering one of the emerging domains—the Internet of Vehicle (IoV) networks, at first, we show that communication in the IoV networks can get benefit from the presence of M-UAVs. However, as the link between the M-UAV and edge server can be intercepted by an R-UAV, the adversary may access the sensitive information from the IoV networks. Followed by this, we propose atiming-basedalgorithm for identifying the presence of rogue APs (or R-UAVs) in the channel. The M-UAV executes the timing-based algorithm, and the detection methoddoes notrequire any auxiliary hardware or any modification to the network protocols for meeting the objective. Supported by an extensive evaluation study, we show that without any rigid restriction on the M-UAV’s speed (e.g., by limiting it to almost static) the proposed approach significantly enhances the detection accuracy (at least by a margin of 29.7% and 16.65%) compared to the state-of-the-art methods.
Nilesh Chakraborty, Yao Chao, Jianqiang Li 0001, Sumit Mishra, Chengwen Luo 0001, Ying He 0006, Jie Chen 0027, Yi Pan 0001
IEEE Internet Things J.4
2022 MMCo-Clus - An Evolutionary Co-clustering Algorithm for Gene Selection
abstract
In the era of Big Data, cluster analysis of high-dimensional data sets often suffers from theCurse of dimensionality. To overcome this problem, the dimensionality reduction throughfeature selectionbecomes inevitable. Co-clustering or two-way clustering is considered to be a more sophisticated tool than conventional one-way clustering. Moreover, the advent of multi-view learning shows that the subjects of a data set can be interpreted in many ways. Interestingly, a minimal number of existing feature selection algorithms take advantage of the co-clustering method and are designed to consider multi-view data. Motivated by this, in the current article, we propose a feature (gene) selection method for high dimensional gene expression (GE) data through amulti-objective optimization basedmulti-viewCo-Clustering algorithm (namedMMCo-Clus). A popular evolutionary technique – Non-dominated Sorting Genetic Algorithm-II (NSGA-II) has been utilized as the proposed method's underlying optimization strategy. First, we construct two views of a chosen data set, utilizing knowledge from two different biological data sources. Next, we develop the MMCo-Clusalgorithm considering the constructed views to identify a set of “good” co-clustering solutions. Finally, based on a concept ofconsensus operationon the co-clustering outcome, a small number of most relevant and non-redundant features are extracted from the original feature-space. The reduced dimension formed by new feature-space causes to decrease the computational burden and noise level of original data. For experimental analysis, we have chosen three benchmark GE data sets. Our feature selection method's effectiveness is evaluated through sample-classification accuracy, accompanied by the cluster profile plot/Eisen plot/t-SNE plot, and biological/statistical significance test. A thorough comparative analysis with existing feature selection algorithms using external and internal evaluation metrics supports our proposed method's potency.
Laizhong Cui, Sudipta Acharya, Sumit Mishra, Yi Pan 0001, Joshua Zhexue Huang
IEEE Trans. Knowl. Data Eng.3
2021 Hypervolume by Slicing Objective Algorithm: An Improved Version
abstract
The hypervolume remains a popular performance indicator in evolutionary multi-objective, mainly because of its nice mathematical properties (i.e., it's the only performance indicator known to be Pareto-compliant). However, its high computational cost (which grows polynomially on the population size but exponentially on the number of objectives) has severely limited its use in many-objective optimization. This has motivated a variety of proposals that attempt to overcome this limitation. One of the most popular proposals currently available is the so-called Hypervolume by Slicing Objectives (HSO) algorithm. Here, we show that the worst-case time complexity of the HSO algorithm, as obtained by its authors, is incorrect. Then, we provide an efficient implementation of the HSO algorithm, which guarantees that unique slices are generated to compute the hypervolume.
Sumit Mishra, Srinibas Swain, Sangita Sarmah, Carlos A. Coello Coello
CEC1
2021 Invalid Scenarios of External Cluster Validity Indices: An Analysis Using Bell Polynomial
abstract
External cluster validity indices (CVIs) are used to evaluate various clustering algorithms by comparing the obtained clustering result with the gold standard. These indices provide some values for the obtained clustering result, indicating how good or bad the obtained clustering result is compared to the gold standard. For an external CVI, it is desirable to always provide the values within its range; otherwise, the clustering result cannot be correctly interpreted. Thus, in this work, our objective is to theoretically obtain the scenarios where these indices are unable to provide valid values. 26 commonly used indices identified in the existing literature are considered for our analysis purposes. For all these CVIs, we are able to determine the scenarios where these indices provide undefined values. Using the Bell number and Bell polynomial notion, we have attempted to represent the number of such scenarios.
Sumit Mishra, Sanjay Moulik, Ved Prakash
SMC1
2021 A parallel naive approach for non-dominated sorting: a theoretical study considering PRAM CREW model
Sumit Mishra, Carlos A. Coello Coello
Soft Comput.1
2020 If unsure, shuffle: deductive sort is Θ(MN3), but O(MN2) in expectation over input permutations
abstract
Despite significant advantages in theory of evolutionary computation, many papers related to evolutionary algorithms still lack proper analysis and limit themselves by rather vague reflections on why making a certain design choice improves the performance. While this seems to be unavoidable when talking about the behavior of an evolutionary algorithm on a practical optimization problem, doing the same for computational complexities of parts of evolutionary algorithms is harmful and should be avoided.
Sumit Mishra, Maxim Buzdalov 0001
GECCO1
2020 Filter Sort Is $\varOmega (N^3)$ in the Worst Case
Sumit Mishra, Maxim Buzdalov 0001
PPSN (2)1
2020 DDA-ENS: Dominance Degree Approach based Efficient Non-dominated Sort
abstract
In Pareto based multi- and many-objective evolutionary algorithms (MOEAs and MaOEAs), non-dominated sorting is an important step that divides the set of solutions into different disjoint non-dominated fronts. In the last 20 years, there have been various approaches proposed for non-dominated sorting. Recently, an approach known as Dominance Degree Approach for Non-Dominated Sorting (DDA-NS) has been proposed. This approach guarantees that the number of floating-point comparisons is always bounded by O(M N log N) for N solutions and M objectives, which is not true for various approaches. However, the worst-case time complexity of DDANS is recently proved to be Θ(MN2+ N3). In this paper, we develop an approach on the top of DDA-NS, which also guarantees (M N log N) floating-point comparisons and its worst-case time complexity is proved to be Θ(MN2) as opposed to Θ(MN2+ N3) of DDA-NS.
Sumit Mishra, Rakesh Senwar
SMC1
2019 An Approach for Non-domination Level Update Problem in Steady-State Evolutionary Algorithms With Parallelism
abstract
One of the bottlenecks in steady-state multiobjective evolutionary algorithms (MOEAs) is non-dominated sorting because it is performed every time whenever a new offspring is generated. The recent literature shows that there is no requirement to perform the complete non-dominated sorting procedure because the entire structure of non-domination level (NDL) does not change. Some approaches have been recently proposed based on this idea. In this paper, we update our previous work where an offspring is inserted into the set of fronts, to further reduce the number of dominance comparisons. Additionally, we also explore parallelism in the updated approach in two different manners considering the PRAM CREW model. Finally, the time and space complexities of two parallel versions is theoretically analyzed.
Sumit Mishra, Carlos A. Coello Coello
CEC1
2019 Parallel Best Order Sort for Non-dominated Sorting: A Theoretical Study Considering the PRAM-CREW Model
abstract
In the current paper we focus on parallelization of non-dominated sorting which is an essential step in Pareto-based multi-objective evolutionary algorithms. The parallel approaches can help to reduce the overall execution time of multi-objective evolutionary algorithms. Although there have been some proposals to parallelize non-dominated sorting algorithms, most of them have focused on the fast non-dominated sort algorithm proposed by Deb et al. This paper explores the scope of parallelism in a recently proposed approach known as Best Order Sort, which was proposed by Roy et al. We focus on two different ways of achieving parallelism in Best Order Sort. The time and space complexity of these two parallel schemes is also analyzed theoretically considering the PRAM CREW model.
Sumit Mishra, Carlos A. Coello Coello
CEC1
2019 A Many Objective Optimization Based Entity Matching Framework for Bibliographic Database
abstract
Entity matching aims at mapping records to different entities where records share the common entity name. The entity matching problem is challenging because several times records do not contain complete information (many of the attributes are missing), unequal distribution of records for different entities, big overlaps between records of different entities. In this paper, we have proposed an unsupervised framework to solve this problem. The aforementioned problem is posed as a partitioning problem. Three unknown artifacts for solving this partitioning problem: optimal partitioning including the optimal number of partitions, suitable distance measure which can be utilized to measure the distance between records and a set of attributes/features which can take part in the distance calculation, are determined automatically using the search capability of a multiobjective optimization technique. Several objective functions which help in measuring the goodness of partitioning are optimized simultaneously by some popular multiobjective optimization techniques, NSGA-II/NSGA-III (Non-dominated Sorting Genetic Algorithm-II/III) to solve the partitioning problem. Total 247 combinations of eight different objective functions are used in the experiments for partitioning fourteen bibliographic datasets. A detailed comparative study of the proposed approach using NSGA-II and NSGA-III.
Sumit Mishra, Sriparna Saha 0001, Samrat Mondal
TENCON1
2019 Worker Selection in Crowd-sourced Platforms using Non-dominated Sorting
abstract
Crowdsourcing has lead to a paradigm shift in the manner commercial houses execute projects by lowering the cost-per-unit of production. A crucial aspect in crowdsourcing is selecting the best set of workers that can perform a task. The environment envisaged in this work is an independent pool of workers, each equipped with a pre-defined set of skills. We assume that these skills do not follow any priority order over each other. Given a task with a set of required skills, our aim is to perform a non-dominated sorting of the workers based on the requirement. From this set of ordered workers, we use domination count to select the best set of workers that can perform the task. Empirical results using real dataset is presented.
Sumit Mishra, Akash Yadav, Ashok Singh Sairam
TENCON1
2018 MBOS: Modified Best Order Sort Algorithm for Performing Non-Dominated Sorting
abstract
The current paper aims to improve an efficient algorithm for solving the problem of non-dominated sorting which is one of the dominant steps of any Pareto based multi/many objective optimization algorithms. Recent years witnessed a large number of attempts in developing some solution frameworks for the problem of non-dominated sorting. One such recent approach is `Best Order Sort' which is efficient with respect to the number of dominance comparisons. However, this approach does not perform well in the presence of duplicate solutions. In this paper attempts have been made to modify the `Best Order Sort' and we call this modified version as `Modified Best Order Sort' to remove the above mentioned limitation. The modified best order sort algorithm has been thoroughly analyzed in different scenarios. Current work shows that `Best Order Sort' can be generalized without affecting its best and the worst case time complexities.
Sumit Mishra, Sriparna Saha 0001, Samrat Mondal
CEC1
2018 P-ENS: Parallelism in Efficient Non-Dominated Sorting
abstract
In recent years, several non-dominated sorting approaches have been proposed. Non-dominated sorting is an essential part of Pareto dominance-based multi-objective evolutionary algorithms (MOEAs) and therefore the relevance of being able to perform such process as efficiently as possible is important. As the use of parallelism has become increasingly popular within MOEAs, there is an evident need for parallel implementations of non-dominated sorting algorithms. In this paper, we have focused on an efficient non-dominated sorting (ENS) approach and explored its parallelization. The time complexity of the parallel version of ENS is theoretically analyzed in four different scenarios.
Sumit Mishra, Carlos A. Coello Coello
CEC1
2018 Towards Obtaining Upper Bound on Sensitivity Computation Process for Cluster Validity Measures
abstract
Cluster validity indices are proposed in the literature to measure the goodness of a clustering result. The validity measure provides a value which shows how good or bad the obtained clustering result is, as compared to the actual clustering result. However, the validity measures are not arbitrarily generated. A validity measure should satisfy some of the important properties. However, there are cases when in-spite of satisfying these properties, a validity measure is not able to differentiate the two clustering results correctly. In this regard, sensitivity as a property of validity measure is introduced to capture the differences between the two clustering results. However, sensitivity computation is a computationally expensive task as it requires to explore all the possible combinations of clustering results which are very large in number and these are growing exponentially. So, it is required to compute the sensitivity efficiently. As the possible combinations of clustering results grow exponentially, so it is required to first obtain an upper bound on this possible number of combinations which will be sufficient to compute the value of the sensitivity. In this paper, we obtain an upper bound on the number of possible combinations of clustering results. For this purpose, a generic approach which is suitable for various validity measures and a specific approach which is applicable for two validity measures are proposed. It is also shown that this upper bound is sufficient to compute the sensitivity of various validity measures. This upper bound is very less as compared to the total number of possible combinations of clustering results.
Sumit Mishra, Samrat Mondal, Sriparna Saha 0001
Fundam. Informaticae1
2017 Unsupervised method to ensemble results of multiple clustering solutions for bibliographic data
abstract
Multiobjective optimization refers to optimization of multiple conflicting objective functions simultaneously. Clustering problem is often formulated as a multiobjective optimization problem where multiple cluster quality measures are simultaneously optimized and Pareto based approaches are popular in solving that Pareto based approaches yield a set of solutions known as Pareto front where all the solutions are non-dominated with respect to each other. A single solution is selected by the decision maker according to his/her preference. But when the number of non-dominated solutions is large in number, then it is difficult for the decision maker to choose the one solution. The selection of a solution from the given Pareto front is known as Post-Pareto optimality analysis. In the past many approaches were proposed for solving the aforementioned problem, but most of these involve the decision maker. In this paper, we have proposed an approach to obtain a single solution from a set of non-dominated solutions by combining these solutions without the intervention of the decision maker. We have evaluated our approach on the set of solutions obtained after application of a newly developed multiobjective based clustering technique on bibliographic databases like DBLP.
Sumit Mishra, Sripama Saha, Samrat Mondal
CEC1
2017 GAEMTBD: Genetic algorithm based entity matching techniques for bibliographic databases
Sumit Mishra, Sriparna Saha 0001, Samrat Mondal
Appl. Intell.1
2016 An automatic framework for entity matching in bibliographic databases
abstract
Entity matching is to map the records to the corresponding entity. It is a well known problem studied by many researchers over the last few years. In bibliographic database, the data evolve over time. For example, the email id of an author in DBLP and ArnetMiner which are two bibliographic databases changes with time. Authors also keep on changing their affiliations. The set of authors with whom they work also changes with time. These types of variations make the entity matching task more difficult. In this paper, we have addressed this problem and proposed the nondominated sorting genetic algorithm-II (NSGA-II) based solution framework. The dissimilarities between different records can be measured using various distance measures. One distance measure can be suitable for one data set while some other distance measure can be suitable for some other data sets. So selecting the appropriate distance measure is also difficult. To address this issue, this paper presents an automatic framework which selects the suitable distance measure along with the appropriate partitioning. To encode the partitions, medoid based encoding is used. Several new mutation operations are used to explore the search space efficiently. Silhouette Index along with Xie-Beni Index are optimized simultaneously during the experiments for three bibliographic data sets. The results of our approach are compared with two existing well known techniques, DBLP and ArnetMiner. From the results it is clear that our proposed approach performs well.
Sumit Mishra, Sriparna Saha 0001, Samrat Mondal
CEC1
2016 Divide and conquer based non-dominated sorting for parallel environment
abstract
Many of the real-life problems involve simultaneous optimization of multiple objectives. In recent years there is an enormous increase in the number of multi-objective optimization problems related to different real-life domains. Evolutionary algorithms are the most popular in solving these types of problems. The non-dominating sorting is one of the steps of any multiobjective evolutionary algorithms. This is used mostly to select the non-dominated set of solutions from a given population. In the past various efficient approaches are proposed in the literature to reduce the complexity of this step. As the evolutionary algorithms inhibit parallelism in it. But not all the existing non-dominating sorting approaches have the parallelism property. So in this paper, we have proposed a new approach named as DCNS (Divide and Conquer based Non-dominating Sorting) which inhibits parallelism in it. It has been shown theoretically and empirically that the proposed approach is computationally efficient than existing state-of-the-art methods.
Sumit Mishra, Sriparna Saha 0001, Samrat Mondal
CEC1
2016 Fast implementation of steady-state NSGA-II
abstract
In steady-state evolutionary algorithms, the parent population is updated each time once a new offspring solution is generated. Due to the updation of the parent population, the non-dominated sorting needs to be applied again and again. The repetition of non-dominated sorting makes steady-state algorithms computationally expensive. But the recent study has identified that the insertion of an offspring solution in the known non-domination level structure of solutions does not change the entire structure. So there is no need to apply the complete non-dominated sorting algorithm again and again. In this regard an efficient non-domination level update approach known as ENLU approach was proposed. In steady-state evolutionary algorithm, the same pair of solutions can be compared multiple times in different generations of the algorithm. In this paper, we have performed the same ENLU approach in a different way so that the same pair of solutions is compared only once if they are in the current population. The worst case time complexity of ENLU approach is O(MN2). So for G generations of the steady-state evolutionary algorithm, the worst case time complexity is GO(MN2). In this paper, we have utilized the same ENLU approach which performs the same number of comparisons all the times. The worst case time complexity of our approach is G(O(MN log N) + O(N2)). However, in terms of space complexity the proposed approach requires O(N2) space as compared to O(N) of the ENLU approach. So we have achieved the speedup at the cost of extra space. At the end, we have explored the possibility of parallelism to make the ENLU approach faster.
Sumit Mishra, Samrat Mondal, Sriparna Saha 0001
CEC1
2016 A multiobjective optimization based entity matching technique for bibliographic databases
Sumit Mishra, Sriparna Saha 0001, Samrat Mondal
Expert Syst. Appl.1
2014 On Validation of Clustering Techniques for Bibliographic Databases
abstract
In entity name disambiguation, performance evaluation of any approach is difficult. This is due to the fact that correct or actual results are often not known. Generally for evaluation purpose, three measures namely precision, recall and f-measure are used. They all are external validity indices because they need golden standard data. But in Bibliographic databases like DBLP, Arnetminer, Scopus, Web of Science, Google Scholar, etc., gold standard data is not easily available and it is very difficult to obtain this due to the overlapping nature of data. So, there is a need to use some other matrices for evaluation purpose. In this paper, some internal cluster validity index based schemes are proposed for evaluating entity name disambiguation algorithms when applied on bibliographic data without using any gold standard datasets. Two new internal validity indices are also proposed in the current paper for this purpose. Experimental results shown on seven bibliographic datasets reveal that proposed internal cluster validity indices are able to compare the results obtained by different methods without prior/gold standard. Thus the present paper demonstrates a novel way of evaluating any entity matching algorithm for bibliographic datasets without using any prior/gold standard information.
Sumit Mishra, Sriparna Saha 0001, Samrat Mondal
ICPR1
2013 Entity Matching Technique for Bibliographic Database
Sumit Mishra, Samrat Mondal, Sriparna Saha 0001
DEXA (2)1