Ji Meng Loh

dblp:78/8749 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 3 since 2021Databases, data management, data science and information retrieval · 5Artificial intelligence and machine learning · 3Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Subsumption, correctness and relative correctness: Implications for software testing
Samia Al Blwi, Imen Marsit, Besma Khaireddine, Amani Ayad, Ji Meng Loh, Ali Mili 0001
Sci. Comput. Program.5
2024 Invariant relations for affine loops
abstract
Abstract Invariant relations are used to analyze while loops; while their primary application is to derive the function of a loop, they can also be used to derive loop invariants, weakest preconditions, strongest postconditions, sufficient conditions of correctness, necessary conditions of correctness, and termination conditions of loops. In this paper we present two generic invariant relations that capture the semantics of loops whose loop body applies affine transformations on numeric variables.
Wided Ghardallou, Hessamaldin Mohammadi, Richard C. Linger, Mark G. Pleszkoch, Ji Meng Loh, Ali Mili 0001
Acta Informatica5
2022 Generalized Mutant Subsumption
Samia Al Blwi, Imen Marsit, Besma Khaireddine, Amani Ayad, Ji Meng Loh, Ali Mili 0001
ICSOFT5
2021 The ratio of equivalent mutants: A key to analyzing mutation equivalence
Imen Marsit, Amani Ayad, Monsour Latif, Ji Meng Loh, Mohamed Nazih Omri, Ali Mili 0001
J. Syst. Softw.5
2019 Quantitative Metrics for Mutation Testing
Amani Ayad, Imen Marsit, Ji Meng Loh, Mohamed Nazih Omri, Ali Mili 0001
ICSOFT3
2018 Impact of Mutation Operators on Mutant Equivalence
Imen Marsit, Mohamed Nazih Omri, Ji Meng Loh, Ali Mili 0001
ICSOFT3
2018 Impact of Mutation Operators on the Ratio of Equivalent Mutants
abstract
Software mutation is a widely used technique of software testing that consists in generating variants of a base program by applying standard modifications to its source code. One of the main obstacles in the use of software mutations is the existence of equivalent mutants, i.e. mutants whose behavior is indistinguishable from the base program, even though their source code is distinct. Despite several decades of research, the identification of equivalent mutants remains an open problem. Rather than attempting to identify individual mutants that are equivalent to the base, we argue that it is often sufficient to estimate the number of equivalent mutants; also, we argue that the number of equivalent mutants depends on two factors that must be considered in the estimation effort, namely the base program and the mutation operators that are used; in this paper, we explore the impact of mutation operators on the number of equivalent mutants.
Imen Marsit, Mohamed Nazih Omri, Ji Meng Loh, Ali Mili 0001
SoMeT3
2015 A masking index for quantifying hidden glitches
Laure Berti-Équille, Ji Meng Loh, Tamraparni Dasu
Knowl. Inf. Syst.2
2014 Empirical glitch explanations
abstract
Data glitches are unusual observations that do not conform to data quality expectations, be they logical, semantic or statistical. By applying data integrity constraints, potentially large sections of data could be flagged as being noncompliant. Ignoring or repairing significant sections of the data could fundamentally bias the results and conclusions drawn from analyses. In the context of Big Data where large numbers and volumes of feeds from disparate sources are integrated, it is likely that significant portions of seemingly noncompliant data are actually legitimate usable data.
Tamraparni Dasu, Ji Meng Loh, Divesh Srivastava
KDD2
2013 A Masking Index for Quantifying Hidden Glitches
abstract
Data glitches are errors in a data set, they are complex entities that often span multiple attributes and records. When they co-occur in data, the presence of one type of glitch can hinder the detection of another type of glitch. This phenomenon is called masking. In this paper, we define two important types of masking, and we propose a novel, statistically rigorous indicator called masking index for quantifying the hidden glitches in four cases of masking: outliers masked by missing values, outliers masked by duplicates, duplicates masked by missing values, and duplicates masked by outliers. The masking index is critical for data quality profiling and data exploration, it enables a user to measure the extent of masking and hence the confidence in the data. In this sense, it is a valuable data quality index for measuring the true cleanliness of the data. It is also an objective and quantitative basis for choosing an anomaly detection method that is best suited for the glitches that are present in any given data set. We demonstrate the utility and effectiveness of the masking index by intensive experiments on synthetic and real-world datasets.
Laure Berti-Équille, Ji Meng Loh, Tamraparni Dasu
ICDM2
2012 Statistical Distortion: Consequences of Data Cleaning
abstract
We introduce the notion of statistical distortion as an essential metric for measuring the effectiveness of data cleaning strategies. We use this metric to propose a widely applicable yet scalable experimental framework for evaluating data cleaning strategies along three dimensions: glitch improvement, statistical distortion and cost-related criteria. Existing metrics focus on glitch improvement and cost, but not on the statistical impact of data cleaning strategies. We illustrate our framework on real world data, with a comprehensive suite of experiments and analyses.
Tamraparni Dasu, Ji Meng Loh
Proc. VLDB Endow.2
2011 Route classification using cellular handoff patterns
abstract
Understanding utilization of city roads is important for urban planners. In this paper, we show how to use handoff patterns from cellular phone networks to identify which routes people take through a city. Specifically, this paper makes three contributions. First, we show that cellular handoff patterns on a given route are stable across a range of conditions and propose a way to measure stability within and between routes using a variant of Earth Mover's Distance. Second, we present two accurate classification algorithms for matching cellular handoff patterns to routes: one requires test drives on the routes while the other uses signal strength data collected by high-resolution scanners. Finally, we present an application of our algorithms for measuring relative volumes of traffic on routes leading into and out of a specific city, and validate our methods using statistics published by a state transportation authority.
Richard A. Becker, Ramón Cáceres, Karrie J. Hanson, Ji Meng Loh, Simon Urbanek, Alexander Varshavsky, Chris Volinsky
UbiComp4
2010 Spatial probabilistic modeling of calls to businesses
abstract
Local search engines allow users to search for entities such as businesses in a particular geographic location. To improve the geographic relevance of search, user feedback data such as logged click locations are traditionally used. In this paper, we use anonymized mobile call log data as an alternate source of data and investigate its relevance to local search. Such data consists of records of anonymized mobile calls made to local businesses along with the locations of celltowers that handled the calls. We model the probability of calls made to particular categories of businesses as a function of distance, using a generalized linear model framework. We provide a detailed comparison between a click log and a mobile call log, showing its relevance to local search. We describe our probabilistic models and apply them to anonymized mobile call logs for New York City and Los Angeles restaurants.
Ramaswamy Hariharan, Ji Meng Loh, James Shanahan, Kenji Yamada
GIS2