Mijung Kim

dblp:83/7296 · DBLP profile ↗
← Back
34ranked-venue papers
16as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 8 first-author · 1 since 2021Software engineering, systems software and programming languages · 11 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 9 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-authorSystems, architecture and hardware · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Lightweight Concolic Testing via Path-Condition Synthesis for Deep Learning Libraries
abstract
Many techniques have been recently developed for testing deep learning (DL) libraries. Although these techniques have effectively improved API and code coverage and detected unknown bugs, they rely on blackbox fuzzing for input generation. Concolic testing (also known as dynamic symbolic execution) can be more effective in exploring diverse execution paths, but applying it to DL libraries is extremely challenging due to their inherent complexity. In this paper, we introduce the first concolic testing technique for DL libraries. Our technique offers a lightweight approach that significantly reduces the heavy overhead associated with traditional concolic testing. While symbolic execution maintains symbolic expressions for every variable with non-concrete values to build a path condition, our technique computes approximate path conditions by inferring branch conditions via inductive program synthesis. Despite potential imprecision from approximation, our method's light overhead allows for effective exploration of diverse execution paths within the complex implementations of DL libraries. We have implemented our tool, Pathfinder, and evaluated it on PyTorch and TensorFlow. Our results show that Pathfinder outperforms existing API-level DL library fuzzers by achieving 67% more branch coverage on average; up to 63% higher than TitanFuzz and 120% higher than FreeFuzz. Pathfinder is also effective in bug detection, uncovering 61 crash bugs, 59 of which were confirmed by developers as previously unknown, with 32 already fixed.
Yonghyeon Kim, Dahyeon Park, Yuseok Jeon, Jooyong Yi, Mijung Kim
ICSE6
2025 How Effective are Large Language Models in Generating Software Specifications?
abstract
Software specifications are essential for many Software Engineering (SE) tasks such as bug detection and test generation. Many existing approaches are proposed to extract the specifications defined in natural language form (e.g., com-ments) into formal machine-readable form (e.g., first-order logic). However, existing approaches suffer from limited generalizability and require manual efforts. The recent emergence of Large Language Models (LLMs), which have been successfully applied to numerous SE tasks, offers a promising avenue for automating this process. In this paper, we conduct the first empirical study to evaluate the capabilities of LLMs for generating software specifications from software comments or documentation. We evaluate LLMs' performance with Few-Shot Learning (FSL) and compare the performance of 13 state-of-the-art LLMs with traditional approaches on three public datasets. In addition, we conduct a comparative diagnosis of the failure cases from both LLMs and traditional methods, identifying their unique strengths and weaknesses. Our study offers valuable insights for future research to improve specification generation.
Danning Xie, Byoung-Joo Yoo, Nan Jiang 0012, Mijung Kim, Lin Tan 0001, Xiangyu Zhang 0001, Judy S. Lee
SANER4
2025 CMASan: Custom Memory Allocator-aware Address Sanitizer
abstract
Custom Memory Allocator (CMA) replaces the standard memory allocator for various purposes, such as improving memory efficiency or enhancing security. However, memory objects allocated by CMA are vulnerable to memory bugs similar to those allocated by the standard memory allocator. Unfortunately, existing memory bug detection approaches, including Address Sanitizer (ASan), do not work properly with these CMAs because existing approaches are mainly designed for the standard memory allocator. This paper presents CMASan, the first CMA-aware address sanitizer designed to effectively detect memory bugs on CMA objects that ASan misses without requiring expert knowledge, manual code modifications, or changing the unique internal logic of CMAs. According to our evaluation, CMASan successfully identifies 19 previously unknown CMA memory bugs undetected by ASan, including some undetected for 9 years. Compared to ASan, CMASan incurs only an additional 9.63% overhead.
Junwha Hong, Wonil Jang, Mijung Kim, Yonghwi Kwon 0001, Yuseok Jeon
SP3
2025 Asymmetric Voltage Latch Type and Ultra-Low Swing Bitline Sense Amplifiers for Low-Power High-Density 1R1W 8T SRAM in 14 nm FinFET
abstract
We propose two innovative sense amplifiers, the asymmetric voltage latched-type sense amplifier (A-VLSA) and the ultra-low swing bitline sense amplifier (ULS-SA), to enhance read performance and reduce power consumption in the non-hierarchical BL 2-port 8-transistor SRAM (2P-SRAM). A-VLSA minimizes offset voltage through a MOS capacitor-based asymmetric operation, while ULS-SA achieves reduced read power by adopting clipped precharging and a charge-sharing read mechanism without incurring delay penalties. Measurement results demonstrate that A-VLSA (ULS-SA) achieves 63% (37%) and 61% (63%) lower energy-delay product (EDP) at VDD= 0.8V and 0.6V, respectively, with a 24% (43%) smaller area compared to the previous pseudo-differential asymmetric current latched-type sense amplifier (A-CLSA). These improvements address key limitations of prior approaches, delivering significant advancements in both power efficiency and read performance, especially under high cell density conditions.
Keon-hee Cho, Ji Sang Oh, Younmee Bae, Mijung Kim, Sangyeop Baeck, Taejoong Song, Seong-Ook Jung
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 Linearly Controllable GAN: Unsupervised Feature Categorization and Decomposition for Image Generation and Manipulation
Sehyung Lee, Mijung Kim, Yeongnam Chae, Björn Stenger
ECCV (4)2
2024 Constrained Sorter Design using Zero-One Principle
abstract
To derive efficient sorting architectures constrained to application-specific input/output conditions, we present in this paper a systematic design methodology that can effectively prune dispensable compare-and-swap (CAS) units. Unlike the previous works resorting to heuristic approaches, the proposed framework exploits the zero-one principle to validate the pruning of a CAS unit at a time, generating the cost-optimized sorter architecture in an iterative manner with a reasonable complexity. In addition to the given input/output constraints, we newly develop the architecture options for the proposed framework, allowing more design spaces for finding the most attractive constrained-sorter design. For 8-list polar decoders, the proposed framework successfully reduces 70% of CAS units in the baseline full sorter, relaxing the area-time complexity by 35% compared with the state-of-the-art solutions.
Sangil Han, Jaehee Kim, Dongyun Kam, Byeong Yong Kong, Mijung Kim, Young-Seok Kim, Youngjoo Lee 0002
ISCAS5
2024 Testing Diverse Geographical Features of Autonomous Driving Systems
abstract
Testing in various driving scenarios is one of the essential methods to enhance the reliability of autonomous driving systems (ADS). Existing ADS testing research has shown effectiveness in detecting safety violations by generating diverse driving scenarios. However, they do not consider the various geographical features and thus have limited ability to find safety violations caused by complex geographical features. Our paper addresses this limitation by analyzing a given high-definition map and collecting its geographical features. We leverage this information and develop a technique for generating corner case scenarios that exercise diverse geographical features such as curves and slopes. Our approach first generates the ego-vehicle’s driving routes so that they achieve full lane coverage on the entire map, then clusters those routes by geographical features, and constructs driving scenarios by adding other objects and environments. In our experiments on Autoware-Universe, we evaluate our technique with six high-definition maps from the Carla simulator. Our results show that driving scenarios generated by our tool effectively exercise more diverse geographical features than existing work. As a result, our tool uncovers new safety violations that are caused by complex geographical features and would not be detected by existing work.
Seongdeok Seo, Judy S. Lee, Mijung Kim
ISSRE3
2024 Studying Versioning in Stack Overflow
abstract
In Stack Overflow (SO), a post consists of multiple components: title, question, answers, question tags, and comments. Developers can create any of these components and make changes, which we call 'edits'. Edits are an important aspect of QA websites to ensure the quality and correctness of the texts. We performed multiple analyses on the revision history of 23 million SO posts from 2008 to 2023, and we gain a more comprehensive understanding of developers' content maintenance behaviors which lay the foundation for further research.
Fuxiang Chen, Mijung Kim, Fatemeh Hendijani Fard
ASE3
2023 Multi-scale Cell-based Layout Representation for Document Understanding
abstract
Deep learning techniques have achieved remarkable progress in document understanding. Most models use co-ordinates to represent absolute or relative spatial information of components, but they are difficult to represent latent rules in the document layout. This makes learning layout representation to be more difficult. Unlike the previous researches which have employed the coordinate system, graph or grid to represent the document layout, we propose a novel layout representation, the cell-based layout, to provide easy-to-understand spatial information for backbone models. In line with human reading habits, it uses cell information, i.e. row and column index, to represent the position of components in a document, and makes the document layout easier to understand. Furthermore, we proposed the multi-scale layout to represent the hierarchical structure of layout, and developed a data augmentation method to improve the performance. Experiment results show that our method achieves the state-of-the-art performance in text-based tasks, including form understanding and receipt understanding, and improves the performance in image-based task such as document image classification. We released the code in the repoa.
Yuzhi Shi, Mijung Kim, Yeongnam Chae
WACV2
2022 DocTer: documentation-guided fuzzing for testing deep learning API functions
abstract
Input constraints are useful for many software development tasks. For example, input constraints of a function enable the generation of valid inputs, i.e., inputs that follow these constraints, to test the function deeper. API functions of deep learning (DL) libraries have DL-specific input constraints, which are described informally in the free-form API documentation. Existing constraint-extraction techniques are ineffective for extracting DL-specific input constraints.
Danning Xie, Mijung Kim, Hung Viet Pham, Lin Tan 0001, Xiangyu Zhang 0001, Michael W. Godfrey
ISSTA3
2021 DEVIATE: A Deep Learning Variance Testing Framework
abstract
Deep learning (DL) training is nondeterministic and such nondeterminism was shown to cause significant variance of model accuracy (up to 10.8%). Such variance may affect the validity of the comparison of newly proposed DL techniques with baselines. To ensure such validity, DL researchers and practitioners must replicate their experiments multiple times with identical settings to quantify the variance of the proposed approaches and baselines. Replicating and measuring DL variances reliably and efficiently is challenging and understudied.We propose a ready-to-deploy framework DEVIATE that (1) measures DL training variance of a DL model with minimal manual efforts, and (2) provides statistical tests of both accuracy and variance. Specifically, DEVIATE automatically analyzes the DL training code and extracts monitored important metrics (such as accuracy and loss). In addition, DEVIATE performs popular statistical tests and provides users with a report of statistical p-values and effect sizes along with various confidence levels when comparing to selected baselines.We demonstrate the effectiveness of DEVIATE by performing case studies with adversarial training. Specifically, for an adversarial training process that uses the Fast Gradient Signed Method to generate adversarial examples as the training data, DEVIATE measures a max difference of accuracy among 8 identical training runs with fixed random seeds to be up to 5.1%.Tool and demo links: https://github.com/lin-tan/DEVIATE
Hung Viet Pham, Mijung Kim, Lin Tan 0001, Yaoliang Yu, Nachiappan Nagappan
ASE2
2018 Computer-Aided Diagnosis and Localization of Glaucoma Using Deep Learning
Mijung Kim, Homin Park, Jasper Zuallaert, Olivier Janssens, Sofie Van Hoecke, Wesley De Neve
BIBM1
2018 Which generated test failures are fault revealing? prioritizing failures based on inferred precondition violations using PAF
abstract
Automated unit testing tools, such as Randoop, have been developed to produce failing tests as means of finding faults. However, these tools often produce false alarms, so are not widely used in practice. The main reason for a false alarm is that the generated failing test violates an implicit precondition of the method under test, such as a field should not be null at the entry of the method. This condition is not explicitly programmed or documented but implicitly assumed by developers. To address this limitation, we propose a technique called PAF to cluster generated test failures due to the same cause and reorder them based on their likelihood of violating an implicit precondition of the method under test. From various test executions, PAF observes their dataflows to the variables whose values are used when the program fails. Based on the dataflow similarity and where these values are originated, PAF clusters failures and determines their likelihood of being fault revealing. We integrated PAF into Randoop. Our empirical results on open-source projects show that PAF effectively clusters fault revealing tests arising from the same fault and successfully prioritizes the fault-revealing ones.
Mijung Kim, Shing-Chi Cheung, Sunghun Kim 0001
ESEC/SIGSOFT FSE1
2018 SpliceRover: interpretable convolutional neural networks for improved splice site prediction
abstract
Motivation: During the last decade, improvements in high-throughput sequencing have generated a wealth of genomic data. Functionally interpreting these sequences and finding the biological signals that are hallmarks of gene function and regulation is currently mostly done using automated genome annotation platforms, which mainly rely on integrated machine learning frameworks to identify different functional sites of interest, including splice sites. Splicing is an essential step in the gene regulation process, and the correct identification of splice sites is a major cornerstone in a genome annotation system. Results: In this paper, we present SpliceRover, a predictive deep learning approach that outperforms the state-of-the-art in splice site prediction. SpliceRover uses convolutional neural networks (CNNs), which have been shown to obtain cutting edge performance on a wide variety of prediction tasks. We adapted this approach to deal with genomic sequence inputs, and show it consistently outperforms already existing approaches, with relative improvements in prediction effectiveness of up to 80.9% when measured in terms of false discovery rate. However, a major criticism of CNNs concerns their 'black box' nature, as mechanisms to obtain insight into their reasoning processes are limited. To facilitate interpretability of the SpliceRover models, we introduce an approach to visualize the biologically relevant information learnt. We show that our visualization approach is able to recover features known to be important for splice site prediction (binding motifs around the splice site, presence of polypyrimidine tracts and branch points), as well as reveal new features (e.g. several types of exclusion patterns near splice sites). Availability and implementation: SpliceRover is available as a web service. The prediction tool and instructions can be found at http://bioit2.irc.ugent.be/splicerover/. Supplementary information: Supplementary data are available at Bioinformatics online.
Jasper Zuallaert, Fréderic Godin, Mijung Kim, Arne Soete, Yvan Saeys, Wesley De Neve
Bioinform.3
2017 Interpretable convolutional neural networks for effective translation initiation site prediction
abstract
Thanks to rapidly evolving sequencing techniques, the amount of genomic data at our disposal is growing increasingly large. Determining the gene structure is a fundamental requirement to effectively interpret gene function and regulation. An important part in that determination process is the identification of translation initiation sites. In this paper, we propose a novel approach for automatic prediction of translation initiation sites, leveraging convolutional neural networks that allow for automatic feature extraction. Our experimental results demonstrate that we are able to improve the state-of-the-art approaches with a decrease of 75.2% in false positive rate and with a decrease of 24.5% in error rate on chosen datasets. Furthermore, an in-depth analysis of the decision-making process used by our predictive model shows that our neural network implicitly learns biologically relevant features from scratch, without any prior knowledge about the problem at hand, such as the Kozak consensus sequence, the influence of stop and start codons in the sequence and the presence of donor splice site patterns. In summary, our findings yield a better understanding of the internal reasoning of a convolutional neural network when applying such a neural network to genomic data.
Jasper Zuallaert, Mijung Kim, Yvan Saeys, Wesley De Neve
BIBM2
2017 Sandpiper: Scaling probabilistic inferencing to large scale graphical models
abstract
We design and implement a scalable version of loopy belief propagation (BP), a widely used algorithm for performing inference on probabilistic graphical models. However, implementations of BP on generic data processing platforms such as Apache Spark do not scale well to very large graphical models containing billions of vertices. To handle such large-scale graphs, we leverage a number of strategies. Our implementation is based on Apache Spark GraphX. We propose a novel graph partitioning strategy to reduce both computation and communication overhead providing a 2x speed-up. We use efficient memory management for storing the graph and shared memory for highspeed communication. To evaluate performance and demonstrate scalability of the approach, we perform a range of experiments including using real-world graphs with billions of vertices, where we achieve an overall 10x speed-up over a vanilla Spark baseline. Further, we apply our BP implementation to infer the probability of a website being malicious by performing inference on a graphical model derived from real, large-scale hyperlinked webcrawl data. We have open sourced our implementation.
Alexander Ulanov, Manish Marwah, Mijung Kim, Roshan Dathathri, Carlos Zubieta, Jun Li 0008
IEEE BigData3
2017 Sparkle: optimizing spark for large memory machines and analytics
abstract
Given the growing availability of affordable scale-up servers, our goal is to bring the performance benefits of in-memory processing on scale-up servers to an increasingly common class of data analytics applications that process small to medium size datasets (up to a few 100GBs) that can easily fit in the memory of a typical scale-up server To achieve this goal, we leverage Spark, an existing memory-centric data analytics framework with wide-spread adoption among data scientists. Bringing Spark's data analytic capabilities to a scale-up system requires rethinking the original design assumptions, which, although effective for a scale-out system, are a poor match to a scale-up system resulting in unnecessary communication and memory inefficiencies.
Mijung Kim, Jun Li 0008, Haris Volos 0001, Manish Marwah, Alexander Ulanov, Kimberly Keeton, Joseph A. Tucek, Ludmila Cherkasova, Pradeep Fernando
SoCC1
2016 Decomposition-by-normalization (DBN): leveraging approximate functional dependencies for efficient CP and tucker decompositions
Mijung Kim, K. Selçuk Candan
Data Min. Knowl. Discov.1
2015 REMI: defect prediction for efficient API testing
abstract
Quality assurance for common APIs is important since the the reliability of APIs affects the quality of other systems using the APIs. Testing is a common practice to ensure the quality of APIs, but it is a challenging and laborious task especially for industrial projects. Due to a large number of APIs with tight time constraints and limited resources, it is hard to write enough test cases for all APIs. To address these challenges, we present a novel technique, REMI that predicts high risk APIs in terms of producing potential bugs. REMI allows developers to write more test cases for the high risk APIs. We evaluate REMI on a real-world industrial project, Tizen-wearable, and apply REMI to the API development process at Samsung Electronics. Our evaluation results show that REMI predicts the bug-prone APIs with reasonable accuracy (0.681 f-measure on average). The results also show that applying REMI to the Tizen-wearable development process increases the number of bugs detected, and reduces the resources required for executing test cases.
Mijung Kim, Jaechang Nam, Jaehyuk Yeon, Soonhwang Choi, Sunghun Kim 0001
ESEC/SIGSOFT FSE1
2014 Efficient Static and Dynamic In-Database Tensor Decompositions on Chunk-Based Array Stores
abstract
As the relevant data sets get large, existing in-memory schemes for tensor decomposition become increasingly ineffective and, instead, memory-independent solutions, such as in-database analytics, are necessitated. In this paper, we present techniques for efficient implementations of in-database tensor decompositions on chunk-based array data stores. The proposed static and incremental in-database tensor decomposition operators and their optimizations address the constraints imposed by the main memory limitations when handling large and high-order tensor data. Firstly, we discuss how to implement alternating least squares operations efficiently on a chunk-based data storage system. Secondly, we consider scenarios with frequent data updates and show that compressed matrix multiplication techniques can be effective in reducing the incremental tensor decomposition maintenance costs. To the best of our knowledge, this paper presents the first attempt to develop efficient and optimized in-database tensor decomposition operations. We evaluate the proposed algorithms on tensor data sets that do not fit into the available memory and results show that the proposed techniques significantly improve the scalability of this core data analysis.
Mijung Kim, K. Selçuk Candan
CIKM1
2014 TensorDB: In-Database Tensor Manipulation with Tensor-Relational Query Plans
abstract
Today's data management systems increasingly need to support both tensor-algebraic operations (for analysis) as well as relational-algebraic operations (for data manipulation and integration). Tensor decomposition techniques are commonly used for discovering underlying structures of multi-dimensional data sets. However, as the relevant data sets get large, existing in-memory schemes for tensor decomposition become increasingly ineffective and, instead, memory-independent solutions, such as in-database analytics, are necessitated. We introduce an in-database analytic system for efficient implementations of in-database tensor decompositions on chunk-based array data stores, so called, TensorDB. TensorDB includes static in-database tensor decomposition and dynamic in-database tensor decomposition operators. TensorDB extends an array database and leverages array operations for data manipulation and integration. TensorDB supports complex data processing plans where multiple relational algebraic and tensor algebraic operations are composed with each other.
Mijung Kim, K. Selçuk Candan
CIKM1
2014 Pushing-Down Tensor Decompositions over Unions to Promote Reuse of Materialized Decompositions
Mijung Kim, K. Selçuk Candan
ECML/PKDD (1)1
2014 Palette: enabling scalable analytics for big-memory, multicore machines
abstract
Hadoop and its variants have been widely used for processing large scale analytics tasks in a cluster environment. However, use of a commodity cluster for analytics tasks needs to be reconsidered based on two key observations: (1) in recent years, large memory, multicore machines have become more affordable; and (2) recent studies show that most analytics tasks in practice are smaller than 100 GB. Thus, replacing a commodity cluster with a large memory, multicore machine can enable in-memory analytics at an affordable cost. However= programming on a big-memory, multicore machine is a challenge. Multi-threaded programming is notoriously difficult. Further, the memory design of most large memory servers follows non-uniform memory access (NUMA) architecture. While NUMA-aware programming often leads to high efficiency in analytics tasks, it is usually done in an ad hoc manner.
Tere Gonzalez, Jun Li 0008, Manish Marwah, Jim Pruyne, Krishnamurthy Viswanathan, Mijung Kim
SIGMOD Conference7
2012 Systematic Modeling, Testing, and Monitoring of Information Integrity in Federated Ontology-driven Data Sources
Mijung Kim, Jake Cobb, Tahsin M. Kurç, Alessandro Orso, Mary Jean Harrold, Andrew R. Post, Shamkant B. Navathe, Joel H. Saltz
AMIA1
2012 Decomposition-by-normalization (DBN): leveraging approximate functional dependencies for efficient tensor decomposition
abstract
For many multi-dimensional data applications, tensor operations as well as relational operations need to be supported throughout the data lifecycle. Although tensor decomposition is shown to be effective for multi-dimensional data analysis, the cost of tensor decomposition is often very high. We propose a novel decomposition-by-normalization scheme that first normalizes the given relation into smaller tensors based on the functional dependencies of the relation and then performs the decomposition using these smaller tensors. The decomposition and recombination steps of the decomposition-by- normalization scheme fit naturally in settings with multiple cores. This leads to a highly efficient, effective, and parallelized decomposition-by-normalization algorithm for both dense and sparse tensors. Experiments confirm the efficiency and effectiveness of the proposed decomposition-by-normalization scheme compared to the conventional nonnegative CP decomposition approach.
Mijung Kim, K. Selçuk Candan
CIKM1
2012 Efficient regression testing of ontology-driven systems
abstract
To manage and integrate information gathered from heterogeneous databases, an ontology is often used. Like all systems, ontology-driven systems evolve over time and must be regression tested to gain confidence in the behavior of the modified system. Because rerunning all existing tests can be extremely expensive, researchers have developed regression-test-selection (RTS) techniques that select a subset of the available tests that are affected by the changes, and use this subset to test the modified system. Existing RTS techniques have been shown to be effective, but they operate on the code and are unable to handle changes that involve ontologies. To address this limitation, we developed and present in this paper a novel RTS technique that targets ontology-driven systems. Our technique creates representations of the old and new ontologies, compares them to identify entities affected by the changes, and uses this information to select the subset of tests to rerun. We also describe in this paper OntoRetest, a tool that implements our technique and that we used to empirically evaluate our approach on two biomedical ontology-driven database systems. The results of our evaluation show that our technique is both efficient and effective in selecting tests to rerun and in reducing the overall time required to perform regression testing.
Mijung Kim, Jake Cobb, Mary Jean Harrold, Tahsin M. Kurç, Alessandro Orso, Joel H. Saltz, Andrew R. Post, Kunal Malhotra, Shamkant B. Navathe
ISSTA1
2012 SBV-Cut: Vertex-cut based graph partitioning using structural balance vertices
Mijung Kim, K. Selçuk Candan
Data Knowl. Eng.1
2011 Approximate tensor decomposition within a tensor-relational algebraic framework
abstract
In this paper, we first introduce a tensor-based relational data model and define algebraic operations on this model. We note that, while in traditional relational algebraic systems the join operation tends to be the costliest operation of all, in the tensor-relational framework presented here, tensor decomposition becomes the computationally costliest operation. Therefore, we consider optimization of tensor decomposition operations within a relational algebraic framework. This leads to a highly efficient, effective, and easy-to-parallelize join-by-decomposition approach and a corresponding KL-divergence based optimization strategy. Experimental results provide evidence that minimizing KL-divergence within the proposed join-by-decomposition helps approximate the conventional join-then-decompose scheme well, without the associated time and space costs.
Mijung Kim, K. Selçuk Candan
CIKM1
2011 Two-stage multinomial logit model
Jin-Hyung Kim, Mijung Kim
Expert Syst. Appl.2
2010 Automated Bug Neighborhood Analysis for Identifying Incomplete Bug Fixes
abstract
Although many static-analysis techniques have been developed for automatically detecting bugs, such as null dereferences, fewer automated approaches have been presented for analyzing whether and how such bugs are fixed. Attempted bug fixes may be incomplete in that a related manifestation of the bug remains unfixed. In this paper, we characterize the “completeness” of attempted bug fixes that involve the flow of invalid values from one program point to another, such as null dereferences, in Java programs. Our characterization is based on the definition of a bug neighborhood, which is a scope of flows of invalid values. We present an automated analysis that, given two versions P and P' of a program, identifies the bugs in P that have been fixed in P', and classifies each fix as complete or incomplete. We implemented our technique for null-dereference bugs and conducted empirical studies using open-source projects. Our results indicate that, for the projects we studied, many bug fixes are not complete, and thus, may cause failures in subsequent executions of the program.
Mijung Kim, Saurabh Sinha 0001, Carsten Görg, Hina Shah, Mary Jean Harrold, Mangala Gowri Nanda
ICST1
2009 Enabling accessible interfaces to digital library content
abstract
Most of the Web interfaces are primarily designed for people with sight, with visually rich features that makes effective use of the tools to enhance visual usability but in process making it impossible for users who are blind or visually impaired to use them. In this work, our goal is to improve participation to NSF's National Science Digital Library (NSDL) by teachers, librarians, and learners who are blind. The middleware for accessible information spaces on NSDL (MAISON) is enhancing the accessibility of NSDL, its internal and external resources and existing services (such as strand maps of educational benchmarks). Relying on cutting-edge, context-aware graph segmentation, filtering and summarization, and concept propagation techniques, the middleware provides information space adaptation, reduction, and preview services through open Web-based service APIs to enable implementation of informative navigation interfaces that are able to reduce the complexity of the information space and provide previews to prevent user disorientation.
Syed Toufeeq Ahmed, K. Selçuk Candan, Suganthi Cidambaram, Shruti Gaur, Jong Wook Kim, Mijung Kim, Hari Sundaram, Renwei Yu
ICME6
2009 Fault localization and repair for Java runtime exceptions
abstract
This paper presents a new approach for locating and repairing faults that cause runtime exceptions in Java programs. The approach handles runtime exceptions that involve a flow of an incorrect value that finally leads to the exception. This important class of exceptions includes exceptions related to dereferences of null pointers, arithmetic faults (e.g., ArithmeticException), and type faults (e.g., ArrayStoreException). Given a statement at which such an exception occurred, the technique combines dynamic analysis (using stack-trace information) with static backward data-flow analysis (beginning at the point where the runtime exception occurred) to identify the source statement at which an incorrect assignment was made; this information is required to locate the fault. The approach also identifies the source statements that may cause this same exception on other executions, along with the reference statements that may raise an exception in other executions because of this incorrect assignment; this information is required to repair the fault. The paper also presents an application of our technique to null pointer exceptions. Finally, the paper describes an implementation of the null-pointer-exception analysis and a set of studies that demonstrate the advantages of our approach for locating and repairing faults in the program.
Saurabh Sinha 0001, Hina Shah, Carsten Görg, Shujuan Jiang, Mijung Kim, Mary Jean Harrold
ISSTA5
2009 Reproducible gene selection algorithm with random effect model in cDNA microarray-based CGH data
Mijung Kim
Expert Syst. Appl.1
2009 Two-stage logistic regression model
Mijung Kim
Expert Syst. Appl.1