Jaehyung Lee 0002

dblp:12/2962-2 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0002-1951-9625ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 100%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution
bug triage
0.612022
A Light Bug Triage Framework for Applying Large Pre-trained Language Model · ASE 2022
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor completion
0.312018
Fast Tucker Factorization for Large-Scale Tensor Completion · ICDM 2018
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor factorization
0.312018
Fast Tucker Factorization for Large-Scale Tensor Completion · ICDM 2018
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis › tensor factorization
tucker decomposition
0.312018
Fast Tucker Factorization for Large-Scale Tensor Completion · ICDM 2018
Natural language and speech › Language models and text generation › pre-trained language model
BERT
0.212022
A Light Bug Triage Framework for Applying Large Pre-trained Language Model · ASE 2022
Natural language and speech › Language models and text generation
pre-trained language model
0.212022
A Light Bug Triage Framework for Applying Large Pre-trained Language Model · ASE 2022

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 1.1fine-tuning · 1.1coordinate descent · 0.3caching · 0.3
YearPublicationVenuePosition
2022 A Light Bug Triage Framework for Applying Large Pre-trained Language Model
abstract
Assigning appropriate developers to the bugs is one of the main challenges in bug triage. Demands for automatic bug triage are increasing in the industry, as manual bug triage is labor-intensive and time-consuming in large projects. The key to the bug triage task is extracting semantic information from a bug report. In recent years, large Pre-trained Language Models (PLMs) including BERT [4] have achieved dramatic progress in the natural language processing (NLP) domain. However, applying large PLMs to the bug triage task for extracting semantic information has several challenges. In this paper, we address the challenges and propose a novel framework for bug triage named LBT-P, standing for Light Bug Triage framework with a Pre-trained language model. It compresses a large PLM into small and fast models using knowledge distillation techniques and also prevents catastrophic forgetting of PLM by introducing knowledge preservation fine-tuning. We also develop a new loss function exploiting representations of earlier layers as well as deeper layers in order to handle the overthinking problem. We demonstrate our proposed framework on the real-world private dataset and three public real-world datasets [11]: Google Chromium, Mozilla Core, and Mozilla Firefox. The result of the experiments shows the superiority of LBT-P.
Jaehyung Lee 0002, Kisun Han, Hwanjo Yu
ASE1
2018 Fast Tucker Factorization for Large-Scale Tensor Completion
abstract
Tensor completion is the task of completing multi-aspect data represented as a tensor by accurately predicting missing entries in the tensor. It is mainly solved by tensor factorization methods, and among them, Tucker factorization has attracted considerable interests due to its powerful ability to learn latent factors and even their interactions. Although several Tucker methods have been developed to reduce the memory and computational complexity, the state-of-the-art method still 1) generates redundant computations and 2) cannot factorize a large tensor that exceeds the size of memory. This paper proposes FTcom, a fast and scalable Tucker factorization method for tensor completion. FTcom performs element-wise updates for factor matrices based on coordinate descent, and adopts a novel caching algorithm which stores frequently-required intermediate data. It also uses a tensor file for disk-based data processing and loads only a small part of the tensor at a time into the memory. Experimental results show that FTcom is much faster and more scalable compared to all other competitors. It significantly shortens the training time of Tucker factorization, especially on real-world tensors, and it can be executed on a billion-scale tensor which is bigger than the memory capacity within a single machine.
Dongha Lee 0003, Jaehyung Lee 0002, Hwanjo Yu
ICDM2