Ji Joong Moon

dblp:36/4610 · also Jijoong Moon · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0003-0888-2143ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 4 · 3 since 2021
YearPublicationVenuePosition
2025 Efficient Compositional Multi-tasking for On-device Large Language Models
abstract
Adapter parameters provide a mechanism to modify the behavior of machine learning models and have gained significant popularity in the context of large language models (LLMs) and generative AI.These parameters can be merged to support multiple tasks via a process known as task merging.However, prior work on merging in LLMs, particularly in natural language processing, has been limited to scenarios where each test example addresses only a single task.In this paper, we focus on on-device settings and study the problem of text-based compositional multi-tasking, where each test example involves the simultaneous execution of multiple tasks.For instance, generating a translated summary of a long text requires solving both translation and summarization tasks concurrently.To facilitate research in this setting, we propose a benchmark comprising four practically relevant compositional tasks.We also present an efficient method (Learnable Calibration) tailored for on-device applications, where computational resources are limited, emphasizing the need for solutions that are both resourceefficient and high-performing.Our contributions lay the groundwork for advancing the capabilities of LLMs in real-world multi-tasking scenarios, expanding their applicability to complex, resource-constrained use cases.Project page:
Ondrej Bohdal, Mete Ozay, Ji Joong Moon, Kyeng-Hun Lee, Hyeonmok Ko, Umberto Michieli
EMNLP3
2025 HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Merging
abstract
Large language models (LLMs) often leverage adapters, such as low-rank-based adapters, to achieve strong performance on downstream tasks.However, storing a separate adapter for each task significantly increases memory requirements, posing a challenge for resourceconstrained environments such as mobile devices.Although model merging techniques can reduce storage costs, they typically result in substantial performance degradation.In this work, we introduce HydraOpt, a new model merging technique that capitalizes on the inherent similarities between the matrices of low-rank adapters.Unlike existing methods that produce a fixed trade-off between storage size and performance, HydraOpt allows us to navigate this spectrum of efficiency and performance.Our experiments show that Hy-draOpt significantly reduces storage size (48% reduction) compared to storing all adapters, while achieving competitive performance (0.2-1.8% drop).Furthermore, it outperforms existing merging techniques in terms of performance at the same or slightly worse storage efficiency.
Taha Ceritli, Ondrej Bohdal, Mete Ozay, Ji Joong Moon, Kyeng-Hun Lee, Hyeonmok Ko, Umberto Michieli
EMNLP4
2025 Controllable Forgetting Mechanism for Few-Shot Class-Incremental Learning
abstract
Class-incremental learning in the context of limited personal labeled samples (few-shot) is critical for numerous real-world applications, such as smart home devices. A key challenge in these scenarios is balancing the trade-off between adapting to new, personalized classes and maintaining the performance of the model on the original, base classes. Fine-tuning the model on novel classes often leads to the phenomenon of catastrophic forgetting, where the accuracy of base classes declines unpredictably and significantly. In this paper, we propose a simple yet effective mechanism to address this challenge by controlling the trade-off between novel and base class accuracy. We specifically target the ultra-low-shot scenario, where only a single example is available per novel class. Our approach introduces a Novel Class Detection (NCD) rule, which adjusts the degree of forgetting a priori while simultaneously enhancing performance on novel classes. We demonstrate the versatility of our solution by applying it to state-of-the-art Few-Shot Class-Incremental Learning (FSCIL) methods, showing consistent improvements across different settings. To better quantify the trade-off between novel and base class performance, we introduce new metrics: NCR@2FOR and NCR@5FOR. Our approach achieves up to a 30% improvement in novel class accuracy on the CIFAR100 dataset (1-shot, 1 novel class) while maintaining a controlled base class forgetting rate of 2%.
Kirill Paramonov, Mete Ozay, Eunju Yang, Ji Joong Moon, Umberto Michieli
ICASSP4
2025 Continual Error Correction on Low-Resource Devices
abstract
The proliferation of AI models in everyday devices has highlighted a critical challenge: prediction errors that degrade user experience. While existing solutions focus on error detection, they rarely provide efficient correction mechanisms, especially for resource-constrained devices. We present a novel system enabling users to correct AI misclassifications through few-shot learning, requiring minimal computational resources and storage. Our approach combines server-side foundation model training with on-device prototype-based classification, enabling efficient error correction through prototype updates rather than model retraining. The system consists of two key components: (1) a server-side pipeline that leverages knowledge distillation to transfer robust feature representations from foundation models to device-compatible architectures, and (2) a device-side mechanism that enables ultra-efficient error correction through prototype adaptation. We demonstrate our system's effectiveness on both image classification and object detection tasks, achieving over 50% error correction in one-shot scenarios on Food-101 and Flowers-102 datasets while maintaining minimal forgetting (less than 0.02%) and negligible computational overhead. Our implementation, validated through an Android demonstration app, proves the system's practicality in real-world scenarios.
Kirill Paramonov, Mete Ozay, Aristeidis Mystakidis, Nikolaos Tsalikidis, Dimitrios Sotos, Anastasios Drosou, Dimitrios Tzovaras, Kiseok Chang, Sangdok Mo, Namwoong Kim, Woojong Yoo, Ji Joong Moon, Umberto Michieli
MMSys13
2024 Deep Neural Network Models Trained with a Fixed Random Classifier Transfer Better Across Domains
abstract
The recently discovered Neural collapse (NC) phenomenon states that the last-layer weights of Deep Neural Networks (DNN), converge to the so-called Equiangular Tight Frame (ETF) simplex, at the terminal phase of their training. This ETF geometry is equivalent to vanishing within-class variability of the last layer activations. Inspired by NC properties, we explore in this paper the transferability of DNN models trained with their last layer weight fixed according to ETF. This enforces class separation by eliminating class covariance information, effectively providing implicit regularization. We show that DNN models trained with such a fixed classifier significantly improve transfer performance, particularly on out-of-domain datasets. On a broad range of fine-grained image classification datasets, our approach outperforms i) baseline methods that do not perform any covariance regularization (up to 22%), as well as ii) methods that explicitly whiten covariance of activations throughout training (up to 19%). Our findings suggest that DNNs trained with fixed ETF classifiers offer a powerful mechanism for improving transfer learning across domains.
Hafiz Tiomoko Ali, Umberto Michieli, Ji Joong Moon, Mete Ozay
ICASSP3
2024 FFT-Based Selection and Optimization of Statistics for Robust Recognition of Severely Corrupted Images
abstract
Improving model robustness in case of corrupted images is among the key challenges to enable robust vision systems on smart devices, such as robotic agents. Particularly, robust test-time performance is imperative for most of the applications. This paper presents a novel approach to improve robustness of any classification model, especially on severely corrupted images. Our method (FROST) employs high-frequency features to detect input image corruption type, and select layer-wise feature normalization statistics. FROST provides the state-of-the-art results for different models and datasets, outperforming competitors on ImageNet-C by up to 37.1% relative gain, improving baseline of 40.9% mCE on severe corruptions.
Elena Camuffo, Umberto Michieli, Ji Joong Moon, Mete Ozay
ICASSP3
2024 Object-Conditioned Bag of Instances for Few-Shot Personalized Instance Recognition
abstract
Nowadays, users demand for increased personalization of vision systems to localize and identify personal instances of objects (e.g., my dog rather than dog) from a few-shot dataset only. Despite outstanding results of deep networks on classical label-abundant benchmarks (e.g., those of the latest YOLOv8 model for standard object detection), they struggle to maintain within-class variability to represent different instances rather than object categories only. We construct an Object-conditioned Bag of Instances (OBoI) based on multiorder statistics of extracted features, where generic object detection models are extended to search and identify personal instances from the OBoI’s metric space, without need for backpropagation. By relying on multi-order statistics, OBoI achieves consistent superior accuracy in distinguishing different instances. In the results, we achieve 77.1% personal object recognition accuracy in case of 18 personal instances, showing about 12% relative gain over the state of the art.
Umberto Michieli, Ji Joong Moon, Mete Ozay
ICASSP2
2024 Cross-Architecture Auxiliary Feature Space Translation for Efficient Few-Shot Personalized Object Detection
abstract
Recent years have seen object detection robotic systems deployed in several personal devices (e.g., home robots and appliances). This has highlighted a challenge in their design, i.e., they cannot efficiently update their knowledge to distinguish between general classes and user-specific instances (e.g., a dog vs. user’s dog). We refer to this challenging task as Instance-level Personalized Object Detection (IPOD). The personalization task requires many samples for model tuning and optimization in a centralized server, raising privacy concerns. An alternative is provided by approaches based on recent large-scale Foundation Models, but their compute costs preclude on-device applications. In our work we tackle both problems at the same time, designing a Few-Shot IPOD strategy called AuXFT. We introduce a conditional coarse-to-fine few-shot learner to refine the coarse predictions made by an efficient object detector, showing that using an off-the-shelf model leads to poor personalization due to neural collapse. Therefore, we introduce a Translator block that generates an auxiliary feature space where features generated by a self-supervised model (e.g., DINOv2) are distilled without impacting the performance of the detector. We validate AuXFT on three publicly available datasets and one in-house benchmark designed for the IPOD task, achieving remarkable gains in all considered scenarios with excellent time-complexity trade-off: AuXFT reaches a performance of 80% its upper bound at just 32% of the inference time, 13% of VRAM and 19% of the model size.
Francesco Barbato, Umberto Michieli, Ji Joong Moon, Pietro Zanuttigh, Mete Ozay
IROS3
2024 Enhanced Model Robustness to Input Corruptions by Per-corruption Adaptation of Normalization Statistics
abstract
Developing a reliable vision system is a fundamental challenge for robotic technologies (e.g., indoor service robots and outdoor autonomous robots) which can ensure reliable navigation even in challenging environments such as adverse weather conditions (e.g., fog, rain), poor lighting conditions (e.g., over/under exposure), or sensor degradation (e.g., blurring, noise), and can guarantee high performance in safety-critical functions. Current solutions proposed to improve model robustness usually rely on generic data augmentation techniques or employ costly test-time adaptation methods. In addition, most approaches focus on addressing a single vision task (typically, image recognition) utilising synthetic data. In this paper, we introduce Per-corruption Adaptation of Normalization statistics (PAN) to enhance the model robustness of vision systems. Our approach entails three key components: (i) a corruption type identification module, (ii) dynamic adjustment of normalization layer statistics based on identified corruption type, and (iii) real-time update of these statistics according to input data. PAN can integrate seamlessly with any convolutional model for enhanced accuracy in several robot vision tasks. In our experiments, PAN obtains robust performance improvement on challenging real-world corrupted image datasets (e.g., OpenLoris, ExDark, ACDC), where most of the current solutions tend to fail. Moreover, PAN outperforms the baseline models by 20-30% on synthetic benchmarks in object recognition tasks.
Elena Camuffo, Umberto Michieli, Simone Milani, Ji Joong Moon, Mete Ozay
IROS4
2024 Swiss DINO: Efficient and Versatile Vision Framework for On-device Personal Object Search
abstract
In this paper, we address a recent trend in robotic home appliances to include vision systems on personal devices, capable of personalizing the appliances on the fly. In particular, we formulate and address an important technical task of personal object search, which involves localization and identification of personal items of interest on images captured by robotic appliances, with each item referenced only by a few annotated images. The task is crucial for robotic home appliances and mobile systems, which need to process personal visual scenes or to operate with particular personal objects (e.g., for grasping or navigation). In practice, personal object search presents two main technical challenges. First, a robot vision system needs to be able to distinguish between many fine-grained classes, in the presence of occlusions and clutter. Second, the strict resource requirements for the on-device system restrict the usage of most state-of-the-art methods for few-shot learning and often prevent on-device adaptation. In this work, we propose Swiss DINO: a simple yet effective framework for one-shot personal object search based on the recent DINOv2 transformer model, which was shown to have strong zero-shot generalization properties. Swiss DINO handles challenging on-device personalized scene understanding requirements and does not require any adaptation training. We show significant improvement (up to 55%) in segmentation and recognition accuracy compared to the common lightweight solutions, and significant footprint reduction of backbone inference time (up to 100×) and GPU consumption (up to 10×) compared to the heavy transformer-based solutions1.
Kirill Paramonov, Jia-Xing Zhong, Umberto Michieli, Ji Joong Moon, Mete Ozay
IROS4
2004 Optimal Blade System Design of a New Concept VTOL Vehicle Using the Departmental Computing Grid System
abstract
The blade system of a new concept VTOL vehicle is designed utilizing high performance and Grid computing technologies. The VTOL vehicle called cyclocopter employs a cycloidal propulsion system to generate the propulsion and lift for VTOL maneuver. The structural design and weight minimization of the composite blade system are critically related to the efficiency of whole cyclocopter system. The structural design is carried out using a hybrid genetic algorithm-based optimization framework on the Departmental Computing Grid(DCG) system, an aggregation of cluster resources installed in the Aerospace department of Seoul National University. High-fidelity simulation is conducted using our parallel finite element code (IPSAP) which employs the domain-wise parallel multifrontal solver, a direct solution method characterized by the solution robustness and the predictability of running time. The optimization results and computational aspects are displayed emphasizing the potential of the high performance Grid computing technology utilized for high-fidelity simulation based aerospace system design.
Jin Woo Park, Si Hyoung Park, In Seong Hwang, Ji Joong Moon, Youngha Yoon, Seung Jo Kim
SC4