VLDB 2026 Research / reviewers in the wild / expert
Taejin Park
dblp:08/5631
· DBLP profile ↗
26ranked-venue papers
12as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 10 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Free Few-Shot Segmentation via Vision-Language Guided PromptingabstractObject segmentation relies heavily on costly pixel-level annotations and struggles to generalize to unseen domains. The recent introduction of the Segment Anything Model (SAM), a foundation model for segmentation, offers a prompt-driven, zero-shot capability that has been applied in various domains (e.g., autonomous driving, satellite imagery, medical imaging) and extended to Few-Shot Segmentation (FSS) tasks. However, existing SAM-based FSS methods typically generate prompts by using a vision encoder to measure support–query image similarity, which often biases towards the support images and fails when there are significant support–query context shifts. To address this limitation, we propose a training-free FSS approach that combines visual and textual cues to generate effective prompts for the target class. By leveraging both vision and language information, our approach bridges the support–query gap and guides SAM to segment novel objects more reliably. Without any additional training, our method outperforms previous state-of-the-art FSS methods on established benchmarks (COCO-20i, Pascal-5i), demonstrating its effectiveness and robust generalization. Our code is publicly available on GitHub. Euihyun Yoon, Taejin Park, Jaekoo Lee |
WACV | 2 |
| 2026 | Recent trends in distant conversational speech recognition: A review of CHiME-7 and 8 DASR challenges
Samuele Cornell, Christoph Böddeker, Taejin Park, He Huang 0012, Desh Raj, Matthew Wiesner, Yoshiki Masuyama, Xuankai Chang, Zhongqiu Wang 0001, Stefano Squartini, L. Paola García-Perera, Shinji Watanabe 0001 |
Comput. Speech Lang. | 3 |
| 2025 | CSRM-LLM: Embracing Multilingual LLMs for Cold-Start Relevance Matching in Emerging E-commerce MarketsabstractAs global e-commerce platforms continue to expand, companies are entering new markets where they encounter cold-start challenges due to limited human labels and user behaviors. In this paper, we share our experiences in Coupang to provide a competitive cold-start performance of relevance matching for emerging e-commerce markets. Specifically, we present a Cold-Start Relevance Matching (CSRM) framework, utilizing a multilingual Large Language Model (LLM) to address three challenges: (1) activating cross-lingual transfer learning abilities of LLMs through machine translation tasks; (2) enhancing query understanding and incorporating e-commerce knowledge by retrieval-based query augmentation; (3) mitigating the impact of training label errors through a multi-round self-distillation training strategy. Our experiments demonstrate the effectiveness of CSRM-LLM and the proposed techniques, resulting in successful real-world deployment and significant online gains, with a 45.8% reduction in defect ratio and a 0.866% uplift in session purchase rate. Yujing Wang 0002, Huoran Li, Chunxu Xu, Yuchong Luo, Xianghui Mao, Cong Li 0021, Lun Du, Chunyang Ma, Qiqi Jiang, Wenting Mo, Pei Wen, Shantanu Kumar, Taejin Park, Yiwei Song, Vijay Rajaram, Sonu Durgia, Pranam Kolari |
CIKM | 16 |
| 2025 | NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing TasksabstractSelf-supervised learning (SSL) has been proved to benefit a wide range of speech processing tasks, such as speech recognition/translation, speaker verification and diarization, etc. However, most of current speech SSL approaches are computationally expensive. In this paper, we introduce a simplified and more efficient SSL framework, termed as NeMo Encoder for Speech Tasks (NEST). Specifically, we adopt the FastConformer architecture with 8x sub-sampling rate, which is faster than Transformer or Conformer architectures. Instead of clusteringbased quantization, we use fixed random projection for its simplicity and effectiveness. We also implement a generalized noisy speech augmentation that teaches the model to disentangle the main speaker from noise or other speakers. Experiments show that NEST improves over existing self-supervised models and achieves new state-of-the-art performance on a variety of speech processing tasks, such as speech recognition/translation, speaker diarization, spoken language understanding, etc. Code and checkpoints are publicly available via NVIDIA NeMo framework123. He Huang 0012, Taejin Park, Kunal Dhawan, Ivan Medennikov, Krishna C. Puvvada, Nithin Rao Koluguri, Jagadeesh Balam, Boris Ginsburg |
ICASSP | 2 |
| 2025 | META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASRabstractWe propose a novel end-to-end multi-talker automatic speech recognition (ASR) framework that enables both multi-speaker (MS) ASR and target-speaker (TS) ASR. Our proposed model is trained in a fully end-to-end manner, incorporating speaker supervision from a pre-trained speaker diarization module. We introduce an intuitive yet effective method for masking ASR encoder activations using output from the speaker supervision module, a technique we term Meta-Cat (meta-information concatenation), that can be applied to both MS-ASR and TS-ASR. Our results demonstrate that the proposed architecture achieves competitive performance in both MS-ASR and TS-ASR tasks, without the need for traditional methods, such as neural mask estimation or masking at the audio or feature level. Furthermore, we demonstrate a glimpse of a unified dual-task model which can efficiently handle both MS-ASR and TS-ASR tasks. Thus, this work illustrates that a robust end-to-end multi-talker ASR framework can be implemented with a streamlined architecture, obviating the need for the complex speaker filtering mechanisms employed in previous studies. Jinhan Wang, Kunal Dhawan, Taejin Park, Myungjong Kim, Ivan Medennikov, He Huang 0012, Nithin Rao Koluguri, Jagadeesh Balam, Boris Ginsburg |
ICASSP | 4 |
| 2025 | Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text SystemsabstractSortformer is an encoder-based speaker diarization model designed for supervising speaker tagging in speech-to-text models. Instead of relying solely on permutation invariant loss (PIL), Sortformer introduces Sort Loss to resolve the permutation problem, either independently or in tandem with PIL. In addition, we propose a streamlined multi-speaker speech-to-text architecture that leverages Sortformer for speaker supervision, embedding speaker labels into the encoder using sinusoidal kernel functions. This design addresses the speaker permutation problem through sorted objectives, effectively bridging timestamps and tokens to supervise speaker labels in the output transcriptions. Experiments demonstrate that Sort Loss can boost speaker diarization performance, and incorporating the speaker supervision from Sortformer improves multi-speaker transcription accuracy. We anticipate that the proposed Sortformer and multi-speaker architecture will enable the seamless integration of speaker tagging capabilities into foundational speech-to-text systems and multimodal large language models (LLMs), offering an easily adoptable and user-friendly mechanism to enhance their versatility and performance in speaker-aware tasks. The code and trained models are made publicly available through the NVIDIA NeMo Framework. Taejin Park, Ivan Medennikov, Kunal Dhawan, He Huang 0012, Nithin Rao Koluguri, Krishna C. Puvvada, Jagadeesh Balam, Boris Ginsburg |
ICML | 1 |
| 2025 | SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
Raymond Grossman, Taejin Park, Kunal Dhawan, Andrew Titus, Sophia Zhi, Yulia Shchadilova, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 2 |
| 2025 | Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
Ivan Medennikov, Taejin Park, He Huang 0012, Kunal Dhawan, Jinhan Wang, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 2 |
| 2025 | Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
Taejin Park, Ivan Medennikov, Jinhan Wang, Kunal Dhawan, He Huang 0012, Nithin Rao Koluguri, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 2 |
| 2024 | Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search ApproachabstractLarge language models (LLMs) have shown great promise for capturing contextual information in natural language processing tasks. We propose a novel approach to speaker diarization that incorporates the prowess of LLMs to exploit contextual cues in human dialogues. Our method builds upon an acoustic-based speaker diarization system by adding lexical information from an LLM in the inference stage. We model the multi-modal decoding process probabilistically and perform joint acoustic and lexical beam searches to incorporate cues from both modalities: audio and text. Our experiments demonstrate that infusing lexical knowledge from the LLM into an acoustics-only diarization system improves the overall speaker-attributed word error rate (SA-WER). The experimental results show that LLMs can provide complementary information to acoustic models for the speaker diarization task via the proposed beam search decoding approach showing up to 39.8% relative delta-SA-WER improvement from the baseline system. Thus, we substantiate that the proposed technique is able to exploit contextual information that is inaccessible to acoustics-only systems which is represented by speaker embeddings. In addition, these findings point to the potential of using LLMs to improve speaker diarization and other speech-processing tasks by capturing semantic and contextual cues. Taejin Park, Kunal Dhawan, Nithin Rao Koluguri, Jagadeesh Balam |
ICASSP | 1 |
| 2024 | Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASRabstractSpeech foundation models have achieved state-of-the-art (SoTA) performance across various tasks, such as automatic speech recognition (ASR) in hundreds of languages. However, multi-speaker ASR remains a challenging task for these models due to data scarcity and sparsity. In this paper, we present approaches to enable speech foundation models to process and understand multi-speaker speech with limited training data. Specifically, we adapt a speech foundation model for the multi-speaker ASR task using only telephonic data. Remarkably, the adapted model also performs well on meeting data without any fine-tuning, demonstrating the generalization ability of our approach. We conduct several ablation studies to analyze the impact of different parameters and strategies on model performance. Our findings highlight the effectiveness of our methods. Results show that less parameters give better overall cpWER, which, although counterintuitive, provides insights into adapting speech foundation models for multi-speaker ASR tasks with minimal annotated data. Kunal Dhawan, Taejin Park, Krishna C. Puvvada, Ivan Medennikov, Somshubra Majumdar, He Huang 0012, Jagadeesh Balam, Boris Ginsburg |
SLT | 3 |
| 2024 | Large Language Model Based Generative Error Correction: A Challenge and Baselines For Speech Recognition, Speaker Tagging, and Emotion RecognitionabstractGiven recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen, pretrained automatic speech recognition (ASR) model. To explore new capabilities in language modeling for speech processing, we introduce the generative speech transcription error correction (GenSEC) challenge. This challenge comprises three post-ASR language modeling tasks: (i) post-ASR transcription correction, (ii) speaker tagging, and (iii) emotion recognition. These tasks aim to emulate future LLM-based agents handling voice-based interfaces while remaining accessible to a broad audience by utilizing open pretrained language models or agent-based APIs. We also discuss insights from baseline evaluations, as well as lessons learned for designing future evaluations. Chao-Han Huck Yang, Taejin Park, Yuan Gong 0001, Yuanchao Li, Zhehuai Chen, Chen Chen 0075, Kunal Dhawan, Piotr Zelasko, Chao Zhang 0031, Yun-Nung Chen, Yu Tsao 0001, Jagadeesh Balam, Boris Ginsburg, Sabato Marco Siniscalchi, Chng Eng Siong, Peter Bell 0001, Catherine Lai, Shinji Watanabe 0001, Andreas Stolcke |
SLT | 2 |
| 2023 | Improving the Quality of MODIS LAI Products by Exploiting Spatiotemporal Correlation InformationabstractThe Moderate Resolution Imaging Spectroradiometer (MODIS) Leaf Area Index (LAI) product is critical for global terrestrial carbon monitoring and ecosystem modeling. However, MODIS LAI is calculated on a pixel-by-pixel and day-by-day basis without using spatial or temporal correlation information, which leads to its high sensitivity of LAI to uncertainties in observed reflectance resulting in an increased noise level in time series. While exploiting prior knowledge is a common practice to fill gaps in observations, little research has been conducted on reducing noisy fluctuations and improving the overall quality of the MODIS LAI product. To address this issue, we proposed a Spatio-Temporal Information Composition Algorithm (STICA), which directly introduces prior Spatio-temporal correlation and Multiple Quality Assessment (MQA) information into the existing MODIS LAI product. STICA reduces the noise level and improves the quality of the product while maintaining the original physically-based (Radiative Transfer Model, RTM) LAI production process. In our analysis, the R2 increased from 0.79 to 0.81, and the RMSE decreased from 0.81 to 0.68 compared to the ground-based LAI reference. The improvement was more pronounced with the degradation of the data quality. STICA reduced noisy fluctuations in the LAI time series to varying degrees among eight biome types. In the Amazon Forest, STICA significantly improved the time-series stability of LAI. Moreover, STICA can effectively eliminate abnormal declines in time series and correct for extreme outliers in LAI. We expect that the MODIS LAI Reanalyzed product generated by this method will better support the application of high-quality LAI datasets. Kai Yan 0001, Jiabin Pu, Jinxiu Liu, Taejin Park, Jian Bi, Eduardo Eiji Maeda, Janne Heiskanen, Yuri Knyazikhin, Ranga B. Myneni |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | TitaNet: Neural Model for Speaker Representation with 1D Depth-Wise Separable Convolutions and Global ContextabstractIn this paper, we propose TitaNet, a novel neural network architecture for extracting speaker representations. We employ 1D depth-wise separable convolutions with Squeeze-and-Excitation (SE) layers with global context followed by channel attention based statistics pooling layer to map variable-length utterances to a fixed-length embedding (t-vector). TitaNet is a scalable architecture and achieves state-of-the-art performance on speaker verification task with an equal error rate (EER) of 0.68% on the VoxCeleb1 trial file and also on speaker diarization tasks with diarization error rate (DER) of 1.73% on AMI-MixHeadset, 1.99% on AMI-Lapel and 1.11% on CH109. Furthermore, we investigate various sizes of TitaNet and present a light TitaNet-S model with only 6M parameters that achieve near state-of-the-art results in diarization tasks. Nithin Rao Koluguri, Taejin Park, Boris Ginsburg |
ICASSP | 2 |
| 2022 | Multi-scale Speaker Diarization with Dynamic Scale Weighting
Taejin Park, Nithin Rao Koluguri, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 1 |
| 2022 | NeMo Open Source Speaker Diarization System
Taejin Park, Nithin Rao Koluguri, Fei Jia, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 1 |
| 2020 | Robust Multi-Channel Speech Recognition Using Frequency Aligned NetworkabstractConventional speech enhancement technique such as beamforming has known benefits for far-field speech recognition. Our own work in frequency-domain multi-channel acoustic modeling has shown additional improvements by training a spatial filtering layer jointly within an acoustic model. In this paper, we further develop this idea and use frequency aligned network for robust multi-channel automatic speech recognition (ASR). Unlike an affine layer in the frequency domain, the proposed frequency aligned component prevents one frequency bin influencing other frequency bins. We show that this modification not only reduces the number of parameters in the model but also significantly and improves the ASR performance. We investigate effects of frequency aligned network through ASR experiments on the real-world far-field data where users are interacting with an ASR system in uncontrolled acoustic environments. We show that our multi-channel acoustic model with a frequency aligned network shows up to 18% relative reduction in word error rate. Taejin Park, Ken'ichi Kumatani, Minhua Wu, Shiva Sundaram |
ICASSP | 1 |
| 2018 | Generating Global Products of LAI and FPAR From SNPP-VIIRS Data: Theoretical Background and ImplementationabstractLeaf area index (LAI) and fraction of photosynthetically active radiation (FPAR) absorbed by vegetation have been successfully generated from the Moderate Resolution Imaging Spectroradiometer (MODIS) data since early 2000. As the Visible Infrared Imaging Radiometer Suite (VIIRS) instrument onboard, the Suomi National Polar-orbiting Partnership (SNPP) has inherited the scientific role of MODIS, and the development of a continuous, consistent, and well-characterized VIIRS LAI/FPAR data set is critical to continue the MODIS time series. In this paper, we build the radiative transfer-based VIIRS-specific lookup tables by achieving minimal difference with the MODIS data set and maximal spatial coverage of retrievals from the main algorithm. The theory of spectral invariants provides the configurable physical parameters, i.e., single scattering albedos (SSAs) that are optimized for VIIRS-specific characteristics. The effort finds a set of smaller red-band SSA and larger near-infrared-band SSA for VIIRS compared with the MODIS heritage. The VIIRS LAI/FPAR is evaluated through comparisons with one year of MODIS product in terms of both spatial and temporal patterns. Further validation efforts are still necessary to ensure the product quality. Current results, however, imbue confidence in the VIIRS data set and suggest that the efforts described here meet the goal of achieving the operationally consistent multisensor LAI/FPAR data sets. Moreover, the strategies of parametric adjustment and LAI/FPAR evaluation applied to SNPP-VIIRS can also be employed to the subsequent Joint Polar Satellite System VIIRS or other instruments. Kai Yan 0001, Taejin Park, Chi Chen 0004, Baodong Xu, Wanjuan Song, Bin Yang 0008, Yelu Zeng, Guangjian Yan, Yuri Knyazikhin, Ranga B. Myneni |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | A Comparative Study of Predicting DBH and Stem Volume of Individual Trees in a Temperate Forest Using Airborne Waveform LiDARabstractUsing airborne full-waveform LiDAR metrics derived by 3-D tree segmentation, this study estimated single tree's diameter at breast height (DBH) and stem volume (STV). Four regression models were used, including multilinear regression and three up-to-date regression models (i.e., least square boosting trees regression, random forest, and ε-support vector regression) from the machine learning field. This study aimed to comparatively evaluate these regression models in predicting DBH and STV at single-tree level and find some clues to regression model's selection. The study sites were located in the Bavarian Forest National Park, Germany, a mixed temperate mountain forest. Our comparisons were performed across different tree species types (coniferous and deciduous) and foliage conditions (leaf-on/leaf-off seasons). The importance of predictor variables was also examined. Experimental results revealed that the best accuracy from machine learning methods outperformed the multilinear model by 1.5 cm for DBH and 0.18 m3for STV in terms of rmse. Through comparative analysis, our work provided some clues to the performance variation of regression models for extracting 3-D tree parameters. Wei Yao 0008, Sungho Choi, Taejin Park, Ranga B. Myneni |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2012 | The Effect of Perceptual Complexity on Affective Picture Processing
Taejin Park, Soodam Park |
CogSci | 1 |
| 2011 | Accumulative sampling for noisy evolutionary multi-objective optimizationabstractObjective evaluation is subject to noise in many real-world problems. The noise can deteriorate the performance of multi-objective evolutionary algorithms, by misleading the population to a local optimum and reducing the convergence rate. This paper proposes three novel noise handling techniques: accumulative sampling, a new ranking method, and a different selection scheme for recombination. The accumulative sampling is basically a kind of dynamic resampling, but it does not explicitly decide the number of samples. Instead, it repeatedly takes additional samples of objectives for the solutions in the archive at every generation, and updates the estimated objectives using all the accumulated samples. The new ranking method combines probabilistic Pareto rank and crowding distance into a single aggregated value to promote the diversity in the archive. Finally, the fitness function and selection method used for recombination are made different from those for the archive to accelerate the convergence rate. Experiments on various benchmark problems have shown that the algorithm adopting all these features performs better than other MOEAs in various performance metrics. Taejin Park, Kwang Ryel Ryu |
GECCO | 1 |
| 2010 | A Dual-Population Genetic Algorithm for Adaptive Diversity ControlabstractA variety of previous works exist on maintaining population diversity of genetic algorithms (GAs). Dual-population GA (DPGA) is a type of multipopulation GA (MPGA) that uses an additional population as a reservoir of diversity. The main population is similar to that of an ordinary GA and evolves to find good solutions. The reserve population evolves to maintain and provide diversity to the main population. While most MPGAs use migration as a means of information exchange between different populations, DPGA uses crossbreeding because the two populations have entirely different fitness functions. The reserve population cannot provide useful diversity to the main population unless the two maintain an appropriate distance. Therefore, DPGA adjusts the distance dynamically to achieve an appropriate balance between exploration and exploitation. The experimental results on various classes of problems using binary, real-valued, and order-based representations show that DPGA quite often outperforms not only the standard GAs but also other GAs having additional mechanisms of diversity preservation. Taejin Park, Kwang Ryel Ryu |
IEEE Trans. Evol. Comput. | 1 |
| 2008 | Dual-population genetic algorithm for nonstationary optimizationabstractIn order to solve nonstationary optimization problems efficiently, evolutionary algorithms need sufficient diversity to adapt to environmental changes. The dual-population genetic algorithm (DPGA) is a novel evolutionary algorithm that uses an extra population called the reserve population to provide additional diversity to the main population through crossbreeding. Preliminary experimental results on various periods and degrees of environmental change have shown that the distance between the two populations of DPGA is one of the most important factors that affect its per-formance. However, it is very difficult to determine the best popu-lation distance without prior knowledge about the given problem. This paper proposes a new DPGA that uses two reserve populations (DPGA2). The reserve populations are at different distances from the main population. The information inflow from the reserve populations is controlled by survival selection. Experimental results show that DPGA2 shows a better performance than other evolutionary algorithms for nonstationary optimization problems without relying on prior knowledge about the problem. Taejin Park, Ri Choe, Kwang Ryel Ryu |
GECCO | 1 |
| 2007 | A dual population genetic algorithm with evolving diversityabstractWe propose a dual population genetic algorithm inspired by the complementary and dominance mechanism prevalent in nature. The proposed algorithm has two distinct populations: a main population and a reserve population. The main population is similar to that of an ordinary genetic algorithm and evolves to find good solutions. The reserve population evolves to maintain and offer diversity to the main population. While most multi-population genetic algorithms use migration as a means of information ex-change between different populations, our algorithm uses crossbreeding and survival selection because the two populations have different evolutionary objectives. The experimental results on various multimodal optimization problems show that the proposed algorithm is better than not only ordinary genetic algorithms but also than the other algorithms based on similar idea. Taejin Park, Kwang Ryel Ryu |
IEEE Congress on Evolutionary Computation | 1 |
| 2006 | Solving a Large-Scaled Crew Pairing Problem by Using a Genetic Algorithm
Taejin Park, Kwang Ryel Ryu |
IEA/AIE | 1 |
| 2004 | Exploiting Unexpressed Genes for Solving Large-Scaled Maximal Covering Problems
Taejin Park, Kwang Ryel Ryu |
PRICAI | 1 |