EDBT 2026 Demo / reviewers in the wild / expert
Samuel Kim
dblp:68/5004
· DBLP profile ↗
32ranked-venue papers
18as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 15 first-author · 1 since 2021Artificial intelligence and machine learning · 16 · 7 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 64% Multimedia analysis and retrieval · 36% | |
| Artificial intelligence
3 papers |
Planning, search and constraint satisfaction · 44% Robot manipulation · 34% Information extraction and text analysis · 16% | |
| Human-computer interaction and pervasive computing
2 papers |
Accessibility and assistive technology · 68% Collaborative and social computing · 32% |
Topics — the 12 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Visualization and visual analytics › visualization literacy › visualization interpretation
chart understanding |
1.0 | 1 | 2026 | VisAnatomy: An SVG Chart Corpus with Fine-Grained Semantic Labels · IEEE Trans. Vis. Comput. Graph. 2026 |
Multimedia analysis and retrieval › multimedia analysis › multimedia content description
semantic annotation |
1.0 | 1 | 2026 | VisAnatomy: An SVG Chart Corpus with Fine-Grained Semantic Labels · IEEE Trans. Vis. Comput. Graph. 2026 |
Robotics › Robot manipulation › manipulation control
collision handling |
0.9 | 1 | 2025 | Work Smarter Not Harder: Simple Imitation Learning with CS-PIBT Outperforms Large-Scale Imitation Learning for MAPF · ICRA 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
multi-agent path finding |
0.9 | 1 | 2025 | Work Smarter Not Harder: Simple Imitation Learning with CS-PIBT Outperforms Large-Scale Imitation Learning for MAPF · ICRA 2025 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.3 | 1 | 2018 | Syntactical Analysis of the Weaknesses of Sentiment Analyzers · EMNLP 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search |
0.3 | 1 | 2025 | Work Smarter Not Harder: Simple Imitation Learning with CS-PIBT Outperforms Large-Scale Imitation Learning for MAPF · ICRA 2025 |
Multimedia analysis and retrieval › multimedia analysis
multimodal behavior analysis |
0.1 | 1 | 2012 | Predicting the conflict level in television political debates: an approach based on crowdsourcing, nonverbal communication and gaussian processes · ACM Multimedia 2012 |
Collaborative and social computing
crowdsourcing |
0.1 | 1 | 2012 | Predicting the conflict level in television political debates: an approach based on crowdsourcing, nonverbal communication and gaussian processes · ACM Multimedia 2012 |
Natural language and speech › Speech recognition and synthesis › speaker diarization
speaker clustering |
0.1 | 1 | 2008 | Strategies to Improve the Robustness of Agglomerative Hierarchical Clustering Under Data Source Variation for Speaker Diarization · IEEE Trans. Speech Audio Process. 2008 |
Natural language and speech › Speech recognition and synthesis
speaker diarization |
0.1 | 1 | 2008 | Strategies to Improve the Robustness of Agglomerative Hierarchical Clustering Under Data Source Variation for Speaker Diarization · IEEE Trans. Speech Audio Process. 2008 |
Query processing and optimization › OLAP
OLAP query processing |
0.0 | 1 | 2001 | The MD-join: An Operator for Complex OLAP · ICDE 2001 |
Data models and query languages
query algebra |
0.0 | 1 | 2001 | The MD-join: An Operator for Complex OLAP · ICDE 2001 |
Methods — techniques the papers use, named apart from their topics
semantic annotation · 2.0imitation learning · 0.9collision shielding · 0.9syntactic analysis · 0.3nonverbal behavioral cue extraction · 0.3gaussian process regression · 0.3crowdsourcing · 0.3selective AHC · 0.1information change rate · 0.1generalized likelihood ratio · 0.1bayesian information criterion · 0.1agglomerative hierarchical clustering · 0.1query optimization · 0.0algebraic transformations · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VisAnatomy: An SVG Chart Corpus with Fine-Grained Semantic LabelsabstractChart corpora, which comprise data visualizations and their semantic labels, are crucial for advancing visualization research. However, the labels in most existing corpora are high-level (e.g., chart types), hindering their utility for broader applications in the era of AI. In this paper, we contribute VisAnatomy, a corpus containing 942 real-world SVG charts produced by over 50 tools, encompassing 40 chart types and featuring structural and stylistic design variations. Each chart is augmented with multi-level fine-grained labels on its semantic components, including each graphical element's type, role, and position, hierarchical groupings of elements, group layouts, and visual encodings. In total, VisAnatomy provides labels for more than 383k graphical elements. We demonstrate the richness of the semantic labels by comparing VisAnatomy with existing corpora. We illustrate its usefulness through four applications: semantic role inference for SVG elements, chart semantic decomposition, chart type classification, and content navigation for accessibility. Finally, we discuss research opportunities to further improve VisAnatomy. Chen Chen 0080, Hannah K. Bako, Peihong Yu, John Hooker, Jeffrey Joyal, Simon C. Wang, Samuel Kim, Jessica Wu, Aoxue Ding, Lara Sandeep, Alex Chen, Chayanika Sinha, Zhicheng Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Work Smarter Not Harder: Simple Imitation Learning with CS-PIBT Outperforms Large-Scale Imitation Learning for MAPFabstractMulti-Agent Path Finding (MAPF) is the problem of effectively finding efficient collision-free paths for a group of agents in a shared workspace. The MAPF community has largely focused on developing high-performance heuristic search methods. Recently, several works have applied various machine learning (ML) techniques to solve MAPF, usually involving sophisticated architectures, reinforcement learning techniques, and set-ups, but none using large amounts of high-quality supervised data. Our initial objective in this work was to show how simple large-scale imitation learning of high-quality heuristic search methods can lead to state-of-the-art ML MAPF performance. However, we find that, at least with our model architecture, simple large-scale (700k examples with hundreds of agents per example) imitation learning does not produce impressive results. Instead, we find that by using prior work that post-processes MAPF model predictions to resolve 1-step collisions (CS-PIBT), we can train a simple ML MAPF policy in minutes that dramatically outperforms existing ML MAPF policies. This has serious implications for all future ML MAPF policies (with local communication) which currently struggle to scale. In particular, this finding implies that future learnt policies should always (1) use smart 1-step collision shields (e.g, CS-PIBT) and (2) include the collision shield with greedy actions as a baseline (e.g. PIBT), as well as (3) motivates future models to focus on longer horizon / more complex planning as 1-step collisions can be efficiently resolved. Rishi Veerapaneni, Arthur Jakobsson, Kevin Ren, Samuel Kim, Jiaoyang Li 0001, Maxim Likhachev |
ICRA | 4 |
| 2024 | Deep Learning and Symbolic Regression for Discovering Parametric EquationsabstractSymbolic regression is a machine learning technique that can learn the equations governing data and thus has the potential to transform scientific discovery. However, symbolic regression is still limited in the complexity and dimensionality of the systems that it can analyze. Deep learning, on the other hand, has transformed machine learning in its ability to analyze extremely complex and high-dimensional datasets. We propose a neural network architecture to extend symbolic regression to parametric systems where some coefficient may vary, but the structure of the underlying governing equation remains constant. We demonstrate our method on various analytic expressions and partial differential equations (PDEs) with varying coefficients and show that it extrapolates well outside of the training domain. The proposed neural-network-based architecture can also be enhanced by integrating with other deep learning architectures such that it can analyze high-dimensional data while being trained end-to-end. To this end, we demonstrate the scalability of our architecture by incorporating a convolutional encoder to analyze 1-D images of varying spring systems. Samuel Kim, Peter Y. Lu, Marin Soljacic |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Multi-Site Clinical Federated Learning Using Recursive and Attentive Models and NVFlareabstractThe prodigious growth of digital health data has precipitated a mounting interest in harnessing machine learning methodologies, such as natural language processing (NLP), to scrutinize medical records, clinical notes, and other text-based health information. Although NLP techniques have exhibited substantial potential in augmenting patient care and informing clinical decision-making, data privacy and adherence to regulations persist as critical concerns. Federated learning (FL) emerges as a viable solution, empowering multiple organizations to train machine learning models collaboratively without disseminating raw data. This paper proffers a pragmatic approach to medical NLP by amalgamating FL, NLP models, and the NVFlare framework, developed by NVIDIA. We introduce two exemplary NLP models, the Long-Short Term Memory (LSTM)-based model and Bidirectional Encoder Representations from Transformers (BERT), which have demonstrated exceptional performance in comprehending context and semantics within medical data. This paper encompasses the development of an integrated framework that addresses data privacy and regulatory compliance challenges while maintaining elevated accuracy and performance, incorporating BERT pretraining, and comprehensively substantiating the efficacy of the proposed approach. Won Joon Yun, Samuel Kim, Joongheon Kim |
ICDCS | 2 |
| 2021 | Integration of Neural Network-Based Symbolic Regression in Deep Learning for Scientific DiscoveryabstractSymbolic regression is a powerful technique to discover analytic equations that describe data, which can lead to explainable models and the ability to predict unseen data. In contrast, neural networks have achieved amazing levels of accuracy on image recognition and natural language processing tasks, but they are often seen as black-box models that are difficult to interpret and typically extrapolate poorly. In this article, we use a neural network-based architecture for symbolic regression called the equation learner (EQL) network and integrate it with other deep learning architectures such that the whole system can be trained end-to-end through backpropagation. To demonstrate the power of such systems, we study their performance on several substantially different tasks. First, we show that the neural network can perform symbolic regression and learn the form of several functions. Next, we present an MNIST arithmetic task where a convolutional network extracts the digits. Finally, we demonstrate the prediction of dynamical systems where an unknown parameter is extracted through an encoder. We find that the EQL-based architecture can extrapolate quite well outside of the training data set compared with a standard neural network-based architecture, paving the way for deep learning to be applied in scientific exploration and discovery. Samuel Kim, Peter Y. Lu, Srijon Mukherjee, Michael Gilbert, Li Jing 0001, Vladimir Ceperic, Marin Soljacic |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | High impedance holographic metasurfaces for conformal and high gain antenna applicationsabstractWe present simulated and experimental results from a conformal meta-antenna that was designed with a holographic-derived surface impedance for placement on a cylindrical surface. This class of antenna relies on the interaction between a surface wave generated by a simple radiator, e.g., a vertical monopole, and a spatially modulated surface impedance to generate high gain and predictable far-field radiation profiles. The surface impedance profile was formed using a distribution of sub-wavelength isotropic metal patches printed on a thin grounded dielectric substrate. An analysis of the meta-antenna radiative performance on both planar and cylindrical substrate geometries illustrates the importance of appropriately tailoring the surface impedance profile to account for the influence of curvature. Samuel Kim, David Shrekenhamer, Jeffrey Will, Ra'id Awadallah, Joseph A. Miragliotta |
CCNC | 1 |
| 2018 | Syntactical Analysis of the Weaknesses of Sentiment AnalyzersabstractWe carry out a syntactic analysis of two stateof-the-art sentiment analyzers, Google Cloud Natural Language and Stanford CoreNLP, to assess their classification accuracy on sentences with negative polarity items.We were motivated by the absence of studies investigating sentiment analyzer performance on sentences with polarity items, a common construct in human language.Our analysis focuses on two sentential structures: downward entailment and non-monotone quantifiers; and demonstrates weaknesses of Google Natural Language and CoreNLP in capturing polarity item information.We describe the particular syntactic phenomenon that these analyzers fail to understand that any ideal sentiment analyzer must.We also provide a set of 150 test sentences that any ideal sentiment analyzer must be able to understand. Rohil Verma, Samuel Kim, David Walter |
EMNLP | 2 |
| 2014 | Detecting pathological speech using contour modeling of harmonic-to-noise ratioabstractThis paper proposes a new feature extraction method for automatically detecting pathological voice in a normal conversation scenario. Unlike conventional approaches that utilize the static harmonic-to-noise ratio (HNR) characteristics of sustained vowel, the proposed method considers the dynamic movements of articulatory organs depending on the types of phonations. Assuming those movements reflect the health status of subjects, the proposed method utilizes the characteristics of HNR contour within a single sentence-level speech signal. Experimental results show that the proposed method reduces the classification error rate by 35.2 % (relative) compared to the conventional method. Jung-Won Lee, Samuel Kim, Hong-Goo Kang |
ICASSP | 2 |
| 2014 | Predicting Continuous Conflict Perceptionwith Bayesian Gaussian ProcessesabstractConflict is one of the most important phenomena of social life, but it is still largely neglected by the computing community. This work proposes an approach that detects common conversational social signals (loudness, overlapping speech, etc.) and predicts the conflict level perceived by human observers in continuous, non-categorical terms. The proposed regression approach is fully Bayesian and it adopts automatic relevance determination to identify the social signals that influence most the outcome of the prediction. The experiments are performed over the SSPNet Conflict Corpus, a publicly available collection of 1,430 clips extracted from televised political debates (roughly 12 hours of material for 138 subjects in total). The results show that it is possible to achieve a correlation close to 0.8 between actual and predicted conflict perception. Samuel Kim, Fabio Valente, Maurizio Filippone, Alessandro Vinciarelli |
IEEE Trans. Affect. Comput. | 1 |
| 2013 | On-line genre classification of TV programs using audio contentabstractAutomatic genre classification of TV programs can benefit users in various ways such as allowing for rapid selection of multimedia content. In this paper, we introduce an on-line method that can classify genres of TV programs using audio content. We deploy an acoustic topic model (ATM) which was originally designed to capture contextual information embedded within audio segments. With a dataset based on RAI content, we perform both on-line and off-line classification; we segment audio signals with a fixed length and feed into the system for on-line classification tasks, while we use whole audio signals for off-line tasks. The off-line experimental results suggest that the proposed method using audio content yields competitive performance with conventional methods using audio-visual features and outperforms conventional audio-based approaches. The on-line results show promising results in classifying genre of TV programs with short segments and also suggest that ATM performs better than conventional GMM method if the length of audio segments is longer (>1 second). Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 1 |
| 2013 | Annotation and classification of Political advertisements
Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan |
INTERSPEECH | 1 |
| 2013 | Annotation and detection of conflict escalation in Political debatesabstractConflict escalation in multi-party conversations refers to an increase \nin the intensity of conflict during conversations. Here we study annotation \nand detection of conflict escalation in broadcast political \ndebates towards a machine-mediated conflict management system. \nIn this regard, we label conflict escalation using crowd-sourced annotations \nand predict it with automatically extracted conversational \nand prosodic features. In particular, to annotate the conflict escalation \nwe deploy two different strategies, i.e., indirect inference and \ndirect assessment; the direct assessment method refers to a way that \nannotators watch and compare two consecutive clips during the annotation \nprocess, while the indirect inference method indicates that \neach clip is independently annotated with respect to the level of \nconflict then the level conflict escalation is inferred by comparing \nannotations of two consecutive clips. Empirical results with 792 \npairs of consecutive clips in classifying three types of conflict escalation, \ni.e., escalation, de-escalation, and constant, show that labels \nfrom direct assessment yield higher classification performance \n(45.3% unweighted accuracy (UA)) than the one from indirect inference \n(39.7% UA), although the annotations from both methods are \nhighly correlated (ρ = 0.74 in continuous values and 63% agreement \nin ternary classes). Samuel Kim, Fabio Valente, Alessandro Vinciarelli |
INTERSPEECH | 1 |
| 2013 | The INTERSPEECH 2013 computational paralinguistics challenge: social signals, conflict, emotion, autismabstractInternational audience Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Klaus R. Scherer, Fabien Ringeval, Mohamed Chetouani, Felix Weninger, Florian Eyben, Erik Marchi, Marcello Mortillaro, Hugues Salamin, Anna Polychroniou, Fabio Valente, Samuel Kim |
INTERSPEECH | 15 |
| 2012 | Automatic detection of conflicts in spoken conversations: Ratings and analysis of broadcast political debatesabstractAutomatic analysis of spoken conversations has recently searched for phenomena like agreement/disagreement in collaborative and non-conflictual discussions (e.g., meetings). This work adds a novel dimension investigating conflicts in spontaneous conversations. The study makes use of broadcasted political debates where conflicts naturally arise between participants. In the first part, an annotation scheme to rate the degree of conflict in conversations is described and applied to 12 hours of recordings. In the second part, the correlation between various prosodic/conversational features and the degree of conflict is investigated. In the third part, we perform automatic detection of the level of conflict based on those features showing an F-measure of 71.6% in three-level classification tasks. Samuel Kim, Fabio Valente, Alessandro Vinciarelli |
ICASSP | 1 |
| 2012 | Automatic detection of conflict escalation in spoken conversationsabstractThis paper investigates the automatic recognition of conflict escalations during spontaneous conversations. In our previous work, we studied if the level of conflict in a segment of conversation can be automatically inferred by means of prosodic and conversational features. This work investigates the possibility of automatically recognizing if the conflict is increasing, i.e., escalating, or not. The dataset used for the study consists of political debates where short clips are classified into escalation, de-escalation and constant labels. Results show a Weighted Accuracy (WA) equals to 69.6% and an Unweighted Accuracy (UA) equals to 49.5% thus revealing lower accuracies compared to the simple conflict detection task (WA 86.1%, UA 71.1%). While the task appears more difficult compared to conflict detection, results are significantly better than chance level showing the feasibility of this approach. Furthermore, the paper investigates the use of a speaker diarization algorithm to extract features in a completely automatic fashion highlighting some limitations of diarization system. Index Terms — Spoken Language Understanding, Conflicts, Paralinguistic, Spontaneous Conversation, Prosodic features, Turn-taking features Samuel Kim, Sree Harsha Yella, Fabio Valente |
INTERSPEECH | 1 |
| 2012 | Annotation and Recognition of Personality Traits in Spoken Conversations from the AMI Meetings CorpusabstractLIDIAP Fabio Valente, Samuel Kim, Petr Motlícek |
INTERSPEECH | 2 |
| 2012 | Predicting the conflict level in television political debates: an approach based on crowdsourcing, nonverbal communication and gaussian processesabstractOne of the most recent trends in multimedia indexing is to represent data in terms of the social and psychological phenomena that users perceive. In such a perspective this article proposes an approach for the automatic detection of conflict level in television political debates. The proposed approach includes the use of crowdsourcing techniques for modeling the perception of data consumers, the extraction of (language independent) nonverbal behavioral cues and the application of regression techniques based on Gaussian Processes. The experiments have been performed over 1430 clips of 30 seconds extracted from 45 political debates (roughly 12 hours of material). The results show that a correlation up to 0.8 can be achieved between the actual and predicted conflict level. Samuel Kim, Maurizio Filippone, Fabio Valente, Alessandro Vinciarelli |
ACM Multimedia | 1 |
| 2010 | Using naïve text queries for robust audio information retrievalabstractThe goal of this work is to build an audio information retrieval system which provides users with flexibility in formulating their queries: from audio examples to naïve text. Specifically, the focus of this paper is on using naïve text to create input queries describing the desired information of the users. Using naïve text queries, however, raises interoperability issues between annotation and retrieval processes due to the wide variety of available audio descriptions. In this paper, we propose an intermediate audio description layer (iADL) to solve the interoperability issues between the annotation and retrieval processes. The iADL comprises two axes corresponding to semantic and onomatopoeic descriptions based on human-to-human communication experiments on how humans express sounds verbally. Various text modeling schemes, such as latent semantic analysis (LSA) and latent topic model, are utilized to transform the naïve text onto the proposd iADL. Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan, Shiva Sundaram |
ICASSP | 1 |
| 2010 | A study of intra-speaker and inter-speaker affective variability using electroglottograph and inverse filtered glottal waveformsabstractIt is well-known that different speakers utilize their vocal instruments in diverse ways to express linguistic intention with some paralinguistic coloring such as emotional quality. The study of voice source features, which describe the action of the vocal folds, is important for a deeper understanding of emotion encoding in speech. In this study we investigate inter and intra-speaker differences in voicing activities as a function of emotion using electroglottography (EGG) and an inverse filtering technique. Results demonstrate that while voice quality features are good indicators of affective state, voice source descriptors vary in affective information across speakers. Glottal ratio measurements taken directly from the EGG signal are more reliable than measurements from the inverse-filtered glottal airflow signal, but the spectral harmonic amplitude differences of EGG are less useful than from inverse filtering. Index Terms: emotion, speech production, voice source, EGG, inverse filtering Daniel Bone, Samuel Kim, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 2 |
| 2010 | An N-gram model for unstructured audio signals toward information retrievalabstractAn N-gram modeling approach for unstructured audio signals is introduced with applications to audio information retrieval. The proposed N-gram approach aims to capture local dynamic information in acoustic words within the acoustic topic model framework which assumes an audio signal consists of latent acoustic topics and each topic can be interpreted as a distribution over acoustic words. Experimental results on classifying audio clips from BBC Sound Effects Library according to both semantic and onomatopoeic labels indicate that the proposed N-gram approach performs better than using only a bag-of-words approach by providing complementary local dynamic information. Samuel Kim, Shiva Sundaram, Panayiotis G. Georgiou, Shri Narayanan |
MMSP | 1 |
| 2009 | A robust harmony structure modeling scheme for classical music opus identificationabstractA robust algorithm to model the harmony structure of a music piece is proposed. The harmony structure is extracted directly from a music audio signal using a second-order statistic of chroma feature vectors. The method is experimentally shown to be robust against the degradation of chroma feature vectors due to noisy pitch estimation in our classical music opus identification evaluation. To analyze the effects of the noisy pitch estimation, we propose a noise model that describes difference between the oracle chroma feature vectors as obtained from a symbolic representation and those extracted from the rendered audio signal. The results suggest that the harmony structure modeling scheme employing the covariance matrix is more robust than the alternative investigated second-order statistics. The results also show that the proposed method obtains 84.3% accuracy with the symbolic representations and 72.0% with the synthesized audio data, which suggest that the proposed harmony structure modeling method has room for further improvement by addressing the signal processing challenges of pitch extraction, or through employing more robust features. Samuel Kim, Panayiotis G. Georgiou, Shri Narayanan |
ICASSP | 1 |
| 2008 | Music fingerprint extraction for classical music cover song identificationabstractAn algorithm for extracting music fingerprints directly from an audio signal is proposed in this paper. The proposed music fingerprint aims to encapsulate various aspects of musical information, such as overall note distribution, harmony structure, and their temporal changes, all in a compact representation. The utility of the proposed music fingerprint to the task of automatic classical music cover song identification is explored through experimental studies; specifically, the goal here is to identify the different versions of the same music through similarity comparisons of the music fingerprints. The results show an improved performance over the state-of-the-art cover song identification systems in terms of both accuracy and speed: the accuracy improved by approximately 40% while the search speed is about 60 times faster than the conventional system. Samuel Kim, Erdem Ünal, Shri Narayanan |
ICME | 1 |
| 2008 | Dynamic chroma feature vectors with applications to cover song identificationabstractA new chroma-based dynamic feature vector is proposed inspired by psychophysical observations that the human auditory system detects reltative pitch changes rather than absolute pitch values. The proposed chroma-based dynamic feature vector describes the relative pitch change intervals. The utility of the proposed feature vector incorporated with a music fingerprint extraction algorithm is experimentally explored within a music cover song identification framework. The results with a classical music database suggest that the proposed biologically plausible dynamic chroma feature vector can be successfully added to the conventional chroma feature vector as a complementary feature; it provides a 5.8% relative performance improvement. Samuel Kim, Shri Narayanan |
MMSP | 1 |
| 2008 | Strategies to Improve the Robustness of Agglomerative Hierarchical Clustering Under Data Source Variation for Speaker DiarizationabstractMany current state-of-the-art speaker diarization systems exploit agglomerative hierarchical clustering (AHC) as their speaker clustering strategy, due to its simple processing structure and acceptable level of performance. However, AHC is known to suffer from performance robustness under data source variation. In this paper, we address this problem. We specifically focus on the issues associated with the widely used clustering stopping method based on Bayesian information criterion (BIC) and the merging-cluster selection scheme based on generalized likelihood ratio (GLR). First, we propose a novel alternative stopping method for AHC based on information change rate (ICR). Through experiments on several meeting corpora, the proposed method is demonstrated to be more robust to data source variation than the BIC-based one. The average improvement obtained in diarization error rate (DER) by this method is 8.76% (absolute) or 35.77% (relative). We also introduce a selective AHC (SAHC) in the paper, which first runs AHC with the ICR-based stopping method only on speech segments longer than 3 s and then classifies shorter speech segments into one of the clusters given by the initial AHC. This modified version of AHC is motivated by our previous analysis that the proportion of short speech turns (or segments) in a data source is a significant factor contributing to the robustness problem arising in the GLR-based merging-cluster selection scheme. The additional performance improvement obtained by SAHC is 3.45% (absolute) or 14.08% (relative) in terms of averaged DER. Kyu Jeong Han, Samuel Kim, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Robust speaker clustering strategies to data source variation for improved speaker diarizationabstractAgglomerative hierarchical clustering (AHC) has been widely used in speaker diarization systems to classify speech segments in a given data source by speaker identity, but is known to be not robust to data source variation. In this paper, we identify one of the key potential sources of this variability that negatively affects clustering error rate (CER), namely short speech segments, and propose three solutions to tackle this issue. Through experiments on various meeting conversation excerpts, the proposed methods are shown to outperform simple AHC in terms of relative CER improvements in the range of 17-32%. Kyu Jeong Han, Samuel Kim, Shri Narayanan |
ASRU | 2 |
| 2007 | Real-time Emotion Detection System using Speech: Multi-modal Fusion of Different Timescale FeaturesabstractThe goal of this work is to build a real-time emotion detection system which utilizes multi-modal fusion of different timescale features of speech. Conventional spectral and prosody features are used for intra-frame and supra-frame features respectively, and a new information fusion algorithm which takes care of the characteristics of each machine learning algorithm is introduced. In this framework, the proposed system can be associated with additional features, such as lexical or discourse information, in later steps. To verify the realtime system performance, binary decision tasks on angry and neutral emotion are performed using concatenated speech signal simulating realtime conditions. Samuel Kim, Panayiotis G. Georgiou, Sungbok Lee, Shri Narayanan |
MMSP | 1 |
| 2005 | A noise-robust pitch synchronous feature extraction algorithm for speaker recognition systemsabstractA noise-robust pitch synchronous feature extraction algorithm for speaker recognition systems is proposed in this paper. Since the pitch synchronous algorithms utilize pitch information, which is meaningful only for periodic segments, we propose a new scheme to deal with non-periodic ones such as unvoiced and noise-corrupted. The experimental results show that the proposed algorithm outperforms the conventional algorithm us- ing fixed length of analysis window in actual identification tasks even in low SNR noisy environments. Samuel Kim, Sung-Wan Yoon, Thomas Eriksson, Hong-Goo Kang, Dae Hee Youn |
INTERSPEECH | 1 |
| 2005 | An information-theoretic perspective on feature selection in speaker recognitionabstractThis letter studies feature selection in speaker recognition from an information-theoretic view. We closely tie the performance, in terms of the expected classification error probability, to the mutual information between speaker identity and features. Information theory can then help us to make qualitative statements about feature selection and performance. We study various common features used for speaker recognition, such as mel-warped cepstrum coefficients and various parameterizations of linear prediction coefficients. The theory and experiments give valuable insights in feature selection and performance of speaker-recognition applications. Thomas Eriksson, Samuel Kim, Hong-Goo Kang, Chungyong Lee |
IEEE Signal Process. Lett. | 2 |
| 2004 | A pitch synchronous feature extraction method for speaker recognitionabstractThe paper presents a novel feature extraction method to improve the performance of speaker identification systems. The proposed feature has the form of a typical conventional feature, Mel frequency cepstral coefficients (MFCC), but a flexible segmentation to reduce spectral mismatch between training and testing processes. Specifically, the length and shift size of the analysis frame are determined by a pitch synchronous method, pitch synchronous MFCC (PSMFCC). To verify the performance of the new feature, we measure the cepstral distortion between training and testing and also perform closed set speaker identification tests. With text-independent and text-dependent experiments, the proposed algorithm provides 44.3% and 26.7% relative improvement, respectively. Samuel Kim, Thomas Eriksson, Hong-Goo Kang, Dae Hee Youn |
ICASSP (1) | 1 |
| 2004 | Theory for speaker recognition over IP
Thomas Eriksson, Samuel Kim, Hong-Goo Kang, Chungyong Lee |
INTERSPEECH | 2 |
| 2004 | On the time variability of vocal tract for speaker recognition
Samuel Kim, Thomas Eriksson, Hong-Goo Kang |
INTERSPEECH | 1 |
| 2001 | The MD-join: An Operator for Complex OLAPabstractOLAP queries (i.e. group-by or cube-by queries with aggregation) have proven to be valuable for data analysis and exploration. Many decision support applications need very complex OLAP queries, requiring a fine degree of control over both the group definition and the aggregates that are computed. For example, suppose that the user has access to a data cube whose measure attribute is Sum(Sales). Then the user might wish to compute the sum of sales in New York and the sum of sales in California for those data cube entries in which Sum(Sales)>$1,000,000. This type of complex OLAP query is often difficult to express and difficult to optimize using standard relational operators (including standard aggregation operators). In this paper, we propose the MD-join operator for complex OLAP queries. The MD-join provides a clean separation between group definition and aggregate computation, allowing great flexibility in the expression of OLAP queries. In addition, the MD-join has a simple and easily optimizable implementation, while the equivalent relational algebra expression is often complex and difficult to optimize. We present several algebraic transformations that allow relational algebra queries that include MD-joins to be optimized. Damianos Chatziantoniou, Michael O. Akinde, Theodore Johnson, Samuel Kim |
ICDE | 4 |