Oshry Ben-Harush

dblp:25/9230 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-authorSystems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 77% Visualization and visual analytics · 23%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
speaker diarization
0.112012
Initialization of Iterative-Based Speaker Diarization Systems for Telephone Conversations · IEEE Trans. Speech Audio Process. 2012
Visualization and visual analytics
clustering
0.012012
Initialization of Iterative-Based Speaker Diarization Systems for Telephone Conversations · IEEE Trans. Speech Audio Process. 2012

Methods — techniques the papers use, named apart from their topics

self-organizing map · 0.1k-means initialization · 0.1gaussian mixture model · 0.1
YearPublicationVenuePosition
2017 Predicting HDD failures from compound SMART attributes
abstract
Hard Disk Drives (HDD s) sometimes fail with no apparent reason; some SMART (Self-Monitoring, Analysis and Reporting Technology) attributes present strong correlations with drive failures[3], yet a drive may also fail without (supposedly) any previous indication. Most of the host systems today utilize alert-methods which are reactive by nature-a drive is indicated to fail when some SMART attribute exceeds its vendor defined threshold for valid operation[4]. This approach does not take the cross correlation between different attributes into account and the fact that thresholds vary across different vendors.
Shiri Gaber, Oshry Ben-Harush, Amihai Savir
SYSTOR2
2012 Initialization of Iterative-Based Speaker Diarization Systems for Telephone Conversations
abstract
Speaker diarization systems attempt to assign temporal segments from a conversation betweenRspeakers to an appropriate speakerr. This task is generally performed when no prior information is given regarding the speakers. The number of speakers is usually unknown and needs to be estimated. However, there are applications where the number of speakers is known in advance. The diarization process generally consists of change detection, clustering and labeling of a given audio stream. Speaker diarization can be performed using an iterative approach that is optimized by the selection of appropriate initial conditions. This study examines the influence of several common initialization algorithms including two variants of a recently proposed, K-means based initialization algorithm over the performance of an iterative-based speaker diarization system applied to two speaker telephone conversations. The suggested speaker diarization system employs either self organizing maps or Gaussian mixture models in order to model the speakers and non-speech in the conversation. The diarization system and initialization algorithms are tuned using 108 telephone conversations taken from LDC CallHome corpus, this is the development set. The evaluation subset is composed of 2048 telephone conversations extracted from the NIST 2005 Rich Transcription corpus. The results obtained show that by initializing the speaker diarization system using the K-means based algorithms provide a relative improvement of 10.4% for the LDC development set and 12.2% for the NIST evaluation subset when compared to random initialization after 12 iterations which are required for the convergence of the diarization process using random initialization. However, when using the K-means based initialization approach, only five iterations are required for the system to converge. Thus, using the new initialization allows us to improve the performances both in terms of diarization error rate and speed of convergence.
Oshry Ben-Harush, Itshak Lapidot, Hugo Guterman
IEEE Trans. Speech Audio Process.1
2010 Incremental diarization of telephone conversations
abstract
Speaker diarization systems attempt segmentation and labeling of a conversation between R speakers, while no prior information is given regarding the conversation. Most state of the art diarization systems require the full body of the conversation data prior to the application of some diarization approach. However, for some applications such as forensics, which handles vast amount of data, an on-line or incremental diarization is of high importance. For that purpose, a two-stage incremental diarization of telephone conversations algorithm is suggested. On the first stage, a fully unsupervised diarization algorithm is applied over an initial training segment from the conversation. The secondstage is composed of time-series clustering of increments of the conversation. Applying incremental diarization over 1802 telephone conversations from NIST 2005 SER generated an increase in diarization error of approximately 2% compared to the diarization error of an off-line diarization system.
Oshry Ben-Harush, Itshak Lapidot, Hugo Guterman
INTERSPEECH1
2009 Entropy based overlapped speech detection as a pre-processing stage for speaker diarization
abstract
One inherent deficiency of most diarization systems is their inability to handle co-channel or overlapped speech. Most of the suggested algorithms perform under singular conditions, require high computational complexity in both time and frequency domains. In this study, frame based entropy analysis of the audio data in the time domain serves as a single feature for an overlapped speech detection algorithm. Identification of overlapped speech segments is performed using Gaussian Mixture Modeling (GMM) along with well known classification algorithms applied on two speaker conversations. By employing this methodology, the proposed method eliminates the need for setting a hard threshold for each conversation or database. LDC CALLHOME American English corpus is used for evaluation of the suggested algorithm. The proposed method successfully detects 63.2% of the frames labeled as overlapped speech by the manual segmentation, while keeping a 5.4% false-alarm rate.
Oshry Ben-Harush, Itshak Lapidot, Hugo Guterman
INTERSPEECH1
2008 Weighted segmental k-means initialization for SOM-based speaker clustering
abstract
A new approach for initial assignment of data in a speaker clustering application is presented. This approach employs Weighted Segmental K-Means clustering algorithm prior to competitive based learning. The clustering system relies on Self-Organizing Maps (SOM) for speaker modeling and likelihood estimation. Performance is evaluated on 108 two speaker conversations taken from LDC CALLHOME American English Speech corpus using NIST criterion and shows an improvement of approximately 48% in Cluster Error Rate (CER) relative to the randomly initialized clustering system. The number of iterations was reduced significantly, which contributes to both speed and efficiency of the clustering system.
Oshry Ben-Harush, Itshak Lapidot, Hugo Guterman
INTERSPEECH1