Roy Hirsch

dblp:296/4438 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Speech recognition and synthesis · 46% Language models and text generation · 23% Kernel, tree and ensemble methods · 15%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 67% Recommender systems · 33%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
speech language model
0.812024
Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM · ICLR 2024
Natural language and speech › Language models and text generation › natural language understanding › question answering
spoken question answering
0.812024
Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM · ICLR 2024
Information retrieval › evaluation
benchmark dataset
0.712023
Efficient Discovery and Effective Evaluation of Visual Perceptual Similarity: A Benchmark and Beyond · ICCV 2023
Information retrieval › similarity measure
image similarity
0.712023
Efficient Discovery and Effective Evaluation of Visual Perceptual Similarity: A Benchmark and Beyond · ICCV 2023
Computer vision › 3D vision › geometric deep learning › set learning
set prediction
0.512021
Trees with Attention for Set Prediction Tasks · ICML 2021
Machine learning › Kernel, tree and ensemble methods
tree-based models
0.512021
Trees with Attention for Set Prediction Tasks · ICML 2021
Recommender systems
cold-start recommendation
0.512021
Cold Item Integration in Deep Hybrid Recommenders via Tunable Stochastic Gates · ICDM 2021
Recommender systems › collaborative filtering
hybrid recommendation
0.112021
Cold Item Integration in Deep Hybrid Recommenders via Tunable Stochastic Gates · ICDM 2021

Methods — techniques the papers use, named apart from their topics

speech encoder · 0.8spectrogram modeling · 0.8large language model adaptation · 0.8cross-modal chain-of-thought · 0.8labeling procedure · 0.7evaluation metrics · 0.7tunable stochastic gates · 0.5set-compatible split criteria · 0.5collaborative filtering · 0.5attention mechanism · 0.5
YearPublicationVenuePosition
2024 Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
abstract
We present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire system is trained end-to-end and operates directly on spectrograms, simplifying our architecture. Key to our approach is a training objective that jointly supervises speech recognition, text continuation, and speech synthesis using only paired speech-text pairs, enabling a `cross-modal' chain-of-thought within a single decoding pass. Our method surpasses existing spoken language models in speaker preservation and semantic coherence. Furthermore, the proposed model improves upon direct initialization in retaining the knowledge of the original LLM as demonstrated through spoken QA datasets. We release our audio samples and spoken QA dataset via our website.
Eliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar, Chulayuth Asawaroengchai, Soroosh Mariooryad, Ehud Rivlin, R. J. Skerry-Ryan, Michelle Tadmor Ramanovich
ICLR3
2024 Random Walks for Temporal Action Segmentation with Timestamp Supervision
abstract
Temporal action segmentation relates to high-level video understanding, commonly formulated as frame-wise classification of untrimmed videos into predefined actions. Fully-supervised deep-learning approaches require dense video annotations which are time and money consuming. Furthermore, the temporal boundaries between consecutive actions typically are not well-defined, leading to inherent ambiguity and interrater disagreement. A promising approach to remedy these limitations is timestamp supervision, requiring only one labeled frame per action instance in a training video. In this work, we reformulate the task of temporal segmentation as a graph segmentation problem with weakly-labeled vertices. We introduce an efficient segmentation method based on random walks on graphs, obtained by solving a sparse system of linear equations. Furthermore, the proposed technique can be employed in any one or combination of the following forms: (1) as a standalone solution for generating dense pseudo-labels from timestamps; (2) as a training loss; (3) as a smoothing mechanism given intermediate predictions. Extensive experiments with three datasets (50Salads, Breakfast, GTEA) show that our method competes with state-of-the-art, and allows the identification of regions of uncertainty around action boundaries.
Roy Hirsch, Regev Cohen, Tomer Golany, Daniel Freedman, Ehud Rivlin
WACV1
2023 Efficient Discovery and Effective Evaluation of Visual Perceptual Similarity: A Benchmark and Beyond
abstract
Visual similarity discovery (VSD) is an important task with broad e-commerce applications. Given an image of a certain object, the goal of VSD is to retrieve images of different objects with high perceptual visual similarity. Although being a highly addressed problem, the evaluation of proposed methods for VSD is often based on a proxy of an identification-retrieval task, evaluating the ability of a model to retrieve different images of the same object. We posit that evaluating VSD methods based on identification tasks is limited, and faithful evaluation must rely on expert annotations. In this paper, we introduce the first large-scale fashion visual similarity benchmark dataset, consisting of more than 110K expert-annotated image pairs. Besides this major contribution, we share insight from the challenges we faced while curating this dataset. Based on these insights, we propose a novel and efficient labeling procedure that can be applied to any dataset. Our analysis examines its limitations and inductive biases, and based on these findings, we propose metrics to mitigate those limitations. Though our primary focus lies on visual similarity, the methodologies we present have broader applications for discovering and evaluating perceptual similarity across various domains.
Oren Barkan, Tal Reiss, Jonathan Weill, Ori Katz, Roy Hirsch, Itzik Malkiel, Noam Koenigstein
ICCV5
2023 Self-supervised Learning for Endoscopic Video Analysis
Roy Hirsch, Mathilde Caron, Regev Cohen, Amir Livne, Ron Shapiro, Tomer Golany, Roman Goldenberg, Daniel Freedman, Ehud Rivlin
MICCAI (5)1
2021 Anchor-based Collaborative Filtering
abstract
Modern-day recommender systems are often based on learning representations in a latent vector space that encode user and item preferences. In these models, each user/item is represented by a single vector and user-item interactions are modeled by some function over the corresponding vectors. This paradigm is common to a large body of collaborative filtering models that repeatedly demonstrated superior results. In this work, we break away from this paradigm and present ACF: Anchor-based Collaborative Filtering. Instead of learning unique vectors for each user and each item, ACF learns a spanning set of anchor-vectors that commonly serve both users and items. In ACF, each anchor corresponds to a unique "taste'' and users/items are represented as a convex combination over the spanning set of anchors. Additionally, ACF employs two novel constraints: (1) exclusiveness constraint on item-to-anchor relations that encourages each item to pick a single representative anchor, and (2) an inclusiveness constraint on anchors-to-items relations that encourages full utilization of all the anchors. We compare ACF with other state-of-the-art alternatives and demonstrate its effectiveness on multiple datasets.
Oren Barkan, Roy Hirsch, Ori Katz, Avi Caciularu, Noam Koenigstein
CIKM2
2021 Cold Start Revisited: A Deep Hybrid Recommender with Cold-Warm Item Harmonization
abstract
Collaborative filtering-based recommender systems are known to suffer from the item cold-start problem. Most recent attempts to mitigate this problem presented parametric approaches, such as deep content based models. In this paper, we show that a straightforward application of parametric models may lead to discrepancies between the cold and warm items’ distributions in the CF space. As a remedy, we propose to combine parametric with non-parametric estimation for robust cold item placement. Extensive evaluation indicates that our method is competitive with other baselines, while producing cold items placement that better resembles the distribution of warm items in the collaborative filtering space.
Oren Barkan, Roy Hirsch, Ori Katz, Avi Caciularu, Yoni Weill, Noam Koenigstein
ICASSP2
2021 Cold Item Integration in Deep Hybrid Recommenders via Tunable Stochastic Gates
abstract
A major challenge in collaborative filtering methods is how to produce recommendations for cold items (items with no ratings), or integrate cold items into an existing catalog. Over the years, a variety of hybrid recommendation models have been proposed to address this problem by utilizing items’ metadata and content along with their ratings or usage patterns. In this work, we wish to revisit the cold start problem in order to draw attention to an overlooked challenge: the ability to integrate and balance between (regular) warm items and completely cold items. In this case, two different challenges arise: (1) preserving high-quality performance on warm items, while (2) learning to promote cold items to relevant users. First, we show that these two objectives are in fact conflicting, and the balance between them depends on the business needs and the application at hand. Next, we propose a novel hybrid recommendation algorithm that bridges these two conflicting objectives and enables a harmonized balance between preserving high accuracy for warm items while effectively promoting completely cold items. We demonstrate the effectiveness of the proposed algorithm on movies, apps, and articles recommendations, and provide an empirical analysis of the cold-warm trade-off.
Oren Barkan, Roy Hirsch, Ori Katz, Avi Caciularu, Jonathan Weill, Noam Koenigstein
ICDM2
2021 Trees with Attention for Set Prediction Tasks
abstract
In many machine learning applications, each record represents a set of items. For example, when making predictions from medical records, the medications prescribed to a patient are a set whose size is not fixed and whose order is arbitrary. However, most machine learning algorithms are not designed to handle set structures and are limited to processing records of fixed size. Set-Tree, presented in this work, extends the support for sets to tree-based models, such as Random-Forest and Gradient-Boosting, by introducing an attention mechanism and set-compatible split criteria. We evaluate the new method empirically on a wide range of problems ranging from making predictions on sub-atomic particle jets to estimating the redshift of galaxies. The new method outperforms existing tree-based methods consistently and significantly. Moreover, it is competitive and often outperforms Deep Learning. We also discuss the theoretical properties of Set-Trees and explain how they enable item-level explainability.
Roy Hirsch, Ran Gilad-Bachrach
ICML1