VLDB 2026 Research / reviewers in the wild / expert
Benjamín Béjar Haro
dblp:118/9633 · also Benjamín Béjar
· DBLP profile ↗
16ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0001-9705-4483ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cross-Channel Unlabeled Sensing over a Union of Signal SubspacesabstractCross-channel unlabeled sensing addresses the problem of recovering a multi-channel signal from measurements that were shuffled across channels. This work expands the cross-channel unlabeled sensing framework to signals that lie in a union of subspaces. The extension allows for handling more complex signal structures and broadens the framework to tasks like compressed sensing. These mismatches between samples and channels often arise in applications such as whole-brain calcium imaging of freely moving organisms or multi-target tracking. We improve over previous models by deriving tighter bounds on the required number of samples for unique reconstruction, while supporting more general signal types. The approach is validated through an application in whole-brain calcium imaging, where organism movements disrupt sample-to-neuron mappings. This demonstrates the utility of our framework in real-world settings with imprecise sample-channel associations, achieving accurate signal reconstruction. Taulant Koka, Manolis C. Tsakiris, Benjamín Béjar Haro, Michael Muma |
ICASSP | 3 |
| 2025 | Language-Guided Instance-Aware Domain-Adaptive Panoptic SegmentationabstractThe increasing relevance of panoptic segmentation is tied to the advancements in autonomous driving and AR/VR applications. However, the deployment of such models has been limited due to the expensive nature of dense data an-notation, giving rise to unsupervised domain adaptation (UDA). A key challenge in panoptic UDA is reducing the domain gap between a labeled source and an unlabeled target domain while harmonizing the subtasks of seman-tic and instance segmentation to limit catastrophic inter-ference. While considerable progress has been achieved, existing approaches mainly focus on the adaptation of se-mantic segmentation. In this work, we focus on incorpo-rating instance-level adaptation via a novel instance-aware cross-domain mixing strategy IMix. IMix significantly en-hances the panoptic quality by improving instance segmen-tation performance. Specifically, we propose inserting high-confidence predicted instances from the target domain onto source images, retaining the exhaustiveness of the resulting pseudo-labels while reducing the injected confirmation bias. Nevertheless, such an enhancement comes at the cost of degraded semantic performance, attributed to catastrophic forgetting. To mitigate this issue, we regu-larize our semantic branch by employing CLIP-based do-main alignment (CDA), exploiting the domain-robustness of natural language prompts. Finally, we present an end-to-end model incorporating these two mechanisms called LIDAPS, achieving state-of-the-art results on all popular panoptic UDA benchmarks. https://github.com/elhamAm/LIDAPS Elham Amin Mansour, Ozan Unal, Suman Saha 0001, Benjamín Béjar Haro, Luc Van Gool |
WACV | 4 |
| 2024 | Semantic-aware Video Representation for Few-shot Action RecognitionabstractRecent work on action recognition leverages 3D features and textual information to achieve state-of-the-art performance. However, most of the current few-shot action recognition methods still rely on 2D frame-level representations, often require additional components to model temporal relations, and employ complex distance functions to achieve accurate alignment of these representations. In addition, existing methods struggle to effectively integrate textual semantics, some resorting to concatenation or addition of textual and visual features, and some using text merely as an additional supervision without truly achieving feature fusion and information transfer from different modalities. In this work, we propose a simple yet effective Semantic-Aware Few-Shot Action Recognition (SAFSAR) model to address these issues. We show that directly leveraging a 3D feature extractor combined with an effective feature-fusion scheme, and a simple cosine similarity for classification can yield better performance without the need of extra components for temporal modeling or complex distance functions. We introduce an innovative scheme to encode the textual semantics into the video representation which adaptively fuses features from text and video, and encourages the visual encoder to extract more semantically consistent features. In this scheme, SAFSAR achieves alignment and fusion in a compact way. Experiments on five challenging few-shot action recognition benchmarks under various settings demonstrate that the proposed SAFSAR model significantly improves the state-of-the-art performance. Yutao Tang, Benjamín Béjar Haro, René Vidal |
WACV | 2 |
| 2024 | Shuffled multi-channel sparse signal recoveryabstractMismatches between samples and their respective channel or target commonly arise in several real-world applications. For instance, whole-brain calcium imaging of freely moving organisms, multiple-target tracking or multi-person contactless vital sign monitoring may be severely affected by mismatched sample-channel assignments. To address this issue systematically, we frame it as a signal reconstruction problem where correspondences between samples and channels are lost. Assuming a sensing matrix for the signals, we show the problem’s equivalence to a highly structured unlabeled sensing problem and establish conditions for unique recovery. This is crucial since existing unlabeled sensing theory is inapplicable and results for reconstructing shuffled multi-channel signals do not yet exist. Our results extend to continuous-time sparse signals, and we derive conditions for reconstructing shuffled sparse signals. For the two-channel case, we provide a first reconstruction method, which combines sparse signal recovery with robust linear regression, outperforming existing unlabeled sensing methods in numerical experiments. Additionally, we showcase its effectiveness in a real-world application involving calcium imaging traces. Our theory marks a significant initial step in addressing this challenging signal reconstruction problem, with potential extensions to diverse signal representations encountered in real-world problems with imprecise measurement or channel assignment. Taulant Koka, Manolis C. Tsakiris, Michael Muma, Benjamín Béjar Haro |
Signal Process. | 4 |
| 2022 | Facial Tic Detection in Untrimmed Videos of Tourette Syndrome PatientsabstractTourette Syndrome (TS) is a behavioral disorder that onsets in childhood and is characterized by the expression of involuntary movements and sounds commonly referred to as tics. Behavioral therapy is the first-line treatment for patients with TS, and it helps patients raise awareness about tic occurrence as well as develop tic inhibition strategies. However, the limited availability of therapists and the difficulties for in-home follow up work limits its effectiveness. An automatic tic detection system that is easy to deploy could alleviate the difficulties of home-therapy by providing feedback to the patients while exercising tic awareness. In this work, we propose a novel architecture (T-Net) for automatic tic detection and classification from untrimmed videos. T-Net combines temporal detection and segmentation and operates on features that are interpretable to a clinician. We compare T-Net to several state-of-the-art systems working on deep features extracted from the raw videos and T-Net achieves comparable performance in terms of average precision while relying on interpretable features needed in clinical practice. Yutao Tang, Benjamín Béjar Haro, Joey K.-Y. Essoe, Joseph F. McGuire, René Vidal |
ICPR | 2 |
| 2022 | The Fastest $\ell _{1, \infty }$ℓ1, ∞ Prox in the WestabstractProximal operators are of particular interest in optimization problems dealing with non-smooth objectives because in many practical cases they lead to optimization algorithms whose updates can be computed in closed form or very efficiently. A well-known example is the proximal operator of the vector L1 norm, which is given by the soft-thresholding operator. In this paper we study the proximal operator of the mixed L1,oo matrix norm and show that it can be computed in closed form by applying the well-known soft-thresholding operator to each column of the matrix. However, unlike the vector L1 norm case where the threshold is constant, in the mixed L1,oo norm case each column of the matrix might require a different threshold and all thresholds depend on the given matrix. We propose a general iterative algorithm for computing these thresholds, as well as two efficient implementations that further exploit easy to compute lower bounds for the mixed norm of the optimal solution. Experiments on large-scale synthetic and real data indicate that the proposed methods can be orders of magnitude faster than state-of-the-art methods. Benjamín Béjar Haro, Ivan Dokmanic, René Vidal |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Representation Learning on Visual-Symbolic Graphs for Video Understanding
Effrosyni Mavroudi, Benjamín Béjar Haro, René Vidal |
ECCV (29) | 2 |
| 2017 | Unlabeled sensing: Reconstruction algorithm and theoretical guaranteesabstractIt often happens that we are interested in reconstructing an unknown signal from partial measurements. Also, it is typically assumed that the location (temporal or spatial) of each sample is known and that the only distortion present in the observations is due to additive measurement noise. However, there are some applications where such location information is lost. In this paper, we consider the situation in which the order of noisy samples, taken from a linear measurement system, is missing. Previous work on this topic has only considered the noiseless case and exhaustive search combinatorial algorithms. We propose a much more efficient algorithm based on a geometrical viewpoint of the problem. We also study the uniqueness of the solution under different choices of the sampling matrix and its robustness to noise for the case of two-dimensional signals. Finally we provide simulation results to confirm the theoretical findings of the paper. Golnoosh Elhami, Adam Scholefield, Benjamín Béjar Haro, Martin Vetterli |
ICASSP | 3 |
| 2017 | Shape from bandwidth: The 2-D orthogonal projection caseabstractCould bandwidth—one of the most classic concepts in signal processing—have a new purpose? In this paper, we investigate the feasibility of using bandwidth to infer shape from a single image. As a first analysis, we limit our attention to orthographic projection and assume a 2-D world. Adam Scholefield, Benjamín Béjar Haro, Martin Vetterli |
ICASSP | 2 |
| 2015 | Enhancing local - Transmitting less - Improving globalabstractSuper-resolving a natural image is an ill-posed problem. The classical approach is based on the registration and subsequent interpolation of a given set of low-resolution images. However, achieving satisfactory results typically requires the combination of a large number of them. Such an approach would be impractical over heterogeneous rate-constrained wireless networks due to the associated communication cost and limited data available. In this paper, we present an approach for local image enhancement following the finite rate of innovation sampling framework, and motivate its application to the super-resolution problem over heterogeneous networks. Local estimates can be exchanged among the nodes of the network in order to regularize the super-resolution problem while, at the same time, reduce data exchange. Benjamín Béjar Haro, Martin Vetterli |
ICASSP | 1 |
| 2014 | Topology optimization for energy-efficient communications in consensus wireless networksabstractOver the past years there has been an increasing interest in developing distributed computation methods over wireless networks. A new communication paradigm has emerged where distributed algorithms such as consensus have played a key role in the development of such networks. A special case are wireless sensor networks (WSN) which have found application in a large variety of problems such as environmental monitoring, surveillance, or localization, to cite a few. One major design issue in WSNs is energy efficiency. Nodes are typically battery-powered devices and thus, it is critical to make a proper use of the scarce energy resources. This fact motivates the search for optimal conditions that favor the communication environment. It is well known that the rate at which the information is spread across the network depends on the topology of the network and that finding the optimal topology is a hard combinatorial problem. However, using convex optimization tools, we propose a method that tries to find the optimal topology in a consensus wireless network that uses broadcast messages. Our results show that exploiting the broadcast nature of the wireless channel leads to more energy efficient configurations than using dedicated unicast messages and that our algorithm performs very close to the optimal solution. Benjamín Béjar Haro, Martin Vetterli |
ICASSP | 1 |
| 2013 | String Motif-Based Description of Tool Motion for Detecting Skill and Gestures in Robotic Surgery
Narges Ahmidi, Benjamín Béjar Haro, S. Swaroop Vedula, Sanjeev Khudanpur, René Vidal, Gregory D. Hager |
MICCAI (1) | 3 |
| 2013 | Surgical gesture classification from video and kinematic data
Luca Zappella, Benjamín Béjar Haro, Gregory D. Hager, René Vidal |
Medical Image Anal. | 2 |
| 2012 | Lifetime maximization for beamforming applications in wireless sensor networksabstractEnergy efficiency is a major design issue in the context of Wireless Sensor Networks (WSN). If data is to be sent to a far-away base station, collaborative beamforming by the sensors may help to distribute the load among the nodes and reduce fast battery depletion. However, collaborative beamforming techniques are far from optimality and in many cases may be wasting more power than required. In this contribution we consider the issue of energy efficiency in beamforming applications. Using a convex optimization framework, we propose the design of a virtual beamformer that maximizes the network's lifetime while satisfying a pre-specified Quality of Service (QoS) requirement. A distributed consensus-based algorithm for the computation of the optimal beamformer is also provided. Benjamín Béjar Haro, Santiago Zazo, Daniel Pérez Palomar |
ICASSP | 1 |
| 2012 | Surgical Gesture Classification from Video Data
Benjamín Béjar Haro, Luca Zappella, René Vidal |
MICCAI (1) | 1 |
| 2011 | A Top-Down Random Generator for the In-Home PLC ChannelabstractWe propose a random channel generator for in-home power line communications (PLC). We follow a statistical top-down approach and we model the multipath propagation effects of the PLC channel in the frequency domain. Then, we introduce the variability into the model, i.e., we let the parameters associated to the reflections be random, according to a certain statistics. Finally, we fit the model to the experimental data. We target the average path loss and root-mean-square (RMS) delay spread of the measured channels. According to this methodology, we show that the randomly generated channels are in good agreement with the experimental ones in terms of the main metrics. Andrea M. Tonello, Fabio Versolatto, Benjamín Béjar Haro |
GLOBECOM | 3 |