Laura Toni

dblp:81/7871 · DBLP profile ↗
← Back
67ranked-venue papers
17as first author
24since 2021 · last 2026
0000-0002-8441-8791ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 11 first-author · 10 since 2021Computer networks · 18 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Theory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GT-MilliNoise: Graph transformer for point-wise denoising of indoor millimetre-wave point clouds
Walter Brescia, Laura Toni, Saverio Mascolo, Luca De Cicco
Signal Process. Image Commun.3
2026 MERINA+: Improving Generalization for Neural Video Adaptation via Information-Theoretic Meta-Reinforcement Learning
abstract
Adaptive bitrate (ABR) streaming is a popular technique used to improve the quality of experience (QoE) for users who watch videos online, which, for example, can provide a smoother video playback by dynamically adjusting the requested video quality with associated bitrate according to the constrained yet diverse network conditions. Recently, learning-based ABR algorithms have achieved a notable performance gain with lower inference overhead than the conventional heuristic or model-based baselines. However, their performance may degrade significantly in an unseen network environment with time-varying and heterogeneous throughput dynamics. For a better generalization, in this paper, we propose a meta-reinforcement learning (meta-RL)-based neural ABR algorithm that is able to quickly adapt its policy to these unseen throughput dynamics. Specifically, we propose a model-free system framework comprising an inference network and a policy network. The inference network infers distribution of the latent representation for underlying dynamics based on the recent throughout context, while the policy network is trained to quickly adapt to the changing throughout dynamics with the sampled latent representation. To effectively learn the inference network and meta-policy on mixed dynamics of the practical ABR scenarios, we further design a variational information bottleneck theory-based loss function for training the inference and policy networks, whose objective is to strike a trade-off between brevity of the latent representation and expressiveness of the meta-policy. We also derive a theoretically necessary condition for the bitrate versions that yield higher long-term QoE, based on which a dynamic action pruning strategy is further developed for practical implementation. This pruning strategy can not only prevent unsafe policy outputs in midst of unseen throughput dynamics, but may also reduce the computational complexity of model-based ABR algorithms. Finally, the meta-training and meta-adaptation procedures of our proposed algorithm are implemented across a range of throughput dynamics. The empirical evaluations on various datasets containing real-world network traces verify that our algorithm surpasses the state-of-the-art ABR algorithms, particularly in terms of the average chunk QoE and fast adaptation across out-of-distribution throughput traces.
Nuowen Kan, Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong, Laura Toni
IEEE Trans. Circuits Syst. Video Technol.7
2025 Heterogeneous Graph Structure Learning through the Lens of Data-generating Processes
abstract
Inferring the graph structure from observed data is a key task in graph machine learning to capture the intrinsic relationship between data entities. While significant advancements have been made in learning the structure of homogeneous graphs, many real-world graphs exhibit heterogeneous patterns where nodes and edges have multiple types. This paper fills this gap by introducing the first approach for heterogeneous graph structure learning (HGSL). To this end, we first propose a novel statistical model for the data-generating process (DGP) of heterogeneous graph data, namely hidden Markov networks for heterogeneous graphs (H2MN). Then we formalize HGSL as a maximum a-posterior estimation problem parameterized by such DGP and derive an alternating optimization method to obtain a solution together with a theoretical justification of the optimization conditions. Finally, we conduct extensive experiments on both synthetic and real-world datasets to demonstrate that our proposed method excels in learning structure on heterogeneous graphs in terms of edge type identification and edge weight recovery.
Keyue Jiang, Bohan Tang, Xiaowen Dong 0001, Laura Toni
AISTATS4
2025 Near-Optimal Sample Complexity in Reward-Free Kernel-based Reinforcement Learning
abstract
Reinforcement Learning (RL) problems are being considered under increasingly more complex structures. While tabular and linear models have been thoroughly explored, the analytical study of RL under non-linear function approximation, especially kernel-based models, has recently gained traction for their strong representational capacity and theoretical tractability. In this context, we examine the question of statistical efficiency in kernel-based RL within the reward-free RL framework, specifically asking: how many samples are required to design a near-optimal policy? Existing work addresses this question under restrictive assumptions about the class of kernel functions. We first explore this question assuming a generative model, then relax this assumption at the cost of increasing the sample complexity by a factor of $H$, the episode length. We tackle this fundamental problem using a broad class of kernels and a simpler algorithm compared to prior work. Our approach derives new confidence intervals for kernel ridge regression, specific to our RL setting, that may be of broader applicability. We further validate our theoretical findings through simulations.
Aya Kayal, Sattar Vakili, Laura Toni, Alberto Bernacchia
AISTATS3
2025 Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
abstract
Bayesian optimization (BO) with preference-based feedback has recently garnered significant attention due to its emerging applications. We refer to this problem as Bayesian Optimization from Human Feedback (BOHF), which differs from conventional BO by learning the best actions from a reduced feedback model, where only the preference between two actions is revealed to the learner at each time step. The objective is to identify the best action using a limited number of preference queries, typically obtained through costly human feedback. Existing work, which adopts the Bradley-Terry-Luce (BTL) feedback model, provides regret bounds for the performance of several algorithms. In this work, within the same framework we develop tighter performance guarantees. Specifically, we derive regret bounds of $\tilde{\mathcal{O}}(\sqrt{\Gamma(T)T})$, where $\Gamma(T)$ represents the maximum information gain—a kernel-specific complexity term—and $T$ is the number of queries. Our results significantly improve upon existing bounds. Notably, for common kernels, we show that the order-optimal sample complexities of conventional BO—achieved with richer feedback models—are recovered. In other words, the same number of preferential samples as scalar-valued samples is sufficient to find a nearly optimal solution.
Aya Kayal, Sattar Vakili, Laura Toni, Da-Shan Shiu, Alberto Bernacchia
ICML3
2025 NAVIX: Scaling MiniGrid Environments with JAX
abstract
As Deep Reinforcement Learning (Deep RL) research moves towards solving large-scale worlds, efficient environment simulations become crucial for rapid experimentation. However, most existing environments struggle to scale to high throughput, setting back meaningful progress. Interactions are typically computed on the CPU, limiting training speed and throughput, due to slower computation and communication overhead when distributing the task across multiple machines. Ultimately, Deep RL training is CPU-bound, and developing batched, fast, and scalable environments has become a frontier for progress. Among the most used Reinforcement Learning (RL) environments, MiniGrid is at the foundation of several studies on exploration, curriculum learning, representation learning, diversity, meta-learning, credit assignment, and language-conditioned RL, and still suffers from the limitations described above. In this work, we introduce NAVIX, a re-implementation of MiniGrid in JAX. NAVIX achieves over $160\,000\times$ speed improvements in batch mode, supporting up to 2048 agents in parallel on a single Nvidia A100 80 GB. This reduces experiment times from one week to 15 minutes, promoting faster design iterations and more scalable RL model development.
Eduardo Pignatelli, Jarek Liesen, Robert T. Lange, Chris Lu 0001, Pablo Samuel Castro, Laura Toni
NeurIPS6
2025 Effects of Dropout on Performance in Long-range Graph Learning Tasks
abstract
Message Passing Neural Networks (MPNNs) are a class of Graph Neural Networks (GNNs) that propagate information across the graph via local neighborhoods. The scheme gives rise to two key challenges: over-smoothing and over-squashing. While several Dropout-style algorithms, such as DropEdge and DropMessage, have successfully addressed over-smoothing, their impact on over-squashing remains largely unexplored. This represents a critical gap in the literature, as failure to mitigate over-squashing would make these methods unsuitable for long-range tasks – the intended use case of deep MPNNs. In this work, we study the aforementioned algorithms, and closely related edge-dropping algorithms – DropNode, DropAgg and DropGNN – in the context of over-squashing. We present theoretical results showing that DropEdge-variants reduce sensitivity between distant nodes, limiting their suitability for long-range tasks. To address this, we introduce DropSens, a sensitivity-aware variant of DropEdge, which is developed following the message-passing scheme of GCN. DropSens explicitly controls the proportion of information lost due to edge-dropping, thereby increasing sensitivity to distant nodes despite dropping the same number of edges. Our experiments on long-range synthetic and real-world datasets confirm the predicted limitations of existing edge-dropping and feature-dropping methods. Moreover, DropSens with GCN consistently outperforms graph rewiring techniques designed to mitigate over-squashing, suggesting that simple, targeted modifications can substantially improve a model’s ability to capture long-range interactions. Our conclusions highlight the need to re-evaluate and re-design existing methods for training deep GNNs, with a renewed focus on modelling long-range interactions. The code for reproducing the results in our this work is available at https://github.com/ignasa007/Dropout-Effects-GNNs.
Jasraj Singh, Keyue Jiang, Brooks Paige, Laura Toni
NeurIPS4
2025 The impact of intrinsic rewards on exploration in Reinforcement Learning
abstract
Abstract One of the open challenges in Reinforcement Learning (RL) is the hard exploration problem in sparse reward environments. Various types of intrinsic rewards have been proposed to address this challenge by pushing toward diversity. This diversity might be imposed at different levels, favoring the agent to explore different states, policies, or behaviors (State, Policy, and Skill level diversity, respectively). However, the impact of diversity on the agent’s behavior remains unclear. In this work, we aim to fill this gap by studying the effect of different levels of diversity imposed by intrinsic rewards on the exploration patterns of RL agents. We select four intrinsic rewards (State Count, Intrinsic Curiosity Module (ICM), Maximum Entropy, and Diversity is All You Need (DIAYN)), each pushing for a different diversity level. We conduct an empirical study on MiniGrid environments to compare their impact on exploration considering various metrics related to the agent’s exploration, namely: episodic return, observation coverage, agent’s position coverage, policy entropy, and timeframes to reach the sparse reward. The main outcome of the study is that State Count leads to the best exploration performance in the case of low-dimensional observations. However, in the case of RGB observations, the performance of State Count is highly degraded mainly due to representation learning challenges. Conversely, Maximum Entropy is less impacted, resulting in a more robust exploration, despite not always being optimal. Lastly, our empirical study revealed that learning diverse skills with DIAYN, often linked to improved robustness and generalization, does not promote exploration in MiniGrid environments. This is because: (i) Learning the skill space itself can be challenging, and (ii) exploration within the skill space prioritizes differentiating between behaviors rather than achieving uniform state visitation.
Aya Kayal, Eduardo Pignatelli, Laura Toni
Neural Comput. Appl.3
2025 A Clustering Approach to Unveil User Similarities in 6 df Extended Reality Applications
abstract
The advent in our daily life of Extended Reality (XR) technologies, such as Virtual and Augmented Reality, has led to the rise of user-centric systems, offering higher level of interaction and presence in virtual environments. In this context, understanding the actual interactivity of users is still an open challenge and a key step to enabling user-centric system. In this work, our goal is to construct an efficient clustering tool for 6 df navigation trajectories by extending the applicability of existing behavioural tool. Specifically, we first compare the navigation in 6 df with its 3 df counterpart, highlighting the main differences and novelties. Then, we investigate new metrics aimed at better modelling behavioural similarities between users in a 6 df system. More concretely, we define and compare 11 similarity metrics which are based on different distance features (i.e., user positions in the 3D space, user viewing directions) and distance measurements (i.e., Euclidean, Geodesic, angular distance). Our solutions are validated and tested on real navigation paths of users interacting with dynamic volumetric media in both 6 df Virtual Reality and Augmented Reality conditions. Results show that metrics based on both user position and viewing direction better perform in detecting user similarity while navigating in a 6 df system. Such easy-to-use but robust metrics allow us to answer a fundamental question for user-centric systems: ‘How do we detect if users look at the same content in 6 df?’, opening the gate to new solutions based on users interactivity, such as viewport prediction, live streaming services optimised based on users behaviour but also for user-based quality assessment methods.
Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César
ACM Trans. Multim. Comput. Commun. Appl.3
2024 MilliNoise: a Millimeter-wave Radar Sparse Point Cloud Dataset in Indoor Scenarios
abstract
Millimeter-wave (mmWave) radar sensors produce Point Clouds (PCs) that are much sparser and noisier than other PC data (e.g., Li-DAR), yet they are more robust in challenging conditions such as in the presence of fog, dust, smoke, or rain. This paper presents MilliNoise, a point cloud dataset captured in indoor scenarios through a mmWave radar sensor installed on a wheeled mobile robot. Each of the 12M points in the MilliNoise dataset is accurately labeled as true/noise point by leveraging known information of the scenes and a motion capture system to obtain the ground truth position of the moving robot. Each frame is carefully pre-processed to produce a fixed number of points for each cloud, enabling classification tools which require data with a fixed shape. Moreover, MilliNoise has been post-processed by labeling each point with the distance to its closest obstacle in the scene, which allows casting the denoising task into the regression framework. Along with the dataset, we provide researchers with the tools to visualize the data and prepare it for statistical and machine learning analysis. MilliNoise is available at: https://github.com/c3lab/MilliNoise
Walter Brescia, Laura Toni, Saverio Mascolo, Luca De Cicco
MMSys3
2024 Information-Theoretic Characterizations of Generalization Error for the Gibbs Algorithm
abstract
Various approaches have been developed to upper bound the generalization error of a supervised learning algorithm. However, existing bounds are often loose and even vacuous when evaluated in practice. As a result, they may fail to characterize the exact generalization ability of a learning algorithm. Our main contributions are exact characterizations of the expected generalization error of the well-known Gibbs algorithm (a.k.a. Gibbs posterior) using different information measures, in particular, the symmetrized KL information between the input training samples and the output hypothesis. Our result can be applied to tighten existing expected generalization errors and PAC-Bayesian bounds. Our information-theoretic approach is versatile, as it also characterizes the generalization error of the Gibbs algorithm with a data-dependent regularizer and that of the Gibbs algorithm in the asymptotic regime, where it converges to the standard empirical risk minimization algorithm. Of particular relevance, our results highlight the role the symmetrized KL information plays in controlling the generalization error of the Gibbs algorithm.
Gholamali Aminian, Yuheng Bu, Laura Toni, Miguel R. D. Rodrigues, Gregory W. Wornell
IEEE Trans. Inf. Theory3
2024 AGAR - Attention Graph-RNN for Adaptative Motion Prediction of Point Clouds of Deformable Objects
abstract
This article focuses on motion prediction for point cloud sequences in the challenging case of deformable 3D objects, such as human body motion. First, we investigate the challenges caused by deformable shapes and complex motions present in this type of representation, with the ultimate goal of understanding the technical limitations of state-of-the-art models. From this understanding, we propose an improved architecture for point cloud prediction of deformable 3D objects. Specifically, to handle deformable shapes, we propose a graph-based approach that learns and exploits the spatial structure of point clouds to extract more representative features. Then, we propose a module able to combine the learned features in aadaptativemanner according to the point cloud movements. The proposed adaptative module controls the composition of local and global motions for each point, enabling the network to model complex motions in deformable 3D objects more effectively. We tested the proposed method on the following datasets: MNIST moving digits, theMixamohuman bodies motions [ 15 ], JPEG [ 5 ] and CWIPC-SXR [ 32 ] real-world dynamic bodies. Simulation results demonstrate that our method outperforms the current baseline methods given its improved ability to model complex movements as well as preserve point cloud shape. Furthermore, we demonstrate the generalizability of the proposed framework for dynamic feature learning by testing the framework for action recognition on the MSRAction3D dataset [ 19 ] and achieving results on par with state-of-the-art methods.
Silvia Rossi 0001, Laura Toni
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Extending 3-DoF Metrics to Model User Behaviour Similarity in 6-DoF Immersive Applications
abstract
Immersive reality technologies, such as Virtual and Augmented Reality, have ushered a new era of user-centric systems, in which every aspect of the coding-delivery-rendering chain is tailored to the interaction of the users. Understanding the actual interactivity and behaviour of the users is still an open challenge and a key step to enabling such a user-centric system. Our main goal is to extend the applicability of existing behavioural methodologies for studying user navigation in the case of 6 Degree-of-Freedom (DoF). Specifically, we first compare the navigation in 6-DoF with its 3-DoF counterpart highlighting the main differences and novelties. Then, we define new metrics aimed at better modelling behavioural similarities between users in a 6-DoF system. We validate and test our solutions on real navigation paths of users interacting with dynamic volumetric media in 6-DoF Virtual Reality conditions. Our results show that metrics that consider both user position and viewing direction better perform in detecting user similarity while navigating in a 6-DoF system. Having easy-to-use but robust metrics that underpin multiple tools and answer the question "how do we detect if two users look at the same content?" open the gate to new solutions for a user-centric system.
Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César
MMSys3
2023 Online Network Source Optimization with Graph-Kernel MAB
Laura Toni, Pascal Frossard
ECML/PKDD (3)1
2023 MiDi: Mixed Graph and 3D Denoising Diffusion for Molecule Generation
Clément Vignac, Nagham Osman, Laura Toni, Pascal Frossard
ECML/PKDD (2)3
2023 Bootstrapped Personalized Popularity for Cold Start Recommender Systems
abstract
Recommender Systems are severely hampered by the well-known Cold Start problem, identified by the lack of information on new items and users. This has led to research efforts focused on data imputation and augmentation models as predominantly data pre-processing strategies, yet their improvement of cold-user performance is largely indirect and often comes at the price of a reduction in accuracy for warmer users. To address these limitations, we propose Bootstrapped Personalized Popularity (B2P), a novel framework that improves performance for cold users (directly) and cold items (implicitly) via popularity models personalized with item metadata. B2P is scalable to very large datasets and directly addresses the Cold Start problem, so it can complement existing Cold Start strategies. Experiments on a real-world dataset from the BBC iPlayer and a public dataset demonstrate that B2P (1) significantly improves cold-user performance, (2) boosts warm-user performance for bootstrapped models by lowering their training sparsity, and (3) improves total recommendation accuracy at a competitive diversity level relative to existing high-performing Collaborative Filtering models. We demonstrate that B2P is a powerful and scalable framework for strongly cold datasets.
Iason Chaimalas, Duncan Martin Walker, Edoardo Gruppi, Benjamin Richard Clark, Laura Toni
RecSys5
2022 An Information-theoretical Approach to Semi-supervised Learning under Covariate-shift
abstract
A common assumption in semi-supervised learning is that the labeled, unlabeled, and test data are drawn from the same distribution. However, this assumption is not satisfied in many applications. In many scenarios, the data is collected sequentially (e.g., healthcare) and the distribution of the data may change over time often exhibiting so-called covariate shifts. In this paper, we propose an approach for semi-supervised learning algorithms that is capable of addressing this issue. Our framework also recovers some popular methods, including entropy minimization and pseudo-labeling. We provide new information-theoretical based generalization error upper bounds inspired by our novel framework. Our bounds are applicable to both general semi-supervised learning and the covariate-shift scenario. Finally, we show numerically that our method outperforms previous approaches proposed for semi-supervised learning under the covariate shift.
Gholamali Aminian, Mahed Abroshan, Mohammad Mahdi Khalili, Laura Toni, Miguel R. D. Rodrigues
AISTATS4
2022 Characterizing and Understanding the Generalization Error of Transfer Learning with Gibbs Algorithm
abstract
We provide an information-theoretic analysis of the generalization ability of Gibbs-based transfer learning algorithms by focusing on two popular empirical risk minimization (ERM) approaches for transfer learning, $\alpha$-weighted-ERM and two-stage-ERM. Our key result is an exact characterization of the generalization behavior using the conditional symmetrized Kullback-Leibler (KL) information between the output hypothesis and the target training samples given the source training samples. Our results can also be applied to provide novel distribution-free generalization error upper bounds on these two aforementioned Gibbs algorithms. Our approach is versatile, as it also characterizes the generalization errors and excess risks of these two Gibbs algorithms in the asymptotic regime, where they converge to the $\alpha$-weighted-ERM and two-stage-ERM, respectively. Based on our theoretical results, we show that the benefits of transfer learning can be viewed as a bias-variance trade-off, with the bias induced by the source distribution and the variance induced by the lack of target samples. We believe this viewpoint can guide the choice of transfer learning algorithms in practice.
Yuheng Bu, Gholamali Aminian, Laura Toni, Gregory W. Wornell, Miguel R. D. Rodrigues
AISTATS3
2022 M4MM '22: 1st International Workshop on Methodologies for Multimedia
abstract
Transversely to all multimedia research, the methods and tools are often shared by several MM topics and applications. M4MM aims to foster discussion around fundamental methodologies and tools that are used in various multimedia research topics. To that aim, the technical program of the workshop consists of two keynote speakers, on complementary and very relevant methods and tools for MM, as well as four technical papers. The complete M4MM'22 workshop proceedings are available at: https://dl.acm.org/doi/proceedings/10.1145/3552487
Xavier Alameda-Pineda, Qin Jin, Vincent Oria, Laura Toni
ACM Multimedia4
2022 Explaining Hierarchical Features in Dynamic Point Cloud Processing
abstract
This paper aims at bringing some light and understanding to the field of deep learning for dynamic point cloud processing. Specifically, we focus on the hierarchical features learning aspect, with the ultimate goal of understanding which features are learned at the different stages of the process and what their meaning is. Last, we bring clarity on how hierarchical components of the network affect the learned features and their importance for a successful learning model. This study is conducted for point cloud prediction tasks, useful for predicting coding applications.
Silvia Rossi 0001, Laura Toni
PCS3
2021 Spatio-Temporal Graph-RNN for Point Cloud Prediction
abstract
In this paper, we propose an end-to-end learning network to predict future frames in a point cloud sequence. As main novelty, an initial layer learns topological information of point clouds as geometric features, to form representative spatiotemporal neighborhoods. This module is followed by multiple Graph-RNN cells. Each cell learns point dynamics (i.e., RNN states) by processing each point jointly with its spatiotemporal neighbours. We tested the network performance with a MNIST dataset of moving digits, a synthetic human bodies motions, and JPEG dynamic bodies datasets. Simulation results demonstrate that our method outperforms baseline ones, which neglect geometry features information.
Silvia Rossi 0001, Laura Toni
ICIP3
2021 A New Challenge: Behavioural Analysis Of 6-DOF User When Consuming Immersive Media
abstract
Thanks to recent advances in computer graphics, wearable technology, and connectivity, Virtual Reality (VR) has landed in our daily life. A key novelty in VR is the role of the user, which has turned from merely passive to entirely active. Thus, improving any aspect of the coding-delivery-rendering chain starts with the need for understanding user behaviour. To do so, we investigate the navigation trajectories of users within a 6-Degrees-of-Freedom (DoF) VR environment. Specifically, we investigate the main differences and similarities between 3 and 6-DoF navigation through existing methodologies adopted to study user behaviour in 3-DoF settings. Our simulation results, based on real navigation paths of users while displaying dynamic volumetric media in 6-DoF conditions, show the limitations of clustering algorithms for 3-DoF in assessing user similarity in 6-DoF. Given these observations, we state the need for developing new solutions for the analysis of 6-DoF trajectories.
Silvia Rossi 0001, Irene Viola 0001, Laura Toni, Pablo César
ICIP3
2021 Information-Theoretic Bounds on the Moments of the Generalization Error of Learning Algorithms
abstract
Generalization error bounds are critical to understanding the performance of machine learning models. In this work, building upon a new bound of the expected value of an arbitrary function of the population and empirical risk of a learning algorithm, we offer a more refined analysis of the generalization behaviour of a machine learning models based on a characterization of (bounds) to their generalization error moments. We discuss how the proposed bounds - which also encompass new bounds to the expected generalization error - relate to existing bounds in the literature. We also discuss how the proposed generalization error moment bounds can be used to construct new generalization error high-probability bounds.
Gholamali Aminian, Laura Toni, Miguel R. D. Rodrigues
ISIT2
2021 An Exact Characterization of the Generalization Error for the Gibbs Algorithm
abstract
Various approaches have been developed to upper bound the generalization error of a supervised learning algorithm. However, existing bounds are often loose and lack of guarantees. As a result, they may fail to characterize the exact generalization ability of a learning algorithm.Our main contribution is an exact characterization of the expected generalization error of the well-known Gibbs algorithm (a.k.a. Gibbs posterior) using symmetrized KL information between the input training samples and the output hypothesis. Our result can be applied to tighten existing expected generalization error and PAC-Bayesian bounds. Our approach is versatile, as it also characterizes the generalization error of the Gibbs algorithm with data-dependent regularizer and that of the Gibbs algorithm in the asymptotic regime, where it converges to the empirical risk minimization algorithm. Of particular relevance, our results highlight the role the symmetrized KL information plays in controlling the generalization error of the Gibbs algorithm.
Gholamali Aminian, Yuheng Bu, Laura Toni, Miguel R. D. Rodrigues, Gregory W. Wornell
NeurIPS3
2020 Laplacian-Regularized Graph Bandits: Algorithms and Theoretical Analysis
abstract
We consider a stochastic linear bandit problem with multiple users, where the relationship between users is captured by an underlying graph and user preferences are represented as smooth signals on the graph. We introduce a novel bandit algorithm where the smoothness prior is imposed via the random-walk graph Laplacian, which leads to a single-user cumulative regret scaling as $\Tilde{\mathcal{O}}(\Psi d \sqrt{T})$ with time horizon $T$, feature dimensionality $d$, and the scalar parameter $\Psi \in (0,1)$ that depends on the graph connectivity. This is an improvement over $\Tilde{\mathcal{O}}(d \sqrt{T})$ in \algo{LinUCB} \Ccite{li2010contextual}, where user relationship is not taken into account.In terms of network regret (sum of cumulative regret over $n$ users), the proposed algorithm leads to a scaling as $\Tilde{\mathcal{O}}(\Psi d\sqrt{nT})$, which is a significant improvement over $\Tilde{\mathcal{O}}(nd\sqrt{T})$ in the state-of-the-art algorithm \algo{Gob.Lin} \Ccite{cesa2013gang}. To improve scalability, we further propose a simplified algorithm with a linear computational complexity with respect to the number of users, while maintaining the same regret. Finally, we present a finite-time analysis on the proposed algorithms, and demonstrate their advantage in comparison with state-of-the-art graph-based bandit algorithms on both synthetic and real-world data.
Kaige Yang, Laura Toni, Xiaowen Dong 0001
AISTATS2
2020 Jensen-Shannon Information Based Characterization of the Generalization Error of Learning Algorithms
abstract
Generalization error bounds are critical to understanding the performance of machine learning models. In this work, we propose a new information-theoretic based generalization error upper bound applicable to supervised learning scenarios. We show that our general bound can specialize in various previous bounds. We also show that our general bound can be specialized under some conditions to a new bound involving the Jensen-Shannon information between a random variable modelling the set of training samples and another random variable modelling the hypothesis. We also prove that our bound can be tighter than mutual information-based bounds under some conditions.
Gholamali Aminian, Laura Toni, Miguel R. D. Rodrigues
ITW2
2020 Large Database Compression Based on Perceived Information
abstract
Lossy compression algorithms trade bits for quality, aiming at reducing as much as possible the bitrate needed to represent the original source (or set of sources), while preserving the source quality. In this letter, we propose a novel paradigm of compression algorithms, aimed at minimizing the information loss perceived by the final user instead of the actual source quality loss, under compression rate constraints. As main contributions, we first introduce the concept of perceived information (PI), which reflects the information perceived by a given user experiencing a data collection, and which is evaluated as the volume spanned by the sources features in a personalized latent space. We then formalize the rate-PI optimization problem and propose an algorithm to solve this compression problem. Finally, we validate our algorithm against benchmark solutions with simulation results, showing the gain in taking into account users' preferences while also maximizing the perceived information in the feature domain.
Thomas Maugey, Laura Toni
IEEE Signal Process. Lett.2
2020 Do Users Behave Similarly in VR? Investigation of the User Influence on the System Design
abstract
With the overarching goal of developing user-centric Virtual Reality (VR) systems, a new wave of studies focused on understanding how users interact in VR environments has recently emerged. Despite the intense efforts, however, current literature still does not provide the right framework to fully interpret and predict users’ trajectories while navigating in VR scenes. This work advances the state-of-the-art on both the study of users’ behaviour in VR and the user-centric system design. In more detail, we complement current datasets by presenting a publicly available dataset that provides navigation trajectories acquired for heterogeneous omnidirectional videos and different viewing platforms—namely, head-mounted display, tablet, and laptop. We then present an exhaustive analysis on the collected data to better understand navigation in VR across users, content, and, for the first time, across viewing platforms. The novelty lies in the user-affinity metric, proposed in this work to investigate users’ similarities when navigating within the content. The analysis reveals useful insights on the effect of device and content on the navigation, which could be precious considerations from the system design perspective. As a case study of the importance of studying users’ behaviour when designing VR systems, we finally propose a user-centric server optimisation. We formulate an integer linear program that seeks the best stored set of omnidirectional content that minimises encoding and storage cost while maximising the user’s experience. This is posed while taking into account network dynamics, type of video content, and also user population interactivity. Experimental results prove that our solution outperforms common company recommendations in terms of experienced quality but also in terms of encoding and storage, achieving a savings up to 70%. More importantly, we highlight a strong correlation between the storage cost and the user-affinity metric, showing the impact of the latter in the system architecture design.
Silvia Rossi 0001, Cagri Ozcinar, Aljoscha Smolic, Laura Toni
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Introduction to the Best Papers from the ACM Multimedia Systems (MMSys) 2019 and Co-Located Workshops
abstract
No abstract available.
Michael Zink, Laura Toni, Ali C. Begen
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Representation Learning on Graphs: A Reinforcement Learning Application
abstract
In this work, we study value function approximation in reinforcement learning (RL) problems with high dimensional state or action spaces via a generalized version of representation policy iteration (RPI). We consider the limitations of proto-value functions (PVFs) at accurately approximating the value function in low dimensions and we highlight the importance of features learning for an improved low-dimensional value function approximation. Then, we adopt different representation learning algorithms on graphs to learn the basis functions that best represent the value function. We empirically show that node2vec, an algorithm for scalable feature learning in networks, and Graph Auto-Encoder constantly outperform the commonly used smooth proto-value functions in low-dimensional feature space.
Sephora Madjiheurem, Laura Toni
AISTATS2
2019 Spherical Clustering of Users Navigating 360° Content
abstract
In Virtual Reality (VR) applications, understanding how users explore the omnidirectional content is important to optimize content creation, to develop user-centric services, or even to detect disorders in medical applications. Clustering users based on their common navigation patterns is a first direction to understand users behavior. However, classical clustering techniques fail in identifying this common paths, since they are usually focused on minimizing a simple distance metric. In this paper, we argue that minimizing the distance metric does not necessarily guarantee to identify users that experience similar navigation path in the VR domain. Therefore, we propose a graph-based method to identify clusters of users who are attending the same portion of the spherical content over time. The proposed solution takes into account the spherical geometry of the content and aims at clustering users based on the actual overlap of displayed content among users. Our method is tested on real VR user navigation patterns. Results show that our solution leads to clusters in which at least 85% of the content displayed by one user is shared among the other users belonging to the same cluster.
Silvia Rossi 0001, Francesca De Simone, Pascal Frossard, Laura Toni
ICASSP4
2019 Streaming from a Moving Platform with Real-Time and Playback Distortion Constraints
abstract
Video streaming from remotely controlled moving platforms such as drones have stringent constraints in terms of delay. In some applications such videos have to provide real-time visual feedback to the pilot with an acceptable distortion while satisfying high-quality requirements at playback. Furthermore the output rate of the source encoder required to achieve a target distortion depends on the speed of the platform. Motivated by this, we consider a novel source model which takes the source speed into account and derive its rate-distortion region. A transmission strategy based on successive joint encoding, which efficiently takes the source correlation into account, is then considered for transmission over a block fading channel. Our numerical results show that such scheme largely enhances over an independent coding scheme in terms of on-line distortion while approaching the playback distortion performance of an optimal encoder as the group of pictures size grows.
Giuseppe Cocco, Laura Toni
ICC2
2019 Adaptive Streaming in Interactive Multiview Video Systems
abstract
Multiview applications endow final users with the possibility to freely navigate within 3D scenes with minimum-delay. A real feeling of scene navigation is enabled by transmitting multiple high-quality camera views, which can be used to synthesize additional virtual views to offer a smooth navigation. However, when network resources are limited, not all camera views can be sent at high quality. It is therefore important, yet challenging, to find the right tradeoff between coding artifacts (reducing the quality of camera views) and virtual synthesis artifacts (reducing the number of camera views sent to users). To this aim, we propose an optimal transmission strategy for interactive multiview HTTP adaptive streaming. We propose a problem formulation to select the optimal set of camera views that the client requests for downloading, such that the navigation quality experienced by the user is optimized while the bandwidth constraints are satisfied. We show that our optimization problem is NP-hard, and we therefore develop an optimal solution based on the dynamic programming algorithm with polynomial time complexity. To further simplify the deployment, we present a suboptimal greedy algorithm with effective performance and lower complexity. The proposed controller is evaluated in theoretical and realistic settings characterized by realistic network statistics estimation, buffer management, and server-side representation optimization. Simulation results show significant improvement in terms of navigation quality compared with alternative baseline multiview adaptation logic solutions.
Xue Zhang 0008, Laura Toni, Pascal Frossard, Yao Zhao 0001, Chunyu Lin
IEEE Trans. Circuits Syst. Video Technol.2
2018 Delay-Power-Rate-Distortion Optimization of Video Representations for Dynamic Adaptive Streaming
abstract
Dynamic adaptive streaming addresses user heterogeneity by providing multiple encoded representations at different rates and/or resolutions for the same video content. For delay-sensitive applications, such as live streaming, there is however a stringent requirement on the encoding delay, and usually the encoding power (or rate) budget is also limited by the computational (or storage) capacity of the server. It is therefore important, yet challenging, to optimally select the source coding parameters for each encoded representation in order to minimize the resource consumption while maintaining a high quality of experience for the users. To address this, we propose an optimization framework with an optimal representation selection problem for delay, power, and rate constrained adaptive video streaming. Then, by the optimal selection of source coding parameters for each selected representation, we maximize the overall expected user satisfaction, subject not only to the encoding rate constraint, but also to the delay and power constraints at the server. We formulate the proposed optimization problem as an integer linear program formulation to provide the performance upper bound, and as a submodular maximization problem with two knapsack constraints to develop a practically feasible algorithm. Simulation results show that the proposed weighted rate and power cost benefit greedy algorithm is able to achieve a near-optimal performance with very low time complexity. In addition, it can strike the best tradeoff both between the rate and power cost, and between the algorithm's performance and the delay requirements proposed by delay sensitive applications.
Laura Toni, Junni Zou, Hongkai Xiong, Pascal Frossard
IEEE Trans. Circuits Syst. Video Technol.2
2018 QoE-Driven Mobile Edge Caching Placement for Adaptive Video Streaming
abstract
Caching at mobile edge servers can smooth temporal traffic variability and reduce the service load of base stations in mobile video delivery. However, the assignment of multiple video representations to distributed servers is still a challenging question in the context of adaptive streaming, since any two representations from different videos or even from the same video will compete for the limited caching storage. Therefore, it is important, yet challenging, to optimally select the cached representations for each edge server in order to effectively reduce the service load of base station while maintaining a high quality of experience (QoE) for users. To address this, we study a QoE-driven mobile edge caching placement optimization problem for dynamic adaptive video streaming that properly takes into account the different rate-distortion (R-D) characteristics of videos and the coordination among distributed edge servers. Then, by the optimal caching placement of representations for multiple videos, we maximize the aggregate average video distortion reduction of all users while minimizing the additional cost of representation downloading from the base station, subject not only to the storage capacity constraints of the edge servers, but also to the transmission and initial startup delay constraints of the users. We formulate the proposed optimization problem as an integer linear program to provide the performance upper bound, and as a submodular maximization problem with a set of knapsack constraints to develop a practically feasible cost benefit greedy algorithm. The proposed algorithm has polynomial computational complexity and a theoretical lower bound on its performance. Simulation results further show that the proposed algorithm is able to achieve a near-optimal performance with very low time complexity. Therefore, the proposed optimization framework reveals the caching performance upper bound for general adaptive video streaming systems, while the proposed algorithm provides some design guidelines for the edge servers to select the cached representations in practice based on both the video popularity and content information.
Laura Toni, Junni Zou, Hongkai Xiong, Pascal Frossard
IEEE Trans. Multim.2
2017 Optimization of Joint Progressive Source and Channel Coding for MIMO Systems
abstract
The optimization of joint source and channel coding for a sequence of numerous progressive packets is a challenging problem. Further, the problem becomes more complicated if the space- time coding is also involved with the optimization in a multiple-input multiple-output (MIMO) system. This is because the number of ways of jointly assigning channel codes and space-time codes to progressive packets is much larger than that of solely assigning channel codes to the packets. This paper applies a parametric approach to address that complex joint optimization problem in a MIMO system. Employing the parametric distortion-rate function, the joint assignment of channel codes and space-time codes to the packets can be optimized in a packet-by- packet manner. As a result, the computational complexity of the optimization is exponentially reduced, compared to the exhaustive search. The numerical results show that the proposed method significantly improves the peak-signal-to-noise ratio performance of the rate-based optimal solution in a MIMO system.
Meesue Shin, Sang-Hyo Kim, Laura Toni, Seok-Ho Chang
GLOBECOM3
2017 Finite length performance of random MAC strategies
abstract
The Internet of Things (IoT) is fueling innovation in nearly every part of our lives. From smart homes, cars, and cities, the Internet of Things is creating a more convenient, secure, intelligent, and personalized experience. While for any final user this IoT vision is a substantial innovation step, for communication providers is a compelling thread with massive number of devices connected to the Internet. Multiple connected devices sharing common wireless resources might create interference if they access the channel simultaneously. Medium access control protocols generally regulate the access of the devices to the shared channel to limit signal interference. In particular, irregular repetition slotted ALOHA (IRSA) techniques can achieve high-throughput performance when interference cancellation methods are adopted to recover from collisions. In this work, we study the finite length performance of IRSA schemes by building on the analogy between successive interference cancellation and iterative belief-propagation on erasure channels. We use a novel combinatorial derivation based on the matrix-occupancy theory to compute the error probability and we validate our method with simulation results.
Konstantinos Dovelos, Laura Toni, Pascal Frossard
ICC2
2017 Optimized receiver control in interactive multiview video streaming systems
abstract
Multiview applications endow final users with the possibility to freely navigate within 3D scenes with minimum-delay. High-quality rendering of the scene is enabled by transmitting multiple high-quality camera views, which can be used to synthesize additional virtual views to offer a smooth navigation in the scene. When network resources are limited, the set of camera views needs to be properly selected by the client. The right tradeoff between coding artifacts (reducing the quality of camera views) and virtual synthesis artifacts (reducing the number of camera views sent to users) has to be optimized. Existing client adaptation logic strategies usually fail to properly consider the content characteristics and the client navigation properties in the view selection problem. We therefore propose an optimal representation selection for interactive multiview HTTP adaptive streaming (HAS), with a complete problem formulation to select the optimal set of camera views that optimize the navigation quality experienced by the user while satisfying the bandwidth constraints. We show that our optimization problem is NP-hard and develop an effective solution based on a dynamic programming algorithm with polynomial time complexity. Simulation results show significant navigation quality improvement compared to two baseline multiview adaptation logic solutions. This confirms that adaptation logics have to consider both video content and interactivity level of the user in the representation selection strategy.
Xue Zhang 0008, Laura Toni, Pascal Frossard, Yao Zhao 0001, Chunyu Lin
ICC2
2017 Navigation-aware adaptive streaming strategies for omnidirectional video
abstract
Virtual reality (VR) applications target high-quality and zero-latency scene navigation to provide users with a full-immersion sensation within a scene. From a network perspective, this requires transmission of the omnidirectional content in its entirety, at a high resolution, which is not always feasible in bandwidth-limited networks. In this work, we propose an optimal transmission strategy for virtual reality applications able to fulfill the bandwidth requirements, while optimizing the end-user quality experienced in the navigation. In further detail, we consider a tile-based coded content for adaptive streaming systems, and we propose a navigation-aware transmission strategy at the clientside (i.e., adaptation logic), which is able to optimize the rate at which each tile is downloaded. First, we introduce the viewport- quality as metric that reflects the quality of any portion of the sphere displayed by the end-user. Then, we cast the tile-rate optimization as an integer linear programming problem and show that the proposed solution achieves substantial quality gains when compared to state-of-the-art adaptation logic methods.
Silvia Rossi 0001, Laura Toni
MMSP2
2017 Optimal Representations for Adaptive Streaming in Interactive Multiview Video Systems
abstract
Interactive multiview video streaming (IMVS) services permit to remotely navigate within a 3D scene with an immersive experience. This is possible by transmitting a set of reference camera views (anchor views), which are used by the clients to freely navigate in the scene and possibly synthesize additional viewpoints of interest. From a networking perspective, the big challenge in IMVS systems is to deliver to each client the best set of anchor views that maximizes the navigation quality, minimizes the view-switching delay and yet satisfies the network constraints. Integrating adaptive streaming solutions in free-viewpoint systems offers a promising solution to deploy IMVS in large and heterogeneous scenarios, as long as the multiview video representations on the server are properly selected. Therefore, we propose to optimize the multiview data at the server by minimizing the overall resource requirements while offering a good navigation quality to the different users. We propose a representation set optimization problem for multiview adaptive streaming systems, and we show that it is NP-hard. Therefore, we introduce the concept of multiview navigation segment that permits to cast the video representation set selection as an integer linear programming problem with a bounded computational complexity. We then show that the proposed solution reduces the computational complexity, while preserving optimality in most of the 3D scenes. We finally provide simulation results for different classes of users and show the gain offered by an optimal multiview video representation selection compared to recommended representation sets (e.g., Netflix and Apple ones) or to a baseline representation selection algorithm, where the encoding parameters are decided a priori for all the camera views.
Laura Toni, Pascal Frossard
IEEE Trans. Multim.1
2017 Improved Utility-Based Congestion Control for Delay-Constrained Communication
abstract
Due to the presence of buffers in the inner network nodes, each congestion event leads to buffer queueing and thus to an increasing end-to-end delay. In the case of delay sensitive applications, a large delay might not be acceptable and a solution to properly manage congestion events while maintaining a low end-to-end delay is required. Delay-based congestion algorithms are a viable solution as they target to limit the experienced end-to-end delay. Unfortunately, they do not perform well when sharing the bandwidth with congestion control algorithms not regulated by delay constraints (e.g., loss-based algorithms). Our target is to fill this gap, proposing a novel congestion control algorithm for delay-constrained communication over best effort packet switched networks. The proposed algorithm is able to maintain a bounded queueing delay when competing with other delay-based flows, and avoid starvation when competing with loss-based flows. We adopt the well-known price-based distributed mechanism as congestion control, but: 1) we introduce a novel non-linear mapping between the experienced delay and the price function and 2) we combine both delay and loss information into a single price term based on packet interarrival measurements. We then provide a stability analysis for our novel algorithm and we show its performance in the simulation results carried out in the NS3 framework. Simulation results demonstrate that the proposed algorithm is able to: achieve good intra-protocol fairness properties, control efficiently the end-to-end delay, and finally, protect the flow from starvation when other flows cause the queuing delay to grow excessively.
Stefano D'Aronco, Laura Toni, Sergio Mena, Pascal Frossard
IEEE/ACM Trans. Netw.2
2016 Price-Based Controller for Quality-Fair HTTP Adaptive Streaming
abstract
HTTP adaptive streaming (HAS) has become the universal technology for video streaming over the Internet. Many HAS system designs aim at sharing the network bandwidth in a rate-fair manner. However, rate fairness is in general not equivalent to quality fairness as different video sequences might have different characteristics and resource requirements. In this work, we focus on this limitation and propose a novel controller for HAS clients that is able to reach quality fairness while preserving the main characteristics of HAS systems and with a limited support from the network devices. In particular, we adopt a price-based mechanism in order to build a controller that maximizes the aggregate video quality for a set of HAS clients that share a common bottleneck. When network resources are scarce, the clients with simple video sequences reduce the requested bitrate in favor of users that subscribe to more complex video sequences, leading to a more efficient network usage. The proposed controller has been implemented in a network simulator, and the simulation results demonstrate its ability to share the available bandwidth among the HAS users in a quality-fair manner.
Stefano D'Aronco, Laura Toni, Pascal Frossard
ISM2
2016 Online learning adaptation strategy for DASH clients
abstract
In this work, we propose an online adaptation logic for Dynamic Adaptive Streaming over HTTP (DASH) clients, where each client selects the representation that maximize the long term expected reward. The latter is defined as a combination of the decoded quality, the quality fluctuations and the rebuffering events experienced by the user during the playback. To solve this problem, we cast a Markov Decision Process (MDP) optimization for the selection of the optimal representations. System dynamics required in the MDP model are a priori unknown and are therefore learned through a Reinforcement Learning (RL) technique. The developed learning process exploits a parallel learning technique that improves the learning rate and limits sub-optimal choices, leading to a fast and yet accurate learning process that quickly converges to high and stable rewards. Therefore, the efficiency of our controller is not sacrificed for fast convergence. Simulation results show that our algorithm achieves a higher QoE than existing RL algorithms in the literature as well as heuristic solutions, as it is able to increase average QoE and reduce quality fluctuations.
Federico Chiariotti, Stefano D'Aronco, Laura Toni, Pascal Frossard
MMSys3
2016 A comparative study of DASH representation sets using real user characteristics
abstract
Adaptive streaming strategies over HTTP allow to serve heterogeneous video users with varying demands. By providing different encoded versions (representations) of each video sequence on the server, clients have the freedom to select a representation that best fits their needs. While the topic of selecting a representation based on a pre-defined set is covered very well in the literature, the problem of how to properly select the representation set stored at the main server is usually an overlooked challenge. In this work, we provide an analysis on how the choice of representations on the server impacts the clients' quality. This is achieved by conducting NS-3 based simulations with a total of 10k users and up to 300 concurrent DASH clients for several recommended sets (e.g., Netflix, YouTube, and Apple), and measuring the experienced quality over a timespan of 24 hours. The results show that under heavy load (at peak hours) there is still room for improvement.
Christian Kreuzberger, Benjamin Rainer, Hermann Hellwagner, Laura Toni, Pascal Frossard
NOSSDAV4
2016 Complexity constrained representation selection for dynamic adaptive streaming
abstract
In this paper, we propose a representation selection optimization problem for complexity constrained adaptive video streaming that properly takes into account the different complexity-rate-distortion (C-R-D) characteristics of the videos when implementing rate control for desired representations. Our objective is to maximize the expected video distortion reduction of users, subject not only to encoding rate constraints, but also to complexity constraints. We prove that our optimization problem is a submodular maximization problem with two knapsack constraints. A weighted rate and complexity cost benefit greedy algorithm is then developed to obtain an approximate solution with polynomial time complexity and good approximation performance in simulations.
Laura Toni, Pascal Frossard, Hongkai Xiong, Junni Zou
VCIP2
2016 In-Network View Synthesis for Interactive Multiview Video Systems
abstract
In multiview applications, camera views can be used as reference views to synthesize additional virtual viewpoints, allowing users to freely navigate within a 3D scene. However, bandwidth constraints may restrict the number of reference views sent to clients, limiting the quality of the synthesized viewpoints. In this work, we study the problem of in-network reference view synthesis aimed at improving the navigation quality at the clients. We consider a distributed cloud network architecture, where data stored in a main cloud is delivered to end users with the help of cloudlets, i.e., resource-rich proxies close to the users. We argue that, in case of limited bandwidth from the cloudlet to the users, re-sampling at the couldlet the viewpoints of the 3D scene (i.e., synthesizing novel virtual views in the cloudlets to be used as new references to the decoder) is beneficial compared to mere subsampling of the original set of camera views. We therefore cast a new reference view selection problem that seeks the subset of views minimizing the distortion over a view navigation window defined by the user under bandwidth constraints. We prove that the problem is NP-hard, and we propose an effective polynomial time algorithm using dynamic programming to solve the optimization problem under general assumptions that cover most of the multiview scenarios in practice. Simulation results confirm the performance gain offered by virtual view synthesis in the network.
Laura Toni, Gene Cheung, Pascal Frossard
IEEE Trans. Multim.1
2015 In-network view re-sampling for interactive free viewpoint video streaming
abstract
Interactive free viewpoint video offers the possibility for each user to independently choose the views of a 3D scene to be displayed at the decoder. The visual content is commonly represented by N texture and depth map pairs that capture different viewpoints. A server selects an appropriate subset of M ≤ N views for transmission, so that the user can freely navigate in the corresponding window of viewpoints without being affected by network delay. During navigation, a user can synthesize any intermediate virtual view image in the navigation window via depth-image-based rendering (DIBR) using two nearby camera views as references. When the available bandwidth is too small to transmit all camera views typically used to synthesize views in the navigation window, we propose to synthesize intermediate virtual views as new references for transmission - a resampling of viewpoints for the 3D scene - so that the synthesized view distortion within the navigation window is minimized. We formulate a combinatorial optimization problem to find the best set of M virtual views to synthesize as new references, and show that the problem is NP-hard. We approximate the original problem with a new reference view equivalence model and derive in this case an optimal dynamic programming algorithm to determine the best set of M views to be transmitted to each user. Experimental results show that synthesizing virtual views as new references for client-side view synthesis can outperform simple selection from camera views by up to 0.73dB in synthesized view quality.
Laura Toni, Gene Cheung, Pascal Frossard
ICIP1
2015 Optimal layered representation for adaptive interactive multiview video streaming
Ana De Abreu, Laura Toni, Nikolaos Thomos, Thomas Maugey, Fernando Pereira 0001, Pascal Frossard
J. Vis. Commun. Image Represent.2
2015 Prioritized Random MAC Optimization Via Graph-Based Analysis
abstract
Motivated by the analogy between successive interference cancellation and iterative belief-propagation on erasure channels, irregular repetition slotted ALOHA (IRSA) strategies have received a lot of attention in the design of medium access control protocols. In this work, we consider generic systems where sources in different importance classes compete for a common channel. We propose a new prioritized IRSA algorithm and derive the probability to correctly resolve collisions for data from each source class. We then make use of our theoretical analysis to formulate a new optimization problem for selecting the transmission strategies of heterogenous sources. We optimize both the replication probability per class and the source rate per class, in such a way that the overall system utility is maximized. We then propose a heuristic-based algorithm for the selection of the transmission strategy, which is built on intrinsic characteristics of the iterative decoding methods adopted for recovering from collisions. Experimental results validate the accuracy of the theoretical study and show the gain of well-chosen prioritized transmission strategies for transmission of data from heterogenous classes over shared wireless channels.
Laura Toni, Pascal Frossard
IEEE Trans. Commun.1
2015 Optimized Packet Scheduling in Multiview Video Navigation Systems
abstract
We study coding and transmission strategies in multicamera systems, where correlated sources send data through a bottleneck channel to a central server, which eventually transmits views to different interactive users. We propose a dynamic navigation -path aware packet scheduling optimization under delay, bandwidth, and interactivity constraints aimed at optimizing the quality-of-experience of interactive users. In particular , the scene distortion is minimized jointly with the distortion variations along most likely navigation paths. The optimization relies both on a novel rate-distortion model, which captures the importance of each view in the scene reconstruction , and on an objective function that optimizes resources based on a client navigation model. The latter takes into account the distortion experienced by interactive clients as well as the distortion variations that might be observed by clients during multiview navigation. We solve the scheduling problem with a novel trellis-based solution, which permits to formally decompose the multivariate optimization problem, thereby significantly reducing the computation complexity. Simulation results show the PSNR quality gain offered by the proposed algorithm compared to baseline scheduling policies. Finally, we show that the best scheduling policy consistently adapts to the most likely user navigation path and that it minimizes distortion variations that can be very disturbing for users in traditional navigation systems.
Laura Toni, Thomas Maugey, Pascal Frossard
IEEE Trans. Multim.1
2015 Optimal Selection of Adaptive Streaming Representations
abstract
Adaptive streaming addresses the increasing and heterogeneous demand of multimedia content over the Internet by offering several encoded versions for each video sequence. Each version (or representation) is characterized by a resolution and a bit rate, and it is aimed at a specific set of users, like TV or mobile phone clients. While most existing works on adaptive streaming deal with effective playout-buffer control strategies on the client side, in this article we take a providers' perspective and propose solutions to improve user satisfaction by optimizing the set of available representations. We formulate an integer linear program that maximizes users' average satisfaction, taking into account network dynamics, type of video content, and user population characteristics. The solution of the optimization is a set of encoding parameters corresponding to the representations set that maximizes user satisfaction. We evaluate this solution by simulating multiple adaptive streaming sessions characterized by realistic network statistics, showing that the proposed solution outperforms commonly used vendor recommendations, in terms of user satisfaction but also in terms of fairness and outage probability. The simulation results show that video content information as well as network constraints and users' statistics play a crucial role in selecting proper encoding parameters to provide fairness among users and to reduce network resource usage. We finally propose a few theoretical guidelines that can be used, in realistic settings, to choose the encoding parameters based on the user characteristics, the network capacity and the type of video content.
Laura Toni, Ramon Aparicio-Pardo, Karine Pires, Gwendal Simon, Alberto Blanc, Pascal Frossard
ACM Trans. Multim. Comput. Commun. Appl.1
2015 Resource Allocation and Performance Analysis for Multiuser Video Transmission Over Doubly Selective Channels
abstract
We consider an uplink multicarrier system with multiple video users who want to send compressed video data to the base station. In the time domain, we model the time-varying channel using Jakes' model, and in the frequency domain, each subcarrier is assumed to be independently fading. The video is scalably coded in units of a group of pictures (GOP), and users have different video rate distortion (RD) functions. At the beginning of the GOP, the base station collects both the RD information and the instantaneous channel state information (CSI) for subcarrier allocation purposes. We design a cross-layer resource allocation algorithm to assign subcarriers to users based on both the demand of the video and the quality of the channel. Once the resource allocation decision is made, the users then periodically adapt the modulation format of the subcarriers allocated according to the evolution of the CSI for the duration of the GOP. We show that our cross-layer resource allocation robustly outperforms two baseline algorithms, each of which uses only one layer of information for resource allocation.
Dawei Wang 0010, Laura Toni, Pamela C. Cosman, Laurence B. Milstein
IEEE Trans. Wirel. Commun.2
2014 Optimal set of video representations in adaptive streaming
abstract
Adaptive streaming addresses the increasing and heterogenous demand of multimedia content over the Internet by offering several streams for each video. Each stream has a different resolution and bit rate, aimed at a specific set of users, e.g., TV, mobile phone. While most existing works on adaptive streaming deal with optimal playout-control strategies at the client side, in this paper we concentrate on the providers' side, showing how to improve user satisfaction by optimizing the encoding parameters. We formulate an integer linear program that maximizes users' average satisfaction, taking into account the network characteristics, the type of video content, and the user population. The solution of the optimization is a set of encoding parameters that outperforms commonly used vendor recommendations, in terms of user satisfaction and total delivery cost. Results show that video content information as well as network constraints and users' statistics play a crucial role in selecting proper encoding parameters to provide fairness among users and reduce network usage. By combining patterns common to several representative cases, we propose a few practical guidelines that can be used to choose the encoding parameters based on the user base characteristics, the network capacity and the type of video content.
Laura Toni, Ramon Aparicio-Pardo, Gwendal Simon, Alberto Blanc, Pascal Frossard
MMSys1
2014 Multiview video representations for quality-scalable navigation
abstract
Interactive multiview video (IMV) applications offer to users the freedom of selecting their preferred viewpoint. Usually, in these systems texture and depth maps of captured views are available at the user side, as they permit the rendering of intermediate virtual views. However, the virtual views' quality depends on the distance to the available views used as references and on their quality, which is generally constrained by the heterogeneous capabilities of the users. In this context, this work proposes an IMV scalable system, where views are optimally organized in layers, each one offering an incremental improvement in the interactive navigation quality. We propose a distortion model for the rendered virtual views and an algorithm that selects the optimal views' subset per layer. Simulation results show the efficiency of the proposed distortion model, and that the careful choice of reference cameras permits to have a graceful quality degradation for clients with limited capabilities.
Ana De Abreu, Laura Toni, Thomas Maugey, Nikolaos Thomos, Pascal Frossard, Fernando Pereira 0001
VCIP2
2014 Packet scheduling in multicamera capture systems
abstract
In multiview video services, multiple cameras acquire the same scene from different perspectives, which results in correlated video streams. This generates large amounts of highly redundant data, which need to be properly handled during encoding and transmission of the multi-view data. In this work, we study coding and transmission strategies in multicamera sets, where correlated sources need to be sent to a central server through a bottleneck channel, and eventually delivered to interactive clients. We propose a dynamic correlation-aware packet scheduling optimization under delay, bandwidth, and interactivity constraints. A novel trellis-based solution permits to formally decompose the multivariate optimization problem, thereby significantly reducing the computation complexity. Simulation results show the gain of the proposed algorithm compared to baseline scheduling policies.
Laura Toni, Thomas Maugey, Pascal Frossard
VCIP1
2014 Correlation-Aware Packet Scheduling in Multi-Camera Networks
abstract
In multiview applications, multiple cameras acquire the same scene from different viewpoints and generally produce correlated video streams. This results in large amounts of highly redundant data. In order to save resources, it is critical to handle properly this correlation during encoding and transmission of the multiview data. In this work, we propose a correlation-aware packet scheduling algorithm for multi-camera networks, where information from all cameras are transmitted over a bottleneck channel to clients that reconstruct the multiview images. The scheduling algorithm relies on a new rate-distortion model that captures the importance of each view in the scene reconstruction. We propose a problem formulation for the optimization of the packet scheduling policies, which adapt to variations in the scene content. Then, we design a low complexity scheduling algorithm based on a trellis search that selects the subset of candidate packets to be transmitted towards effective multiview reconstruction at clients. Extensive simulation results confirm the gain of our scheduling algorithm when inter-source correlation information is used in the scheduler, compared to scheduling policies with no information about the correlation or non-adaptive scheduling policies. We finally show that increasing the optimization horizon in the packet scheduling algorithm improves the transmission performance, especially in scenarios where the level of correlation rapidly varies with time.
Laura Toni, Thomas Maugey, Pascal Frossard
IEEE Trans. Multim.1
2013 Interactive free viewpoint video streaming using prioritized network coding
abstract
In free viewpoint applications, the images are captured by an array of cameras that acquire a scene of interest from different perspectives. Any intermediate viewpoint not included in the camera array can be virtually synthesized by the decoder, at a quality that depends on the distance between the virtual view and the camera views available at decoder. Hence, it is beneficial for any user to receive camera views that are close to each other for synthesis. This is however not always feasible in bandwidth-limited overlay networks, where every node may ask for different camera views. In this work, we propose an optimized delivery strategy for free viewpoint streaming over overlay networks. We introduce the concept of layered quality-of-experience (QoE), which describes the level of interactivity offered to clients. Based on these levels of QoE, camera views are organized into layered subsets. These subsets are then delivered to clients through a prioritized network coding streaming scheme, which accommodates for the network and clients heterogeneity and effectively exploit the resources of the overlay network. Simulation results show that, in a scenario with limited bandwidth or channel reliability, the proposed method outperforms baseline network coding approaches, where the different levels of QoE are not taken into account in the delivery strategy optimization.
Laura Toni, Nikolaos Thomos, Pascal Frossard
MMSP1
2013 Uplink Resource Management for Multiuser OFDM Video Transmission Systems: Analysis and Algorithm Design
abstract
We consider a multiuser OFDM system in which users want to transmit videos via a base station. The base station knows the channel state information (CSI) as well as the rate distortion (RD) information of the video streams and tries to allocate power and spectrum resources to the users according to both physical layer CSI and application layer RD information. We derive and analyze a condition for the optimal resource allocation solution in a continuous frequency response setting. The optimality condition for this cross layer optimization scenario is similar to the equal slope condition for conventional video multiplexing resource allocation. Based on our analysis, we design an iterative subcarrier assignment and power allocation algorithm for an uplink system, and provide numerical performance analysis with different numbers of users. Comparing to systems with either only physical layer or only application layer information available at the base station, our results show that the user capacity and the video PSNR performance can be increased significantly by using cross layer design. Bit-level simulations which take into account the imperfection of the video coding rate control, the variation of RD curve fitting, as well as channel errors, are presented.
Dawei Wang 0010, Laura Toni, Pamela C. Cosman, Laurence B. Milstein
IEEE Trans. Commun.2
2012 Channel Coding Optimization Based on Slice Visibility for Transmission of Compressed Video over OFDM Channels
abstract
Optimization of multimedia transmissions over wireless channels should be aimed at maximizing the video quality perceived by the final user. For transmission of video sequences over an orthogonal frequency division multiplexing (OFDM) system in a slowly varying Rayleigh faded environment, we develop a cross-layer technique, based on a slice loss visibility (SLV) model used to evaluate the visual importance of each slice. In particular, taking into account the visibility scores available from the bitstream, depending on the scenario, we optimize the mapping of video slices within a 2-D time-frequency resource block and/or the channel code rates, in order to better protect more visually important slices. The proposed algorithm is investigated for several scenarios, with different levels of information about the channel available in the optimization process. Results demonstrate that, for different physical environments and different video sequences, the proposed algorithm outperforms baseline ones which do not take into account either the SLV or the CSI in the video transmission.
Laura Toni, Pamela C. Cosman, Laurence B. Milstein
IEEE J. Sel. Areas Commun.1
2011 On the Performances of a Target Detection Algorithm in Underwater Sensor Networks with Aqua-Sim
abstract
This paper is aimed at investigating the performance of a behavior-based controller for area coverage and target detection missions, in which multiple Autonomous Underwater Vehicles (AUVs) are employed. The main contribution of this work is the analysis of the performance improvement when a realistic underwater communication network is considered among robots, exploited in a realistic mission. In our model, we integrate the algorithm proposed in [1] and implemented in MATLAB with the Aqua-Sim simulator, a network simulator for underwater sensor network based on NS-2. Results show that the performance improvement is deeply related to the communication parameters, which have to be rightly chosen in order to find a good trade-off between energy consumption and algorithm efficiency.
Laura Sorbi, Graziano Pio De Capua, El Hadi Cherkaoui, Laura Toni, Jean-Guy Fontaine
EUC4
2011 Unequal error protection based on slice visibility for transmission of compressed video over OFDM channels
abstract
We address channel code rate optimization for transmission of non-scalable coded video sequences over orthogonal frequency division multiplexing networks. A slice loss visibility (SLV) model is used to evaluate the visual importance of each H.264 slice. Based on both the SLV model and the frequency diversity order available from the channel, we propose a cross-layer technique to allocate video slices within a 2-D time-frequency resource block, and optimize the unequal channel code rate profile, in order to better protect more visually important slices. The proposed algorithm outperforms baseline ones which do not take into account the SLV.
Laura Toni, Pamela C. Cosman, Laurence B. Milstein
ICME1
2011 Taking advantage of the diversity in wireless access networks: On the simulation of a user centric approach
abstract
“Always Best Connected” or simply ABC concept has been introduced to express the possibility for mobile users to experience with smartphone/computer a continuity of service at any place any time. In this context, the aim of the fourth generation of wireless networks is to not only support high speed connection but also implement ABC taking benefit of the numerous underlying wireless technologies. For that, smart-phones should implement sophisticated access network selection mechanism to take benefit of this diversity. In our previous works, we have used the utility theory to propose several utility functions that measures the value of each access network vs. the preferences of the end users and we have shown how these preferences can be used by the user terminal to select the most appropriate access network. In this paper, we extend that work with the implementation of the solution in a simulator of heterogeneous access networks and perform a set of simulations to highlight the value added of the proposed solution. The obtained results show similar results as those obtained analytically and confirm the validity of the approach for the end users and the operators.
El Hadi Cherkaoui, Nazim Agoulmine, Thinh Nguyen, Laura Toni, Jean-Guy Fontaine
Integrated Network Management4
2011 Joint source channel coding optimization for heterogeneous access networks in multiuser scenario
abstract
In a multiple access technologies scenario, aimed at providing the “always best connected” (ABC) service, the future fourth-generation wireless networks will be characterized by an increasing heterogeneity. When a multimedia transmission is considered, the ABC can be improved by the optimization of the source coding. In this work, we propose a joint source channel coding (JSCC) optimization technique, able to adapt the system parameters to channel variations, providing at the same time a high level of scalability of the encoded video bitstream. With the proposed JSCC method, we address the ABC issue in multi-user scenario affected by multiple interference. The results show that, despite the simplicity of the algorithm, the adaptive system experiences an improvement of the performance compared to fixed coding schemes.
Laura Toni, Lorenzo Rossi 0002, Nazim Agoulmine, Jean-Guy Fontaine
Integrated Network Management1
2011 Threshold Evaluation in Link Adaptation Schemes for Progressive Images Transmission
abstract
Link adaptation (LA) has been demonstrated to be a very effective technique to improve transmission reliability in wireless communication systems subject to fading channel fluctuations. Due to the diffusion in commercial systems, LA techniques should take into account realistic channel models, low complexity algorithms, and application requirements. In this work, after analyzing the adaptive modulation and coding (AMC) performance for progressive images transmission, we propose a modified AMC technique that jointly takes into account both the channel variations and the progressive nature of the encoded bitstream. Since different priorities are assigned to the bits by the source encoder, an unequal error protection AMC method, that enhances transmission reliability for high priority bits, is necessary. Rather than increasing the complexity of the channel coding and modulation techniques, we adopt classical AMC methods optimizing the thresholds for the selection of opportunist modulation and coding schemes. The results show that the proposed technique is robust even with non-ideal channel state information.
Laura Toni, Barbara M. Masini
VTC Fall1
2011 Does Fast Adaptive Modulation Always Outperform Slow Adaptive Modulation?
abstract
Link adaptation techniques are important modern and future wireless communication systems to cope with quality of service fluctuations in fading channels. These techniques require the knowledge of the channel state obtained with a portion of resources devoted to channel estimation instead of data and updated every coherence time of the process to be tracked. In this paper, we analyze fast and slow adaptive modulation systems with diversity and non-ideal channel estimation under energy constraints. The framework enables to address the following questions: (i) What is the impact of non-ideal channel estimation on fast and slow adaptive modulation systems? (ii) How to define a proper figure of merit which considers both resources dedicated to data and those to channel estimation? (iii) Does fast adaptive always outperform slow adaptive techniques? Our analysis shows that, despite the lower complexity and feedback rate, slow adaptive modulation (SAM) can achieve higher spectral efficiency than fast adaptive modulation (FAM) in the presence of energy constraint, diversity, and non-ideal channel estimation. In addition, SAM satisfies bit error outage requirements also in FAM-denied region.
Laura Toni, Andrea Conti 0001
IEEE Trans. Wirel. Commun.1
2010 Communication in a Behavior-based Approach to Target Detection and Tracking with Autonomous Underwater Vehicles
abstract
In this paper, we address the challenging topic of target detection and tracking in underwater environments, proposing a behavior-based algorithm suitable for both fixed and moving targets. Taking into account the stringent constraints of the acoustic underwater channel, we apply this method to fleets of collaborative AUVs. The resulting control strategy is able to benefit from the behavior-based approach and the cooperation among robots. Results show that the performance of the algorithm can be improved by tuning the range of vehicles communication, meeting the tradeoff between energy consumption and behavior efficiency.
Laura Sorbi, Laura Toni, Graziano Pio De Capua, Lorenzo Rossi 0002
EUC2
2009 Channel Coding for Progressive Images in a 2-D Time-Frequency OFDM Block With Channel Estimation Errors
abstract
Coding and diversity are very effective techniques for improving transmission reliability in a mobile wireless environment. The use of diversity is particularly important for multimedia communications over fading channels. In this work, we study the transmission of progressive image bitstreams using channel coding in a 2-D time-frequency resource block in an OFDM network, employing time and frequency diversities simultaneously. In particular, in the frequency domain, based on the order of diversity and the correlation of individual subcarriers, we construct symmetric n -channel FEC-based multiple descriptions using channel erasure codes combined with embedded image coding. In the time domain, a concatenation of RCPC codes and CRC codes is employed to protect individual descriptions. We consider the physical channel conditions arising from various coherence bandwidths and coherence times, leading to a range of orders of diversities available in the time and frequency domains. We investigate the effects of different error patterns on the delivered image quality due to various fade rates. We also study the tradeoffs and compare the relative effectiveness associated with the use of erasure codes in the frequency domain and convolutional codes in the time domain under different physical environments. Both the effects of intercarrier interference and channel estimation errors are included in our study. Specifically, the effects of channel estimation errors, frequency selectivity and the rate of the channel variations are taken into consideration for the construction of the 2-D time-frequency block. We provide results showing the gain that the proposed model achieves compared to a system without temporal coding. In one example, for a system experiencing flat fading, low Doppler, and imperfect CSI, we find that the increase in PSNR compared to a system without time diversity is as much as 9.4 dB.
Laura Toni, Yee Sin Chan, Pamela C. Cosman, Laurence B. Milstein
IEEE Trans. Image Process.1