EDBT 2026 Demo / reviewers in the wild / expert
Mattia Rossi
dblp:96/6402
· DBLP profile ↗
23ranked-venue papers
8as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Computer networks · 4 · 1 first-authorSystems, architecture and hardware · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Streamlined Attention-Based Network for Descriptor ExtractionabstractWe introduce SANDesc a Streamlined Attention-based Network for Descriptor extraction that aims to improve on existing architectures for keypoint description. Our descriptor network learns to compute descriptors that improve matching without modifying the underlying keypoint detector. We employ a revised U-Net-like architecture enhanced with Convolutional Block Attention Modules and residual paths, enabling effective local representation while maintaining computational efficiency. We refer to the building blocks of our model as Residual U-Net Blocks with Attention. The model is trained using a modified triplet loss in combination with a curriculum learning-inspired hard negative mining strategy, which improves training stability. Extensive experiments on HPatches, MegaDepth-1500, and the Image Matching Challenge 2021 show that training SANDESC on top of existing keypoint detectors leads to improved results on multiple matching tasks compared to the original keypoint descriptors. At the same time, SANDesc has a model complexity of just 2.4 million parameters. As a further contribution, we introduce a new urban dataset featuring 4K images and pre-calibrated intrinsics, designed to evaluate feature extractors. On this benchmark, SANDESC achieves substantial performance gains over the existing descriptors while operating with limited computational resources. Mattia D'Urso, Emanuele Santellani, Christian Sormann, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer |
3DV | 4 |
| 2025 | MultimodalStudio: A Heterogeneous Sensor Dataset and Framework for Neural Rendering across Multiple Imaging ModalitiesabstractNeural Radiance Fields (NeRF) have shown impressive performances in the rendering of 3D scenes from arbitrary viewpoints. While RGB images are widely preferred for training volume rendering models, the interest in other radiance modalities is also growing. However, the capability of the underlying implicit neural models to learn and transfer information across heterogeneous imaging modalities has seldom been explored, mostly due to the limited training data availability. For this purpose, we present MultimodalStudio (MMS): it encompasses MMS-DATA and MMS-FW. MMS-DATA is a multimodal multi-view dataset containing 32 scenes acquired with 5 different imaging modalities: RGB, monochrome, near-infrared, polarization and multi-spectral. MMS-FW is a novel modular multimodal NeRF framework designed to handle multimodal raw data and able to support an arbitrary number of multi-channel devices. Through extensive experiments, we demonstrate that MMS-FW trained on MMS-DATA can transfer information between different imaging modalities and produce higher quality renderings than using single modalities alone. We publicly release the dataset and the framework, to promote the research on multimodal volume rendering and beyond. Federico Lincetto, Gianluca Agresti, Mattia Rossi, Pietro Zanuttigh |
CVPR | 3 |
| 2024 | GMM-IKRS: Gaussian Mixture Models for Interpretable Keypoint Refinement and Scoring
Emanuele Santellani, Martin Zach, Christian Sormann, Mattia Rossi, Andreas Kuhn 0002, Friedrich Fraundorfer |
ECCV (77) | 4 |
| 2023 | Exploiting Multiple Priors for Neural 3D Indoor Reconstruction
Federico Lincetto, Gianluca Agresti, Mattia Rossi, Pietro Zanuttigh |
BMVC | 3 |
| 2023 | S-TREK: Sequential Translation and Rotation Equivariant Keypoints for local feature extractionabstractIn this work we introduce S-TREK, a novel local feature extractor that combines a deep keypoint detector, which is both translation and rotation equivariant by design, with a lightweight deep descriptor extractor. We train the S-TREK keypoint detector within a framework inspired by reinforcement learning, where we leverage a sequential procedure to maximize a reward directly related to keypoint repeatability. Our descriptor network is trained following a "detect, then describe" approach, where the descriptor loss is evaluated only at those locations where keypoints have been selected by the already trained detector. Extensive experiments on multiple benchmarks confirm the effectiveness of our proposed method, with S-TREK often outperforming other state-of-the-art methods in terms of repeatability and quality of the recovered poses, especially when dealing with in-plane rotations. Emanuele Santellani, Christian Sormann, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer |
ICCV | 3 |
| 2023 | DELS-MVS: Deep Epipolar Line Search for Multi-View StereoabstractWe propose a novel approach for deep learning-based Multi-View Stereo (MVS). For each pixel in the reference image, our method leverages a deep architecture to search for the corresponding point in the source image directly along the corresponding epipolar line. We denote our method DELS-MVS: Deep Epipolar Line Search Multi-View Stereo. Previous works in deep MVS select a range of interest within the depth space, discretize it, and sample the epipolar line according to the resulting depth values: this can result in an uneven scanning of the epipolar line, hence of the image space. Instead, our method works directly on the epipolar line: this guarantees an even scanning of the image space and avoids both the need to select a depth range of interest, which is often not known a priori and can vary dramatically from scene to scene, and the need for a suitable discretization of the depth space. In fact, our search is iterative, which avoids the building of a cost volume, costly both to store and to process. Finally, our method performs a robust geometry-aware fusion of the estimated depth maps, leveraging a confidence predicted alongside each depth. We test DELS-MVS on the ETH3D, Tanks and Temples and DTU benchmarks and achieve competitive results with respect to state-of-the-art approaches. Christian Sormann, Emanuele Santellani, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer |
WACV | 3 |
| 2022 | MD-Net: Multi-Detector for Local Feature ExtractionabstractEstablishing a sparse set of keypoint correspondences between images is a fundamental task in many computer vision pipelines. Often, this translates into a computationally expensive nearest neighbor search, where every keypoint descriptor at one image must be compared with all the descriptors at the others. In order to lower the computational cost of the matching phase, we propose a deep feature extraction network capable of detecting a predefined number of complementary sets of keypoints at each image. Since only the descriptors within the same set need to be compared across the different images, the matching phase computational complexity decreases with the number of sets. We train our network to predict the keypoints and compute the corresponding descriptors jointly. In particular, in order to learn complementary sets of keypoints, we introduce a novel unsupervised loss which penalizes intersections among the different sets. Additionally, we propose a novel descriptor-based weighting scheme meant to penalize the detection of keypoints with non-discriminative descriptors. With extensive experiments we show that our feature extraction network, trained only on synthetically warped images and in a fully unsupervised manner, achieves competitive results on 3D reconstruction and re-localization tasks at a reduced matching complexity. Emanuele Santellani, Christian Sormann, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer |
ICPR | 3 |
| 2021 | IB-MVS: An Iterative Algorithm for Deep Multi-View Stereo based on Binary Decisions
Christian Sormann, Mattia Rossi, Andreas Kuhn 0005, Friedrich Fraundorfer |
BMVC | 2 |
| 2020 | DeepC-MVS: Deep Confidence Prediction for Multi-View Stereo ReconstructionabstractDeep Neural Networks (DNNs) have the potential to improve the quality of image-based 3D reconstructions. However, the use of DNNs in the context of 3D reconstruction from large and high-resolution image datasets is still an open challenge, due to memory and computational constraints. We propose a pipeline which takes advantage of DNNs to improve the quality of 3D reconstructions while being able to handle large and high-resolution datasets. In particular, we propose a confidence prediction network explicitly tailored for Multi-View Stereo (MVS) and we use it for both depth map outlier filtering and depth map refinement within our pipeline, in order to improve the quality of the final 3D reconstructions. We train our confidence prediction network on (semi-)dense ground truth depth maps from publicly available real world MVS datasets. With extensive experiments on popular benchmarks, we show that our overall pipeline can produce state-of-the-art 3D reconstructions, both qualitatively and quantitatively. Andreas Kuhn 0005, Christian Sormann, Mattia Rossi, Oliver Erdler, Friedrich Fraundorfer |
3DV | 3 |
| 2020 | BP-MVSNet: Belief-Propagation-Layers for Multi-View-StereoabstractIn this work, we propose BP-MVSNet, a convolutional neural network (CNN)-based Multi-View-Stereo (MVS) method that uses a differentiable Conditional Random Field (CRF) layer for regularization. To this end, we propose to extend the BP layer [16] and add what is necessary to successfully use it in the MVS setting. We therefore show how we can calculate a normalization based on the expected 3D error, which we can then use to normalize the label jumps in the CRF. This is required to make the BP layer invariant to different scales in the MVS setting. In order to also enable fractional label jumps, we propose a differentiable interpolation step, which we embed into the computation of the pairwise term. These extensions allow us to integrate the BP layer into a multi-scale MVS network, where we continuously improve a rough initial estimate until we get high quality depth maps as a result. We evaluate the proposed BP-MVSNet in an ablation study and conduct extensive experiments on the DTU, Tanks and Temples and ETH3D data sets. The experiments show that we can significantly outperform the baseline and achieve state-of-the-art results. Christian Sormann, Patrick Knöbelreiter, Andreas Kuhn 0005, Mattia Rossi, Thomas Pock, Friedrich Fraundorfer |
3DV | 4 |
| 2020 | Joint Graph-Based Depth Refinement and Normal EstimationabstractDepth estimation is an essential component in understanding the 3D geometry of a scene, with numerous applications in urban and indoor settings. These scenarios are characterized by a prevalence of human made structures, which in most of the cases are either inherently piece-wise planar or can be approximated as such. With these settings in mind, we devise a novel depth refinement framework that aims at recovering the underlying piece-wise planarity of those inverse depth maps associated to piece-wise planar scenes. We formulate this task as an optimization problem involving a data fidelity term, which minimizes the distance to the noisy and possibly incomplete input inverse depth map, as well as a regularization, which enforces a piece-wise planar solution. As for the regularization term, we model the inverse depth map pixels as the nodes of a weighted graph, with the weight of the edge between two pixels capturing the likelihood that they belong to the same plane in the scene. The proposed regularization fits a plane at each pixel automatically, avoiding any a priori estimation of the scene planes, and enforces that strongly connected pixels are assigned to the same plane. The resulting optimization problem is solved efficiently with the ADAM solver. Extensive tests show that our method leads to a significant improvement in depth refinement, both visually and numerically, with respect to state-of-the-art algorithms on the Middlebury, KITTI and ETH3D multi-view datasets. Mattia Rossi, Mireille El Gheche, Andreas Kuhn 0005, Pascal Frossard |
CVPR | 1 |
| 2019 | A Practical Solution for Torsional Vibrations Evasion in Variable Speed DrivesabstractThe paper presents a practical solution for an existing drive-train to avoid torsional vibrations for variable speed drives under SPWM modulation technique. Torsional natural frequencies can be a problem because either ill defined in design stage or one of the mechanical parameters has changed after the initial design stage. The proposed method comprises of dividing the frequency range into several partitions where in each partition a value of modulation frequency ratio mfis chosen. A criterion is developed to choose mfbased on avoiding torsional excitations with the minimum converter losses. Torque harmonic distortion and converter losses are analyzed. Matlab/Simulink is used to validate the approach. Khaled ElShawarby, Roberto Perini, Gian Maria Foglia, Antonino Di Gerlando, Mattia Rossi, Francesco Castelli-Dezza |
IECON | 5 |
| 2018 | A Nonsmooth Graph-Based Approach to Light Field Super-ResolutionabstractWe propose a new super-resolution algorithm tailored for light field cameras, which suffer by design from a limited spatial resolution. In particular, we cast light field super-resolution into an optimization problem, where the particular structure of the light field data is captured by a nonsmooth graph-based regularizer, and where all the light field views are super-resolved jointly. Our experiments show that the proposed method compares favorably to the state-of-the-art light field super-resolution algorithms in terms of PSNR and visual quality. In particular, the nonsmooth graph-based regularizer leads to sharper images while preserving fine details. Mattia Rossi, Mireille El Gheche, Pascal Frossard |
ICIP | 1 |
| 2018 | Voltage Control Comparison for Low-Power DC-DC Converters in EVs: PI and Explicit MPCabstractMultilevel converters are used to improve the traction converter of modern electric vehicles (EV). Such topologies can be applied even to low-power converter such as dc-dc auxiliary converters, to improve the EV overall efficiency. The low voltage rating of auxiliary circuits implies the usage of low-cost hardware both for power as well for control platforms. A cost-effective explicit model predictive control (eMPC) is designed for a three-level neutral point clamped step-down dc-dc auxiliary converter and compared to an eMPC for a two-level solution. Both eMPCs are compared to a classical PI-type voltage control. Mattia Rossi, Luigi Piegari, Francesco Castelli-Dezza, Marco Mauri, Maria Stefania Carmeli |
IECON | 1 |
| 2018 | Optical Responses on Multiple Spatial Scales for Assessing Vegetation Dynamics - A Case Study for Alpine GrasslandsabstractVegetation growth is highly dynamic over space and time. Especially mountainous vegetation is regionally affected by climatic variability and human impacts. Tracking optical reflectance provides a unique possibility to analyze vegetation across scales from a small plot up to an extensive area. This study compares the NDVI index of grasslands across four different spatial scales during the growing period in 2017. These scales are covered by measurements from (i) ground spectrometer; (ii) station-based reflectance sensors; (iii) fixed installed Phenocam and (iv) Remote Sensing Sentinel-2 MSI images. We compared the reflectance on single points within a grassland and among different grassland sites in order to assess the strengths and weaknesses of each sensor. Secondly, we analyzed the detectability of anthropogenic management activities using the NDVI index. First results show clear differences in the optical response among scales and in the detectability of management (e.g. fertilization, harvesting) activities. Mattia Rossi, Georg Niedrist, Sarah Asam, Giustino Tonon, Marc Zebisch |
IGARSS | 1 |
| 2018 | Geometry-Consistent Light Field Super-Resolution via Graph-Based RegularizationabstractLight field cameras capture the 3D information in a scene with a single exposure. This special feature makes light field cameras very appealing for a variety of applications: from post-capture refocus to depth estimation and image-based rendering. However, light field cameras suffer by design from strong limitations in their spatial resolution. Off-the-shelf super-resolution algorithms are not ideal for light field data, as they do not consider its structure. On the other hand, the few super-resolution algorithms explicitly tailored for light field data exhibit significant limitations, such as the need to carry out a costly disparity estimation procedure with sub-pixel precision. We propose a new light field super-resolution algorithm meant to address these limitations. We use the complementary information in the different light field views to augment the spatial resolution of the whole light field at once. In particular, we show that coupling the multi-view approach with a graph-based regularizer, which enforces the light field geometric structure, permits to avoid the need of a precise and costly disparity estimation step. Extensive experiments show that the new algorithm compares favorably to the state-of-the-art methods for light field super-resolution, both in terms of visual quality and in terms of reconstruction error. Mattia Rossi, Pascal Frossard |
IEEE Trans. Image Process. | 1 |
| 2017 | Graph-based light field super-resolutionabstractLight field cameras can capture the 3D information in a scene with a single exposure. This special feature makes light field cameras very appealing for a variety of applications: from post capture refocus, to depth estimation and image-based rendering. However, light field cameras exhibit a very limited spatial resolution, which should therefore be increased by computational methods. Off-the-shelf single-frame and multi-frame super-resolution algorithms are not ideal for light field data, as they ignore its particular structure. A few super-resolution algorithms explicitly devised for light field data exist, but they exhibit significant limitations, such as the need to carry out an explicit disparity estimation step for one or several light field views. In this work we present a new light field super-resolution algorithm meant to address these limitations. We adopt a multi-frame alike super-resolution approach, where the information in the different light field views is used to augment the spatial resolution of the whole light field. In particular, we show that coupling the multi-frame paradigma with a graph regularizer that enforces the light field structure permits to avoid the costly and challenging disparity estimation step. Our experiments show that the proposed method compares favorably to the state-of-the-art for light field super-resolution algorithms, both in terms of PSNR and visual quality. Mattia Rossi, Pascal Frossard |
MMSP | 1 |
| 2016 | Torsional issues related to variable frequency control of elastic drive systemsabstractSystems with variable frequency drives may experience torsional vibration. One reason is the pulsating nature of motor torque due to converter supply, another one is the electro-mechanical interaction and the effects of closed loop control. This paper deals with the speed control of elastic drive systems, mainly from torsional performance point of view. Modeling of elastic shaft string is being explored. The importance of control engineering to minimize the torsional vibration is explained. The paper also discusses different control strategies and their practical use in large resonant drive systems. Key factors in this part are complexity of the controller, its robustness and availability of required parameters. The practical aspects of the implementation in real control hardware are being considered. Attention is also paid to detection and monitoring of vibrations. Finally a verification method for correct parameterization on site is proposed. Different aspects are illustrated on simulations and calculations as well as on a real 6.1 MW compressor drive. Martin Bruha, Miroslav Byrtus, Kai Pietiläinen, Mattia Rossi, Marco Mauri |
IECON | 4 |
| 2014 | Luminance driven sparse representation based demosaickingabstractTypical consumer cameras sense at each pixel only one out of the three color components the representation of a color image requires. Then the missing components are estimated via a procedure referred to as demosaicking. The recent spread of sparse regularization approaches for signal reconstruction purposes has extended to demosaicking algorithms too. In this paper, starting from a sparse representation based demosaicking algorithm recently appeared in the literature, a new one is devised. The proposed algorithm effectively estimates the original luminance component from the acquired data and then uses it to guide the sparsity based reconstruction of the full-resolution image. This hybrid approach to demosaicking allows the new algorithm to outperform the past one and to compete with leading demosaicking algorithms, both in terms of PSNR measure and visual quality. Mattia Rossi, Giancarlo Calvagno |
ICIP | 1 |
| 2011 | Improving HTTP performance using "stateless" TCPabstractTCP is quite a heavyweight protocol when serving very small web pages. We introduce a server-side kernel modification which enables a web server to perform HTTP over a UDP socket while the kernel provides a regular TCP interface 'on the wire' to remote clients. We show that our 'stateless' TCP modification can greatly reduce a server's CPU usage (>20%) and TCP related memory requirements(>90%), potentially enabling it to serve small web pages even under extreme overload conditions. David A. Hayes, Michael Welzl, Grenville J. Armitage, Mattia Rossi |
NOSSDAV | 4 |
| 2011 | Inferring the time-zones of prefixes and autonomous systems by monitoring game server discovery trafficabstractGeolocation of IP addresses is used for determining authenticity of webpages, delivering specific country or location related content and advertisements, or to add security for online transactions. Although IP geolocation databases exist, it is sometimes useful to validate their entries or create new, independent databases using independent sources of information. We propose and demonstrate a method whereby collecting and analyzing online game server discovery traffic over short periods of time can allow us to detect in which timezone a certain prefix or AS is located. Our method provides very good estimates of various AS timezones which we verify using publicly available IP geolocation databases. Mattia Rossi, Philip Branch, Grenville J. Armitage |
NOSSDAV | 1 |
| 2010 | A Technique for Reducing BGP Update Announcements through Path Exploration DampingabstractThis paper defines and evaluates Path Exploration Damping (PED) - a router-level mechanism for reducing the volume of propagation of likely transient update messages within a BGP network and decreasing average time to restore reachability compared to current BGP Update damping practices. PED selectively delays and suppresses the propagation of BGP updates that either lengthen an existing AS Path or vary an existing AS Path without shortening its length. We show how PED impacts on convergence time compared to currently deployed mechanisms like Route Flap Damping (RFD), Minimum Route Advertisement Interval (MRAI) and Withdrawal Rate Limiting (WRATE). We replay Internet BGP update traffic captured at two Autonomous Systems to observe that a PED-enabled BGP speaker can reduce the total number of BGP announcements by up to 32% and reduce Path Exploration by 77% compared to conventional use of MRAI. We also describe how PED can be incrementally deployed in the Internet, as it interacts well with prevailing MRAI deployment, and enables restoration of reachability more quickly than MRAI. Geoff Huston, Mattia Rossi, Grenville J. Armitage |
IEEE J. Sel. Areas Commun. | 2 |
| 2008 | TCP/IP over IEEE 802.11b WLAN: the Challenge of Harnessing Known-Corrupt DataabstractThe two transport protocols DCCP and UDP-Lite can make use of data that are known to be erroneous, provided that the link layer hands over such data. A similar functionality has been suggested for TCP. In order to investigate the potential of these mechanisms in WiFi networks, we carried out a measurement study where we examined how often information about corrupt data reaches the transport layer when the corruption control at the link layer is disabled. Our results suggest that this may be a rare occurrence in certain scenarios. Michael Welzl, Mattia Rossi, Andrea Fumagalli, Marco Tacca |
ICC | 2 |