Titus Zaharia

dblp:45/100 · also Titus Bogdan Zaharia · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-6589-1241ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Multi-modal BLIP-2 Approach for Video Captioning
Antoine Brimont, Titus Zaharia, Ruxandra Tapu
ICPR (4)2
2026 A Survey on Video Captioning in the Era of Large Language Models
abstract
Recent technological advancements have led to the widespread presence of video in our daily lives, particularly through social media, where it has become the dominant medium of communication. As a result, the ability to analyze and understand videos has become essential for managing such a massive volume of content. Within this context, video captioning (VC) is one of the most effective ways to achieve this understanding. Recent advances in Natural Language Generation, combined with breakthroughs in Computer Vision, has facilitated the development of robust VC models. However, despite these advancements, selecting the right model for the right application remains challenging due to limited evaluation methods, and significant limitations still hinder the widespread adoption of these models. This survey aims to address these issues by proposing a comprehensive overview of the VC models landscape, along with a comparative experimental study, carried out on two benchmark datasets and involving eight representative, high quality models.
Antoine Brimont, Titus Zaharia, Ruxandra Tapu
ACM Trans. Multim. Comput. Commun. Appl.2
2025 MI-Cap: A Multi-Modal Interpretable Model for Video Captioning
abstract
Video captioning aims to describing video content in natural language, which implies understanding and interpreting scenes, objects, actions and events that are present in the video. Current approaches have mostly concentrated on visual cues, often neglecting the rich information available from other important modalities, such as the audio or textual (e.g. subtitles) channels. In this paper, we introduce a novel video captioning method trained with a multi-modal contrastive loss function that emphasizes both multi-modal integration and interpretability. Our approach is designed to capture the rich dependencies between the different modalities, resulting in more accurate and pertinent captions. Concerning the interpretability issues, by exploiting multiple attention mechanisms, the model is able to provide explanations of the results proposed. The experimental evaluation, caried out on widely used benchmark datasets such as MSR-VTT and VATEX, demonstrate that the proposed method performs favorably against state-of the-art models.
Antoine Hanna-Asaad, Decky Aspandi, Titus Zaharia
CBMI3
2025 A 3D Mesh Convolution-Based Autoencoder for Geometry Compression
abstract
In this paper, we introduce a novel 3D mesh convolution-based autoencoder for geometry compression, able to deal with irregular mesh data without requiring neither preprocessing nor manifold/watertightness conditions. The proposed approach extracts meaningful latent representations by learning features directly from the mesh faces, while preserving connectivity through dedicated pooling and unpooling operations. The encoder compresses the input mesh into a compact base mesh space, which ensures that the latent space remains comparable. The decoder reconstructs the original connectivity and restores the compressed geometry to its full resolution. Extensive experiments on multi-class datasets demonstrate that our method outperforms state-of-the-art approaches in both 3D mesh geometry reconstruction and latent space classification tasks.
Germain Bregeon, Marius Preda, Radu Ispas, Titus Zaharia
ICIP4
2024 MeshConv3D: Efficient Convolution and Pooling Operators for Triangular 3D Meshes
abstract
Convolutional neural networks (CNNs) have been pivotal in various 2D image analysis tasks, including computer vision, image indexing and retrieval or semantic classification. Extending CNNs to 3D data such as point clouds and 3D meshes raises significant challenges since the very basic convolution and pooling operators need to be completely revisited and re-defined in an appropriate manner to tackle irregular connectivity issues. In this paper, we introduce MeshConv3D, a 3D mesh-dedicated methodology integrating specialized convolution and face collapse-based pooling operators. MeshConv3D operates directly on meshes of arbitrary topology, without any need of prior remeshing/conversion techniques. In order to validate our approach, we have considered a semantic classification task. The experimental results obtained on three distinct benchmark datasets show that the proposed approach makes it possible to achieve equivalent or superior classification results, while minimizing the related memory footprint and computational load.
Germain Bregeon, Marius Preda, Radu Ispas, Titus Zaharia
CBMI4
2023 Multimodal emotion recognition using cross modal audio-video fusion with attention and deep metric learning
Bogdan Cosmin Mocanu, Ruxandra Tapu, Titus Zaharia
Image Vis. Comput.3
2022 Affine Transformation-Based Color Compression For Dynamic 3D Point Clouds
abstract
Recently, high-quality, humanoid-like 3D point clouds have become extensively used in various use cases related to VR/AR applications. Such high-density point clouds, represented by a huge number of points (e.g., 1 million points) carrying various photometric attributes, require efficient compression techniques for storage and transmission. However, most research works in the literature mainly focus on geometry compression, while only a few consider the spatio-temporal compression of color attributes. In this paper, we propose a novel color attribute prediction method, which exploits a skeleton-based affine motion estimation technique. The skeleton and the motion parameters are compressed in a lossless manner, to preserve accurate color prediction. The color residuals are lossy compressed using a video-based coding solution. Our proposal has been integrated into the Video-based Point Cloud Compression (V-PCC) test model of MPEG. The experimental results demonstrate that the proposed method outperforms the reference V-PCC test model, notably in low bitrate conditions.
Marius Preda, Titus Zaharia
ICIP3
2022 One-Cycle Pruning: Pruning Convnets With Tight Training Budget
abstract
Introducing sparsity in a convnet has been an efficient way to reduce its complexity while keeping its performance almost intact. Most of the time, sparsity is introduced using a three-stage pipeline: 1) training the model to convergence, 2) pruning the model, 3) fine-tuning the pruned model to recover performance. The last two steps are often performed iteratively, leading to reasonable results but also to a time-consuming process. In our work, we propose to remove the first step of the pipeline and to combine the two others in a single training-pruning cycle, allowing the model to jointly learn the optimal weights while being pruned. We do this by introducing a novel pruning schedule, named One-Cycle Pruning (OCP), which starts pruning from the beginning of the training, and until its very end. Experiments conducted on a variety of combinations between architectures (VGG-16, ResNet-18), datasets (CIFAR-10, CIFAR-100, Caltech-101), and sparsity values (80%, 90%, 95%) show that not only OCP consistently outperforms common pruning schedules such as One-Shot, Iterative and Automated Gradual Pruning, but also that it drastically reduces the required training budget. More-over, experiments following the Lottery Ticket Hypothesis show that OCP allows to find higher quality and more stable pruned networks.
Nathan Hubens, Matei Mancas, Bernard Gosselin, Marius Preda, Titus Zaharia
ICIP5
2022 ATOFIS, an AR Training System for Manual Assembly: A Full Comparative Evaluation against Guides
abstract
This paper reports on a user study to comparatively evaluate two AR training systems designed for step-by-step manual operations: ATOFIS - recently proposed in the literature, and Microsoft Dynamics 365 Guides (hereinafter Guides) - one of the most relevant state-of-the-art commercial solutions. The user study (N=16) was conducted in two stages- i.e., training and authoring, on a partial replica of a real-world assembly workstation. During training, the participant learns a sequence of manual operations by performing two assembly cycles, guided by each of the two AR training systems. During authoring, the participant creates the two sets of AR work instructions used in the next training session, one set with each of the two authoring systems. We bound the authoring and training procedures during the experiment to comparatively assess the AR systems overall, and address at the same time an evaluation gap observed in the literature. The experimental results demonstrated advantages of the authoring approach proposed by ATOFIS (i.e., low-cost, formalized, in-situ, immersive and on-the-fly), proved the usability and effectiveness of the AR instructions authored with ATOFIS and validated a set of hypotheses formulated by the authors of the system. ATOFIS authoring was $1.72\times$ faster and unanimously preferred by the participants; ATOFIS training reported zero assembly errors and was 13% faster than Guides. ATOFIS reported excellent system usability (i.e., SUS) and mental workload (i.e., NASA-TLX) scores for both authoring and training, outperforming Guides on all dimensions.
Traian Lavric, Emmanuel Bricard, Marius Preda, Titus Zaharia
ISMAR4
2021 Exploring Low-Cost Visual Assets for Conveying Assembly Instructions in AR
abstract
Augmented Reality (AR) is an emerging technology offering a great potential in assisting humans in a wide range of industrial processes, from manufacturing to validation and maintenance. However, very few AR solutions have been adopted so far in industrial sectors, mainly because of technical and acceptability issues. This paper has three main contributions: (1) it identifies potential barriers for AR adoption in manufacturing environments; (2) it proposes an AR training methodology that overcomes these challenges and finally, (3) it evaluates the proposed AR training approach in a concrete, real-world use case by conducting a field experiment with 12 participants. Our findings indicate that low-cost assistive visual assets (i.e. text, images, videos and arrows) can be sufficient for effectively delivering complex manual assembly instructions through AR head-mounted displays (i.e. Hololens 2). We finally discuss how the proposed methodology simplifies considerably the authoring of the AR instructions and how this technique can potentially be generalized to other similar industrial scenarios.
Traian Lavric, Emmanuel Bricard, Marius Preda, Titus Zaharia
INISTA4
2021 Active learning to measure opinion and violence in French newspapers
abstract
News articles analysis may be oversimplified when restricted to detecting classes of interest already benefiting from trustworthy labeled datasets, like political affiliation or fakeness. Behind an apparent neutrality, an editorial slant may be embodied by favoring one-sided interviews, avoiding topics or choosing oriented illustrations. These challenges, seen as machine learning problems, would require a tedious annotation task. We introduce ReALMS, an active learning framework capable of quickly elaborating models which detect arbitrary classes in multi-modal text and image documents. Evidence of this capability is given by a case study on French news outlets: the detection of subjectivity, demonstrations and violence.
Paul Guélorget, Guillaume Gadek, Titus Zaharia, Bruno Grilhères
KES3
2021 Compression of Sparse and Dense Dynamic Point Clouds - Methods and Standards
abstract
In this article, a survey of the point cloud compression (PCC) methods by organizing them with respect to the data structure, coding representation space, and prediction strategies is presented. Two paramount families of approaches reported in the literature-the projection- and octree-based methods-are proven to be efficient for encoding dense and sparse point clouds, respectively. These approaches are the pillars on which the Moving Picture Experts Group Committee developed two PCC standards published as final international standards in 2020 and early 2021, respectively, under the names: video-based PCC and geometry-based PCC. After surveying the current approaches for PCC, the technologies underlying the two standards are described in detail from an encoder perspective, providing guidance for potential standard implementors. In addition, experiment evaluations in terms of compression performances for both solutions are provided.
Marius Preda, Vladyslav Zakharchenko, Euee S. Jang, Titus Zaharia
Proc. IEEE5
2020 Skeleton-based motion estimation for Point Cloud Compression
abstract
With the rapid development of point cloud acquisition technologies, high-quality human-shape point clouds are more and more used in VR/AR applications and in general in 3D Graphics. To achieve near-realistic quality, such content usually contains an extremely high number of points (over 0.5 million points per 3D object per frame) and associated attributes (such as color). For this reason, disposing of efficient, dedicated 3D Point Cloud Compression (3DPCC) methods becomes mandatory. This requirement is even stronger in the case of dynamic content, where the coordinates and attributes of the 3D points are evolving over time. In this paper, we propose a novel skeleton-based 3DPCC approach, dedicated to the specific case of dynamic point clouds representing humanoid avatars. The method relies on a multi-view 2D human pose estimation of 3D dynamic point clouds. By using the DensePose neural network, we first extract the body parts from projected 2D images. The obtained 2D segmentation information is back-projected and aggregated into the 3D space. This procedure makes it possible to partition the 3D point cloud into a set of 3D body parts. For each part, a 3D affine transform is estimated between every two consecutive frames and used for 3D motion compensation. The proposed approach has been integrated into the Video-based Point Cloud Compression (V-PCC) test model of MPEG. Experimental results show that the proposed method, in the particular case of body motion with small amplitudes, outperforms the V-PCC test mode in the lossy inter-coding condition by up to 83% in terms of bitrate reduction in low bit rate conditions. Meanwhile, the proposed framework holds the potential of supporting various features such as regions of interests and level of details.
Christian Tulvan, Marius Preda, Titus Zaharia
MMSP4
2020 Wearable assistive devices for visually impaired: A state of the art survey
Ruxandra Tapu, Bogdan Cosmin Mocanu, Titus Zaharia
Pattern Recognit. Lett.3
2017 Object Tracking Using Deep Convolutional Neural Networks and Visual Appearance Models
Bogdan Cosmin Mocanu, Ruxandra Tapu, Titus Zaharia
ACIVS3
2017 A computer vision-based perception system for visually impaired
Ruxandra Tapu, Bogdan Cosmin Mocanu, Titus Zaharia
Multim. Tools Appl.3
2016 Automatic Segmentation of TV News into Stories Using Visual and Temporal Information
Bogdan Cosmin Mocanu, Ruxandra Tapu, Titus Zaharia
ACIVS3
2016 Scribble-based object segmentation with modified gaussian mixture models
Raluca-Diana Sambra-Petre, Titus Zaharia
Pattern Anal. Appl.2
2016 Laban descriptors for gesture recognition and emotional analysis
Arthur Truong, Hugo Boujut, Titus Zaharia
Vis. Comput.3
2012 Retrieval of Multiple Instances of Objects in Videos
Andrei Bursuc, Titus Zaharia, Françoise J. Prêteux
MMM2
2012 OVIDIUS: A Web Platform for Video Browsing and Search
Andrei Bursuc, Titus Zaharia, Françoise J. Prêteux
MMM2
2010 Mobile video browsing and retrieval with the OVIDIUS platform
abstract
This paper describes a mobile video browsing and retrievalapproach, based on the so-called OVIDIUS (On-line VIDeo Indexing Universal System) platform. In contrast with traditional and commercial video retrieval platforms, where video content is treated in a more or less monolithic manner (i.e. with global descriptions associated with the whole document), the proposed approach makes it possible to browse and access video content in a finer, per-segment basis. The hierarchical metadata structure exploits the MPEG-7 approach for structural description of video content. The MPEG-7 description schemes have been here enriched with both semantic and content-based metadata. The developed approach shows all its pertinence within a multiterminal context and in particular for video access from mobile devices. The platform has been recently (February, 2010) validated within the framework of the Mé[email protected] French national project.
Andrei Bursuc, Titus Zaharia, Françoise J. Prêteux
ACM Multimedia2
2009 A Triangle-Fan-based approach for low complexity 3D mesh compression
abstract
This paper proposes a novel approach for mono-resolution 3D mesh compression, called TFAN (Triangle Fan-based compression). TFAN treats in a unified manner meshes of arbitrary topologies, i.e. manifold or not, oriented or not, while supporting real-time decoding. The proposed approach has been evaluated on both manifold and non-manifold databases, including more than 7000 3D mesh models. Experiments show that the TFAN approach outperforms existing techniques such as MPEG-4 3DMC or Tourna & Gotsman, with decoding times lower by an order of magnitude at equivalent or even better levels of compression efficiency (+/-10% in bitrate). When applied to non-manifold 3D data, the compression performances are significantly enhanced (6% to 30% gain in bitrate). Due to its high compression performances TFAN has been recently retained for ISO MPEG-4 standardization.
Khaled Mamou, Titus Zaharia, Françoise J. Prêteux
ICIP2
2009 TFAN: A low complexity 3D mesh compression algorithm
abstract
Abstract This paper proposes a novel approach for mono‐resolution 3D mesh compression, called TFAN (Triangle Fan‐based compression). TFAN treats in a unified manner meshes of arbitrary topologies, i.e., manifold or not, oriented or not, while offering a linear computational complexity (with respect to the number of mesh vertices) for both encoding and decoding algorithms. In addition, the TFAN compressed representation is optimized for real‐time decoding applications. In order to validate the proposed approach, two databases have been considered for experimentations. The first is the MPEG‐4 test set, which includes over 3500 general purpose manifold meshes. The second, related to the French national project SEMANTIC‐3D, includes over 4000 computer assisted design (CAD) meshes of highly irregular, non‐manifold topologies. In both cases, the TFAN approach outperforms existing techniques such as MPEG‐4/3DMC (3D Mesh Coding) or Touma and Gotsman, with decoding times lower by an order of magnitude at equivalent or even better levels of compression efficiency (±10% in bitrate). In addition, when applied to non‐manifold 3D data, the compression performances are significantly enhanced (6–30% gain in bitrate). Due to its high compression performances the TFAN approach has been recently retained for ISO standardization, within the framework of the MPEG‐4/AFX standard. Copyright © 2009 John Wiley & Sons, Ltd.
Khaled Mamou, Titus Zaharia, Françoise J. Prêteux
Comput. Animat. Virtual Worlds2
2008 FAMC: The MPEG-4 standard for Animated Mesh Compression
abstract
This paper presents a new compression technique for 3D dynamic meshes, referred to as FAMC - frame-based animated mesh compression, recently promoted within the MPEG-4 standard as Amendment 2 of part 16 AFX (Animation Framework extension). The heart of the method is a skinning model optimally computed from a frame-based representation and exploited for compression purposes within the framework of a motion compensation strategy. The proposed encoder offers high compression performances (gains in bitrate of 60% with respect to the previous MPEG-4 technique and of 20 to 40% with respect to state-of-the-art approaches) and is well suited for compressing both geometric and photometric attributes.
Khaled Mamou, Titus Zaharia, Françoise J. Prêteux
ICIP2
2008 Two optimizations of the MPEG-4 FAMC standard for enhanced compression of animated 3D meshes
abstract
The MPEG-4 standard adopted a novel technology for compression of dynamic 3D meshes with constant connectivity and time-varying geometry, referred to as FAMC - frame-based animated mesh compression. In this paper, we propose two optimizations of the FAMC approach, aiming at improving the compression efficiency. The first one is based on a PCA (principal component analysis) decomposition of the motion compensation error residuals. The second improves the bi-orthogonal (4-2) wavelet coding approach supported by the standard, by introducing an optimal bit allocation procedure, combined with an adapted quantization of wavelet coefficients. Experimental results show that both optimizations lead to significant gains in compression rate (about 20-30%) at low bitrates.
Khaled Mamou, Titus Zaharia, Françoise J. Prêteux, Ayman Kamoun, Frédéric Payan, Marc Antonini
ICIP2
2008 Frame-based compression of animated meshes in MPEG-4
abstract
This paper presents a new compression technique for 3D dynamic meshes, referred to as FAMC - frame-based animated mesh compression, promoted within the MPEG-4 standard as amendment 2 of part 16 AFX (animation framework extension). The FAMC approach combines a model-based motion compensation strategy, with transform/predictive coding of residual errors. First, a skinning motion compensation model is automatically computed from a frame-based representation and then encoded. Subsequently, either 1) DCT/lifting wavelets or 2) layer-based predictive coding is employed to exploit remaining spatio-temporal correlations in the residual signal. The proposed encoder offers high compression performances (gains in bit rate of 60% with respect to the previous MPEG-4 technique and of 20% to 40% with respect to state-of-the-art approaches) and is well suited for compressing both geometric and photometric (normal vectors, colors...) attributes. In addition, the FAMC method supports a rich set of functionalities including streaming, scalability (spatial, temporal and quality) and progressive transmission.
Khaled Mamou, Titus Zaharia, Françoise J. Prêteux, Nikolce Stefanoski, Jörn Ostermann
ICME2
2006 A skinning approach for dynamic 3D mesh compression
abstract
Abstract This paper proposes a novel approach for 3D mesh compression, based on a skinning animation technique. The core of the proposed method is a piecewise affine predictor coupled with a skinning model and a DCT representation of the residuals errors. The experimental evaluation shows that the proposed skinning‐based encoder outperforms (with bitrates gains from 47% to 67%) GV, RT, MPEG‐4/AFX‐IC, D3DMC, PCA and Dynapack techniques. Copyright © 2006 John Wiley & Sons, Ltd.
Khaled Mamou, Titus Zaharia, Françoise J. Prêteux
Comput. Animat. Virtual Worlds2
2002 Shape-based retrieval of 3D mesh models
abstract
This paper addresses the issue of 3D mesh indexation by using shape descriptors (SDs) under constraints of geometric and topological invariance. A new shape descriptor, the canonical 3D Hough transform descriptor (C3DHTD) is here proposed. Intrinsically topologically stable, the C3DHTD is not invariant to geometric transformations. Nevertheless, we show mathematically how the C3DHTD can be optimally associated (in terms of compactness of representation and computational complexity) with a spatial alignment procedure which leads to a geometric invariant behavior. Experiments carried out upon the categorized MPEG-7 3D model database objectively show that the C3DHTD outperforms both the MPEG-7 3D SD and classic EGIs descriptor.
Titus Zaharia, Françoise J. Prêteux
ICME (1)1