EDBT 2026 Demo / reviewers in the wild / expert
Bin Ren 0005
dblp:02/6854-5
· DBLP profile ↗
25ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0002-9790-1504ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Masked Clustering Prediction for Unsupervised Point Cloud Pre-trainingabstractVision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative semantic features from point clouds via standard ViTs remains underexplored. We propose MaskClu, a novel unsupervised pre-training method for ViTs on 3D point clouds that integrates masked point modeling with clustering-based learning. MaskClu is designed to reconstruct both cluster assignments and cluster centers from masked point clouds, thus encouraging the model to capture dense semantic information. Additionally, we introduce a global contrastive learning mechanism that enhances instance-level feature learning by contrasting different masked views of the same point cloud. By jointly optimizing these complementary objectives, i.e., dense semantic reconstruction, and instance-level contrastive learning. MaskClu enables ViTs to learn richer and more semantically meaningful representations from 3D point clouds. We validate the effectiveness of MaskClu via multiple 3D tasks, including part segmentation, semantic segmentation, object detection, and classification, setting new competitive results. Bin Ren 0005, Xiaoshui Huang, Mengyuan Liu 0001, Hong Liu 0008, Fabio Poiesi, Nicu Sebe, Guofeng Mei |
AAAI | 1 |
| 2026 | Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression MethodsabstractChenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng, Yiyu Wang, Haocong He, Yuanhuiyi Lyu, Lutao Jiang, Xin Zou, Yuqian Fu, Bin Ren, Linfeng Zhang, Xuming Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng 0002, Haocong He, Yuanhuiyi Lyu, Lutao Jiang, Xin Zou 0001, Yuqian Fu, Bin Ren 0005, Linfeng Zhang 0001, Xuming Hu |
ACL (1) | 11 |
| 2026 | A Unified Masked Jigsaw Puzzle Framework for Vision and Language ModelsabstractIn federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has been revealed that the gradient of Position Embeddings (PEs) in Transformer contains sufficient information, which can be used to reconstruct the input data. To mitigate this issue, we introduce a Masked Jigsaw Puzzle (MJP) framework. MJP starts with random token shuffling to break the token order, and then a learnable unknown (unk) position embedding is used to mask out the PEs of the shuffled tokens. In this manner, the local spatial information which is encoded in the position embeddings is disrupted, and the models are forced to learn feature representations that are less reliant on the local spatial information. Notably, with the careful use of MJP, we can not only improve models' robustness against gradient attacks, but also boost their performance in both vision and text application scenarios, such as classification for images (e.g., ImageNet-1 K) and sentiment analysis for text (e.g., Yelp and Amazon). Experimental results suggest that MJP is a unified framework for different Transformer-based models in both vision and language tasks. Weixin Ye, Wei Wang 0108, Yue Song 0002, Bin Ren 0005, Wei Bi, Rita Cucchiara, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | MXPose: Multiplex interactive learning for multi-view 3D Human Pose Estimation
Wanruo Zhang, Zhou Guan, Mengyuan Liu 0001, Bin Ren 0005, Hong Liu 0008 |
Pattern Recognit. | 4 |
| 2026 | Uncertainty-Aware Testing-Time Optimization for 3D Human Pose EstimationabstractAlthough data-driven methods have achieved success in 3D human pose estimation, they often suffer from domain gaps and exhibit limited generalization. In contrast, optimization-based methods excel in fine-tuning for specific cases but are generally inferior to data-driven methods in overall performance. We observe that previous optimization-based methods commonly rely on projection constraint, which only ensures alignment in 2D space, potentially leading to the overfitting problem. To address this, we propose an Uncertainty-Aware testing-time Optimization (UAO) framework, which keeps the prior information of pre-trained model and alleviates the overfitting problem using the uncertainty of joints. Specifically, during the training phase, we design an effective 2D-to-3D network for estimating the corresponding 3D pose while quantifying the uncertainty of each 3D joint. For optimization during testing, the proposed optimization framework freezes the pre-trained model and optimizes only a latent state. Projection loss is then employed to ensure the generated poses are well aligned in 2D space for high-quality optimization. Furthermore, we utilize the uncertainty of each joint to determine how much each joint is allowed for optimization. The effectiveness and superiority of the proposed framework are validated through extensive experiments on challenging datasets: Human3.6M, MPI-INF-3DHP, and 3DPW. Notably, our approach outperforms the previous best result by a large margin of 5.5% on Human3.6M. Ti Wang, Mengyuan Liu 0001, Hong Liu 0008, Bin Ren 0005, Yingxuan You, Wenhao Li 0002, Nicu Sebe, Xia Li 0005 |
IEEE Trans. Multim. | 4 |
| 2025 | A Large-Scale Dataset of Gaussian Splats and Their Self-Supervised Pretrainingabstract3D Gaussian Splatting (3DGS) has become the de facto method of 3D representation in many vision tasks. This calls for the 3D understanding directly in this representation space. To facilitate the research in this direction, we first build a large-scale dataset of 3DGS using the commonly used ShapeNet and ModelNet datasets. Our dataset ShapeSplat consists of 65K objects from 87 unique categories, whose labels are in accordance with the respective datasets. The creation of this dataset utilized the computing equivalent of 2 GPU years on a TITAN XP GPU. We utilize our dataset for unsupervised pretraining and supervised finetuning for classification and segmentation tasks. To this end, we introduce Gaussian-MAE, which highlights the unique benefits of representation learning from Gaussian parameters. Through exhaustive experiments, we provide several valuable insights. In particular, we show that (1) the distribution of the optimized GS centroids significantly differs from the uniformly sampled point cloud (used for initialization) counterpart; (2) this change in distribution results in degradation in classification but improvement in segmentation tasks when using only the centroids; (3) to leverage additional Gaussian parameters, we propose Gaussian feature grouping in a normalized feature space, along with splats pooling layer, offering a tailored solution to effectively group and embed similar Gaussians, which leads to notable improvement in finetuning tasks. Our dataset and model are publicly available at ShapeSplat. Yue Li 0036, Bin Ren 0005, Nicu Sebe, Ender Konukoglu, Theo Gevers, Luc Van Gool, Danda Pani Paudel |
3DV | 3 |
| 2025 | Multi-Stage Multimodal Distillation for Audio-Visual Speaker TrackingabstractSpeaker tracking plays a crucial role in various human-robot interaction applications. Recently, leveraging multimodal information, such as audio and visual signals, has become an important strategy for enhancing the robustness of the tracking system. However, current methods face challenges in effectively exploring the complementarity between audio and visual modalities. To this end, we propose an Audio-Visual Tracker based on Multi-Stage Multimodal Distillation (MSMD-AVT), which utilizes an audio-visual knowledge distillation framework to facilitate audio-visual information fusion over multiple stages progressively. MSMD-AVT is constructed based on an audio-visual teacher-student model incorporating three distinct distillation losses. During the feature extraction stage, the feature alignment distillation is designed to ensure that the feature representations from the student network remain consistent with the teacher encoding feature. Moreover, during the feature fusion stage, the fusion guidance distillation is proposed, using deep teacher features to guide the multimodal fusion process in the student network, optimizing the complementary benefits of audio-visual fusion. Finally, the logits distillation is applied during the position estimation stage to help the student model better capture localization features through knowledge transfer and output alignment. Additionally, we present a multimodal fusion module based on a bidirectional cross-attention mechanism in the student network, dynamically adjusting the effectiveness of different modal features for the tracking task by extracting complementary audio-visual contextual information. Extensive experimental results on the widely used AV16.3 dataset indicate that MSMD-AVT significantly outperforms existing state-of-the-art methods in terms of accuracy and robustness. Our code is publicly available at https://github.com/moyitech/MSMD-AVT. Yidi Li 0001, Wenkai Zhao, Zhenhuan Xu, Bin Ren 0005, Nicu Sebe |
ICASSP | 5 |
| 2025 | ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
Yuqian Fu, Bin Ren 0005, Guolei Sun, Biao Gong, Yanwei Fu 0001, Danda Pani Paudel, Xuanjing Huang 0001, Luc Van Gool |
ICCV | 3 |
| 2025 | SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language PretrainingabstractRecognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training or together at inference. This highlights the clear absence of a model capable of processing 3D data alone for learning semantics end-to-end, along with the necessary data to train such a model. Meanwhile, 3D Gaussian Splatting (3DGS) has emerged as the de facto standard for 3D scene representation across various vision tasks. However, effectively integrating semantic reasoning into 3DGS in a generalizable manner remains an open challenge. To address these limitations, we introduce SceneSplat, to our knowledge the first large-scale 3D indoor scene understanding approach that operates natively on 3DGS. Furthermore, we propose a self-supervised learning scheme that unlocks rich 3D feature learning from unlabeled scenes. To power the proposed methods, we introduce SceneSplat-7K, the first large-scale 3DGS dataset for indoor scenes, comprising 7916 scenes derived from seven established datasets, such as ScanNet and Matterport3D. Generating SceneSplat-7K required computational resources equivalent to 150 GPU days on an L4 GPU, enabling standardized benchmarking for 3DGS-based reasoning for indoor scenes. Our exhaustive experiments on SceneSplat-7K demonstrate the significant benefit of the proposed method over the established baselines. Yue Li 0036, Runyi Yang, Huapeng Li, Mengjiao Ma, Bin Ren 0005, Nikola Popovic 0001, Nicu Sebe, Ender Konukoglu, Theo Gevers, Luc Van Gool, Martin R. Oswald, Danda Pani Paudel |
ICCV | 6 |
| 2025 | SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splattingabstract3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven to be an effective strategy for 3D scene understanding. Current Language Gaussian Splatting line of work fall into three main groups: (i) per-scene optimization-based, (ii) per-scene optimization-free, and (iii) generalizable approach. However, most of them are evaluated only on rendered 2D views of a handful of scenes and viewpoints close to the training views, limiting ability and insight into holistic 3D understanding. To address this gap, we propose the first large-scale benchmark that systematically assesses these three groups of methods directly in 3D space, evaluating on 1060 scenes across three indoor datasets and one outdoor dataset. Benchmark results demonstrate a clear advantage of the generalizable paradigm, particularly in relaxing the scene-specific limitation, enabling fast feed-forward inference on novel scenes, and achieving superior segmentation performance. We further introduce SceneSplat-49K -- a carefully curated 3DGS dataset comprising of around 49K diverse indoor and outdoor scenes trained from multiple sources, with which we demonstrate generalizable approach could harness strong data priors. Our codes, benchmark, and datasets are available. Mengjiao Ma, Yue Li 0036, Jiahuan Cheng, Runyi Yang, Bin Ren 0005, Nikola Popovic 0001, Mingqiang Wei, Nicu Sebe, Ender Konukoglu, Luc Van Gool, Theo Gevers, Martin R. Oswald, Danda Pani Paudel |
NeurIPS | 6 |
| 2025 | DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
Jinqiao Wei, Xionghui Zhao, Yidi Li 0001, Shaoyi Du, Bin Ren 0005, Nicu Sebe |
PRCV (17) | 6 |
| 2025 | PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object DetectionabstractThe integration of point and voxel representations is becoming more common in Light Detection and Ranging (LiDAR)-based 3D object detection. However, existing fusion strategies suffer from ineffective semantic alignment and contextual information loss, while relying solely on point features within regions of interest leads to geometric detail degradation and limited local–global feature integration. To tackle these challenges, we propose the Point-Voxel Attention Fusion Network (PVAFN), a novel two-stage 3D object detector that introduces a point-voxel attention fusion module based on dual-gated cross-modal interaction and a multi-pooling strategy based on density-space awareness. During the feature extraction and fusion stage, a dual-gated hierarchical attention mechanism is proposed to dynamically fuse three heterogeneous modalities—keypoint-based geometric details, voxel-wise local regularity, and Bird’s-Eye-View (BEV)-level global semantics—through learnable gating functions. In the refinement stage, a density-spatial-aware multi-pooling enhancement module is designed to synergize density-aware cluster pooling and multi-scale spatial-aware pyramid pooling, efficiently capturing key geometric details and fine-grained shape structures. This design enhances the integration of local and global features while enabling adaptive multi-scale context modeling and spatially sensitive feature aggregation. Extensive experiments on the KITTI and Waymo benchmark datasets demonstrate that PVAFN achieves promising detection accuracy in 3D mean Average Precision. • PVAFN: A novel network for 3D object detection, fusing keypoint, voxel, and BEV features dynamically. • Stage-I: Dual-gated hierarchical attention fusion for bidirectional point-voxel feature calibration. • Stage-II: Density-spatial-aware multi-pooling module enhances local and global geometric perception. • SOTA Performance: PVAFN achieves highest AP on KITTI and Waymo benchmarks for autonomous driving. Yidi Li 0001, Bin Ren 0005, Wenhao Li 0002, Hong Liu 0008, Nicu Sebe |
Expert Syst. Appl. | 4 |
| 2025 | A pure MLP-Mixer-based GAN framework for guided image translation
Hao Tang 0005, Bin Ren 0005, Nicu Sebe |
Pattern Recognit. | 2 |
| 2025 | HYRE: Hybrid Regressor for 3D Human Pose and Shape EstimationabstractRegression-based 3D human pose and shape estimation often fall into one of two different paradigms. Parametric approaches, which regress the parameters of a human body model, tend to produce physically plausible but image-mesh misalignment results. In contrast, non-parametric approaches directly regress human mesh vertices, resulting in pixel-aligned but unreasonable predictions. In this paper, we consider these two paradigms together for a better overall estimation. To this end, we propose a novel HYbrid REgressor (HYRE) that greatly benefits from the joint learning of both paradigms. The core of our HYRE is a hybrid intermediary across paradigms that provides complementary clues to each paradigm at the shared feature level and fuses their results at the part-based decision level, thereby bridging the gap between the two. We demonstrate the effectiveness of the proposed method through both quantitative and qualitative experimental analyses, resulting in improvements for each approach and ultimately leading to better hybrid results. Our experiments show that HYRE outperforms previous methods on challenging 3D human pose and shape benchmarks. Wenhao Li 0002, Mengyuan Liu 0001, Hong Liu 0008, Bin Ren 0005, Xia Li 0005, Yingxuan You, Nicu Sebe |
IEEE Trans. Image Process. | 4 |
| 2025 | Hierarchical Cross-Attention Network for Virtual Try-OnabstractIn this article, we present an innovative solution tailored for the intricate challenges of the virtual try-on task—our novel Hierarchical Cross-Attention Network, HCANet. HCANet is meticulously crafted with two primary stages: geometric matching and try-on, each playing a crucial role in delivering realistic and visually convincing virtual try-on outcomes. A distinctive feature of HCANet is the incorporation of a novel Hierarchical Cross-Attention (HCA) block into both stages, enabling the effective capture of long-range correlations between individual and clothing modalities. The HCA block functions as a cornerstone, enhancing the depth and robustness of the network. By adopting a hierarchical approach, it facilitates a nuanced representation of the interaction between the person and clothing, capturing intricate details essential for an authentic virtual try-on experience. Our extensive set of experiments establishes the prowess of HCANet. The results showcase its cutting-edge performance across both objective quantitative metrics and subjective evaluations of visual realism. HCANet stands out as a state-of-the-art solution, demonstrating its capability to generate virtual try-on results that not only excel in accuracy but also satisfy subjective criteria of realism. This marks a significant step forward in advancing the field of virtual try-on technologies. Hao Tang 0005, Bin Ren 0005, Nicu Sebe |
IEEE Trans. Multim. | 2 |
| 2024 | Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-supervised Learning
Bin Ren 0005, Guofeng Mei, Danda Pani Paudel, Weijie Wang 0002, Yawei Li 0001, Mengyuan Liu 0001, Rita Cucchiara, Luc Van Gool, Nicu Sebe |
ACCV (7) | 1 |
| 2024 | Denoising Diffusion Probabilistic Models for Action-Conditioned 3D Motion GenerationabstractDiffusion-based generative models have proven to be highly effective in various domains of synthesis. In this work, we propose a conditional paradigm utilizing the denoising diffusion probabilistic model (DDPM) to address the challenge of realistic and diverse action-conditioned 3D skeleton-based motion generation. The proposed method leverages bidirectional Markov chains to generate samples by inferring the reversed Markov chain based on the learned distribution mapping during the forward diffusion process. To the best of our knowledge, our work is the first to employ DDPM to synthesize a variable number of motion sequences conditioned on a categorical action. The proposed method is evaluated on the NTU RGB+D dataset and the NTU RGB+D two-person dataset, showing significant improvements over state-of-the-art motion generation methods. Mengyi Zhao, Mengyuan Liu 0001, Bin Ren 0005, Shuling Dai, Nicu Sebe |
ICASSP | 3 |
| 2024 | Sharing Key Semantics in Transformer Makes Efficient Image RestorationabstractImage Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergence of Vision Transformers (ViTs) has further propelled these advancements. When computing, the self-attention mechanism, a cornerstone of ViTs, tends to encompass all global cues, even those from semantically unrelated objects or regions. This inclusivity introduces computational inefficiencies, particularly noticeable with high input resolution, as it requires processing irrelevant information, thereby impeding efficiency. Additionally, for IR, it is commonly noted that small segments of a degraded image, particularly those closely aligned semantically, provide particularly relevant information to aid in the restoration process, as they contribute essential contextual cues crucial for accurate reconstruction. To address these challenges, we propose boosting IR's performance by sharing the key semantics via Transformer for IR (i.e., SemanIR) in this paper. Specifically, SemanIR initially constructs a sparse yet comprehensive key-semantic dictionary within each transformer stage by establishing essential semantic connections for every degraded patch. Subsequently, this dictionary is shared across all subsequent transformer blocks within the same stage. This strategy optimizes attention calculation within each block by focusing exclusively on semantically related components stored in the key-semantic dictionary. As a result, attention calculation achieves linear computational complexity within each window. Extensive experiments across 6 IR tasks confirm the proposed SemanIR's state-of-the-art performance, quantitatively and qualitatively showcasing advancements. The visual results, code, and trained models are available at: https://github.com/Amazingren/SemanIR. Bin Ren 0005, Yawei Li 0001, Jingyun Liang, Mengyuan Liu 0001, Rita Cucchiara, Luc Van Gool, Ming-Hsuan Yang 0001, Nicu Sebe |
NeurIPS | 1 |
| 2024 | Cloth Interactive Transformer for Virtual Try-OnabstractThe 2D image-based virtual try-on has aroused increased interest from the multimedia and computer vision fields due to its enormous commercial value. Nevertheless, most existing image-based virtual try-on approaches directly combine the person-identity representation and the in-shop clothing items without taking their mutual correlations into consideration. Moreover, these methods are commonly established on pure convolutional neural networks (CNNs) architectures which are not simple to capture the long-range correlations among the input pixels. As a result, it generally results in inconsistent results. To alleviate these issues, in this article, we propose a novel two-stage cloth interactive transformer (CIT) method for the virtual try-on task. During the first stage, we design a CIT matching block, aiming at precisely capturing the long-range correlations between the cloth-agnostic person information and the in-shop cloth information. Consequently, it makes the warped in-shop clothing items look more natural in appearance. In the second stage, we put forth a CIT reasoning block for establishing global mutual interactive dependencies among person representation, the warped clothing item, and the corresponding warped cloth mask. The empirical results, based on mutual dependencies, demonstrate that the final try-on results are more realistic. Substantial empirical results on a public fashion dataset illustrate that the suggested CIT attains competitive virtual try-on performance. Bin Ren 0005, Hao Tang 0005, Fanyang Meng, Runwei Ding, Philip Torr 0001, Nicu Sebe |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Spatio-Temporal Graph Diffusion for Text-Driven Human Motion Generation
Chang Liu 0030, Mengyi Zhao, Bin Ren 0005, Mengyuan Liu 0001, Nicu Sebe |
BMVC | 3 |
| 2023 | Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision TransformersabstractPosition Embeddings (PEs), an arguably indispensable component in Vision Transformers (ViTs), have been shown to improve the performance of ViTs on many vision tasks. However, PEs have a potentially high risk of privacy leakage since the spatial information of the input patches is exposed. This caveat naturally raises a series of interesting questions about the impact of PEs on accuracy, privacy, prediction consistency, etc. To tackle these issues, we propose a Masked Jigsaw Puzzle (MJP) position embedding method. In particular, MJP first shuffles the selected patches via our block-wise random jigsaw puzzle shuffle algorithm, and their corresponding PEs are occluded. Meanwhile, for the non-occluded patches, the PEs remain the original ones but their spatial relation is strengthened via our dense absolute localization regressor. The experimental results reveal that 1) PEs explicitly encode the 2D spatial relationship and lead to severe privacy leakage problems under gradient inversion attack; 2) Training ViTs with the naively shuffled patches can alleviate the problem, but it harms the accuracy; 3) Under a certain shuffle ratio, the proposed MJP not only boosts the performance and robustness on large-scale datasets (i.e., ImageNet-1K and ImageNet-C, -A/O) but also improves the privacy preservation ability under typical gradient attacks by a large margin. The source code and trained models are available at https://github.com/yhlleo/MJP. Bin Ren 0005, Yue Song 0002, Wei Bi, Rita Cucchiara, Nicu Sebe, Wei Wang 0108 |
CVPR | 1 |
| 2023 | PI-Trans: Parallel-Convmlp and Implicit-Transformation Based Gan for Cross-View Image TranslationabstractFor semantic-guided cross-view image translation, it is crucial to learn where to sample pixels from the source view image and where to reallocate them guided by the target view semantic map, especially when there is little overlap or drastic view difference between the source and target images. Hence, one not only needs to encode the long- range dependencies among pixels in both the source view image and target view semantic map but also needs to translate these learned dependencies. To this end, we propose a novel generative adversarial network, PI-Trans, which mainly consists of a novel Parallel-ConvMLP module and an Implicit Transformation module at multiple semantic levels. Extensive experimental results show that PI-Trans achieves the best qualitative and quantitative performance by a large margin compared to the state-of-the-art methods on two challenging datasets. The source code is available at https://github.com/Amazingren/PI-Trans. Bin Ren 0005, Hao Tang 0005, Yiming Wang 0002, Xia Li 0005, Wei Wang 0108, Nicu Sebe |
ICASSP | 1 |
| 2023 | Deep Unsupervised Key Frame Extraction for Efficient Video ClassificationabstractVideo processing and analysis have become an urgent task, as a huge amount of videos (e.g., YouTube, Hulu) are uploaded online every day. The extraction of representative key frames from videos is important in video processing and analysis since it greatly reduces computing resources and time. Although great progress has been made recently, large-scale video classification remains an open problem, as the existing methods have not well balanced the performance and efficiency simultaneously. To tackle this problem, this work presents an unsupervised method to retrieve the key frames, which combines the convolutional neural network and temporal segment density peaks clustering. The proposed temporal segment density peaks clustering is a generic and powerful framework, and it has two advantages compared with previous works. One is that it can calculate the number of key frames automatically. The other is that it can preserve the temporal information of the video. Thus, it improves the efficiency of video classification. Furthermore, a long short-term memory network is added on the top of the convolutional neural network to further elevate the performance of classification. Moreover, a weight fusion strategy of different input networks is presented to boost performance. By optimizing both video classification and key frame extraction simultaneously, we achieve better classification performance and higher efficiency. We evaluate our method on two popular datasets (i.e., HMDB51 and UCF101), and the experimental results consistently demonstrate that our strategy achieves competitive performance and efficiency compared with the state-of-the-art approaches. Hao Tang 0005, Lei Ding 0008, Songsong Wu, Bin Ren 0005, Nicu Sebe, Paolo Rota |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Cascaded Cross MLP-Mixer GANs for Cross-View Image Translation
Bin Ren 0005, Hao Tang 0005, Nicu Sebe |
BMVC | 1 |
| 2020 | Grouped Temporal Enhancement Module for Human Action RecognitionabstractTemporal information is a significant cue for recognizing human actions from videos. Different from 2D CNN which can only capture spatial information in an efficient way, 3D CNN is good at capturing both spatial and temporal information at the expense of high computational cost. Beyond both methods, this paper presents a Grouped Temporal Enhancement (GTE) module which even outperforms 3D CNN, meanwhile only needs similar low computational cost as 2D CNN. The GTE module firstly decomposes an input video into spatial and temporal groups along channel dimension, and then uses a learnable temporal shift (LTS) operation for efficient temporal modeling. Finally, a 2D convolution filter is used to enhance the ability of LTS for spatial modeling. Extensive experiments on three benchmark datasets validate the effect of our method. Hong Liu 0008, Bin Ren 0005, Mengyuan Liu 0001, Runwei Ding |
ICIP | 2 |