VLDB 2026 Research / reviewers in the wild / expert
Bin Zhao 0001
dblp:73/4325-1
· DBLP profile ↗
69ranked-venue papers
12as first author
59since 2021 · last 2026
0000-0002-0294-8538ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 10 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 4 first-author · 26 since 2021Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow DerivativesabstractReconstructing controllable Gaussian splats for articulated objects from monocular video is especially challenging due to its inherently insufficient constraints. Existing methods address this by relying on dense masks and manually defined control signals, limiting their real-world applications. In this paper, we propose an annotation-free method, FreeGaussian, which mathematically disentangles camera egomotion and articulated movements via flow derivatives. By establishing a connection between 2D flows and 3D Gaussian dynamic flow, our method enables optimization and continuity of dynamic Gaussian motions from flow priors without any control signals. Furthermore, we introduce a 3D spherical vector controlling scheme, which represents the state as a 3D Gaussian trajectory, thereby eliminating the need for complex 1D control signal calculations and simplifying controllable Gaussian modeling. Extensive experiments on articulated objects demonstrate the state-of-the-art visual performance and precise, part-aware controllability of our method. Delin Qu, Junli Liu, Haoming Song, Dong Wang 0028, Yuan Yuan 0001, Bin Zhao 0001 |
AAAI | 8 |
| 2026 | Semi-LLIE: Semi-supervised contrastive learning with Mamba-based low-light enhancement
Ke Zhang 0014, Bin Zhao 0001, Xuelong Li 0001 |
Neural Networks | 5 |
| 2026 | Image harmonization in complex degradation scenes
Bin Zhao 0001, Xuelong Li 0001 |
Pattern Recognit. | 2 |
| 2025 | Learning 2D Invariant Affordance Knowledge for 3D Affordance Groundingabstract3D Object Affordance Grounding aims to predict the functional regions on a 3D object and has laid the foundation for a wide range of applications in robotics. Recent advances tackle this problem via learning a mapping between 3D regions and a single human-object interaction image. However, the geometric structure of the 3D object and the object in the human-object interaction image are not always consistent, leading to poor generalization. To address this issue, we propose to learn generalizable invariant affordance knowledge from multiple human-object interaction images within the same affordance category. Specifically, we introduce the Multi-Image Guided Invariant-Feature-Aware 3D Affordance Grounding (MIFAG) framework. It grounds 3D object affordance regions by identifying common interaction patterns across multiple human-object interaction images. First, the Invariant Affordance Knowledge Extraction Module (IAM) utilizes an iterative updating strategy to gradually extract aligned affordance knowledge from multiple images and integrate it into an affordance dictionary. Then, the Affordance Dictionary Adaptive Fusion Module (ADM) learns comprehensive point cloud representations that consider all affordance candidates in multiple images. Besides, the Multi-Image and Point Affordance (MIPA) benchmark is constructed and our method outperforms existing state-of-the-art methods on various experimental comparisons. Xianqiang Gao 0001, Pingrui Zhang, Delin Qu, Dong Wang 0028, Zhigang Wang 0002, Yan Ding 0002, Bin Zhao 0001 |
AAAI | 7 |
| 2025 | Efficient Diffusion as Low Light EnhancerabstractThe computational burden of the iterative sampling process remains a major challenge in diffusion-based LowLight Image Enhancement (LLIE). Current acceleration methods, whether training-based or training-free, often lead to significant performance degradation, highlighting the trade-off between performance and efficiency. In this paper, we identify two primary factors contributing to performance degradation: fitting errors and the inference gap. Our key insight is that fitting errors can be mitigated by linearly extrapolating the incorrect score functions, while the inference gap can be reduced by shifting the Gaussian flow to a reflectance-aware residual space. Based on the above insights, we design Reflectance-Aware Trajectory Refinement (RATR) module, a simple yet effective module to refine the teacher trajectory using the reflectance component of images. Following this, we introduce Reflectance-aware Diffusion with Distilled Trajectory (ReDDiT), an efficient and flexible distillation framework tailored for LLIE. Our framework achieves comparable performance to previous diffusion-based methods with redundant steps in just 2 steps while establishing new state-of-the-art (SOTA) results with 8 or 4 steps. Comprehensive experimental evaluations on 10 benchmark datasets validate the effectiveness of our method, consistently outperforming existing SOTA methods. Code is available at project page. Guanzhou Lan, Qianli Ma 0008, Zhigang Wang 0002, Dong Wang 0028, Xuelong Li 0001, Bin Zhao 0001 |
CVPR | 7 |
| 2025 | Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot ManipulationabstractBuilding a lifelong robot that can effectively leverage prior knowledge for continuous skill acquisition remains significantly challenging. Despite the success of experience replay and parameter-efficient methods in alleviating catastrophic forgetting problem, naively applying these methods causes a failure to leverage the shared primitives between skills. To tackle these issues, we propose Primitive Prompt Learning (PPL), to achieve lifelong robot manipulation via reusable and extensible primitives. Within our two stage learning scheme, we first learn a set of primitive prompts to represent shared primitives through multi-skills pre-training stage, where motion-aware prompts are learned to capture semantic and motion shared primitives across different skills. Secondly, when acquiring new skills in lifelong span, new prompts are concatenated and optimized with frozen pretrained prompts, boosting the learning via knowledge transfer from old skills to new ones. For evaluation, we construct a large-scale skill dataset and conduct extensive experiments in both simulation and real-world tasks, demonstrating PPL’s superior performance over state-of-the-art methods. Yuanqi Yao, Siao Liu, Haoming Song, Delin Qu, Yan Ding 0002, Bin Zhao 0001, Zhigang Wang 0002, Xuelong Li 0001, Dong Wang 0028 |
CVPR | 7 |
| 2025 | AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional RelationsabstractVisual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual grounding from aerial views. Compared to traditional VG, AerialVG poses new challenges, \emph{e.g.}, appearance-based grounding is insufficient to distinguish among multiple visually similar objects, and positional relations should be emphasized. Besides, existing VG models struggle when applied to aerial imagery, where high-resolution images cause significant difficulties. To address these challenges, we introduce the first AerialVG dataset, consisting of 5K real-world aerial images, 50K manually annotated descriptions, and 103K objects. Particularly, each annotation in AerialVG dataset contains multiple target objects annotated with relative spatial relations, requiring models to perform comprehensive spatial reasoning. Furthermore, we propose an innovative model especially for the AerialVG task, where a Hierarchical Cross-Attention is devised to focus on target regions, and a Relation-Aware Grounding module is designed to infer positional relations. Experimental results validate the effectiveness of our dataset and method, highlighting the importance of spatial reasoning in aerial visual grounding. The code and dataset will be released. Junli Liu, Zhigang Wang 0002, Chi Yan, Dong Wang 0028, Xuelong Li 0001, Bin Zhao 0001 |
ICCV | 9 |
| 2025 | Open-Vocabulary Octree-Graph for 3D Scene UnderstandingabstractOpen-vocabulary 3D scene understanding is indispensable for embodied agents. Recent works leverage pretrained vision-language models (VLMs) for object segmentation and project them to point clouds to build 3D maps. Despite progress, a point cloud is a set of unordered coordinates that requires substantial storage space and does not directly convey occupancy information or spatial relation, making existing methods inefficient for downstream tasks, e.g., path planning and text-based object retrieval. To address these issues, we propose \textbf{Octree-Graph}, a novel scene representation for open-vocabulary 3D scene understanding. Specifically, a Chronological Group-wise Segment Merging (CGSM) strategy and an Instance Feature Aggregation (IFA) algorithm are first designed to get 3D instances and corresponding semantic features. Subsequently, an adaptive-octree structure is developed that stores semantics and depicts the occupancy of an object adjustably according to its shape. Finally, the Octree-Graph is constructed where each adaptive-octree acts as a graph node, and edges describe the spatial relations among nodes. Extensive experiments on various tasks are conducted on several widely-used datasets, demonstrating the versatility and effectiveness of our method. Code is available \href{https://github.com/yifeisu/OV-Octree-Graph}{here}. Zhigang Wang 0002, Yifei Su, Dong Wang 0028, Xuelong Li 0001, Bin Zhao 0001 |
ICCV | 7 |
| 2025 | MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile ManipulationabstractIn mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it effectively. Many navigation approaches primarily define success by proximity to the target, often overlooking the necessity for optimal positioning that facilitates subsequent manipulation. To address this, we introduce MoMa-Kitchen, a benchmark dataset comprising over 100k samples that provide training data for models to learn optimal final navigation positions for seamless transition to manipulation. Our dataset includes affordance-grounded floor labels collected from diverse kitchen environments, in which robotic mobile manipulators of different models attempt to grasp target objects amidst clutter. Using a fully automated pipeline, we simulate diverse real-world scenarios and generate affordance labels for optimal manipulation positions. Visual data are collected from RGB-D inputs captured by a first-person view camera mounted on the robotic arm, ensuring consistency in viewpoint during data collection. We also develop a lightweight baseline model, NavAff, for navigation affordance grounding that demonstrates promising performance on the MoMa-Kitchen benchmark. Our approach enables models to learn affordance-based final positioning that accommodates different arm types and platform heights, thereby paving the way for more robust and generalizable integration of navigation and manipulation in embodied AI. Project page: \href{https://momakitchen.github.io/}{https://momakitchen.github.io/}. Pingrui Zhang, Xianqiang Gao 0001, Kehui Liu, Dong Wang 0028, Zhigang Wang 0002, Bin Zhao 0001, Yan Ding 0002, Xuelong Li 0001 |
ICCV | 7 |
| 2025 | Cocube: a Tabletop Modular Multi-Robot Platform for Education and ResearchabstractThis paper presents CoCube a tabletop modular robotics platform designed for robotics education and multirobot algorithm research. CoCube is characterized by its low cost low floors high ceilings and wide walls offering flexibility and broad applicability across various use cases. The platform comprises four key components: CoCube robots which integrate wireless communication movement and interaction; CoModules which provide versatile external functionality; CoMaps which enable high-precision localization via microdot patterns on regular printed paper; and CoTags for interaction. CoCube operates on MicroBlocks a blocks programming language for physical computing inspired by MIT Scratch a widely-used coding language with a simple visual interface that makes programming accessible to young learners. It offers users both flexibility and ease of use with advanced API support for more complex applications. This paper details the design of the CoCube platform and demonstrates its potential in both educational and research contexts. Songyi Zhu, Zhonghan Tang, Jialing Han, Zemin Lin, Zhongrui You, John Maloney, Bernat Romagosa, Bin Zhao 0001, Zhigang Wang 0002, Zhinan Zhang, Xuelong Li 0001 |
ICRA | 11 |
| 2025 | COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language ModelsabstractLeveraging the powerful reasoning capabilities of large language models (LLMs), recent LLM-based robot task planning methods yield promising results. However, they mainly focus on single or multiple homogeneous robots on simple tasks. Practically, complex long-horizon tasks always require collaboration among multiple heterogeneous robots especially with more complex action spaces, which makes these tasks more challenging. To this end, we propose COHERENT, a novel LLM-based task planning framework for collaboration of heterogeneous multi-robot systems including quadrotors, robotic dogs, and robotic arms. Specifically, a Proposal-Execution-Feedback-Adjustment (PEFA) mechanism is designed to decompose and assign actions for individual robots, where a centralized task assigner makes a task planning proposal to decompose the complex task into subtasks, and then assigns subtasks to robot executors. Each robot executor selects a feasible action to implement the assigned subtask and reports self-reflection feedback to the task assigner for plan adjustment. The PEFA loops until the task is completed. Moreover, we create a challenging heterogeneous multi-robot task planning benchmark encompassing 100 complex long-horizon tasks. The experimental results show that our work surpasses the previous methods by a large margin in terms of success rate and execution efficiency. The experimental videos, code, and benchmark are released at https://github.com/MrKeee/COHERENT. Kehui Liu, Dong Wang 0028, Zhigang Wang 0002, Xuelong Li 0001, Bin Zhao 0001 |
ICRA | 6 |
| 2025 | AlignBot: Aligning VLM-Powered Customized Task Planning with User Reminders Through Fine-Tuning for Household RobotsabstractThis paper presents AlignBot, a novel framework designed to optimize VLM-powered customized task planning for household robots by effectively aligning with user reminders. In domestic settings, aligning task planning with user reminders poses significant challenges due to the limited quantity, diversity, and multimodal nature of the reminders. To address these challenges, AlignBot employs a fine-tuned LLaVA-7B model, functioning as an adapter for GPT-40. This adapter model internalizes diverse forms of user reminders-such as personalized preferences, corrective guidance, and contextual assistance-into structured instruction-formatted cues that prompt GPT-40 in generating customized task plans. Additionally, AlignBot integrates a dynamic retrieval mechanism that selects task-relevant historical successes as prompts for GPT-40, further enhancing task planning accuracy. To validate the effectiveness of AlignBot, experiments are conducted in real-world household environments, which are constructed within the laboratory to replicate typical household settings. A multimodal dataset with over 1,500 entries derived from volunteer reminders is used for training and evaluation. The results demonstrate that AlignBot significantly improves customized task planning, outperforming existing LLM- and VLM-powered planners by interpreting and aligning with user reminders, achieving 86.8 % success rate compared to the vanilla GPT-40 baseline at 21.6%, reflecting a 65% improvement and over four times greater effectiveness. Supplementary materials are available at: https://yding25.com/AlignBot/ Zhaxizhuoma, Pengan Chen, Ziniu Wu, Dong Wang 0028, Peng Zhou 0018, Nieqing Cao, Yan Ding 0002, Bin Zhao 0001, Xuelong Li 0001 |
ICRA | 9 |
| 2025 | Representation discrepancy bridging method for remote sensing image-text retrieval
Hailong Ning, Siying Wang 0012, Tao Lei 0003, Xiaopeng Cao, Huanmin Dou, Bin Zhao 0001, Asoke K. Nandi, Petia Radeva |
Neurocomputing | 6 |
| 2025 | Scale-Adaptive Aerial Object Tracking Network via Location EstimationabstractThe objective in aerial object tracking with the goal of accurately capturing and tracking the dynamic and variable characteristics of objects in various complex environments. However, existing tracking methods often encounter performance bottlenecks when dealing with rapid object movement, occlusions, and changes in appearance. Summarizing the aforementioned issues, we introduce a Scale-adaptive Aerial object Tracking Network (SATNet), which can not only effectively handles scale changes and occlusions of tracking objects from an aerial perspective, but also accurately estimates the positions of tracking objects against complex backgrounds, ensuring the stability and precision for the tracking model. Specifically, SATNet incorporates a Scale-adaptive Feature Fusion Enhancement Module (SFFEM) to integrate multi-scale detailed and semantic features, strengthening object feature representation while mitigating interference from similar objects, thereby improving robustness to occlusions and enabling accurate detection of small or distant objects. In addition, an object search strategy based on Location Estimation Module (LEM) is designed to analyze classification and regression information, achieving precise object localization in complex environments and significantly enhancing the performance of the SATNet tracking model in processing video sequences for reliable and effective object tracking. Widespread experiments proving the efficacy and supremacy of the proposed SATNet against numerous cutting-edge competitors across three public datasets. Chuangye Guo, Mengyao Dong, Bin Zhao 0001, Xuelong Li 0001 |
IEEE Internet Things J. | 5 |
| 2025 | On the Value of Myopic Behavior in Policy ReuseabstractLeveraging learned strategies in unfamiliar scenarios is fundamental to human intelligence. In reinforcement learning, rationally reusing the policies acquired from other tasks or human experts is critical for tackling problems that are difficult to learn from scratch. In this work, we present a framework called Selective Myopic bEhavior Control (SMEC), which results from the insight that the short-term behaviors of prior policies are sharable across tasks. By evaluating the behaviors of prior policies via a hybrid value function architecture, SMEC adaptively aggregates the sharable short-term behaviors of prior policies and the long-term behaviors of the task policy, leading to coordinated decisions. Empirical results on a collection of manipulation and locomotion tasks demonstrate that SMEC outperforms existing methods, and validate the ability of SMEC to leverage related prior policies. Chenjia Bai, Haoran He, Bin Zhao 0001, Zhen Wang 0004, Wei Li 0055, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Color Event Enhanced Single-Exposure HDR ImagingabstractSingle-exposure high dynamic range (HDR) imaging aims to reconstruct the wide-range intensities of a scene by using its single low dynamic range (LDR) image, thus providing significant efficiency. Existing methods pay high attention to restoring the luminance by inversing the tone-mapping process, while the color in the over-/under-exposed area cannot be well restored due to the information loss of the single LDR image. To address this issue, we introduce color events into the imaging pipeline, which record asynchronous pixel-wise color changes in a high dynamic range, enabling edge-like scene perception under challenging lighting conditions. Specifically, we propose a joint framework that incorporates color events and a single LDR image to restore both content and color of an HDR image, where an exposureaware transformer (EaT) module is designed to propagate the informative hints, provided by the normal-exposed LDR regions and the event streams, to the missing areas. In this module, an exposure-aware mask is estimated to suppress distractive information and strengthen the restoration of the over-/under-exposed regions. To our knowledge, we are the first to use color events to enhance single-exposure HDR imaging. We also contribute corresponding datasets, consisting of synthesized datasets and a real-world dataset collected by a DAVIS346-color camera. The datasets can be found at https://www.kaggle.com/datasets/mengyaocui/ce-hdr. Extensive experiments demonstrate the effectiveness of the proposed method. Zhigang Wang 0002, Dong Wang 0028, Bin Zhao 0001, Xuelong Li 0001 |
AAAI | 4 |
| 2024 | X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge TransferabstractThe field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning temporal information within video sequences. To address these issues, we propose a novel cross-modal knowledge transfer framework, called X4D-SceneFormer. This framework enhances 4D-Scene understanding by transferring texture priors from RGB sequences using a Transformer architecture with temporal relationship mining. Specifically, the framework is designed with a dual-branch architecture, consisting of an 4D point cloud transformer and a Gradient-aware Image Transformer (GIT). The GIT combines visual texture and temporal correlation features to offer rich semantics and dynamics for better point cloud representation. During training, we employ multiple knowledge transfer techniques, including temporal consistency losses and masked self-attention, to strengthen the knowledge transfer between modalities. This leads to enhanced performance during inference using single-modal 4D point cloud inputs. Extensive experiments demonstrate the superior performance of our framework on various 4D point cloud video understanding tasks, including action recognition, action segmentation and semantic segmentation. The results achieve 1st places, i.e., 85.3% (+7.9%) accuracy and 47.3% (+5.0%) mIoU for 4D action segmentation and semantic segmentation, on the HOI4D challenge, outperforming previous state-of-the-art by a large margin. We release the code at https://github.com/jinglinglingling/X4D. Linglin Jing, Ying Xue 0003, Xu Yan 0005, Chaoda Zheng, Dong Wang 0028, Ruimao Zhang, Zhigang Wang 0002, Hui Fang 0003, Bin Zhao 0001, Zhen Li 0026 |
AAAI | 9 |
| 2024 | Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained ModelsabstractThe popularity of pre-trained large models has revolutionized downstream tasks across diverse fields, such as language, vision, and multi-modality. To minimize the adaption cost for downstream tasks, many Parameter-Efficient Fine-Tuning (PEFT) techniques are proposed for language and 2D image pre-trained models. However, the specialized PEFT method for 3D pre-trained models is still under-explored. To this end, we introduce Point-PEFT, a novel framework for adapting point cloud pre-trained models with minimal learnable parameters. Specifically, for a pre-trained 3D model, we freeze most of its parameters, and only tune the newly added PEFT modules on downstream tasks, which consist of a Point-prior Prompt and a Geometry-aware Adapter. The Point-prior Prompt adopts a set of learnable prompt tokens, for which we propose to construct a memory bank with domain-specific knowledge, and utilize a parameter-free attention to enhance the prompt tokens. The Geometry-aware Adapter aims to aggregate point cloud features within spatial neighborhoods to capture fine-grained geometric information through local interactions. Extensive experiments indicate that our Point-PEFT can achieve better performance than the full fine-tuning on various downstream tasks, while using only 5% of the trainable parameters, demonstrating the efficiency and effectiveness of our approach. Code is released at https://github.com/Ivan-Tang-3D/Point-PEFT. Ray Zhang 0002, Zoey Guo, Xianzheng Ma, Bin Zhao 0001, Zhigang Wang 0002, Dong Wang 0028, Xuelong Li 0001 |
AAAI | 5 |
| 2024 | HPL-ESS: Hybrid Pseudo-Labeling for Unsupervised Event-based Semantic SegmentationabstractEvent-based semantic segmentation has gained popularity due to its capability to deal with scenarios under high-speed motion and extreme lighting conditions, which cannot be addressed by conventional RGB cameras. Since it is hard to annotate event data, previous approaches rely on event-to-image reconstruction to obtain pseudo labels for training. However, this will inevitably introduce noise, and learning from noisy pseudo labels, especially when generated from a single source, may reinforce the errors. This drawback is also called confirmation bias in pseudo-labeling. In this paper, we propose a novel hybrid pseudo-labeling framework for unsupervised event-based semantic segmentation, HPL-ESS, to alleviate the influence of noisy pseudo labels. Specifically, we first employ a plain unsupervised domain adaptation framework as our baseline, which can generate a set of pseudo labels through self-training. Then, we incorporate offline event-to-image re-construction into the framework, and obtain another set of pseudo labels by predicting segmentation maps on the re-constructed images. A noisy label learning strategy is designed to mix the two sets of pseudo labels and enhance the quality. Moreover, we propose a soft prototypical alignment (SPA) module to further improve the consistency of target domain features. Extensive experiments show that the proposed method outperforms existing state-of-the-art methods by a large margin on benchmarks (e.g., +5.88% accuracy, +10.32% mIoU on DSEC-Semantic dataset), and even surpasses several supervised methods. Linglin Jing, Zhigang Wang 0002, Xu Yan 0005, Dong Wang 0028, Gerald Schaefer, Hui Fang 0003, Bin Zhao 0001, Xuelong Li 0001 |
CVPR | 9 |
| 2024 | Cyclic Learning for Binaural Audio Generation and LocalizationabstractBinaural audio is obtained by simulating the biological structure of human ears, which plays an important role in artificial immersive spaces. A promising approach is to utilize mono audio and corresponding vision to synthesize binaural audio, thereby avoiding expensive binaural audio recording. However, most existing methods di-rectly use the entire scene as a guide, ignoring the corre-spondence between sounds and sounding objects. In this paper, we advocate generating binaural audio using fine-grained raw waveform and object-level visual information as guidance. Specifically, we propose a Cyclic Locating-and-Ul'mixing (CLUP) framework that jointly learns vi-sual sounding object localization and binaural audio generation. Visual sounding object localization establishes the correspondence between specific visual objects and sound modalities, which provides object-aware guidance to improve binaural generation performance. Meanwhile, the spatial information contained in the generated binaural au-dio can further improve the performance of sounding object localization. In this case, visual sounding object localization and binaural audio generation can achieve cyclic learning and benefit from each other. Experimental re-sults demonstrate that on the FAIR-Play benchmark dataset, our method is significantly ahead of the existing baselines in multiple evaluation metrics (STFTJ↓: 0.787 vs. 0.851, ENVJ↑: 0.128 vs. 0.134, WAVJ↓: 5.244 vs. 5.684, SNR↑: 7.546 vs. 7.044). Zhaojian Li 0002, Bin Zhao 0001, Yuan Yuan 0001 |
CVPR | 2 |
| 2024 | Implicit Event-RGBD Neural SLAMabstractImplicit neural SLAM has achieved remarkable progress recently. Nevertheless, existing methods face significant challenges in non-ideal scenarios, such as motion blur or lighting variation, which often leads to issues like convergence failures, localization drifts, and distorted mapping. To address these challenges, we propose EN-SLAM, the first event-RGBD implicit neural SLAM framework, which effectively leverages the high rate and high dynamic range advantages of event data for tracking and mapping. Specif-ically, EN-SLAM proposes a differentiable CRF (Camera Response Function) rendering technique to generate dis-tinct RGB and event camera data via a shared radiance field, which is optimized by learning a unified implicit representation with the captured event and RGBD supervision. Moreover, based on the temporal difference property of events, we propose a temporal aggregating optimization strategy for the event joint tracking and global bundle adjustment, capitalizing on the consecutive difference constraints of events, significantly enhancing tracking accuracy and robustness. Finally, we construct the simulated dataset DEV-Indoors and real captured dataset DEV-Reals containing 6 scenes, 17 sequences with practical motion blur and lighting changes for evaluations. Experimental results show that our method outperforms the SOTA methods in both tracking ATE and mapping ACC with a real-time 17 FPS in various challenging environments. Project page: https://delinqu.github.io/EN-SLAM. Delin Qu, Chi Yan, Dong Wang 0028, Dan Xu 0002, Bin Zhao 0001, Xuelong Li 0001 |
CVPR | 8 |
| 2024 | GS-SLAM: Dense Visual SLAM with 3D Gaussian SplattingabstractIn this paper, we introduce GS-SLAM that first utilizes 3D Gaussian representation in the Simultaneous Localization and Mapping (SLAM) system. It facilitates a better bal-ance between efficiency and accuracy. Compared to recent SLAM methods employing neural implicit representations, our method utilizes a real-time differentiable splatting ren-dering pipeline that offers significant speedup to map opti-mization and RGB-D rendering. Specifically, we propose an adaptive expansion strategy that adds new or deletes noisy 3D Gaussians in order to efficiently reconstruct new observed scene geometry and improve the mapping of pre-viously observed areas. This strategy is essential to ex-tend 3D Gaussian representation to reconstruct the whole scene rather than synthesize a static object in existing meth-ods. Moreover, in the pose tracking process, an effective coarse-to-fine technique is designed to select reliable 3D Gaussian representations to optimize camera pose, resulting in runtime reduction and robust estimation. Our method achieves competitive performance compared with existing state-of-the-art real-time methods on the Replica, TUM-RGBD datasets. Project page: https://gs-slam.github.io/. Chi Yan, Delin Qu, Dan Xu 0002, Bin Zhao 0001, Zhigang Wang 0002, Dong Wang 0028, Xuelong Li 0001 |
CVPR | 4 |
| 2024 | Any2Point: Empowering Any-Modality Large Models for Efficient 3D Understanding
Ray Zhang 0002, Jiaming Liu 0003, Zoey Guo, Bin Zhao 0001, Zhigang Wang 0002, Peng Gao 0007, Hongsheng Li 0001, Dong Wang 0028, Xuelong Li 0001 |
ECCV (36) | 5 |
| 2024 | A Coarse-to-Fine Reconstruction Framework for Non-Lambertian Photometric StereoabstractPhotometric stereo aims to regress object surface normal from a set of images observed under varying illuminations. Although existing methods have achieved promising results, the irregular high-frequency detail is ignored, especially in complex and tiny surface folds. To address this problem, a coarse-to-fine reconstruction framework is proposed for non-Lambertian photometric stereo. Specifically, a coarse network is designed to roughly predict object surface normal, which learns the mapping from observed images to coarse surface normal. Then, to deal with the high-frequency information loss, we introduce a fine network to extract high-frequency information by leveraging both coarse surface normal and observation images. Meanwhile, to provide more supervision, we design a reconstruction module to reconstruct observed images from predicted surface normal and illuminations. Extensive experiments have demonstrated that the proposed method outperforms existing works and restores high-frequency detail effectively. In addition, the proposed method promotes the robustness under sparse illuminations. Zhigang Wang 0002, Peipei Gu, Bin Zhao 0001, Xuelong Li 0001 |
ICME | 5 |
| 2024 | SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied ManipulationabstractAcquiring a multi-task imitation policy in 3D manipulation poses challenges in terms of scene understanding and action prediction. Current methods employ both 3D representation and multi-view 2D representation to predict the poses of the robot’s end-effector. However, they still require a considerable amount of high-quality robot trajectories, and suffer from limited generalization in unseen tasks and inefficient execution in long-horizon reasoning. In this paper, we propose **SAM-E**, a novel architecture for robot manipulation by leveraging a vision-foundation model for generalizable scene understanding and sequence imitation for long-term action reasoning. Specifically, we adopt Segment Anything (SAM) pre-trained on a huge number of images and promptable masks as the foundation model for extracting task-relevant features, and employ parameter-efficient fine-tuning on robot data for a better understanding of embodied scenarios. To address long-horizon reasoning, we develop a novel multi-channel heatmap that enables the prediction of the action sequence in a single pass, notably enhancing execution efficiency. Experimental results from various instruction-following tasks demonstrate that SAM-E achieves superior performance with higher execution efficiency compared to the baselines, and also significantly improves generalization in few-shot adaptation to new tasks. Chenjia Bai, Haoran He, Zhigang Wang 0002, Bin Zhao 0001, Xiu Li 0001, Xuelong Li 0001 |
ICML | 5 |
| 2024 | Robust Quadrupedal Locomotion via Risk-Averse Policy LearningabstractThe robustness of legged locomotion is crucial for quadrupedal robots in challenging terrains. Recently, Reinforcement Learning (RL) has shown promising results in legged locomotion and various methods try to integrate privileged distillation, scene modeling, and external sensors to improve the generalization and robustness of locomotion policies. However, these methods are hard to handle uncertain scenarios such as abrupt terrain changes or unexpected external forces. In this paper, we consider a novel risk-sensitive perspective to enhance the robustness of legged locomotion. Specifically, we employ a distributional value function learned by quantile regression to model the aleatoric uncertainty of environments, and perform risk-averse policy learning by optimizing the worst-case scenarios via a risk distortion measure. Extensive experiments in both simulation environments and a real Aliengo robot demonstrate that our method is efficient in handling various external disturbances, and the resulting policy exhibits improved robustness in harsh and uncertain situations in legged locomotion. Jiyuan Shi, Chenjia Bai, Haoran He, Lei Han 0001, Dong Wang 0008, Bin Zhao 0001, Mingguo Zhao, Xiu Li 0001, Xuelong Li 0001 |
ICRA | 6 |
| 2024 | Kinematic-aware Prompting for Generalizable Articulated Object Manipulation with LLMsabstractGeneralizable articulated object manipulation is essential for home-assistant robots. Recent efforts focus on imitation learning from demonstrations or reinforcement learning in simulation, however, due to the prohibitive costs of real-world data collection and precise object simulation, it still remains challenging for these works to achieve broad adaptability across diverse articulated objects. Recently, many works have tried to utilize the strong in-context learning ability of Large Language Models (LLMs) to achieve generalizable robotic manipulation, but most of these researches focus on high-level task planning, sidelining low-level robotic control. In this work, building on the idea that the kinematic structure of the object determines how we can manipulate it, we propose a kinematic-aware prompting framework that prompts LLMs with kinematic knowledge of objects to generate low-level motion trajectory waypoints, supporting various object manipulation. To effectively prompt LLMs with the kinematic structure of different objects, we design a unified kinematic knowledge parser, which represents various articulated objects as a unified textual description containing kinematic joints and contact location. Building upon this unified description, a kinematic-aware planner model is proposed to generate precise 3D manipulation waypoints via a designed kinematic-aware chain-of-thoughts prompting method. Our evaluation spanned 48 instances across 16 distinct categories, revealing that our framework not only outperforms traditional methods on 8 seen categories but also shows a powerful zero-shot capability for 8 unseen articulated object categories with only 17 demonstrations. Moreover, the real-world experiments on 7 different object categories prove our framework’s adaptability in practical scenarios. Code is released at https://github.com/GeWu-Lab/LLM_articulated_object_manipulation. Wenke Xia, Dong Wang 0028, Xincheng Pang, Zhigang Wang 0002, Bin Zhao 0001, Di Hu 0001, Xuelong Li 0001 |
ICRA | 5 |
| 2024 | Depth Helps: Improving Pre-trained RGB-based Policy with Depth Information Injectionabstract3D perception ability is crucial for generalizable robotic manipulation. While recent foundation models have made significant strides in perception and decision-making with RGB-based input, their lack of 3D perception limits their effectiveness in fine-grained robotic manipulation tasks. To address these limitations, we propose a Depth Information Injection (DI2) framework that leverages the RGB-Depth modality for policy fine-tuning, while relying solely on RGB images for robust and efficient deployment. Concretely, we introduce the Depth Completion Module (DCM) to extract the spatial prior knowledge related to depth information and generate virtual depth information from RGB inputs to aid policy deployment. Further, we propose the Depth-Aware Codebook (DAC) to eliminate noise and reduce the cumulative error from the depth prediction. In the inference phase, this framework employs RGB inputs and accurately predicted depth data to generate the manipulation action. We conduct experiments on simulated LIBERO environments and real-world scenarios, and the experiment results prove that our method could effectively enhance the pre-trained RGB-based policy with 3D perception ability for robotic manipulation. The website is released at https://gewu-lab.github.io/DepthHelps-IROS2024. Xincheng Pang, Wenke Xia, Zhigang Wang 0002, Bin Zhao 0001, Di Hu 0001, Dong Wang 0028, Xuelong Li 0001 |
IROS | 4 |
| 2024 | TAS: Personalized Text-guided Audio Spatialization
Zhaojian Li 0002, Bin Zhao 0001, Yuan Yuan 0001 |
ACM Multimedia | 2 |
| 2024 | Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-TrainingabstractLearning a generalist embodied agent capable of completing multiple tasks poses challenges, primarily stemming from the scarcity of action-labeled robotic datasets. In contrast, a vast amount of human videos exist, capturing intricate tasks and interactions with the physical world. Promising prospects arise for utilizing actionless human videos for pre-training and transferring the knowledge to facilitate robot policy learning through limited robot demonstrations. However, it remains a challenge due to the domain gap between humans and robots. Moreover, it is difficult to extract useful information representing the dynamic world from human videos, because of its noisy and multimodal data structure. In this paper, we introduce a novel framework to tackle these challenges, which leverages a unified discrete diffusion to combine generative pre-training on human videos and policy fine-tuning on a small number of action-labeled robot videos. We start by compressing both human and robot videos into unified video tokens. In the pre-training stage, we employ a discrete diffusion model with a mask-and-replace diffusion strategy to predict future video tokens in the latent space. In the fine-tuning stage, we harness the imagined future videos to guide low-level action learning with a limited set of robot data. Experiments demonstrate that our method generates high-fidelity future videos for planning and enhances the fine-tuned policies compared to previous state-of-the-art approaches with superior performance. Haoran He, Chenjia Bai, Ling Pan, Weinan Zhang 0001, Bin Zhao 0001, Xuelong Li 0001 |
NeurIPS | 5 |
| 2024 | LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Control and RenderingabstractThis paper scales object-level reconstruction to complex scenes, advancing interactive scene reconstruction. We introduce two datasets, OmniSim and InterReal, featuring 28 scenes with multiple interactive objects. To tackle the challenge of inaccurate interactive motion recovery in complex scenes, we propose LiveScene, a scene-level language-embedded interactive radiance field that efficiently reconstructs and controls multiple objects. By decomposing the interactive scene into local deformable fields, LiveScene enables separate reconstruction of individual object motions, reducing memory consumption. Additionally, our interaction-aware language embedding localizes individual interactive objects, allowing for arbitrary control using natural language. Our approach demonstrates significant superiority in novel view synthesis, interactive scene control, and language grounding performance through extensive experiments. Project page: https://livescenes.github.io. Delin Qu, Pingrui Zhang, Xianqiang Gao 0001, Bin Zhao 0001, Zhigang Wang 0002, Dong Wang 0028, Xuelong Li 0001 |
NeurIPS | 5 |
| 2024 | Pessimistic value iteration for multi-task data sharing in Offline Reinforcement Learning
Chenjia Bai, Lingxiao Wang 0003, Jianye Hao, Zhuoran Yang, Bin Zhao 0001, Zhen Wang 0004, Xuelong Li 0001 |
Artif. Intell. | 5 |
| 2024 | Optics-driven drone
Xuelong Li 0001, Zhigang Wang 0002, Bin Zhao 0001 |
Sci. China Inf. Sci. | 4 |
| 2024 | Motion-Aware Video Frame Interpolation
Fuhua Zhang 0002, Bin Zhao 0001, Xuelong Li 0001 |
Neural Networks | 3 |
| 2024 | Image harmonization with Simple Hybrid CNN-Transformer Network
Bin Zhao 0001, Xuelong Li 0001 |
Neural Networks | 2 |
| 2024 | Vehicle Perception From SatelliteabstractSatellites are capable of capturing high-resolution videos. It makes vehicle perception from satellite become possible. Compared to street surveillance, drive recorder or other equipments, satellite videos provide a much broader city-scale view, so that the global dynamic scene of the traffic are captured and displayed. Traffic monitoring from satellite is a new task with great potential applications, including traffic jams prediction, path planning, vehicle dispatching, etc. Practically, limited by the resolution and view, the captured vehicles are very tiny (a few pixels) and move slowly. Worse still, these satellites are in Low Earth Orbit (LEO) to capture such high-resolution videos, so the background is also moving. Under this circumstance, traffic monitoring from the satellite view is an extremely challenging task. To attract more researchers into this field, we build a large-scale benchmark for traffic monitoring from satellite. It supports several tasks, including tiny object detection, counting and density estimation. The dataset is constructed based on 12 satellite videos and 14 synthetic videos recorded from GTA-V. They are separated into 408 video clips, which contain 7,336 real satellite images and 1,960 synthetic images. 128,801 vehicles are annotated totally, and the number of vehicles in each image varies from 0 to 101. Several classic and state-of-the-art approaches in traditional computer vision are evaluated on the datasets, so as to compare the performance of different approaches, analyze the challenges in this task, and discuss the future prospects. Bin Zhao 0001, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Progressive Feature Interleaved Fusion Network for Remote-Sensing Image Salient Object DetectionabstractSalient object detection (SOD) has made significant strides in natural scene images (NSIs) in the span of the past few decades. However, extending these approaches for remote-sensing images (RSIs) faces challenges due to their complex backgrounds, complicated edges, irregular topology, and multiscale object variations, which hinder performance. Existing RSI-SOD techniques are unable to accurately detect salient objects while preserving detailed boundaries, and their computational inefficiency limits their practicality. To overcome these challenges, we entail the development of a progressive feature interleaved framework (PROFILE) in RSI-SOD. In particular, we leverage the interleaved association of the convolutional neural network (CNN) and Transformer (IACTer) to obtain global semantic relations and spatial details. To handle object scale variation, we design a lightweight plug-and-play multiscale hierarchical channel-spatial collaborative feature enhancement module (MHCCF), which can boost the representation of features regarding the relevant region, while identifying the precise location details about the salient region. Finally, a bi-directional consistency constraint module (BCCM) is developed, which can be integrated into the training of arbitrary SOD and segmentation networks to efficiently locate salient regions with refined structures and clear demarcations. Experiments demonstrate that our PROFILE surpasses 20 cutting-edge SOD methods, proving its ability to enhance the accuracy and integrity of SOD in complex backgrounds, such as illumination and shadows. Bin Zhao 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Low-Light Image Enhancement With SAM-Based Structure Priors and GuidanceabstractLow-light images often suffer from severe detail lost in darker areas and non-uniform illumination distribution across distinct regions. Thus, structure modeling and region-specific illumination manipulation are crucial for high-quality enhanced image generation. However, previous methods encounter limitations in exploring robust structure priors and lack adequate modeling of illumination relationships among different regions, resulting in structure artifacts and color deviations. To alleviate this limitation, we propose a Segmentation-Guided Framework (SGF) which integrates the constructed robust segmentation priors to guide the enhancement process. Specifically, SGF first constructs a robust image-level edge prior based on the segmentation results of the Segment Anything Model (SAM) in a zero-shot manner. Then, we generate lighted-up region-aware feature-level prior by incorporating region-aware dynamic convolution. To adequately model long-distance illumination interactions across distinct regions, we design a segmentation-guided transformer block (SGTB), which utilizes the lighted-up region-aware feature-level prior to guide self-attention calculation. By arranging the SGTBs in a symmetric hierarchical structure, we derive a segmentation-guided enhancement module that operates under the guidance of both the image and feature-level priors. Comprehensive experimental results show that our SGF performs remarkably in both quantitative evaluation and visual comparison. Bin Zhao 0001, Xuelong Li 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Weather Translation via Weather-Cue TransferringabstractIn this article, the weather translation task is proposed, which aims to transfer the weather type of the image from one category to another. Weather translation is a complicated image weather editing task that changes the weather cue of an image across multiple weather types, and it is related to image restoration, image editing, and photographic style transfer tasks. Although lots of approaches have been developed for traditional image translation and restoration tasks, only few of them are capable of handling the multicategory weather types problem with a single network due to the rich categories and highly complicated semantic structures of weather images. Especially, it is difficult to change the weather cue while preserving the weather-invariant area. To solve these issues, we developed a weather-cue guided multidomain translation approach based on StarGAN v2, termed WeatherGAN. In the proposed model, the core generator is redesigned to transfer the weather cue according to the target weather type. The weather segmentation module is first introduced to acquire the weather semantic structure of images in a weakly supervised multitask manner. In addition, a weather clues module is presented to reprocess the weather segmentation into a weather-specific clues map, which identifies the weather-invariant and weather-cue areas clearly. Extensive studies and evaluations show that our approach outperforms the state of the art. The data and source code will be publicly available soon after the manuscript is accepted. Xuelong Li 0001, Chen Li 0063, Kai Kou, Bin Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Edge-Aware Network for Flow-Based Video Frame InterpolationabstractVideo frame interpolation can up-convert the frame rate and enhance the video quality. In recent years, although interpolation performance has achieved great success, image blur usually occurs at object boundaries owing to the large motion. It has been a long-standing problem and has not been addressed yet. In this brief, we propose to reduce the image blur and get the clear shape of objects by preserving the edges in the interpolated frames. To this end, the proposed edge-aware network (EA-Net) integrates the edge information into the frame interpolation task. It follows an end-to-end architecture and can be separated into two stages, i.e., edge-guided flow estimation and edge-protected frame synthesis. Specifically, in the flow estimation stage, three edge-aware mechanisms are developed to emphasize the frame edges in estimating flow maps, so that the edge maps are taken as auxiliary information to provide more guidance to boost the flow accuracy. In the frame synthesis stage, the flow refinement module is designed to refine the flow map, and the attention module is carried out to adaptively focus on the bidirectional flow maps when synthesizing the intermediate frames. Furthermore, the frame and edge discriminators are adopted to conduct the adversarial training strategy, so as to enhance the reality and clarity of synthesized frames. Experiments on three benchmarks, including Vimeo90k, UCF101 for single-frame interpolation, and Adobe240-fps for multiframe interpolation, have demonstrated the superiority of the proposed EA-Net for the video frame interpolation task. Bin Zhao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance FieldabstractTalking head generation aims to generate faces that maintain the identity information of the source image and imitate the motion of the driving image. Most pioneering methods rely primarily on 2D representations and thus will inevitably suffer from face distortion when large head rotations are encountered. Recent works instead employ explicit 3D structural representations or implicit neural rendering to improve performance under large pose changes. Nevertheless, the fidelity of identity and expression is not so desirable, especially for novel-view synthesis. In this paper, we propose HiDe-NeRF, which achieves high-fidelity and free-view talking-head synthesis. Drawing on the recently proposed Deformable Neural Radiance Fields, HiDe-NeRF represents the 3D dynamic scene into a canonical appearance field and an implicit deformation field, where the former comprises the canonical source face and the latter models the driving pose and expression. In particular, we improve fidelity from two aspects: (i) to enhance identity expressiveness, we design a generalized appearance module that leverages multi-scale volume features to preserve face shape and details; (ii) to improve expression preciseness, we propose a lightweight deformation module that explicitly decouples the pose and expression to enable precise expression modeling. Extensive experiments demonstrate that our proposed approach can generate better results than previous works. Project page: https://www.waytron.net/hidenerf/ Weichuang Li, Longhao Zhang, Dong Wang 0028, Bin Zhao 0001, Zhigang Wang 0002, Mulin Chen, Bang Zhang, Zhongjian Wang, Liefeng Bo, Xuelong Li 0001 |
CVPR | 4 |
| 2023 | Fully Self-Supervised Depth Estimation from Defocus ClueabstractDepth-from-defocus (DFD), modeling the relationship between depth and defocus pattern in images, has demonstrated promising performance in depth estimation. Recently, several self-supervised works try to overcome the difficulties in acquiring accurate depth ground-truth. However, they depend on the all-in-focus (AIF) images, which cannot be captured in real-world scenarios. Such limitation discourages the applications of DFD methods. To tackle this issue, we propose a completely self-supervised framework that estimates depth purely from a sparse focal stack. We show that our framework circumvents the needs for the depth and AIF image ground-truth, and receives superior predictions, thus closing the gap between the theoretical success of DFD works and their applications in the real world. In particular, we propose (i) a more realistic setting for DFD tasks, where no depth or AIF image ground-truth is available; (ii) a novel self- supervision framework that provides reliable predictions of depth and AIF image under the challenging setting. The proposed framework uses a neural model to predict the depth and AIF image, and utilizes an optical model to validate and refine the prediction. We verify our framework on three benchmark datasets with rendered focal stacks and real focal stacks. Qualitative and quantitative evaluations show that our method provides a strong baseline for self- supervised DFD tasks. The source code is publicly avail- able at https://github.com/Ehzoahis/DEReD. Haozhe Si, Bin Zhao 0001, Dong Wang 0028, Mulin Chen, Zhigang Wang 0002, Xuelong Li 0001 |
CVPR | 2 |
| 2023 | Propagate and Calibrate: Real-Time Passive Non-Line-of-Sight TrackingabstractNon-line-of-sight (NLOS) tracking has drawn increasing attention in recent years, due to its ability to detect object motion out of sight. Most previous works on NLOS tracking rely on active illumination, e.g., laser, and suffer from high cost and elaborate experimental conditions. Besides, these techniques are still far from practical application due to oversimplified settings. In contrast, we propose a purely passive method to track a person walking in an invisible room by only observing a relay wall, which is more in line with real application scenarios, e.g., security. To excavate imperceptible changes in videos of the relay wall, we introduce difference frames as an essential carrier of temporal-local motion messages. In addition, we propose PAC-Net, which consists of alternating propagation and calibration, making it capable of leveraging both dynamic and static messages on a frame-level granularity. To evaluate the proposed method, we build and publish the first dynamic passive NLOS tracking dataset, NLOS-Track, which fills the vacuum of realistic NLOS datasets. NLOS-Track contains thousands of NLOS video clips and corresponding trajectories. Both real-shot and synthetic data are included. Our codes and dataset are available at https://againstentropy.github.io/NLOS-Track/. Zhigang Wang 0002, Bin Zhao 0001, Dong Wang 0028, Mulin Chen, Xuelong Li 0001 |
CVPR | 3 |
| 2023 | ViewRefer: Grasp the Multi-view Knowledge for 3D Visual GroundingabstractUnderstanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view cues embedded in the text modality and fail to weigh the relative importance of different views. In this paper, we propose ViewRefer, a multi-view framework for 3D visual grounding exploring how to grasp the view knowledge from both text and 3D modalities. For the text branch, ViewRefer leverages the diverse linguistic knowledge of large-scale language models, e.g., GPT, to expand a single grounding text to multiple geometry-consistent descriptions. Meanwhile, in the 3D modality, a transformer fusion module with inter-view attention is introduced to boost the interaction of objects across views. On top of that, we further present a set of learnable multi-view prototypes, which memorize scene-agnostic knowledge for different views, and enhance the framework from two perspectives: a view-guided attention module for more robust text features, and a view-guided scoring strategy during the final prediction. With our designed paradigm, ViewRefer achieves superior performance on three benchmarks and surpasses the second-best by +2.8%, +1.5%, and +1.35% on Sr3D, Nr3D, and ScanRefer. Zoey Guo, Ray Zhang 0002, Dong Wang 0028, Zhigang Wang 0002, Bin Zhao 0001, Xuelong Li 0001 |
ICCV | 6 |
| 2023 | Towards Nonlinear-Motion-Aware and Occlusion-Robust Rolling Shutter CorrectionabstractThis paper addresses the problem of rolling shutter correction in complex nonlinear and dynamic scenes with extreme occlusion. Existing methods suffer from two main drawbacks. Firstly, they face challenges in estimating the accurate correction field due to the uniform velocity assumption, leading to significant image correction errors under complex motion. Secondly, the drastic occlusion in dynamic scenes prevents current solutions from achieving better image quality because of the inherent difficulties in aligning and aggregating multiple frames. To tackle these challenges, we model the curvilinear trajectory of pixels analytically and propose a geometry-based Quadratic Rolling Shutter (QRS) motion solver, which precisely estimates the high-order correction field of individual pixels. Besides, to reconstruct high-quality occlusion frames in dynamic scenes, we present a 3D video architecture that effectively Aligns and Aggregates multi-frame context, namely, RSA2-Net. We evaluate our method across a broad range of cameras and video sequences, demonstrating its significant superiority. Specifically, our method surpasses the state-of-the-art by +4.98, +0.77, and +4.33 of PSNR on CarlaRS, Fastec-RS, and BS-RSC datasets, respectively. Code is available at https://github.com/DelinQu/qrsc. Delin Qu, Yizhen Lao, Zhigang Wang 0002, Dong Wang 0028, Bin Zhao 0001, Xuelong Li 0001 |
ICCV | 5 |
| 2023 | Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior RefinementabstractThe popularity of Contrastive Language-Image Pretraining (CLIP) has propelled its application to diverse downstream vision tasks. To improve its capacity on downstream tasks, few-shot learning has become a widelya-dopted technique. However, existing methods either exhibit limited performance or suffer from excessive learnable parameters. In this paper, we propose APE, an Adaptive Prior rEfinement method for CLIP’s pre-trained knowledge, which achieves superior accuracy with high computational efficiency. Via a prior refinement module, we analyze the inter-class disparity in the downstream data and decouple the domain-specific knowledge from the CLIP-extracted cache model. On top of that, we introduce two model variants, a training-free APE and a training-required APE-T. We explore the trilateral affinities between the test image, prior cache model, and textual representations, and only enable a lightweight category-residual module to be trained. For the average accuracy over 11 benchmarks, both APE and APE-T attain state-of-the-art and respectively outperform the second-best by +1.59% and +1.99% under 16 shots with ×30 less learnable parameters. Code is available at https://github.com/yangyangyang127/APE. Renrui Zhang, Bowei He, Aojun Zhou, Dong Wang 0004, Bin Zhao 0001, Peng Gao 0007 |
ICCV | 6 |
| 2023 | Behavior Contrastive Learning for Unsupervised Skill DiscoveryabstractIn reinforcement learning, unsupervised skill discovery aims to learn diverse skills without extrinsic rewards. Previous methods discover skills by maximizing the mutual information (MI) between states and skills. However, such an MI objective tends to learn simple and static skills and may hinder exploration. In this paper, we propose a novel unsupervised skill discovery method through contrastive learning among behaviors, which makes the agent produce similar behaviors for the same skill and diverse behaviors for different skills. Under mild assumptions, our objective maximizes the MI between different behaviors based on the same skill, which serves as an upper bound of the previous MI objective. Meanwhile, our method implicitly increases the state entropy to obtain better state coverage. We evaluate our method on challenging mazes and continuous control tasks. The results show that our method generates diverse and far-reaching skills, and also obtains competitive performance in downstream tasks compared to the state-of-the-art methods. Rushuai Yang, Chenjia Bai, Hongyi Guo, Siyuan Li 0003, Bin Zhao 0001, Zhen Wang 0004, Peng Liu 0008, Xuelong Li 0001 |
ICML | 5 |
| 2023 | Bio-Inspired Audiovisual Multi-Representation Integration via Self-Supervised LearningabstractAudiovisual self-supervised representation learning has made significant strides in various audiovisual tasks. Existing methods mostly focus on single representation modeling between audio and visual modalities, ignoring the complex correspondence between them, resulting in the inability to execute cross-modal understanding in a more natural audiovisual scene. Several biological studies have shown that human learning is influenced by multi-layered synchronization of perception. To this end, inspired by biology, we argue to exploit the naturally existing relationships in audio and visual modalities to learn audiovisual representations under multilayer perceptual integration. Firstly, we introduce an audiovisual multi-representation pretext task that integrates semantic consistency, temporal alignment, and spatial correspondence. Secondly, we propose a self-supervised audiovisual multi-representation learning approach, which simultaneously learns the perceptual relationship between visual and audio modalities at semantic, temporal, and spatial levels. To establish fine-grained correspondence between visual objects and sounds, an audiovisual object detection module is proposed, which detects potential sounding objects by combining unsupervised knowledge at multiple levels. In addition, we propose a modality-wise loss and a task-wise loss to learn a subspace-orthogonal representation space that makes representation relations more discriminative. Finally, experimental results demonstrate that collectively understanding the semantic, temporal, and spatial correspondence between audiovisual modalities enables the model to perform better on downstream tasks such as sound separation, sound spatialization, and audiovisual segmentation. Zhaojian Li 0002, Bin Zhao 0001, Yuan Yuan 0001 |
ACM Multimedia | 2 |
| 2023 | Diffusion Model is an Effective Planner and Data Synthesizer for Multi-Task Reinforcement LearningabstractDiffusion models have demonstrated highly-expressive generative capabilities in vision and NLP. Recent studies in reinforcement learning (RL) have shown that diffusion models are also powerful in modeling complex policies or trajectories in offline datasets. However, these works have been limited to single-task settings where a generalist agent capable of addressing multi-task predicaments is absent. In this paper, we aim to investigate the effectiveness of a single diffusion model in modeling large-scale multi-task offline data, which can be challenging due to diverse and multimodal data distribution. Specifically, we propose Multi-Task Diffusion Model (\textsc{MTDiff}), a diffusion-based method that incorporates Transformer backbones and prompt learning for generative planning and data synthesis in multi-task offline settings. \textsc{MTDiff} leverages vast amounts of knowledge available in multi-task data and performs implicit knowledge sharing among tasks. For generative planning, we find \textsc{MTDiff} outperforms state-of-the-art algorithms across 50 tasks on Meta-World and 8 maps on Maze2D. For data synthesis, \textsc{MTDiff} generates high-quality data for testing tasks given a single demonstration as a prompt, which enhances the low-quality datasets for even unseen tasks. Haoran He, Chenjia Bai, Zhuoran Yang, Weinan Zhang 0001, Dong Wang 0028, Bin Zhao 0001, Xuelong Li 0001 |
NeurIPS | 7 |
| 2023 | Cross-Domain Policy Adaptation via Value-Guided Data FilteringabstractGeneralizing policies across different domains with dynamics mismatch poses a significant challenge in reinforcement learning. For example, a robot learns the policy in a simulator, but when it is deployed in the real world, the dynamics of the environment may be different. Given the source and target domain with dynamics mismatch, we consider the online dynamics adaptation problem, in which case the agent can access sufficient source domain data while online interactions with the target domain are limited. Existing research has attempted to solve the problem from the dynamics discrepancy perspective. In this work, we reveal the limitations of these methods and explore the problem from the value difference perspective via a novel insight on the value consistency across domains. Specifically, we present the Value-Guided Data Filtering (VGDF) algorithm, which selectively shares transitions from the source domain based on the proximity of paired value targets across the two domains. Empirical results on various environments with kinematic and morphology shifts demonstrate that our method achieves superior performance compared to prior approaches. Chenjia Bai, Xiaoteng Ma, Dong Wang 0028, Bin Zhao 0001, Zhen Wang 0004, Xuelong Li 0001, Wei Li 0235 |
NeurIPS | 5 |
| 2023 | Edge-Guided Remote-Sensing Image CompressionabstractUsing high-fidelity image compression makes it possible to transmit remote-sensing images in real-time. Nevertheless, existing lossy remote-sensing image compression (RSIC) methods have some inherent potential issues, including blocking and blurring effects, which are particularly problematic in low-compression-ratio (CR) settings. Although numerous methods have been studied to address the aforementioned issue, the majority of them exploit the prior of local smoothness in images, which usually induces the over-smoothing of regions with noticeable structure (i.e., edges and textures). During this task, we developed an innovative end-to-end framework that enables high-fidelity RSIC while retaining sharp edge and texture information. Initially, we put forth an edge-guided adversarial network (EGA-Net) for simultaneously restoring edge structures and generating texture details. Second, we impose an edge fidelity constraint to direct our network to optimize image content and structural information jointly. In addition, to facilitate this task, we have constructed a large-scale RSIC dataset named NWPU-RS-Compression (NWPU-RSC). This dataset contains over 300000 images of 30 categories, all with a fixed resolution of 600 × 600. Finally, a new quantitative metric for full reference image quality that takes into account signal statistics and the characteristics of the human visual system (HVS) has been developed, which helps evaluate reconstructed remote-sensing images more objectively and accurately. Experimental evidence has demonstrated that the EGA-Net surpasses several representative compression approaches regarding quality metrics on the NWPU-RSC, AID, and ISPR Vaihingen datasets. Code, dataset, and more experimental results can be accessed at https: //github.com/Chenxi1510/Remote-sensing-Image-Compression. Bin Zhao 0001, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | AudioVisual Video SummarizationabstractAudio and vision are two main modalities in video data. Multimodal learning, especially for audiovisual learning, has drawn considerable attention recently, which can boost the performance of various computer vision tasks. However, in video summarization, most existing approaches just exploit the visual information while neglecting the audio information. In this brief, we argue that the audio modality can assist vision modality to better understand the video content and structure and further benefit the summarization process. Motivated by this, we propose to jointly exploit the audio and visual information for the video summarization task and develop an audiovisual recurrent network (AVRN) to achieve this. Specifically, the proposed AVRN can be separated into three parts: 1) the two-stream long-short term memory (LSTM) is used to encode the audio and visual feature sequentially by capturing their temporal dependency; 2) the audiovisual fusion LSTM is used to fuse the two modalities by exploring the latent consistency between them; and 3) the self-attention video encoder is adopted to capture the global dependency in the video. Finally, the fused audiovisual information and the integrated temporal and global dependencies are jointly used to predict the video summary. Practically, the experimental results on the two benchmarks, i.e., SumMe and TVsum, have demonstrated the effectiveness of each part and the superiority of AVRN compared with those approaches just exploiting visual information for video summarization. Bin Zhao 0001, Maoguo Gong, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingabstractMasked Autoencoders (MAE) have shown great potentials in self-supervised pre-training for language and 2D image transformers. However, it still remains an open question on how to exploit masked autoencoding for learning 3D representations of irregular point clouds. In this paper, we propose Point-M2AE, a strong Multi-scale MAE pre-training framework for hierarchical self-supervised learning of 3D point clouds. Unlike the standard transformer in MAE, we modify the encoder and decoder into pyramid architectures to progressively model spatial geometries and capture both fine-grained and high-level semantics of 3D shapes. For the encoder that downsamples point tokens by stages, we design a multi-scale masking strategy to generate consistent visible regions across scales, and adopt a local spatial self-attention mechanism during fine-tuning to focus on neighboring patterns. By multi-scale token propagation, the lightweight decoder gradually upsamples point tokens with complementary skip connections from the encoder, which further promotes the reconstruction from a global-to-local perspective. Extensive experiments demonstrate the state-of-the-art performance of Point-M2AE for 3D representation learning. With a frozen encoder after pre-training, Point-M2AE achieves 92.9% accuracy for linear SVM on ModelNet40, even surpassing some fully trained methods. By fine-tuning on downstream tasks, Point-M2AE achieves 86.43% accuracy on ScanObjectNN, +3.36% to the second-best, and largely benefits the few-shot classification, part segmentation and 3D object detection with the hierarchical pre-training scheme. Code is available at https://github.com/ZrrSkywalker/Point-M2AE. Renrui Zhang, Peng Gao 0007, Rongyao Fang, Bin Zhao 0001, Dong Wang 0004, Yu Qiao 0001, Hongsheng Li 0001 |
NeurIPS | 5 |
| 2022 | Hierarchical multimodal transformer to summarize videos
Bin Zhao 0001, Maoguo Gong, Xuelong Li 0001 |
Neurocomputing | 1 |
| 2022 | Audio-visual collaborative representation learning for Dynamic Saliency Prediction
Hailong Ning, Bin Zhao 0001, Zhanxuan Hu, Ercheng Pei |
Knowl. Based Syst. | 2 |
| 2022 | Reconstructive Sequence-Graph Network for Video SummarizationabstractExploiting the inner-shot and inter-shot dependencies is essential for key-shot based video summarization. Current approaches mainly devote to modeling the video as a frame sequence by recurrent neural networks. However, one potential limitation of the sequence models is that they focus on capturing local neighborhood dependencies while the high-order dependencies in long distance are not fully exploited. In general, the frames in each shot record a certain activity and vary smoothly over time, but the multi-hop relationships occur frequently among shots. In this case, both the local and global dependencies are important for understanding the video content. Motivated by this point, we propose a reconstructive sequence-graph network (RSGN) to encode the frames and shots as sequence and graph hierarchically, where the frame-level dependencies are encoded by long short-term memory (LSTM), and the shot-level dependencies are captured by the graph convolutional network (GCN). Then, the videos are summarized by exploiting both the local and global dependencies among shots. Besides, a reconstructor is developed to reward the summary generator, so that the generator can be optimized in an unsupervised manner, which can avert the lack of annotated data in video summarization. Furthermore, under the guidance of reconstruction loss, the predicted summary can better preserve the main video content and shot-level dependencies. Practically, the experimental results on three popular datasets (i.e., SumMe, TVsum and VTW) have demonstrated the superiority of our proposed approach to the summarization task. Bin Zhao 0001, Haopeng Li 0001, Xiaoqiang Lu, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Low-Light Hyperspectral Image EnhancementabstractDue to inadequate energy captured by the hyperspectral camera sensor in poor illumination conditions, low-light hyperspectral images (HSIs) usually suffer from low visibility, spectral distortion, and various noises. A range of HSI restoration methods have been developed, yet their effectiveness in enhancing low-light HSIs is constrained. This work focuses on the low-light HSI enhancement task, which aims to reveal the spatial-spectral information hidden in darkened areas. To facilitate the development of low-light HSI processing, we collect a low-light HSI (LHSI) dataset of both indoor and outdoor scenes. Based on Laplacian pyramid decomposition and reconstruction, we developed an end-to-end data-driven low-light HSI enhancement (HSIE) approach trained on the LHSI dataset. With the observation that illumination is related to the low-frequency component of HSI, while textural details are closely correlated to the high-frequency component, the proposed HSIE is designed to have two branches. The illumination enhancement branch is adopted to enlighten the low-frequency component with reduced resolution. The high-frequency refinement branch is utilized for refining the high-frequency component via a predicted mask. In addition, to improve information flow and boost performance, we introduce an effective channel attention block (CAB) with residual dense connection, which served as the basic block of the illumination enhancement branch. The effectiveness and efficiency of HSIE both in quantitative assessment measures and visual effects are demonstrated by experimental results on the LHSI dataset. According to the classification performance on the remote sensing Indian Pines dataset, downstream tasks benefit from the enhanced HSI. Datasets and codes are available: https://github.com/guanguanboy/HSIE. Xuelong Li 0001, Bin Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Semantics-Consistent Representation Learning for Remote Sensing Image-Voice RetrievalabstractWith the development of earth observation technology, massive amounts of remote sensing (RS) images are acquired. To find useful information from these images, cross-modal RS image–voice retrieval provides a new insight. This article aims to study the task of RS image–voice retrieval so as to search effective information from massive amounts of RS data. Existing methods for RS image–voice retrieval rely primarily on the pairwise relationship to narrow the heterogeneous semantic gap between images and voices. However, apart from the pairwise relationship included in the data sets, the intramodality and nonpaired intermodality relationships should also be considered simultaneously since the semantic consistency among nonpaired representations plays an important role in the RS image–voice retrieval task. Inspired by this, a semantics-consistent representation learning (SCRL) method is proposed for RS image–voice retrieval. The main novelty is that the proposed method takes the pairwise, intramodality, and nonpaired intermodality relationships into account simultaneously, thereby improving the semantic consistency of the learned representations for the RS image–voice retrieval. The proposed SCRL method consists of two main steps: 1) semantics encoding and 2) SCRL. First, an image encoding network is adopted to extract high-level image features with a transfer learning strategy, and a voice encoding network with dilated convolution is devised to obtain high-level voice features. Second, a consistent representation space is conducted by modeling the three kinds of relationships to narrow the heterogeneous semantic gap and learn semantics-consistent representations across two modalities. Extensive experimental results on three challenging RS image–voice data sets, including Sydney, UCM, and RSICD image–voice data sets, show the effectiveness of the proposed method. Hailong Ning, Bin Zhao 0001, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Video Crowd Localization With Multifocus Gaussian Neighborhood Attention and a Large-Scale BenchmarkabstractVideo crowd localization is a crucial yet challenging task, which aims to estimate exact locations of human heads in the given crowded videos. To model spatial-temporal dependencies of human mobility, we propose a multi-focus Gaussian neighborhood attention (GNA), which can effectively exploit long-range correspondences while maintaining the spatial topological structure of the input videos. In particular, our GNA can also capture the scale variation of human heads well using the equipped multi-focus mechanism. Based on the multi-focus GNA, we develop a unified neural network called GNANet to accurately locate head centers in video clips by fully aggregating spatial-temporal information via a scene modeling module and a context cross-attention module. Moreover, to facilitate future researches in this field, we introduce a large-scale crowd video benchmark named VSCrowd (https://github.com/HopLee6/VSCrowd), which consists of 60K+ frames captured in various surveillance scenes and 2M+ head annotations. Finally, we conduct extensive experiments on three datasets including our VSCrowd, and the experiment results show that the proposed method is capable to achieve state-of-the-art performance for both video crowd localization and counting. Haopeng Li 0001, Lingbo Liu, Shinan Liu, Junyu Gao 0001, Bin Zhao 0001, Rui Zhang 0003 |
IEEE Trans. Image Process. | 6 |
| 2020 | Property-Constrained Dual Learning for Video SummarizationabstractVideo summarization is the technique to condense large-scale videos into summaries composed of key-frames or key-shots so that the viewers can browse the video content efficiently. Recently, supervised approaches have achieved great success by taking advantages of recurrent neural networks (RNNs). Most of them focus on generating summaries by maximizing the overlap between the generated summary and the ground truth. However, they neglect the most critical principle, i.e., whether the viewer can infer the original video content from the summary. As a result, existing approaches cannot preserve the summary quality well and usually demand large amounts of training data to reduce overfitting. In our view, video summarization has two tasks, i.e., generating summaries from videos and inferring the original content from summaries. Motivated by this, we propose a dual learning framework by integrating the summary generation (primal task) and video reconstruction (dual task) together, which targets to reward the summary generator under the assistance of the video reconstructor. Moreover, to provide more guidance to the summary generator, two property models are developed to measure the representativeness and diversity of the generated summary. Practically, experiments on four popular data sets (SumMe, TVsum, OVP, and YouTube) have demonstrated that our approach, with compact RNNs as the summary generator, using less training data, and even in the unsupervised setting, can get comparable performance with those supervised ones adopting more complex summary generators and trained on more annotated data. Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Weather recognition via classification labels and weather-cue maps
Bin Zhao 0001, Lulu Hua, Xuelong Li 0001, Xiaoqiang Lu, Zhigang Wang 0002 |
Pattern Recognit. | 1 |
| 2019 | CAM-RNN: Co-Attention Model Based RNN for Video CaptioningabstractVideo captioning is a technique that bridges vision and language together, for which both visual information and text information are quite important. Typical approaches are based on the recurrent neural network (RNN), where the video caption is generated word by word, and the current word is predicted based on the visual content and previously generated words. However, in the prediction of the current word, there is much uncorrelated visual content, and some of the previously generated words provide little information, which may cause interference in generating a correct caption. Based on this point, we attempt to exploit the visual and text features that are most correlated with the caption. In this paper, a co-attention model based recurrent neural network (CAM-RNN) is proposed, where the CAM is utilized to encode the visual and text features, and the RNN works as the decoder to generate the video caption. Specifically, the CAM is composed of a visual attention module, a text attention module, and a balancing gate. During the generation procedure, the visual attention module is able to adaptively attend to the salient regions in each frame and the frames most correlated with the caption. The text attention module can automatically focus on the most relevant previously generated words or phrases. Moreover, between the two attention modules, a balancing gate is designed to regulate the influence of visual features and text features when generating the caption. In practice, the extensive experiments are conducted on four popular datasets, including MSVD, Charades, MSR-VTT, and MPII-MD, which have demonstrated the effectiveness of the proposed approach. Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu |
IEEE Trans. Image Process. | 1 |
| 2018 | HSA-RNN: Hierarchical Structure-Adaptive RNN for Video SummarizationabstractAlthough video summarization has achieved great success in recent years, few approaches have realized the influence of video structure on the summarization results. As we know, the video data follow a hierarchical structure, i.e., a video is composed of shots, and a shot is composed of several frames. Generally, shots provide the activity-level information for people to understand the video content. While few existing summarization approaches pay attention to the shot segmentation procedure. They generate shots by some trivial strategies, such as fixed length segmentation, which may destroy the underlying hierarchical structure of video data and further reduce the quality of generated summaries. To address this problem, we propose a structure-adaptive video summarization approach that integrates shot segmentation and video summarization into a Hierarchical Structure-Adaptive RNN, denoted as HSA-RNN. We evaluate the proposed approach on four popular datasets, i.e., SumMe, TVsum, CoSum and VTW. The experimental results have demonstrated the effectiveness of HSA-RNN in the video summarization task. Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu |
CVPR | 1 |
| 2018 | Video Captioning with Tube FeaturesabstractVisual feature plays an important role in the video captioning task. Considering that the video content is mainly composed of the activities of salient objects, it has restricted the caption quality of current approaches which just focus on global frame features while paying less attention to the salient objects. To tackle this problem, in this paper, we design an object-aware feature for video captioning, denoted as tube feature. Firstly, Faster-RCNN is employed to extract object regions in frames, and a tube generation method is developed to connect the regions from different frames but belonging to the same object. After that, an encoder-decoder architecture is constructed for video caption generation. Specifically, the encoder is a bi-directional LSTM, which is utilized to capture the dynamic information of each tube. The decoder is a single LSTM extended with an attention model, which enables our approach to adaptively attend to the most correlated tubes when generating the caption. We evaluate our approach on two benchmark datasets: MSVD and Charades. The experimental results have demonstrated the effectiveness of tube feature in the video captioning task. Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu |
IJCAI | 1 |
| 2018 | A CNN-RNN architecture for multi-label weather recognition
Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu, Zhigang Wang 0002 |
Neurocomputing | 1 |
| 2018 | Key Frame Extraction in the Summary SpaceabstractKey frame extraction is an efficient way to create the video summary which helps users obtain a quick comprehension of the video content. Generally, the key frames should be representative of the video content, meanwhile, diverse to reduce the redundancy. Based on the assumption that the video data are near a subspace of a high-dimensional space, a new approach, named as key frame extraction in the summary space, is proposed for key frame extraction in this paper. The proposed approach aims to find the representative frames of the video and filter out similar frames from the representative frame set. First of all, the video data are mapped to a high-dimensional space, named as summary space. Then, a new representation is learned for each frame by analyzing the intrinsic structure of the summary space. Specifically, the learned representation can reflect the representativeness of the frame, and is utilized to select representative frames. Next, the perceptual hash algorithm is employed to measure the similarity of representative frames. As a result, the key frame set is obtained after filtering out similar frames from the representative frame set. Finally, the video summary is constructed by assigning the key frames in temporal order. Additionally, the ground truth, created by filtering out similar frames from human-created summaries, is utilized to evaluate the quality of the video summary. Compared with several traditional approaches, the experimental results on 80 videos from two datasets indicate the superior performance of our approach. Xuelong Li 0001, Bin Zhao 0001, Xiaoqiang Lu |
IEEE Trans. Cybern. | 2 |
| 2017 | MAM-RNN: Multi-level Attention Model Based RNN for Video CaptioningabstractVisual information is quite important for the task of video captioning. However, in the video, there are a lot of uncorrelated content, which may cause interference to generate a correct caption. Based on this point, we attempt to exploit the visual features which are most correlated to the caption. In this paper, a Multi-level Attention Model based Recurrent Neural Network (MAM-RNN) is proposed, where MAM is utilized to encode the visual feature and RNN works as the decoder to generate the video caption. During generation, the proposed approach is able to adaptively attend to the salient regions in the frame and the frames correlated to the caption. Practically, the experimental results on two benchmark datasets, i.e., MSVD and Charades, have shown the excellent performance of the proposed approach. Xuelong Li 0001, Bin Zhao 0001, Xiaoqiang Lu |
IJCAI | 2 |
| 2017 | Hierarchical Recurrent Neural Network for Video SummarizationabstractExploiting the temporal dependency among video frames or subshots is very important for the task of video summarization. Practically, RNN is good at temporal dependency modeling, and has achieved overwhelming performance in many video-based tasks, such as video captioning and classification. However, RNN is not capable enough to handle the video summarization task, since traditional RNNs, including LSTM, can only deal with short videos, while the videos in the summarization task are usually in longer duration. To address this problem, we propose a hierarchical recurrent neural network for video summarization, called H-RNN in this paper. Specifically, it has two layers, where the first layer is utilized to encode short video subshots cut from the original video, and the final hidden state of each subshot is input to the second layer for calculating its confidence to be a key subshot. Compared to traditional RNNs, H-RNN is more suitable to video summarization, since it can exploit long temporal dependency among frames, meanwhile, the computation operations are significantly lessened. The results on two popular datasets, including the Combined dataset and VTW dataset, have demonstrated that the proposed H-RNN outperforms the state-of-the-arts. Bin Zhao 0001, Xuelong Li 0001, Xiaoqiang Lu |
ACM Multimedia | 1 |
| 2017 | A General Framework for Edited Video and Raw Video SummarizationabstractIn this paper, we build a general summarization framework for both of edited video and raw video summarization. Overall, our work can be divided into three folds. 1) Four models are designed to capture the properties of video summaries, i.e., containing important people and objects (importance), representative to the video content (representativeness), no similar key-shots (diversity), and smoothness of the storyline (storyness). Specifically, these models are applicable to both edited videos and raw videos. 2) A comprehensive score function is built with the weighted combination of the aforementioned four models. Note that the weights of the four models in the score function, denoted as property-weight, are learned in a supervised manner. Besides, the property-weights are learned for edited videos and raw videos, respectively. 3) The training set is constructed with both edited videos and raw videos in order to make up the lack of training data. Particularly, each training video is equipped with a pair of mixing-coefficients, which can reduce the structure mess in the training set caused by the rough mixture. We test our framework on three data sets, including edited videos, short raw videos, and long raw videos. Experimental results have verified the effectiveness of the proposed framework. Xuelong Li 0001, Bin Zhao 0001, Xiaoqiang Lu |
IEEE Trans. Image Process. | 2 |