EDBT 2026 Demo / reviewers in the wild / expert
Yunbo Wang
dblp:84/3894
· DBLP profile ↗
68ranked-venue papers
19as first author
44since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 9 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 9 first-author · 14 since 2021Computer networks · 4 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BioDPP: Dynamic Prompt Policy Learning for Biomedical Vision-Language ModelsabstractFoundational vision-language models (VLMs), such as CLIP, are emerging as a promising paradigm in vision tasks due to their strong generalization ability. Nevertheless, adapting them to downstream tasks remains challenging, especially in biomedical imaging, where scarce annotations, low-contrast features and complex patterns hinder model adaptation. Thus, prompt tuning is employed to facilitate the adaptation of VLMs. However, current prompt tuning methods like Context Optimization (CoOp) mainly learn a single yet static prompt which is applied to all images, and such one-size-fits-all prompt cannot describe the case-specific diagnostic cues in biomedical data, compromising the adaptation of VLMs. To this end, we propose a Dynamic Prompt Policy learning method that enables efficient adaptation of Biomedical VLMs (BioDPP) for accurate and highly generalizable few-shot biomedical image classification. Specifically, we conceptualize the learnable context as an agent, and present a paradigm of learning a dynamic prompting policy, rather than obtaining a single yet static prompt. Wherein, a dual-reward mechanism is developed to guide policy learning via the feedback on both classification decision and the consistency between the prompt and the context, steering the agent to generate context-aware prompts. Moreover, we devise adaptive baseline stabilization to dynamically regulate reward advantage value throughout the training process, enabling policy refinement in a complex reward space tailored to biomedical VLMs. Extensive experiments are conducted on 10 biomedical datasets, and the results reveal that our BioDPP achieves superior performance, demonstrating more efficient prompt optimization in biomedical VLMs. Pingyi Miao, Xianlai Chen, Yunbo Wang, Ying An |
AAAI | 4 |
| 2026 | MetaTrader: Learning to Generalize RL Trading Policies Beyond Offline DataabstractReinforcement learning (RL) has shown significant promise in sequential portfolio optimization. A typical solution involves optimizing cumulative returns using historical offline data. However, it may produce less generalizable policies that merely ''memorize'' optimal buying and selling actions from the offline data while neglecting the non-stationary nature of the financial market. We frame portfolio optimization of stock data as a specific type of offline RL problem. Our method, MetaTrader, presents two key contributions. First, it introduces a novel bilevel RL algorithm that operates on both the original stock data and its transformations. The core idea is that a robust policy should generalize effectively to out-of-distribution data. Second, we propose a new temporal difference (TD) method that leverages a transformation-based conservative TD target to address value overestimation under limited offline data. Empirical results on two publicly available datasets demonstrate that MetaTrader outperforms existing methods, including both traditional stock prediction models and RL-based trading approaches. Minting Pan, Yunbo Wang, Siyu Gao, Xiaokang Yang 0001 |
AAAI | 3 |
| 2026 | IMM-GNN: An Integrative Multi-hop and Multi-scale Graph Neural Network for Molecular Property Prediction
Xianlai Chen, Yunbo Wang, Ying An |
ISBRA (2) | 4 |
| 2026 | EADR: Efficient runtime detection and recovery of actuator attacks on UAVsabstractAbstract The actuator is the critical component of the unmanned aerial vehicle (UAV). The interference and suppression to the signal of UAV’s actuators are challenging to directly detect or physically mitigate, thus posing a significant threat to UAV flight safety. Under the assumption of sensor integrity, the state-of-the-art physics-based attack detection approaches can identify the actuator attacks at runtime. However, when both sensor and actuator attacks are allowed simultaneously, such physics-based attack detection approaches cannot differentiate between the two physical attacks, thereby failing to locate the specific compromised actuators or maintain the UAV’s resilience to the actuator attack at runtime. This paper presents EADR , an efficient runtime framework for detecting and recovering from UAV actuator attacks. By leveraging the existing signal-characteristic-based sensor attack detection mechanism, EADR prevents potential sensor attacks from impacting the resilience of actuator attacks. In response to typical attack scenarios, we implement actuator attack detection based on the nonlinear dynamic model combined with the cumulative sum (CUSUM) detection algorithm. We further locate the specific compromised actuators, determine the required compensations for the signal of these actuators at runtime, and apply the compensations to the actuators to effectively recover the UAV system state. The experimental results demonstrate that the time to detection (TTD) of EADR ’s detector is significantly reduced compared with the state-of-the-art approaches. EADR ’s recovery mechanism can reduce the flight positional error by approximately 56.3 – 77.6%. The average runtime overhead of EADR is less than 2%, ensuring the real-time performance required for real-world UAV flight. Cong Sun 0001, Penghao He, Yunbo Wang, Zongxu Zhang, Xiaomin Wei |
Cybersecur. | 3 |
| 2025 | Noisy Correspondence Rectification via Asymmetric Similarity LearningabstractCross-modal matching shows enormous potential to recognize objects across different sensory modalities, which is fundamental to numerous visual-language tasks like image-text retrieval and visual captioning. Existing works generally rely on massive and well-aligned data pairs for model training. Unfortunately, multimodal datasets are extremely difficult to annotate and collect. As an alternative, the co-occurred data pairs collected from the internet have been widely exploited to train a cross-modal matching model. However, the cheaply-collected dataset unavoidably contains mismatched pairs (i.e., noisy correspondence), which are detrimental to the matching model. In this paper, we propose an alternative method termed noisy correspondence rectification via Asymmetric Similarity Learning (ASL), and it allows for dealing with insufficient learning of positive and negative pairs caused by the popular triplet-based symmetric learning fashion. Specifically, the learning of positive or negative pairs within a triplet is conducted in an asymmetric fashion, and the self-paced weighting boundary is imposed on positive pairs to mitigate the effect of noise. Meanwhile, the optimization of negative samples will not be affected in the process of punishing potentially-noisy positive samples. To verify the effectiveness of our proposed approach, a series of experiments are conducted on three widely-used benchmarks (i.e., Flick30k, MS-COCO and CC152k), and the results show superior performance compared to the state-of-the-art methods. Yunbo Wang, YuJie Wu, Zhien Dai, Can Tian, Jianhai Chen |
AAAI | 1 |
| 2025 | HiGraphSum: Hierarchical Hypergraph-Based Dual-Stage Automatic Summarization of Discharge SummariesabstractAutomatic summarization of discharge summaries aims to generate a concise yet informative overview from complex electronic health records (EHRs), thus alleviating the documentation burden on clinicians. Existing methods usually employ pre-trained Transformer architectures to capture textual dependency via self-attention mechanism. However, these methods fail to utilize the information of multi-level hierarchical structure in EHR, compressing the understanding of intricate semantic interaction among different structure levels. In this work, we propose HiGraphSum, a hierarchical hypergraphbased dual-stage model for automatic summarization of discharge summaries. In the first stage, HiGraphSum constructs a hierarchical hypergraph that integrates heterogeneous wordsentence graphs and sentence-based hypergraphs to capture hierarchical and contextual information, facilitating global document structure and intricate clinical semantics learning. In the second stage, we introduce a probability concentration strategy (PCS) to refine summary generation by sharpening the output distribution, thereby enhancing the factual consistency of the generated summaries. Extensive experiments on discharge summarise datasets demonstrate that HiGraphSum outperforms state-of-the-art baselines. Xianlai Chen, Taixiang Li, Ying An, Yunbo Wang |
BIBM | 6 |
| 2025 | Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningabstractTraining visual reinforcement learning (RL) in practical scenarios presents a significant challenge, $\textit{i.e.,}$ RL agents suffer from low sample efficiency in environments with variations. While various approaches have attempted to alleviate this issue by disentangled representation learning, these methods usually start learning from scratch without prior knowledge of the world. This paper, in contrast, tries to learn and understand underlying semantic variations from distracting videos via offline-to-online latent distillation and flexible disentanglement constraints. To enable effective cross-domain semantic knowledge transfer, we introduce an interpretable model-based RL framework, dubbed Disentangled World Models (DisWM). Specifically, we pretrain the action-free video prediction model offline with disentanglement regularization to extract semantic knowledge from distracting videos. The disentanglement capability of the pretrained model is then transferred to the world model through latent distillation. For finetuning in the online environment, we exploit the knowledge from the pretrained model and introduce a disentanglement constraint to the world model. During the adaptation phase, the incorporation of actions and rewards from online environment interactions enriches the diversity of the data, which in turn strengthens the disentangled representation learning. Experimental results validate the superiority of our approach on various benchmarks. Qi Wang 0080, Baao Xie, Xin Jin 0014, Yunbo Wang, Liaomo Zheng, Xiaokang Yang 0001, Wenjun Zeng 0001 |
ICCV | 5 |
| 2025 | Open-World Reinforcement Learning over Long Short-Term ImaginationabstractTraining visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be “short-sighted”, as they are typically trained on short snippets of imagined experiences. We argue that the primary challenge in open-world decision-making is improving the exploration efficiency across a vast state space, especially for tasks that demand consideration of long-horizon payoffs. In this paper, we present LS-Imagine, which extends the imagination horizon within a limited number of state transition steps, enabling the agent to explore behaviors that potentially lead to promising long-term feedback. The foundation of our approach is to build a $\textit{long short-term world model}$. To achieve this, we simulate goal-conditioned jumpy state transitions and compute corresponding affordance maps by zooming in on specific areas within single images. This facilitates the integration of direct long-term values into behavior learning. Our method demonstrates significant improvements over state-of-the-art techniques in MineDojo. Jiajian Li, Qi Wang 0080, Yunbo Wang, Xin Jin 0014, Yang Li 0041, Wenjun Zeng 0001, Xiaokang Yang 0001 |
ICLR | 3 |
| 2025 | EvoMesh: Adaptive Physical Simulation with Hierarchical Graph EvolutionsabstractGraph neural networks have been a powerful tool for mesh-based physical simulation. To efficiently model large-scale systems, existing methods mainly employ hierarchical graph structures to capture multi-scale node relations. However, these graph hierarchies are typically manually designed and fixed, limiting their ability to adapt to the evolving dynamics of complex physical systems. We propose EvoMesh, a fully differentiable framework that jointly learns graph hierarchies and physical dynamics, adaptively guided by physical inputs. EvoMesh introduces anisotropic message passing, which enables direction-specific aggregation of dynamic features between nodes within each hierarchy, while simultaneously learning node selection probabilities for the next hierarchical level based on physical context. This design creates more flexible message shortcuts and enhances the model’s capacity to capture long-range dependencies. Extensive experiments on five benchmark physical simulation datasets show that EvoMesh outperforms recent fixed-hierarchy message passing networks by large margins. The project page is available at https://hbell99.github.io/evo-mesh/. Huayu Deng, Xiangming Zhu 0002, Yunbo Wang, Xiaokang Yang 0001 |
ICML | 3 |
| 2025 | Video-Enhanced Offline Reinforcement Learning: A Model-Based ApproachabstractOffline reinforcement learning (RL) enables policy optimization using static datasets, avoiding the risks and costs of extensive real-world exploration. However, it struggles with suboptimal offline behaviors and inaccurate value estimation due to the lack of environmental interaction. We present Video-Enhanced Offline RL (VeoRL), a model-based method that constructs an interactive world model from diverse, unlabeled video data readily available online. Leveraging model-based behavior guidance, our approach transfers commonsense knowledge of control policy and physical dynamics from natural videos to the RL agent within the target domain. VeoRL achieves substantial performance gains (over 100% in some cases) across visual control tasks in robotic manipulation, autonomous driving, and open-world video games. Project page: https://panmt.github.io/VeoRL.github.io. Minting Pan, Yitao Zheng, Jiajian Li, Yunbo Wang, Xiaokang Yang 0001 |
ICML | 4 |
| 2025 | MetaGS: A Meta-Learned Gaussian-Phong Model for Out-of-Distribution 3D Scene RelightingabstractOut-of-distribution (OOD) 3D relighting requires novel view synthesis under unseen lighting conditions that differ significantly from the observed images. Existing relighting methods, which assume consistent light source distributions between training and testing, often degrade in OOD scenarios. We introduce **MetaGS** to tackle this challenge from two perspectives. First, we propose a meta-learning approach to train 3D Gaussian splatting, which explicitly promotes learning generalizable Gaussian geometries and appearance attributes across diverse lighting conditions, even with biased training data. Second, we embed fundamental physical priors from the *Blinn-Phong* reflection model into Gaussian splatting, which enhances the decoupling of shading components and leads to more accurate 3D scene reconstruction. Results on both synthetic and real-world datasets demonstrate the effectiveness of MetaGS in challenging OOD relighting tasks, supporting efficient point-light relighting and generalizing well to unseen environment lighting maps. Yunbo Wang |
NeurIPS | 2 |
| 2025 | Continual Visual Reinforcement Learning with A Life-Long World Model
Minting Pan, Wendong Zhang 0002, Xiangming Zhu 0002, Siyu Gao, Yunbo Wang, Xiaokang Yang 0001 |
ECML/PKDD (6) | 6 |
| 2025 | An intelligent fault diagnosis for rotating machine under strong noise based on cross-attention-driven spatial-temporal feature fusion and duplexing time sequence convolution optimization
Yongpeng Li, Guanglei Meng, Yunbo Wang |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | A survey on biomedical automatic text summarization with large language modelsabstractAutomatic text summarization in the biomedical field can support efficient literature screening, medical knowledge management, and innovative medical research. In recent years, Large Language Models (LLMs), as a disruptive technology in natural language processing, have shown great potential for Biomedical Automatic Text Summarization (BATS). This technology helps to better understand the terminology of biomedical texts, track medical hotspots, and generate personalized diagnoses and treatment plans. This paper provides an in-depth discussion on the development of BATS, and the opportunities as well as challenges brought by applying LLMs to biomedical automatic text summarization. Firstly, the development of BATS is reviewed, where traditional text summarization, neural network-based summarization, and LLMs-based summarization are analyzed systematically. Meanwhile, the applications of various LLMs (e.g., BERT and GPT series) in three types of BATS are presented in detail, including extractive summarization, abstractive summarization, and hybrid summarization. Next, the relevant datasets are introduced, such as PubMed, COVID-19 and MIMIC-Ⅲ. Then, traditional, emerging, and auxiliary metrics for evaluating the performance of BATS are shown, and the performance evaluation of different models is elaborated. Finally, the opportunities brought by applying LLMs to BATS are described, and the potential challenges along with the corresponding solutions are discussed. Xianlai Chen, Yunbo Wang, Jincai Huang 0002 |
Inf. Process. Manag. | 3 |
| 2025 | Dynamic Scene Understanding Through Object-Centric Voxelization and Neural RenderingabstractLearning object-centric representations from unsupervised videos is challenging. Unlike most previous approaches that focus on decomposing 2D images, we present a 3D generative model named DynaVol-S for dynamic scenes that enables object-centric learning within a differentiable volume rendering framework. The key idea is to perform object-centric voxelization to capture the 3D nature of the scene, which infers per-object occupancy probabilities at individual spatial locations. These voxel features evolve through a canonical-space deformation function and are optimized in an inverse rendering pipeline with a compositional NeRF. Additionally, our approach integrates 2D semantic features to create 3D semantic grids, representing the scene through multiple disentangled voxel grids. DynaVol-S significantly outperforms existing models in both novel view synthesis and unsupervised decomposition tasks for dynamic scenes. By jointly considering geometric structures and semantic features, it effectively addresses challenging real-world scenarios involving complex object interactions. Furthermore, once trained, the explicitly meaningful voxel features enable additional capabilities that 2D scene decomposition methods cannot achieve, such as novel scene generation through editing geometric shapes or manipulating the motion trajectories of objects. Yanpeng Zhao, Yiwei Hao, Siyu Gao, Yunbo Wang, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | VIMU: Effective Physics-based Realtime Detection and Recovery against Stealthy Attacks on UAVsabstractSensor attacks on robotic vehicles have become pervasive and manipulative. Their latest advancements exploit sensor and detector characteristics to bypass detection. Recent security efforts have leveraged the physics-based model to detect or mitigate sensor attacks. However, these approaches are only resilient to a few sensor attacks and still need improvement in detection effectiveness. We present VIMU, an efficient sensor attack detection and resilience system for unmanned aerial vehicles. We propose a detection algorithm, CS-EMA, that leverages low-pass filtering to identify stealthy gyroscope attacks while achieving an overall effective sensor attack detection. We develop a fine-grained nonlinear physical model with precise aerodynamic and propulsion wrench modeling. We also augment the state estimation with a FIFO buffer safeguard to mitigate the impact of high-rate IMU attacks. The proposed physical model and buffer safeguard provide an effective system state recovery toward maintaining flight stability. We implement VIMU on PX4 autopilot. The evaluation results demonstrate the effectiveness of VIMU in detecting and mitigating various realistic sensor attacks, especially stealthy attacks. Yunbo Wang, Cong Sun 0001, Qiaosen Liu, Bingnan Su, Zongxu Zhang, Michael Norris, Gang Tan, Jianfeng Ma 0001 |
ACSAC | 1 |
| 2024 | Latent Intuitive Physics: Learning to Transfer Hidden Physics from A 3D VideoabstractWe introduce latent intuitive physics, a transfer learning framework for physics simulation that can infer hidden properties of fluids from a single 3D video and simulate the observed fluid in novel scenes. Our key insight is to use latent features drawn from a learnable prior distribution conditioned on the underlying particle states to capture the invisible and complex physical properties. To achieve this, we train a parametrized prior learner given visual observations to approximate the visual posterior of inverse graphics, and both the particle states and the visual posterior are obtained from a learned neural renderer. The converged prior learner is embedded in our probabilistic physics engine, allowing us to perform novel simulations on unseen geometries, boundaries, and dynamics without knowledge of the true physical parameters. We validate our model in three ways: (i) novel scene simulation with the learned visual-world physics, (ii) future prediction of the observed fluid dynamics, and (iii) supervised particle simulation. Our model demonstrates strong performance in all three tasks. Xiangming Zhu 0002, Huayu Deng, Yunbo Wang, Xiaokang Yang 0001 |
ICLR | 4 |
| 2024 | DynaVol: Unsupervised Learning for Dynamic Scenes through Object-Centric VoxelizationabstractUnsupervised learning of object-centric representations in dynamic visual scenes is challenging. Unlike most previous approaches that learn to decompose 2D images, we present DynaVol, a 3D scene generative model that unifies geometric structures and object-centric learning in a differentiable volume rendering framework. The key idea is to perform object-centric voxelization to capture the 3D nature of the scene, which infers the probability distribution over objects at individual spatial locations. These voxel features evolve over time through a canonical-space deformation function, forming the basis for global representation learning via slot attention. The voxel features and global features are complementary and are both leveraged by a compositional NeRF decoder for volume rendering. DynaVol remarkably outperforms existing approaches for unsupervised dynamic scene decomposition. Once trained, the explicitly meaningful voxel features enable additional capabilities that 2D scene decomposition methods cannot achieve: it is possible to freely edit the geometric shapes or manipulate the motion trajectories of the objects. Yanpeng Zhao, Siyu Gao, Yunbo Wang, Xiaokang Yang 0001 |
ICLR | 3 |
| 2024 | Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement LearningabstractTraining offline RL models using visual inputs poses two significant challenges, *i.e.*, the overfitting problem in representation learning and the overestimation bias for expected future rewards. Recent work has attempted to alleviate the overestimation bias by encouraging conservative behaviors. This paper, in contrast, tries to build more flexible constraints for value estimation without impeding the exploration of potential advantages. The key idea is to leverage off-the-shelf RL simulators, which can be easily interacted with in an online manner, as the “*test bed*” for offline policies. To enable effective online-to-offline knowledge transfer, we introduce CoWorld, a model-based RL approach that mitigates cross-domain discrepancies in state and reward spaces. Experimental results demonstrate the effectiveness of CoWorld, outperforming existing RL approaches by large margins. Qi Wang 0080, Yunbo Wang, Xin Jin 0014, Wenjun Zeng 0001, Xiaokang Yang 0001 |
NeurIPS | 3 |
| 2024 | Model-Based Reinforcement Learning with Multi-task Offline Pretraining
Minting Pan, Yitao Zheng, Yunbo Wang, Xiaokang Yang 0001 |
ECML/PKDD (7) | 3 |
| 2024 | Model-Based Reinforcement Learning With Isolated ImaginationsabstractWorld models learn the consequences of actions in vision-based interactive systems. However, in practical scenarios like autonomous driving, noncontrollable dynamics that are independent or sparsely dependent on action signals often exist, making it challenging to learn effective world models. To address this issue, we propose Iso-Dream++, a model-based reinforcement learning approach that has two main contributions. First, we optimize the inverse dynamics to encourage the world model to isolate controllable state transitions from the mixed spatiotemporal variations of the environment. Second, we perform policy optimization based on the decoupled latent imaginations, where we roll out noncontrollable states into the future and adaptively associate them with the current controllable state. This enables long-horizon visuomotor control tasks to benefit from isolating mixed dynamics sources in the wild, such as self-driving cars that can anticipate the movement of other vehicles, thereby avoiding potential risks. On top of our previous work (Pan et al. 2022), we further consider the sparse dependencies between controllable and noncontrollable states, address the training collapse problem of state decoupling, and validate our approach in transfer learning setups. Our empirical study demonstrates that Iso-Dream++ outperforms existing reinforcement learning models significantly on CARLA and DeepMind Control. Minting Pan, Xiangming Zhu 0002, Yitao Zheng, Yunbo Wang, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | StockFormer: Learning Hybrid Trading Machines with Predictive CodingabstractTypical RL-for-finance solutions directly optimize trading policies over the noisy market data, such as stock prices and trading volumes, without explicitly considering the future trends and correlations of different investment assets as we humans do. In this paper, we present StockFormer, a hybrid trading machine that integrates the forward modeling capabilities of predictive coding with the advantages of RL agents in policy flexibility. The predictive coding part consists of three Transformer branches with modified structures, which respectively extract effective latent states of long-/short-term future dynamics and asset relations. The RL agent adaptively fuses these states and then executes an actor-critic algorithm in the unified state space. The entire model is jointly trained by propagating the critic's gradients back to the predictive coding module. StockFormer significantly outperforms existing approaches across three publicly available financial datasets in terms of portfolio returns and Sharpe ratios. Siyu Gao, Yunbo Wang, Xiaokang Yang 0001 |
IJCAI | 2 |
| 2023 | Improving Masked Autoencoders by Learning Where to Mask
Haijian Chen, Wendong Zhang 0002, Yunbo Wang, Xiaokang Yang 0001 |
PRCV (8) | 3 |
| 2023 | Out-of-Domain Human Mesh Reconstruction via Dynamic Bilevel Online AdaptationabstractWe consider a new problem of adapting a human mesh reconstruction model to out-of-domain streaming videos, where the performance of existing SMPL-based models is significantly affected by the distribution shift represented by different camera parameters, bone lengths, backgrounds, and occlusions. We tackle this problem through online adaptation, gradually correcting the model bias during testing. There are two main challenges: First, the lack of 3D annotations increases the training difficulty and results in 3D ambiguities. Second, non-stationary data distribution makes it difficult to strike a balance between fitting regular frames and hard samples with severe occlusions or dramatic changes. To this end, we propose the Dynamic Bilevel Online Adaptation algorithm (DynaBOA). It first introduces the temporal constraints to compensate for the unavailable 3D annotations and leverages a bilevel optimization procedure to address the conflicts between multi-objectives. DynaBOA provides additional 3D guidance by co-training with similar source examples retrieved efficiently despite the distribution shift. Furthermore, it can adaptively adjust the number of optimization steps on individual frames to fully fit hard samples and avoid overfitting regular frames. DynaBOA achieves state-of-the-art results on three out-of-domain human mesh reconstruction benchmarks. Shanyan Guan, Jingwei Xu 0005, Michelle Zhang He, Yunbo Wang, Bingbing Ni, Xiaokang Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | PredRNN: A Recurrent Neural Network for Spatiotemporal Predictive LearningabstractThe predictive learning of spatiotemporal sequences aims to generate future images by learning from the historical context, where the visual dynamics are believed to have modular structures that can be learned with compositional subsystems. This paper models these structures by presenting PredRNN, a new recurrent network, in which a pair of memory cells are explicitly decoupled, operate in nearly independent transition manners, and finally form unified representations of the complex environment. Concretely, besides the original memory cell of LSTM, this network is featured by a zigzag memory flow that propagates in both bottom-up and top-down directions across all layers, enabling the learned visual dynamics at different levels of RNNs to communicate. It also leverages a memory decoupling loss to keep the memory cells from learning redundant features. We further propose a new curriculum learning strategy to force PredRNN to learn long-term dynamics from context frames, which can be generalized to most sequence-to-sequence models. We provide detailed ablation studies to verify the effectiveness of each component. Our approach is shown to obtain highly competitive results on five datasets for both action-free and action-conditioned predictive learning scenarios. Yunbo Wang, Haixu Wu, Jianjin Zhang, Zhifeng Gao, Jianmin Wang 0001, Philip S. Yu, Mingsheng Long |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | ModeRNN: Harnessing Spatiotemporal Mode Collapse in Unsupervised Predictive LearningabstractLearning predictive models for unlabeled spatiotemporal data is challenging in part because visual dynamics can be highly entangled, especially in real scenes. In this paper, we refer to the multi-modal output distribution of predictive learning as spatiotemporal modes. We find an experimental phenomenon named spatiotemporal mode collapse (STMC) on most existing video prediction models, that is, features collapse into invalid representation subspaces due to the ambiguous understanding of mixed physical processes. We propose to quantify STMC and explore its solution for the first time in the context of unsupervised predictive learning. To this end, we present ModeRNN, a decoupling-aggregation framework that has a strong inductive bias of discovering the compositional structures of spatiotemporal modes between recurrent states. We first leverage a set of dynamic slots with independent parameters to extract individual building components of spatiotemporal modes. We then perform a weighted fusion of slot features to adaptively aggregate them into a unified hidden representation for recurrent updates. Through a series of experiments, we show high correlation between STMC and the fuzzy prediction results of future video frames. Besides, ModeRNN is shown to better mitigate STMC and achieve the state of the art on five video prediction datasets. Zhiyu Yao, Yunbo Wang, Haixu Wu, Jianmin Wang 0001, Mingsheng Long |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Fully context-aware image inpainting with a learned semantic pyramid
Wendong Zhang 0002, Yunbo Wang, Bingbing Ni, Xiaokang Yang 0001 |
Pattern Recognit. | 2 |
| 2023 | An efficient communication strategy for massively parallel computation in CFD
Yunbo Wang, Lei He 0012, HaoYuan Zhang |
J. Supercomput. | 1 |
| 2022 | Continual Predictive Learning from VideosabstractPredictive learning ideally builds the world model of physical processes in one or more given environments. Typical setups assume that we can collect data from all environments at all times. In practice, however, different prediction tasks may arrive sequentially so that the environments may change persistently throughout the training procedure. Can we develop predictive learning algorithms that can deal with more realistic, non-stationary physical environments? In this paper, we study a new continual learning problem in the context of video prediction, and observe that most existing methods suffer from severe catastrophic forgetting in this setup. To tackle this problem, we propose the continual predictive learning (CPL) approach, which learns a mixture world model via predictive experience replay and performs test-time adaptation with non-parametric task inference. We construct two new benchmarks based on RoboNet and KTH, in which different tasks correspond to different physical robotic environments or human actions. Our approach is shown to effectively mitigate forgetting and remarkably outperform the naíve combinations of previous art in video prediction and continual learning. Wendong Zhang 0002, Han Lu 0004, Siyu Gao, Yunbo Wang, Mingsheng Long, Xiaokang Yang 0001 |
CVPR | 5 |
| 2022 | NeuroFluid: Fluid Dynamics Grounding with Particle-Driven Neural Radiance FieldsabstractDeep learning has shown great potential for modeling the physical dynamics of complex particle systems such as fluids. Existing approaches, however, require the supervision of consecutive particle properties, including positions and velocities. In this paper, we consider a partially observable scenario known as fluid dynamics grounding, that is, inferring the state transitions and interactions within the fluid particle systems from sequential visual observations of the fluid surface. We propose a differentiable two-stage network named NeuroFluid. Our approach consists of (i) a particle-driven neural renderer, which involves fluid physical properties into the volume rendering function, and (ii) a particle transition model optimized to reduce the differences between the rendered and the observed images. NeuroFluid provides the first solution to unsupervised learning of particle-based fluid dynamics by training these two models jointly. It is shown to reasonably estimate the underlying physics of fluids with different initial shapes, viscosity, and densities. Shanyan Guan, Huayu Deng, Yunbo Wang, Xiaokang Yang 0001 |
ICML | 3 |
| 2022 | Iso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World ModelsabstractWorld models learn the consequences of actions in vision-based interactive systems. However, in practical scenarios such as autonomous driving, there commonly exists noncontrollable dynamics independent of the action signals, making it difficult to learn effective world models. Naturally, therefore, we need to enable the world models to decouple the controllable and noncontrollable dynamics from the entangled spatiotemporal data. To this end, we present a reinforcement learning approach named Iso-Dream, which expands the Dream-to-Control framework in two aspects. First, the world model contains a three-branch neural architecture. By solving the inverse dynamics problem, it learns to factorize latent representations according to the responses to action signals. Second, in the process of behavior learning, we estimate the state values by rolling-out a sequence of noncontrollable states (less related to the actions) into the future and associate the current controllable state with them. In this way, the isolation of mixed dynamics can greatly facilitate long-horizon decision-making tasks in realistic scenes, such as avoiding potential future risks by predicting the movement of other vehicles in autonomous driving. Experiments show that Iso-Dream is effective in decoupling the mixed dynamics and remarkably outperforms existing approaches in a wide range of visual control and prediction domains. Minting Pan, Xiangming Zhu 0002, Yunbo Wang, Xiaokang Yang 0001 |
NeurIPS | 3 |
| 2022 | Cross-modal image-text search via Efficient Discrete Class Alignment Hashing
Song Wang 0016, Huan Zhao 0003, Yunbo Wang, Jing Huang 0012, Keqin Li 0001 |
Inf. Process. Manag. | 3 |
| 2022 | Semantic association enhancement transformer with relative position for image captioning
Yunbo Wang, Yuxin Peng 0001, Shengyong Chen |
Multim. Tools Appl. | 2 |
| 2022 | Source Data-Absent Unsupervised Domain Adaptation Through Hypothesis Transfer and Labeling TransferabstractUnsupervised domain adaptation (UDA) aims to transfer knowledge from a related but different well-labeled source domain to a new unlabeled target domain. Most existing UDA methods require access to the source data, and thus are not applicable when the data are confidential and not shareable due to privacy concerns. This paper aims to tackle a realistic setting with only a classification model available trained over, instead of accessing to, the source data. To effectively utilize the source model for adaptation, we propose a novel approach called Source HypOthesis Transfer (SHOT), which learns the feature extraction module for the target domain by fitting the target data features to the frozen source classification module (representing classification hypothesis). Specifically, SHOT exploits both information maximization and self-supervised learning for the feature extraction module learning to ensure the target features are implicitly aligned with the features of unseen source data via the same hypothesis. Furthermore, we propose a new labeling transfer strategy, which separates the target data into two splits based on the confidence of predictions (labeling information), and then employ semi-supervised learning to improve the accuracy of less-confident predictions in the target domain. We denote labeling transfer as SHOT++ if the predictions are obtained by SHOT. Extensive experiments on both digit classification and object recognition tasks show that SHOT and SHOT++ achieve results surpassing or comparable to the state-of-the-arts, demonstrating the effectiveness of our approaches for various visual domain adaptation problems. Code will be available at https://github.com/tim-learn/SHOT-plus. Jian Liang 0001, Dapeng Hu, Yunbo Wang, Ran He 0001, Jiashi Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | VideoDG: Generalizing Temporal Relations in Videos to Novel DomainsabstractThis paper introduces video domain generalization where most video classification networks degenerate due to the lack of exposure to the target domains of divergent distributions. We observe that the global temporal features are less generalizable, due to the temporal domain shift that videos from other unseen domains may have an unexpected absence or misalignment of the temporal relations. This finding has motivated us to solve video domain generalization by effectively learning the local-relation features of different timescales that are more generalizable, and exploiting them along with the global-relation features to maintain the discriminability. This paper presents the VideoDG framework with two technical contributions. The first is a new deep architecture named the Adversarial Pyramid Network, which improves the generalizability of video features by capturing the local-relation, global-relation, and cross-relation features progressively. On the basis of pyramid features, the second contribution is a new and robust approach of adversarial data augmentation that can bridge different video domains by improving the diversity and quality of augmented data. We construct three video domain generalization benchmarks in which domains are divided according to different datasets, different consequences of actions, or different camera views, respectively. VideoDG consistently outperforms the combinations of previous video classification models and existing domain generalization methods on all benchmarks. Zhiyu Yao, Yunbo Wang, Jianmin Wang 0001, Philip S. Yu, Mingsheng Long |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | MARS: Learning Modality-Agnostic Representation for Scalable Cross-Media RetrievalabstractCross-media retrieval (CMR) offers a flexible retrieval experience across multiple modalities. Existing CMR approaches are constrained by the assumption that the paired modalities are available in training, and they leverage the data of all modalities to obtain a common representation. However, as dealing with the data from new modality, the previous all modalities need to be re-trained, compromising the flexibility and practicality of CMR. In this paper, we propose an approach termed learning Modality-Agnostic Representation for Scalable cross-media retrieval (MARS), which allows each modality to be trained independently. To be specific, MARS treats the label information as a distinct modality, and introduces a label parsing module LabNet to generate semantic representation for correlating different modalities. Meanwhile, MARS constructs the modality-specific representation module DataNet to obtain the modality-shared representation and modality-exclusive representation equipped with unbiased semantic classification. Technically, for the first modality, we jointly train the LabNet and its DataNet to preserve the semantic similarity between the Label-derived representation and the modality-shared representation. For new modalities, MARS employs the well-learned LabNet to extract the representation in labels, and then such representation is served as the privilege to guide the associated DataNet training via the same objective. Furthermore, we assign the same classifier to the representation module of all modalities for better semantic alignment. With the above schema, the obtained modality-shared representation is considered to be modality-agnostic. Extensive experiments on several benchmark multi-modality datasets demonstrate that the proposed MARS achieves better results than existing methods. Yunbo Wang, Yuxin Peng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Dual-View 3D Reconstruction via Learning Correspondence and Dependency of Point Cloud RegionsabstractMulti-view 3D reconstruction generally adopts the feature fusion strategy to guide the generation of 3D shape for objects with different views. Empirically, the correspondence learning of object regions across different views enables better feature fusion. However, such idea has not been fully exploited in existing methods. Furthermore, current methods fail to explore the intrinsic dependency among regions within a 3D shape, leading to a rough reconstruction result. To address the above issues, we propose a Dual-View 3D Point Cloud reconstruction architecture named DVPC, which takes two views images as inputs, and progressively generates a refined 3D point cloud. First, a point cloud generation network is assigned to generate a coarse point cloud for each input view. Second, a dual-view point clouds synthesis network is presented in DVPC. It constructs a regional attention mechanism to learn a high-quality correspondence among regions across two coarse point clouds in different views, so that our DVPC can achieve feature fusion accurately. And then it develops a point cloud deformation module to produce a relatively-precise point cloud via establishing the communication between the coarse point cloud and the fused feature. Lastly, a point-region transformer network is devised to model the dependency among regions within the relatively-precise point cloud. With the dependency, the relatively-precise point cloud is refined into a desirable 3D point cloud with rich details. Qualitative and quantitative experiments on the ShapeNet and Pix3D datasets demonstrate that the proposed DVPC outperforms the state-of-the-art methods in terms of reconstruction quality. Jianxin Wang 0001, Shourui Yang, Yunbo Wang, Jianhua Zhang 0002, Yuxin Peng 0001, Shengyong Chen |
IEEE Trans. Image Process. | 3 |
| 2022 | Verifiable Semantic-Aware Ranked Keyword Search in Cloud-Assisted Edge ComputingabstractRanked keyword search has gained Ranked keyword search has gained traction due to its attractive properties such as flexibility and accessibility. However, most existing ranked keyword search schemes ignore the semantic associations between the documents and queries. To solve this challenging issue in cloud-assisted edge computing, we first design theSemantic-awareRankedMulti-keywordSearch (SRMS) scheme by adopting the Latent Dirichlet Allocation (LDA) topic model and the Chinese Remainder Theorem (CRT)-based secret sharing mechanism. Considering that the cloud server may be malicious, we implement a basic verification mechanism in SRMS to verify the correctness and completeness of search results and extend this verification mechanism in cloud-assisted edge computing scenarios. Formal security analysis proves that SRMS and extended result verification mechanisms are secure in both the known ciphertext model and the known background model. Extensive experiments using the real-world dataset demonstrate that SRMS is efficient and practical. Jianfeng Ma 0001, Yinbin Miao, Lei Chen 0029, Yunbo Wang, Ximeng Liu, Kim-Kwang Raymond Choo |
IEEE Trans. Serv. Comput. | 5 |
| 2021 | Bilevel Online Adaptation for Out-of-Domain Human Mesh ReconstructionabstractThis paper considers a new problem of adapting a pretrained model of human mesh reconstruction to out-of-domain streaming videos. However, most previous methods based on the parametric SMPL model [36] underperform in new domains with unexpected, domain-specific attributes, such as camera parameters, lengths of bones, backgrounds, and occlusions. Our general idea is to dynamically fine-tune the source model on test video streams with additional temporal constraints, such that it can mitigate the domain gaps without over-fitting the 2D information of individual test frames. A subsequent challenge is how to avoid conflicts between the 2D and temporal constraints. We propose to tackle this problem using a new training algorithm named Bilevel Online Adaptation (BOA), which divides the optimization process of overall multi-objective into two steps of weight probe and weight update in a training iteration. We demonstrate that BOA leads to state-of-the-art results on two human mesh reconstruction benchmarks1. Shanyan Guan, Jingwei Xu 0005, Yunbo Wang, Bingbing Ni, Xiaokang Yang 0001 |
CVPR | 3 |
| 2021 | MetaSets: Meta-Learning on Point Sets for Generalizable RepresentationsabstractDeep learning techniques for point clouds have achieved strong performance on a range of 3D vision tasks. However, it is costly to annotate large-scale point sets, making it critical to learn generalizable representations that can transfer well across different point sets. In this paper, we study a new problem of 3D Domain Generalization (3DDG) with the goal to generalize the model to other unseen domains of point clouds without any access to them in the training process. It is a challenging problem due to the substantial geometry shift from simulated to real data, such that most existing 3D models underperform due to overfitting the complete geometries in the source domain. We propose to tackle this problem via MetaSets, which meta-learns point cloud representations from a group of classification tasks on carefully-designed transformed point sets containing specific geometry priors. The learned representations are more generalizable to various unseen domains of different geometries. We design two benchmarks for Sim-to-Real transfer of 3D point clouds. Experimental results show that MetaSets outperforms existing 3D deep learning methods by large margins. Zhangjie Cao, Yunbo Wang, Jianmin Wang 0001, Mingsheng Long |
CVPR | 3 |
| 2021 | Context-Aware Image Inpainting with Learned Semantic PriorsabstractRecent advances in image inpainting have shown impressive results for generating plausible visual details on rather simple backgrounds. However, for complex scenes, it is still challenging to restore reasonable contents as the contextual information within the missing regions tends to be ambiguous. To tackle this problem, we introduce pretext tasks that are semantically meaningful to estimating the missing contents. In particular, we perform knowledge distillation on pretext models and adapt the features to image inpainting. The learned semantic priors ought to be partially invariant between the high-level pretext task and low-level image inpainting, which not only help to understand the global context but also provide structural guidance for the restoration of local textures. Based on the semantic priors, we further propose a context-aware image inpainting model, which adaptively integrates global semantics and local features in a unified image generator. The semantic learner and the image generator are trained in an end-to-end manner. We name the model SPL to highlight its ability to learn and leverage semantic priors. It achieves the state of the art on Places2, CelebA, and Paris StreetView datasets Wendong Zhang 0002, Ying Tai, Yunbo Wang, Wenqing Chu, Bingbing Ni, Chengjie Wang 0001, Xiaokang Yang 0001 |
IJCAI | 4 |
| 2021 | Learning Transferable Features for Point Cloud Detection via 3D Contrastive Co-trainingabstractMost existing point cloud detection models require large-scale, densely annotated datasets. They typically underperform in domain adaptation settings, due to geometry shifts caused by different physical environments or LiDAR sensor configurations. Therefore, it is challenging but valuable to learn transferable features between a labeled source domain and a novel target domain, without any access to target labels. To tackle this problem, we introduce the framework of 3D Contrastive Co-training (3D-CoCo) with two technical contributions. First, 3D-CoCo is inspired by our observation that the bird-eye-view (BEV) features are more transferable than low-level geometry features. We thus propose a new co-training architecture that includes separate 3D encoders with domain-specific parameters, as well as a BEV transformation module for learning domain-invariant features. Second, 3D-CoCo extends the approach of contrastive instance alignment to point cloud detection, whose performance was largely hindered by the mismatch between the fictitious distribution of BEV features, induced by pseudo-labels, and the true distribution. The mismatch is greatly reduced by 3D-CoCo with transformed point clouds, which are carefully designed by considering specific geometry priors. We construct new domain adaptation benchmarks using three large-scale 3D datasets. Experimental results show that our proposed 3D-CoCo effectively closes the domain gap and outperforms the state-of-the-art methods by large margins. Yihan Zeng, Chunwei Wang, Yunbo Wang, Hang Xu 0004, Chaoqiang Ye, Chao Ma 0004 |
NeurIPS | 3 |
| 2021 | Computation offloading over multi-UAV MEC network: A distributed deep reinforcement learning approach
Dawei Wei, Jianfeng Ma 0001, Linbo Luo 0001, Yunbo Wang, Lei He 0012, Xinghua Li 0001 |
Comput. Networks | 4 |
| 2021 | Deep Semantic Reconstruction Hashing for Similarity RetrievalabstractHashing has shown enormous potentials in preserving semantic similarity for large-scale data retrieval. Existing methods widely retain the similarity within two binary codes towards their discrete semantic affinity, i.e., 1 or -1. However, such a discrete reconstruction approach has obvious drawbacks. First, two unrelated dissimilar samples would have similar binary codes when both of them are the most dissimilar with an anchor sample. Second, the fine-grained semantic similarity cannot be shown in the generated binary codes among data with multiple semantic concepts. Furthermore, existing approaches generally adopt a point-wise error-minimizing strategy to enforce the real-valued codes close to its associated discrete codes, resulting in the well-learned paired semantic similarity being unintentionally damaged when performing quantization. To address these issues, we propose a novel deep hashing method with pairwise similarity-preserving quantization constraint, termed Deep Semantic Reconstruction Hashing (DSRH), which defines a high-level semantic affinity within each data pair to learn compact binary codes. Specifically, DSRH is expected to learn the specific binary codes whose similarity can reconstruct their high-level semantic similarity. Besides, we adopt a pairwise similarity-preserving quantization constraint instead of the traditional point-wise quantization technique, which is conducive to maintain the well-learned paired semantic similarity when performing quantization. Extensive experiments are conducted on four representative image retrieval benchmarks, and the proposed DSRH outperforms the state-of-the-art deep-learning methods with respect to different evaluation metrics. Yunbo Wang, Xianfeng Ou, Jian Liang 0001, Zhenan Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Probabilistic Video Prediction From Noisy Data With a Posterior ConfidenceabstractWe study a new research problem of probabilistic future frames prediction from a sequence of noisy inputs, which is useful because it is difficult to guarantee the quality of input frames in practical spatiotemporal prediction applications. It is also challenging because it involves two levels of uncertainty: the perceptual uncertainty from noisy observations and the dynamics uncertainty in forward modeling. In this paper, we propose to tackle this problem with an end-to-end trainable model named Bayesian Predictive Network (BP-Net). Unlike previous work in stochastic video prediction that assumes spatiotemporal coherence and therefore fails to deal with perceptual uncertainty, BP-Net models both levels of uncertainty in an integrated framework. Furthermore, unlike previous work that can only provide unsorted estimations of future frames, BP-Net leverages a differentiable sequential importance sampling (SIS) approach to make future predictions based on the inference of underlying physical states, thereby providing sorted prediction candidates in accordance with the SIS importance weights, i.e., the confidences. Our experiment results demonstrate that BP-Net remarkably outperforms existing approaches on predicting future frames from noisy data. Yunbo Wang, Jiajun Wu 0001, Mingsheng Long, Josh Tenenbaum |
CVPR | 1 |
| 2020 | Progressive Adversarial Networks for Fine-Grained Domain AdaptationabstractFine-grained visual categorization has long been considered as an important problem, however, its real application is still restricted, since precisely annotating a large fine-grained image dataset is a laborious task and requires expert-level human knowledge. A solution to this problem is applying domain adaptation approaches to fine-grained scenarios, where the key idea is to discover the commonality between existing fine-grained image datasets and massive unlabeled data in the wild. The main technical bottleneck lies in that the large inter-domain variation will deteriorate the subtle boundaries of small inter-class variation during domain alignment. This paper presents the Progressive Adversarial Networks (PAN) to align fine-grained categories across domains with a curriculum-based adversarial learning framework. In particular, throughout the learning process, domain adaptation is carried out through all multi-grained features, progressively exploiting the label hierarchy from coarse to fine. The progressive learning is applied upon both category classification and domain alignment, boosting both the discriminability and the transferability of the fine-grained features. Our method is evaluated on three benchmarks, two of which are proposed by us, and it outperforms the state-of-the-art domain adaptation methods. Xinyang Chen 0001, Yunbo Wang, Mingsheng Long, Jianmin Wang 0001 |
CVPR | 3 |
| 2020 | A Balanced and Uncertainty-Aware Approach for Partial Domain Adaptation
Jian Liang 0001, Yunbo Wang, Dapeng Hu, Ran He 0001, Jiashi Feng |
ECCV (11) | 2 |
| 2020 | A Multi-Player Minimax Game for Generative Adversarial NetworksabstractWhile multi-discriminators have been recently exploited to enhance the discriminability and diversity of Generative Adversarial Networks (GANs), these independent discriminators may not collaborate harmoniously to learn diverse and complementary decision boundaries. This paper extends the original two-player adversarial game of GANs by introducing a new multi-player objective named Discriminator Discrepancy Loss (DDL) for diversifying the multi-discriminators. Besides the competition between the generator and each discriminator, there are also competitions between the discriminators: 1) When training multi-discriminators, we simultaneously minimize the original GAN loss and maximize DDL, seeking a good trade-off between the accuracy and diversity. This yields diversified multi-discriminators that fit the generated data distribution to the real data distribution from more comprehensive perspectives. 2) When training the generator, we minimize DDL to encourage the generator to confuse all discriminators. This enhances the diversity of the generated data distribution. Further, we propose a layer-sharing network architecture for the multi-discriminators, which allows them to learn from distinct perspectives about the shared low-level features through better collaboration. It also makes our model more lightweight than existing multi-discriminators approaches. Our DDL-GAN remarkably outperforms other GANs over five standard datasets for image generation tasks. Yunbo Wang, Mingsheng Long, Jianmin Wang 0001, Philip S. Yu, Jia-Guang Sun 0001 |
ICME | 2 |
| 2020 | Multi-Task Learning of Generalizable Representations for Video Action RecognitionabstractIn classic video action recognition, labels may not contain enough information about the diverse video appearance and dynamics, thus, existing models that are trained under the standard supervised learning paradigm may extract less generalizable features. We evaluate these models under a cross-dataset experiment setting, as the above label bias problem in video analysis is even more prominent across different data sources. We find that using the optical flows as model inputs harms the generalization ability of most video recognition models.Based on these findings, we present a multi-task learning paradigm for video classification. Our key idea is to avoid label bias and improve the generalization ability by taking data as its own supervision or supervising constraints on the data. First, we take the optical flows and the RGB frames by taking them as auxiliary supervisions, and thus naming our model as Reversed Two-Stream Networks (Rev2Net). Further, we collaborate the auxiliary flow prediction task and the frame reconstruction task by introducing a new training objective to Rev2Net, named Decoding Discrepancy Penalty (DDP), which constraints the discrepancy of the multi-task features in a self-supervised manner. Rev2Net is shown to be effective on the classic action recognition task. It specifically shows a strong generalization ability in the cross-dataset experiments. Zhiyu Yao, Yunbo Wang, Mingsheng Long, Jianmin Wang 0001, Philip S. Yu, Jia-Guang Sun 0001 |
ICME | 2 |
| 2020 | Unsupervised Transfer Learning for Spatiotemporal Predictive NetworksabstractThis paper explores a new research problem of unsupervised transfer learning across multiple spatiotemporal prediction tasks. Unlike most existing transfer learning methods that focus on fixing the discrepancy between supervised tasks, we study how to transfer knowledge from a zoo of unsupervisedly learned models towards another predictive network. Our motivation is that models from different sources are expected to understand the complex spatiotemporal dynamics from different perspectives, thereby effectively supplementing the new task, even if the task has sufficient training samples. Technically, we propose a differentiable framework named transferable memory. It adaptively distills knowledge from a bank of memory states of multiple pretrained RNNs, and applies it to the target network via a novel recurrent structure called the Transferable Memory Unit (TMU). Compared with finetuning, our approach yields significant improvements on three benchmarks for spatiotemporal prediction, and benefits the target task even from less relevant pretext ones. Zhiyu Yao, Yunbo Wang, Mingsheng Long, Jianmin Wang 0001 |
ICML | 2 |
| 2020 | DualSMC: Tunneling Differentiable Filtering and Planning under Continuous POMDPsabstractA major difficulty of solving continuous POMDPs is to infer the multi-modal distribution of the unobserved true states and to make the planning algorithm dependent on the perceived uncertainty. We cast POMDP filtering and planning problems as two closely related Sequential Monte Carlo (SMC) processes, one over the real states and the other over the future optimal trajectories, and combine the merits of these two parts in a new model named the DualSMC network. In particular, we first introduce an adversarial particle filter that leverages the adversarial relationship between its internal components. Based on the filtering results, we then propose a planning algorithm that extends the previous SMC planning approach [Piche et al., 2018] to continuous POMDPs with an uncertainty-dependent policy. Crucially, not only can DualSMC handle complex observations such as image input but also it remains highly interpretable. It is shown to be effective in three continuous POMDP domains: the floor positioning domain, the 3D light-dark navigation domain, and a modified Reacher domain. Yunbo Wang, Bo Liu 0042, Jiajun Wu 0001, Yuke Zhu, Simon S. Du, Li Fei-Fei 0001, Josh Tenenbaum |
IJCAI | 1 |
| 2020 | Vision and force based autonomous coating with rollersabstractCoating rollers are widely popular in structural painting, in comparison with brushes and sprayers, due to thicker paint layer, better color consistency, and effortless customizability of holder frame and naps. In this paper, we introduce a cost-effective method to employ a general purpose robot (Sawyer, Rethink Robotics) for autonomous coating. To sense the position and the shape of the target object to be coated, the robot is combined with an RGB-Depth camera. The combined system autonomously recognizes the number of faces of the object as well as their position and surface normal. Unlike related work based on two-dimensional RGB-based image processing, all the analyses and algorithms here employ three-dimensional point cloud data (PCD). The object model learned from the PCD is then autonomously analyzed to achieve optimal motion planning to avoid collision between the robot arm and the object. To achieve human-level performance in terms of the quality of coating using the bare minimum ingredients, a combination of our own passive and builtin active impedance control is implemented. The former is realized by installing an ultrasonic sensor at the end-effector of robot working with a customized compliant mass-spring-damper roller to keep a precise distance between the end-effector and surface to be coated, maintaining a fixed force. Altogether, the control approach mimics human painting as evidenced by experimental measurements on the thickness of the coating. Coating on two different polyhedral objects is also demonstrated to test the overall method. Yayun Du, Zhaoxing Deng, Zicheng Fang, Yunbo Wang, Taiki Nagata, Karan Bansal, Mohiuddin Quadir, Mohammad K. Jawed |
IROS | 4 |
| 2020 | Global Context Enhanced Multi-modal Fusion for Referring Image Segmentation
Yan Huang 0008, Linjiang Huang, Yunbo Wang, Zhanyu Ma, Liang Wang 0001 |
PRCV (1) | 4 |
| 2020 | IMShell-Dec: Pay More Attention to External Links in PowerShell
Ruidong Han, Chao Yang 0016, Jianfeng Ma 0001, Siqi Ma 0001, Yunbo Wang, Feng Li 0041 |
SEC | 5 |
| 2019 | Memory in Memory: A Predictive Neural Network for Learning Higher-Order Non-Stationarity From Spatiotemporal DynamicsabstractNatural spatiotemporal processes can be highly non-stationary in many ways, e.g. the low-level non-stationarity such as spatial correlations or temporal dependencies of local pixel values; and the high-level variations such as the accumulation, deformation or dissipation of radar echoes in precipitation forecasting. From Cramer's Decomposition, any non-stationary process can be decomposed into deterministic, time-variant polynomials, plus a zero-mean stochastic term. By applying differencing operations appropriately, we may turn time-variant polynomials into a constant, making the deterministic component predictable. However, most previous recurrent neural networks for spatiotemporal prediction do not use the differential signals effectively, and their relatively simple state transition functions prevent them from learning too complicated variations in spacetime. We propose the Memory In Memory (MIM) networks and corresponding recurrent blocks for this purpose. The MIM blocks exploit the differential signals between adjacent recurrent states to model the non-stationary and approximately stationary properties in spatiotemporal dynamics with two cascaded, self-renewed memory modules. By stacking multiple MIM blocks, we could potentially handle higher-order non-stationarity. The MIM networks achieve the state-of-the-art results on four spatiotemporal prediction tasks across both synthetic and real-world datasets. We believe that the general idea of this work can be potentially applied to other time-series forecasting tasks. Yunbo Wang, Jianjin Zhang, Mingsheng Long, Jianmin Wang 0001, Philip S. Yu |
CVPR | 1 |
| 2019 | Towards Joint Multiply Semantics Hashing for Visual Search
Yunbo Wang, Zhenan Sun |
ICIG (3) | 1 |
| 2019 | Eidetic 3D LSTM: A Model for Video Prediction and Beyond
Yunbo Wang, Lu Jiang 0004, Ming-Hsuan Yang 0001, Li-Jia Li 0001, Mingsheng Long, Li Fei-Fei 0001 |
ICLR (Poster) | 1 |
| 2019 | Z-Order Recurrent Neural Networks for Video PredictionabstractWe present a Z-Order RNN (Znet) for predicting future video frames given historical observations. There are two main contributions respectively in deterministic and stochastic modeling perspective. First, we propose a new RNN architecture for modeling the deterministic dynamics, which updates hidden states along a z-order curve to enhance the consistency of the features of mirrored layers. Second, we introduce an adversarial training approach to two-stream Znet for modeling the stochastic variations, which forces the Znet-Predictor to imitate the behavior of the Znet-Probe. This two-stream architecture enables the adversarial training to be conducted in the feature space instead of the image space. Our model achieves the state-of-the-art prediction accuracy on two video datasets. Jianjin Zhang, Yunbo Wang, Mingsheng Long, Jianmin Wang 0001, Philip S. Yu |
ICME | 2 |
| 2019 | Local Semantic-Aware Deep Hashing With Hamming-Isometric QuantizationabstractHashing has attracted increasing attention due to its tremendous potential for efficient image retrieval and data storage. Compared with conventional hashing methods with a handcrafted feature, emerging deep hashing approaches employ deep neural networks to learn feature representations as well as hash functions, which have already been proved to be more powerful and robust in real-world applications. Currently, most of the existing deep hashing methods construct pairwise or triplet-wise constraint to obtain similar binary codes between similar data pair or relative similar binary codes within a triplet. However, some critical local structures of the data are lack of exploiting, thus the effectiveness of hash learning is not fully shown. To address this limitation, we propose a novel deep hashing method named local semantic-aware deep hashing with Hamming-isometric quantization (LSDH), where local similarity of the data is intentionally integrated into hash learning. Specifically, in the Hamming space, we exploit the potential semantic relation of the data to robustly preserve their local similarity. In addition to reducing the error introduced by binary quantizing, we further develop a Hamming-isometric objective to maximize the consistency of similarity between the pairwise binary-like feature and its binary codes pair, which is shown to be able to enhance the quality of binary codes. Extensive experimental results on several benchmark datasets, including three singlelabel datasets (i.e., CIFAR-10, CIFAR-20, and SUN397) and one multi-label dataset (NUS-WIDE), demonstrate that the proposed LSDH achieves superior performance over the latest state-of-theart hashing methods. Yunbo Wang, Jian Liang 0001, Zhenan Sun |
IEEE Trans. Image Process. | 1 |
| 2018 | PredRNN++: Towards A Resolution of the Deep-in-Time Dilemma in Spatiotemporal Predictive LearningabstractWe present PredRNN++, a recurrent network for spatiotemporal predictive learning. In pursuit of a great modeling capability for short-term video dynamics, we make our network deeper in time by leveraging a new recurrent structure named Causal LSTM with cascaded dual memories. To alleviate the gradient propagation difficulties in deep predictive models, we propose a Gradient Highway Unit, which provides alternative quick routes for the gradient flows from outputs back to long-range previous inputs. The gradient highway units work seamlessly with the causal LSTMs, enabling our model to capture the short-term and the long-term video dependencies adaptively. Our model achieves state-of-the-art prediction results on both synthetic and real video datasets, showing its power in modeling entangled motions. Yunbo Wang, Zhifeng Gao, Mingsheng Long, Jianmin Wang 0001, Philip S. Yu |
ICML | 1 |
| 2018 | PredCNN: Predictive Learning with Cascade ConvolutionsabstractPredicting future frames in videos remains an unsolved but challenging problem. Mainstream recurrent models suffer from huge memory usage and computation cost, while convolutional models are unable to effectively capture the temporal dependencies between consecutive video frames. To tackle this problem, we introduce an entirely CNN-based architecture, PredCNN, that models the dependencies between the next frame and the sequential video inputs. Inspired by the core idea of recurrent models that previous states have more transition operations than future states, we design a cascade multiplicative unit (CMU) that provides relatively more operations for previous video frames. This newly proposed unit enables PredCNN to predict future spatiotemporal data without any recurrent chain structures, which eases gradient propagation and enables a fully paralleled optimization. We show that PredCNN outperforms the state-of-the-art recurrent models for video prediction on the standard Moving MNIST dataset and two challenging crowd flow prediction datasets, and achieves a faster training speed and lower memory footprint. Ziru Xu, Yunbo Wang, Mingsheng Long, Jianmin Wang 0001 |
IJCAI | 2 |
| 2017 | Spatiotemporal Pyramid Network for Video Action RecognitionabstractTwo-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact bilinear operator into video analysis tasks. This operator enables efficient training of bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets. Yunbo Wang, Mingsheng Long, Jianmin Wang 0001, Philip S. Yu |
CVPR | 1 |
| 2017 | PredRNN: Recurrent Neural Networks for Predictive Learning using Spatiotemporal LSTMsabstractThe predictive learning of spatiotemporal sequences aims to generate future images by learning from the historical frames, where spatial appearances and temporal variations are two crucial structures. This paper models these structures by presenting a predictive recurrent neural network (PredRNN). This architecture is enlightened by the idea that spatiotemporal predictive learning should memorize both spatial appearances and temporal variations in a unified memory pool. Concretely, memory states are no longer constrained inside each LSTM unit. Instead, they are allowed to zigzag in two directions: across stacked RNN layers vertically and through all RNN states horizontally. The core of this network is a new Spatiotemporal LSTM (ST-LSTM) unit that extracts and memorizes spatial and temporal representations simultaneously. PredRNN achieves the state-of-the-art prediction performance on three video prediction datasets and is a more general framework, that can be easily extended to other predictive learning tasks by integrating with other architectures. Yunbo Wang, Mingsheng Long, Jianmin Wang 0001, Zhifeng Gao, Philip S. Yu |
NIPS | 1 |
| 2012 | Cross-Layer Analysis of the End-to-End Delay Distribution in Wireless Sensor NetworksabstractEmerging applications of wireless sensor networks (WSNs) require real-time quality-of-service (QoS) guarantees to be provided by the network. Due to the nondeterministic impacts of the wireless channel and queuing mechanisms, probabilistic analysis of QoS is essential. One important metric of QoS in WSNs is the probability distribution of the end-to-end delay. Compared to other widely used delay performance metrics such as the mean delay, delay variance, and worst-case delay, the delay distribution can be used to obtain the probability to meet a specific deadline for QoS-based communication in WSNs. To investigate the end-to-end delay distribution, in this paper, a comprehensive cross-layer analysis framework, which employs a stochastic queueing model in realistic channel environments, is developed. This framework is generic and can be parameterized for a wide variety of MAC protocols and routing protocols. Case studies with the CSMA/CA MAC protocol and an anycast protocol are conducted to illustrate how the developed framework can analytically predict the distribution of the end-to-end delay. Extensive test-bed experiments and simulations are performed to validate the accuracy of the framework for both deterministic and random deployments. Moreover, the effects of various network parameters on the distribution of end-to-end delay are investigated through the developed framework. To the best of our knowledge, this is the first work that provides a generic, probabilistic cross-layer analysis of end-to-end delay in WSNs. Yunbo Wang, Mehmet Can Vuran, Steve Goddard |
IEEE/ACM Trans. Netw. | 1 |
| 2011 | Analysis of event detection delay in wireless sensor networksabstractEmerging applications of wireless sensor networks (WSNs) require real-time event detection to be provided by the network. In a typical event monitoring WSN, multiple reports are generated by several nodes when a physical event occurs, and are then forwarded through multi-hop communication to a sink that detects the event. To improve the event detection reliability, usually timely delivery of a certain number of packets is required. Traditional timing analysis of WSNs are, however, either focused on individual packets or traffic flows from individual nodes. In this paper, a spatio-temporal fluid model is developed to capture the delay characteristics of event detection in large-scale WSNs. More specifically, the distribution of delay in event detection from multiple reports is modeled. Accordingly, metrics such as mean delay and soft delay bounds are analyzed for different network parameters. Motivated by the fact that queue build up in WSNs with low-rate traffic is negligible, a lower-complexity model is also developed. Testbed experiments and simulations are used to validate the accuracy of both approaches. The resulting framework can be utilized to analyze the effects of network and protocol parameters on event detection delay to realize real-time operation in WSNs. To the best of our knowledge, this is the first approach that provides a transient analysis of event detection delay when multiple reports via multi-hop communication are needed. Yunbo Wang, Mehmet Can Vuran, Steve Goddard |
INFOCOM | 1 |
| 2010 | Stochastic Analysis of Energy Consumption in Wireless Sensor NetworksabstractLimited energy resources in wireless sensor networks (WSNs) call for a comprehensive cross-layer analysis of energy consumption in a multi-hop network. In this paper, we provide a stochastic analysis of the energy consumption in a random network environment. Accordingly, a comprehensive cross-layer analysis framework, which employs a stochastic queueing model in realistic channel environments, is developed. This framework accurately predicts the distribution of energy consumption for nodes in WSNs during a given time period. We show that when the time duration is long, the energy consumption asymptotically approaches a Normal distribution. Using the distribution of energy consumption, the distribution of node lifetime is also investigated. With the help of this probabilistic model, a case study with an anycast protocol is conducted to show how the developed framework can analytically predict the distribution of energy consumption and lifetime. Comprehensive simulations and testbed experiments are provided to validate the developed model. The cross-layer framework is also used to identify relationships between the distribution of energy consumption and network parameters, such as network density, duty cycle, and traffic rate. To the best of our knowledge, this is the first work to investigate probabilistic distribution of energy consumption in WSNs. Yunbo Wang, Mehmet Can Vuran, Steve Goddard |
SECON | 1 |
| 2009 | Cross-Layer Analysis of the End-to-End Delay Distribution in Wireless Sensor NetworksabstractEmerging applications of wireless sensor networks (WSNs) require real-time quality of service (QoS) guarantees to be provided by the network. However, designing real-time scheduling and communication solutions for these networks is challenging since the characteristics of QoS metrics in WSNs are not well known yet. Due to the nature of wireless connectivity, it is infeasible to satisfy worst-case QoS requirements in WSNs. Instead, probabilistic QoS guarantees should be provided, which requires the definition of probabilistic QoS metrics. To provide an analytical tool for the development of real-time solutions, in this paper, the distribution of end-to-end delay in multi-hop WSNs is investigated. Accordingly, a comprehensive and accurate cross-layer analysis framework, which employs a stochastic queueing model in realistic channel environments, is developed. This framework captures the heterogeneity in WSNs in terms of channel quality, transmit power, queue length, and communication protocols. A case study with the TinyOS CSMA/CA MAC protocol is conducted to show how the developed framework can analytically predict the distribution of end-to-end delay. Testbed experiments are provided to validate the developed model. The cross-layer framework can be used to identify the relationships between network parameters and the distribution of end-to-end delay and accordingly, to design real-time solutions for WSNs. Our ongoing work suggests that this framework can be easily extended to model additional QoS metrics such as energy consumption distribution. To the best of our knowledge, this is the first work to investigate probabilistic QoS guarantees in WSNs. Yunbo Wang, Mehmet Can Vuran, Steve Goddard |
RTSS | 1 |
| 2007 | High Resolution Animated Scenes from StillsabstractCurrent techniques for generating animated scenes involve either videos (whose resolution is limited) or a single image (which requires a significant amount of user interaction). In this paper, we describe a system that allows the user to quickly and easily produce a compelling-looking animation from a small collection of high resolution stills. Our system has two unique features. First, it applies an automatic partial temporal order recovery algorithm to the stills in order to approximate the original scene dynamics. The output sequence is subsequently extracted using a second-order Markov Chain model. Second, a region with large motion variation can be automatically decomposed into semiautonomous regions such that their temporal orderings are softly constrained. This is to ensure motion smoothness throughout the original region. The final animation is obtained by frame interpolation and feathering. Our system also provides a simple-to-use interface to help the user to fine-tune the motion of the animated scene. Using our system, an animated scene can be generated in minutes. We show results for a variety of scenes. Zhouchen Lin, Lifeng Wang 0001, Yunbo Wang, Sing Bing Kang, Tian Fang |
IEEE Trans. Vis. Comput. Graph. | 3 |