Dongdong Yu

dblp:156/2054 · DBLP profile ↗
← Back
28ranked-venue papers
1as first author
23since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 TrackGo: A Flexible and Efficient Method for Controllable Video Generation
abstract
Recent years have seen substantial progress in diffusion-based controllable video generation. However, achieving precise control in complex scenarios, including fine-grained object parts, sophisticated motion trajectories, and coherent background movement, remains a challenge. In this paper, we introduce *TrackGo*, a novel approach that leverages free-form masks and arrows for conditional video generation. This method offers users with a flexible and precise mechanism for manipulating video content. We also propose the *TrackAdapter* for control implementation, an efficient and lightweight adapter designed to be seamlessly integrated into the temporal self-attention layers of a pretrained video generation model. This design leverages our observation that the attention map of these layers can accurately activate regions corresponding to motion in videos. Our experimental results demonstrate that our new approach, enhanced by the TrackAdapter, achieves state-of-the-art performance on key metrics such as FVD, FID, and ObjMC scores.
Haitao Zhou, Chuang Wang 0008, Jinlin Liu, Dongdong Yu, Qian Yu 0002, Changhu Wang
AAAI5
2025 LaVin-DiT: Large Vision Diffusion Transformer
abstract
This paper presents the Large Vision Diffusion Transformer (LaVin-DiT), a scalable and unified foundation model designed to tackle over 20 computer vision tasks in a generative framework. Unlike existing large vision models directly adapted from natural language processing architectures, which rely on less efficient autoregressive techniques and disrupt spatial relationships essential for vision data, LaVin-DiT introduces key innovations to optimize generative performance for vision tasks. First, to address the high dimensionality of visual data, we incorporate a spatial-temporal variational autoencoder that encodes data into a continuous latent space. Second, for generative modeling, we develop a joint diffusion transformer that progressively produces vision outputs. Third, for unified multitask training, in-context learning is implemented. Input-target pairs serve as task context, which guides the diffusion transformer to align outputs with specific tasks within the latent space. During inference, a task-specific context set and test data as queries allow LaVin-DiT to generalize across tasks without fine-tuning. Trained on extensive vision datasets, the model is scaled from 0.1B to 3.4B parameters, demonstrating substantial scalability and state-of-the-art performance across diverse vision tasks. This work introduces a novel pathway for large vision foundation models, underscoring the promising potential of diffusion transformers. The code and models are available at https://derrickwang005.github.io/LaVin-DiT/.
Zhaoqing Wang, Xiaobo Xia, Runnan Chen, Dongdong Yu, Changhu Wang, Mingming Gong, Tongliang Liu
CVPR4
2025 MedSoft-Diffusion: Medical Semantic-Guided Diffusion Model with Soft Mask Conditioning for Vertebral Disease Diagnosis
Shidan He, Enyuan Hu, Zixuan Tang, Dongdong Yu, Yuan Hong 0004, Zhenzhong Liu, Mengtang Li
MICCAI (15)5
2025 MIBF-Net: Multi-Modal Information Balanced Fusion Network for Clinical Diagnosis via Patient Narratives and Lesion Image
Zixuan Tang, Bai Sun, Shidan He, Yuan Hong 0004, Dongdong Yu, Zhenzhong Liu, Mengtang Li
MICCAI (1)5
2025 A Hybrid Framework Based on Bio-Signal and Built-in Force Sensor for Human-Robot Active Co-Carrying
abstract
Human-robot collaboration represents a promising avenue for applications in future factory scenarios. Existing work predominantly concentrates on passive assistance based on real-time sensing sensor like force sensors, while only a few studies have ventured into the exploration and realization of active assistance provided by robots with the assistance of predictive sensors. This paper proposes an innovative hybrid framework that combines human bio-signal information with built-in robot force sensors to implement human-robot active collaboration. First, a Hill-type muscle-skeleton model is adopted and calibrated through partial swarm optimization (PSO). With this model, surface electromyography(sEMG) is used to estimate human limb stiffness. Then, an Radial Basis Function neural network (RBFNN) compensator is developed to account for the uncertainty in human-object-robot dynamics. Subsequently, we propose an adaptive variable impedance controller, incorporating a global bias into the neural network architecture. This innovative modification serves to augment the system’s robustness, streamline the network configuration by curtailing the number of hidden neurons, and consequently, facilitate more consistent and efficient human-robot interaction behavior. Finally, we substantiate the effectiveness of the proposed methodology through a two-link robotic simulation experiment and a real-world co-carrying task employing with the Baxter robot and human partner. These rigorous evaluations unveil a significant alleviation of task-related human workload attributed to our proposed framework.Note to Practitioners—This framework aims to address the existing research gap in human-robot collaboration, particularly involving bio-signal utilization, to facilitate perceptive active assistance within a typical industrial assembly scenario. In such representative tasks, many studies primarily employ real-time sensors such as force, position. These sensors, while essential, are limited by their detection principles and require collaborative operation to ascertain stiffness and realize passive assistance. Conversely, bio-signals intrinsically contain stiffness information and exhibit prospective characteristics that can be leveraged for stiffness prediction. In this typical task, we design a human-robot co-transport system with two crucial characteristics: first, the robot is capable of detecting human stiffness tendencies and comprehending human intent, leading to self-adjusting robotic behavior that provides enhanced protection for the transported object. Secondly, the newly proposed controller can manage sudden disturbances and execute self-repairs, thus increasing the task success rate and ensuring worker safety.
Leyun Hu, Dihua Zhai, Dongdong Yu, Yuanqing Xia
IEEE Trans Autom. Sci. Eng.3
2024 Global-To-Pixel Regression for Human Mesh Recovery
Yabo Xiao, Mingshu He, Dongdong Yu
ECCV (16)3
2024 VertFound: Synergizing Semantic and Spatial Understanding for Fine-Grained Vertebrae Classification via Foundation Models
Yinhao Wu, Jinzhou Tang, Zequan Yao, Yuan Hong 0004, Dongdong Yu, Zhifan Gao
MICCAI (12)6
2023 Box-Level Active Detection
abstract
Active learning selects informative samples for annotation within budget, which has proven efficient recently on object detection. However, the widely used active detection benchmarks conduct image-level evaluation, which is unrealistic in human workload estimation and biased towards crowded images. Furthermore, existing methods still perform image-level annotation, but equally scoring all targets within the same image incurs waste of budget and redundant labels. Having revealed above problems and limitations, we introduce a box-level active detection framework that controls a box-based budget per cycle, prioritizes informative targets and avoids redundancy for fair comparison and efficient application. Under the proposed box-level setting, we devise a novel pipeline, namely Complementary Pseudo Active Strategy (ComPAS). It exploits both human annotations and the model intelligence in a complementary fashion: an efficient input-end committee queries labels for informative objects only; meantime well-learned targets are identified by the model and compensated with pseudo-labels. ComPAS consistently outperforms 10 competitors under 4 settings in a unified codebase. With supervision from labeled data only, it achieves 100% supervised performance of VOC0712 with merely 19% box annotations. On the COCO dataset, it yields up to 4.3% mAP improvement over the second-best method. ComPAS also supports training with the unlabeled pool, where it surpasses 90% COCO supervised performance with 85% label reduction. Our source code is publicly available at https://github.com/lyumengyao/blad.
Mengyao Lyu, Jundong Zhou, Hui Chen 0013, Dongdong Yu, Yandong Guo, Liuyu Xiang, Guiguang Ding
CVPR5
2023 Transformer-based Open-world Instance Segmentation with Cross-task Consistency Regularization
abstract
Open-World Instance Segmentation (OWIS) is an emerging research topic that aims to segment class-agnostic object instances from images. The mainstream approaches use a two-stage segmentation framework, which first locates the candidate object bounding boxes and then performs instance segmentation. In this work, we instead promote a single-stage transformer-based framework for OWIS. We argue that the end-to-end training process in the single-stage framework can be more convenient for directly regularizing the localization of class-agnostic object pixels. Based on the transformer-based instance segmentation framework, we propose a regularization model to predict foreground pixels and use its relation to instance segmentation to construct a cross-task consistency loss. We show that such a consistency loss could alleviate the problem of incomplete instance annotation - a common problem in the existing OWIS datasets. We also show that the proposed loss lends itself to an effective solution to semi-supervised OWIS that could be considered an extreme case that all object annotations are absent for some images. Our extensive experiments demonstrate that the proposed method achieves impressive results in both fully-supervised and semi-supervised settings. Compared to SOTA methods, the proposed method significantly improves the AP_100 score by 4.75% in UVO dataset →UVO dataset setting and 4.05% in COCO dataset →UVO dataset setting.
Xizhe Xue, Dongdong Yu, Lingqiao Liu, Yu Liu 0015, Satoshi Tsutsui, Ying Li 0017, Zehuan Yuan, Zheng Shou 0001
ACM Multimedia2
2023 Trimap-guided feature mining and fusion network for natural image matting
Dongdong Yu, Zhaozhi Xie, Yaoyi Li, Zehuan Yuan, Hongtao Lu 0001
Comput. Vis. Image Underst.2
2023 DMRNet++: Learning Discriminative Features With Decoupled Networks and Enriched Pairs for One-Step Person Search
abstract
Person search aims at localizing and recognizing query persons from raw video frames, which is a combination of two sub-tasks, i.e., pedestrian detection and person re-identification. The dominant fashion is termed as the one-step person search that jointly optimizes detection and identification in a unified network, exhibiting higher efficiency. However, there remain major challenges: (i) conflicting objectives of multiple sub-tasks under the shared feature space, (ii) inconsistent memory bank caused by the limited batch size, (iii) underutilized unlabeled identities during the identification learning. To address these issues, we develop an enhanced decoupled and memory-reinforced network (DMRNet++). First, we simplify the standard tightly coupled pipelines and establish a task-decoupled framework (TDF). Second, we build a memory-reinforced mechanism (MRM), with a slow-moving average of the network to better encode the consistency of the memorized features. Third, considering the potential of unlabeled samples, we model the recognition process as semi-supervised learning. An unlabeled-aided contrastive loss (UCL) is developed to boost the identification feature learning by exploiting the aggregation of unlabeled identities. Experimentally, the proposed DMRNet++ obtains the mAP of 94.5% and 52.1% on CUHK-SYSU and PRW datasets, which exceeds most existing methods.
Chuchu Han, Zhedong Zheng, Dongdong Yu, Zehuan Yuan, Changxin Gao, Nong Sang, Yi Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 MCIBI++: Soft Mining Contextual Information Beyond Image for Semantic Segmentation
abstract
Co-occurrent visual pattern makes context aggregation become an essential paradigm for semantic segmentation. The existing studies focus on modeling the contexts within image while neglecting the valuable semantics of the corresponding category beyond image. To this end, we propose a novel soft mining contextual information beyond image paradigm named MCIBI++ to further boost the pixel-level representations. Specifically, we first set up a dynamically updated memory module to store the dataset-level distribution information of various categories and then leverage the information to yield the dataset-level category representations during network forward. After that, we generate a class probability distribution for each pixel representation and conduct the dataset-level context aggregation with the class probability distribution as weights. Finally, the original pixel representations are augmented with the aggregated dataset-level and the conventional image-level contextual information. Moreover, in the inference phase, we additionally design a coarse-to-fine iterative inference strategy to further boost the segmentation results. MCIBI++ can be effortlessly incorporated into the existing segmentation frameworks and bring consistent performance improvements. Also, MCIBI++ can be extended into the video semantic segmentation framework with considerable improvements over the baseline. Equipped with MCIBI++, we achieved the state-of-the-art performance on seven challenging image or video semantic segmentation benchmarks.
Zhenchao Jin, Dongdong Yu, Zehuan Yuan, Lequan Yu
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 AdaptivePose: Human Parts as Adaptive Points
abstract
Multi-person pose estimation methods generally follow top-down and bottom-up paradigms, both of which can be considered as two-stage approaches thus leading to the high computation cost and low efficiency. Towards a compact and efficient pipeline for multi-person pose estimation task, in this paper, we propose to represent the human parts as points and present a novel body representation, which leverages an adaptive point set including the human center and seven human-part related points to represent the human instance in a more fine-grained manner. The novel representation is more capable of capturing the various pose deformation and adaptively factorizes the long-range center-to-joint displacement thus delivers a single-stage differentiable network to more precisely regress multi-person pose, termed as AdaptivePose. For inference, our proposed network eliminates the grouping as well as refinements and only needs a single-step disentangling process to form multi-person pose. Without any bells and whistles, we achieve the best speed-accuracy trade-offs of 67.4% AP / 29.4 fps with DLA-34 and 71.3% AP / 9.1 fps with HRNet-W48 on COCO test-dev dataset.
Yabo Xiao, Dongdong Yu, Guoli Wang 0004, Qian Zhang 0009, Mingshu He
AAAI3
2022 Learning Quality-Aware Representation for Multi-Person Pose Regression
abstract
Off-the-shelf single-stage multi-person pose regression methods generally leverage the instance score (i.e., confidence of the instance localization) to indicate the pose quality for selecting the pose candidates. We consider that there are two gaps involved in existing paradigm: 1) The instance score is not well interrelated with the pose regression quality. 2) The instance feature representation, which is used for predicting the instance score, does not explicitly encode the structural pose information to predict the reasonable score that represents pose regression quality. To address the aforementioned issues, we propose to learn the pose regression quality-aware representation. Concretely, for the first gap, instead of using the previous instance confidence label (e.g., discrete {1,0} or Gaussian representation) to denote the position and confidence for person instance, we firstly introduce the Consistent Instance Representation (CIR) that unifies the pose regression quality score of instance and the confidence of background into a pixel-wise score map to calibrates the inconsistency between instance score and pose regression quality. To fill the second gap, we further present the Query Encoding Module (QEM) including the Keypoint Query Encoding (KQE) to encode the positional and semantic information for each keypoint and the Pose Query Encoding (PQE) which explicitly encodes the predicted structural pose information to better fit the Consistent Instance Representation (CIR). By using the proposed components, we significantly alleviate the above gaps. Our method outperforms previous single-stage regression-based even bottom-up methods and achieves the state-of-the-art result of 71.7 AP on MS COCO test-dev set.
Yabo Xiao, Dongdong Yu, Guoli Wang 0004, Qian Zhang 0009
AAAI2
2022 You Should Look at All Objects
Zhenchao Jin, Dongdong Yu, Luchuan Song, Zehuan Yuan, Lequan Yu
ECCV (9)2
2022 ByteTrack: Multi-object Tracking by Associating Every Detection Box
Peize Sun, Yi Jiang 0009, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo 0002, Wenyu Liu 0001, Xinggang Wang
ECCV (22)4
2022 QueryPose: Sparse Multi-Person Pose Regression via Spatial-Aware Part-Level Query
abstract
We propose a sparse end-to-end multi-person pose regression framework, termed QueryPose, which can directly predict multi-person keypoint sequences from the input image. The existing end-to-end methods rely on dense representations to preserve the spatial detail and structure for precise keypoint localization. However, the dense paradigm introduces complex and redundant post-processes during inference. In our framework, each human instance is encoded by several learnable spatial-aware part-level queries associated with an instance-level query. First, we propose the Spatial Part Embedding Generation Module (SPEGM) that considers the local spatial attention mechanism to generate several spatial-sensitive part embeddings, which contain spatial details and structural information for enhancing the part-level queries. Second, we introduce the Selective Iteration Module (SIM) to adaptively update the sparse part-level queries via the generated spatial-sensitive part embeddings stage-by-stage. Based on the two proposed modules, the part-level queries are able to fully encode the spatial details and structural information for precise keypoint regression. With the bipartite matching, QueryPose avoids the hand-designed post-processes. Without bells and whistles, QueryPose surpasses the existing dense end-to-end methods with 73.6 AP on MS COCO mini-val set and 72.7 AP on CrowdPose test set. Code is available at https://github.com/buptxyb666/QueryPose.
Yabo Xiao, Dongdong Yu, Lei Jin 0003, Mingshu He, Zehuan Yuan
NeurIPS4
2022 Conditional Hyper-Network for Blind Super-Resolution With Multiple Degradations
abstract
Although the single-image super-resolution (SISR) methods have achieved great success on the single degradation, they still suffer performance drop with multiple degrading effects in real scenarios. Recently, some blind and non-blind models for multiple degradations have been explored. However, these methods usually degrade significantly for distribution shifts between the training and test data. Towards this end, we propose a novel conditional hyper-network framework for super-resolution with multiple degradations (named CMDSR), which helps the SR framework learn how to adapt to changes in the degradation distribution of input. We extract degradation prior at the task-level with the proposed ConditionNet, which will be used to adapt the parameters of the basic SR network (BaseNet). Specifically, the ConditionNet of our framework first learns the degradation prior from a support set, which is composed of a series of degraded image patches from the same task. Then the adaptive BaseNet rapidly shifts its parameters according to the conditional features. Moreover, in order to better extract degradation prior, we propose a task contrastive loss to shorten the inner-task distance and enlarge the cross-task distance between task-level features. Without predefining degradation maps, our blind framework can conduct one single parameter update to yield considerable improvement in SR results. Extensive experiments demonstrate the effectiveness of CMDSR over various blind, and even several non-blind methods. The flexible BaseNet structure also reveals that CMDSR can be a general framework for a large series of SISR models. Our code is available at https://github.com/guanghaoyin/CMDSR.
Guanghao Yin, Wei Wang 0009, Zehuan Yuan, Wei Ji 0008, Dongdong Yu, Shouqian Sun, Tat-Seng Chua, Changhu Wang
IEEE Trans. Image Process.5
2021 F2Net: Learning to Focus on the Foreground for Unsupervised Video Object Segmentation
abstract
Although deep learning based methods have achieved great progress in unsupervised video object segmentation, difficult scenarios (e.g., visual similarity, occlusions, and appearance changing) are still no well-handled. To alleviate these issues, we propose a novel Focus on Foreground Network (F2Net), which delves into the intra-inter frame details for the foreground objects and thus effectively improve the segmentation performance. Specifically, our proposed network consists of three main parts: Siamese Encoder Module, Center Guiding Appearance Diffusion Module, and Dynamic Information Fusion Module. Firstly, we take a siamese encoder to extract the feature representations of paired frames (reference frame and current frame). Then, a Center Guiding Appearance Diffusion Module is designed to capture the inter-frame feature (dense correspondences between reference frame and current frame), intra-frame feature (dense correspondences in current frame), and original semantic feature of current frame. Different from the Anchor Diffusion Network, we establish a Center Prediction Branch to predict the center location of the foreground object in current frame and leverage the center point information as spatial guidance prior to enhance the inter-frame and intra-frame feature extraction, and thus the feature representation considerably focus on the foreground objects. Finally, we propose a Dynamic Information Fusion Module to automatically select relatively important features through three aforementioned different level features. Extensive experiments on DAVIS, Youtube-object, and FBMS datasets show that our proposed F2Net achieves the state-of-the-art performance with significant improvement.
Daizong Liu, Dongdong Yu, Changhu Wang, Pan Zhou 0001
AAAI2
2021 Body Meshes as Points
Dongdong Yu, Jun Hao Liew, Xuecheng Nie, Jiashi Feng
CVPR2
2021 Weakly Supervised Person Search with Region Siamese Networks
abstract
Supervised learning is dominant in person search, but it requires elaborate labeling of bounding boxes and identities. Large-scale labeled training data is often difficult to collect, especially for person identities. A natural question is whether a good person search model can be trained without the need of identity supervision. In this paper, we present a weakly supervised setting where only bounding box annotations are available. Based on this new setting, we provide an effective baseline model termed Region Siamese Networks (R-SiamNets). Towards learning useful representations for recognition in the absence of identity labels, we supervise the R-SiamNet with instance-level consistency loss and cluster-level contrastive loss. For instance-level consistency learning, the R-SiamNet is constrained to extract consistent features from each person region with or without out-of-region context. For cluster-level contrastive learning, we enforce the aggregation of closest instances and the separation of dissimilar ones in feature space. Extensive experiments validate the utility of our weakly supervised method. Our model achieves the rank-1 of 87.1% and mAP of 86.0% on CUHK-SYSU benchmark, which surpasses several fully supervised methods, such as OIM [36] and MGTS [4], by a clear margin. More promising performance can be reached by incorporating extra training data. We hope this work could encourage the future research in this field.
Chuchu Han, Dongdong Yu, Zehuan Yuan, Changxin Gao, Nong Sang, Yi Yang 0001, Changhu Wang
ICCV3
2021 Mining Contextual Information Beyond Image for Semantic Segmentation
abstract
This paper studies the context aggregation problem in semantic image segmentation. The existing researches focus on improving the pixel representations by aggregating the contextual information within individual images. Though impressive, these methods neglect the significance of the representations of the pixels of the corresponding class beyond the input image. To address this, this paper proposes to mine the contextual information beyond individual images to further augment the pixel representations. We first set up a feature memory module, which is updated dynamically during training, to store the dataset-level representations of various categories. Then, we learn class probability distribution of each pixel representation under the supervision of the ground-truth segmentation. At last, the representation of each pixel is augmented by aggregating the dataset-level representations based on the corresponding class probability distribution. Furthermore, by utilizing the stored dataset-level representations, we also propose a representation consistent learning strategy to make the classification head better address intra-class compactness and inter-class dispersion. The proposed method could be effortlessly incorporated into existing segmentation frameworks (e.g., FCN, PSPNet, OCRNet and DeepLabV3) and brings consistent performance improvements. Mining contextual information beyond image allows us to report state-of-the-art performance on various benchmarks: ADE20K, LIP, Cityscapes and COCO-Stuff1.
Zhenchao Jin, Dongdong Yu, Qi Chu 0001, Changhu Wang, Jie Shao 0006
ICCV3
2021 Distributed Covariance Intersection Fusion Estimation With Delayed Measurements and Unknown Inputs
abstract
This article is concerned with the distributed covariance intersection (CI) fusion estimation for cyber-physical systems (CPSs) with delayed measurements and unknown inputs. The measurement transmission is subject to random delays described by a set of independent Bernoulli processes. Based on the provided finite-length buffers, the delayed measurements are retrieved within the corresponding buffer length. By modeling the unknown inputs with a noninformative prior distribution, a local minimum mean square error (MMSE) estimator is derived in the Bayesian framework. Then this result is extended to the multiple sensor scenario, where the sequential CI fusion approach is applied to design a recursively distributed fusion estimator. It is proved that the distributed sequential CI fusion estimator is consistent and performs better than each local estimator in state estimation. An illustrative example is provided to demonstrate the effectiveness of the proposed technique.
Dongdong Yu, Yuanqing Xia, Li Li 0050, Zirui Xing, Cui Zhu
IEEE Trans. Syst. Man Cybern. Syst.1
2020 SPCNet: Spatial Preserve and Content-Aware Network for Human Pose Estimation
abstract
Human pose estimation is a fundamental yet challenging task in computer vision. Although deep learning techniques have made great progress in this area, difficult scenarios (e.g., invisible keypoints, occlusions, complex multi-person scenarios, and abnormal poses) are still not well-handled. To alleviate these issues, we propose a novel Spatial Preserve and Content-aware Network(SPCNet), which includes two effective modules: Dilated Hourglass Module(DHM) and Selective Information Module(SIM). By using the Dilated Hourglass Module, we can preserve the spatial resolution along with large receptive field. Similar to Hourglass Network, we stack the DHMs to get the multi-stage and multi-scale information. Then, a Selective Information Module is designed to select relatively important features from different levels under a sufficient consideration of spatial content-aware mechanism and thus considerably improves the performance. Extensive experiments on MPII, LSP and FLIC human pose estimation benchmarks demonstrate the effectiveness of our network. In particular, we exceed previous methods and achieve the state-of-the-art performance on three aforementioned benchmark datasets.
Yabo Xiao, Dongdong Yu, Tianqi Lv, Yiqi Fan, Lingrui Wu
ECAI2
2019 Multi-Person Pose Estimation With Enhanced Channel-Wise and Spatial Information
abstract
Multi-person pose estimation is an important but challenging problem in computer vision. Although current approaches have achieved significant progress by fusing the multi-scale feature maps, they pay little attention to enhancing the channel-wise and spatial information of the feature maps. In this paper, we propose two novel modules to perform the enhancement of the information for the multi-person pose estimation. First, a Channel Shuffle Module (CSM) is proposed to adopt the channel shuffle operation on the feature maps with different levels, promoting cross-channel information communication among the pyramid feature maps. Second, a Spatial, Channel-wise Attention Residual Bottleneck (SCARB) is designed to boost the original residual unit with attention mechanism, adaptively highlighting the information of the feature maps both in the spatial and channel-wise context. The effectiveness of our proposed modules is evaluated on the COCO keypoint benchmark, and experimental results show that our approach achieves the state-of-the-art results.
Dongdong Yu, Zhenqi Xu, Changhu Wang
CVPR2
2019 Remote Nonlinear State Estimation With Stochastic Event-Triggered Sensor Schedule
abstract
This paper concentrates on the remote state estimation problem for nonlinear systems over a communication-limited wireless sensor network. Because of the non-Gaussian property caused by nonlinear transformation, the unscented transformation technique is exploited to obtain approximate Gaussian probability distributions of state and measurement. To reduce excessive data transmission, uncontrollable and controllable stochastic event-triggered scheduling schemes are developed to decide whether the current measurement should be transmitted. Compared with some existing deterministic event-triggered scheduling schemes, the newly developed ones possess a potential superiority in maintaining Gaussian property of innovation process. Under the proposed schemes, two nonlinear state estimators are designed based on the unscented Kalman filter. Stability and convergence conditions of these two estimators are established by analyzing behaviors of estimation error and error covariance. It is shown that an expected compromise between communication rate and estimation quality can be achieved by properly tuning event-triggered parameter matrix. Numerical examples are provided to testify the validity of the proposed results.
Li Li 0050, Dongdong Yu, Yuanqing Xia, Hongjiu Yang
IEEE Trans. Cybern.2
2017 Multi-crop Convolutional Neural Networks for lung nodule malignancy suspiciousness classification
Mu Zhou, Feng Yang 0009, Dongdong Yu, Di Dong, Caiyun Yang, Yali Zang, Jie Tian 0001
Pattern Recognit.4
2017 Adaptive total-variation for non-negative matrix factorization on manifold
Chengcai Leng, Guo-Rong Cai, Dongdong Yu, Zongyue Wang
Pattern Recognit. Lett.3