VLDB 2026 Research / reviewers in the wild / expert
Dongli Wang
dblp:16/6934
· DBLP profile ↗
27ranked-venue papers
13as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSystems, architecture and hardware · 3 · 3 first-authorComputer networks · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep learning-based group activity recognition in videos: A survey
Xiaolin Zhu 0004, Dongli Wang, Yan Zhou 0003, Yongcan Weng |
Neurocomputing | 2 |
| 2026 | Enhanced skeleton-based Group Activity Recognition through spatio-temporal graph convolution with cross-dimensional attention
Dongli Wang, Yongcan Weng, Xiaolin Zhu 0004, Yan Zhou 0003, Richard Irampaye |
Image Vis. Comput. | 1 |
| 2025 | Multi-dimensional convolution transformer for group activity recognition
Dongli Wang, Xiaolin Zhu 0004, Yan Zhou 0003 |
Multim. Tools Appl. | 1 |
| 2025 | Dynamical Attention Hypergraph Convolutional Network for Group Activity RecognitionabstractRecently, group activity recognition (GAR) has drawn growing interests in video analysis and computer vision communities. The current models of GAR tasks are often impractical in that they suppose that all interactions between actors are pairwise, which only models and leverages part of the information in real entire interactions. Motivated by this, we design a distinct dynamical attention hypergraph convolutional network framework, referred to as DAHGCN, for precise GAR, modeling the entire interactions and capturing the high-order relationships among involved actors in a real-life scenario. Specifically, to learn complementary feature representations for fine-grained GAR, a multilevel feature descriptor (MLFD) module is proposed. Furthermore, for learning higher order interaction relationships, we construct a DAHGCN to accommodate complex group interactions, which can dynamically change the topology of the hypergraph and learn these key representations by virtue of the "similarity-based shared nearest-neighbor (SSNN) clustering" and "attention mechanisms" on hypergraph. Finally, a multiscale temporal convolution (MSTC) module is utilized to explore various long-range temporal dynamic correlations across different frames. In addition, comprehensive experiments on three commonly used GAR datasets clearly demonstrate that, when compared with the state-of-the-art methods, our proposed method can achieve the most optimal performance. Xiaolin Zhu 0004, Dongli Wang, Qin Wan 0001, Yan Zhou 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Lunet: an enhanced upsampling fusion network with efficient self-attention for semantic segmentation
Yan Zhou 0003, Haibin Zhou, Yin Yang 0003, Richard Irampaye, Dongli Wang, Zhengpeng Zhang |
Vis. Comput. | 6 |
| 2024 | Multi-object tracking using context-sensitive enhancement via feature fusion
Yan Zhou 0003, Dongli Wang, Xiaolin Zhu 0004 |
Multim. Tools Appl. | 3 |
| 2023 | Hierarchical Spatial-Temporal Transformer with Motion Trajectory for Individual Action and Group Activity RecognitionabstractGroup activity recognition, which aims to simultaneously understand individual action and group activity in video clips, plays a fundamental role in video analysis. In this paper, we propose a novel reasoning network, Hierarchical Spatial-Temporal Transformer termed HSTT, for individual action and group activity recognition, which focuses on capturing the various degrees of spatial-temporal dynamic interactions adaptively and jointly among actors. Specifically, we first design a hierarchical spatial-temporal Transformer by capturing different levels of relationships to deal with unequal interaction relationships among actors. Furthermore, our proposed spatial-temporal Transformer (STT) block is capable of fully mining long-range spatial-temporal interactions with the virtue of the merge function and cross attention mechanism. Besides, we adopt the motion trajectory branch to provide complementary dynamic features for improving recognition performance. Extensive experiments on the two public GAR datasets clearly show that our approach can achieve very competitive performance by comparing them with state-of-the-art works. Xiaolin Zhu 0004, Dongli Wang, Yan Zhou 0003 |
ICASSP | 2 |
| 2023 | Adversarial learning based intermediate feature refinement for semantic segmentation
Dongli Wang, Zhitian Yuan, Wanli Ouyang, Baopu Li, Yan Zhou 0003 |
Appl. Intell. | 1 |
| 2023 | MLST-Former: Multi-Level Spatial-Temporal Transformer for Group Activity RecognitionabstractGroup activity recognition, which aims to simultaneously understand individual action and group activity in video clips, plays a fundamental role in computer vision and video analysis. In this paper, we propose a novel relational inference framework, termed MLST-Former, for individual action and group activity recognition, capturing the various degrees of spatial-temporal dynamic interactions adaptively and jointly among actors and generating reasonable group representations. Specifically, we first design a multi-level spatial-temporal Transformer to capture the miscellaneous actors’ spatial-temporal contextual information to deal with unbalanced interactions between actors. Furthermore, our proposed network is capable of fully mining long-range spatial-temporal dependencies with the virtue of the merge function and cross attention mechanism. We then propose an inter-frame gating fusion mechanism (IGFM) to selectively aggregate the temporal and structural features of the interacting actors. A new multi-task learning strategy, consisting of the classification cost of individual actions and group activities, is also developed. Moreover, we adopt the motion trajectory branch to provide complementary dynamic features for improving recognition performance. A series of ablation studies demonstrate the effectiveness and respective contributions of the different components within the proposed method. Extensive experiments on four public GAR datasets clearly show that our approach can achieve very competitive performance by comparing them with state-of-the-art methods. Xiaolin Zhu 0004, Yan Zhou 0003, Dongli Wang, Wanli Ouyang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Lightweight Self-Attention Network for Semantic SegmentationabstractThe deep neural network model based on self-attention (SA) for obtaining rich contextual information has been widely adopted in semantic segmentation. However, the computational complexity of the standard self-attentive module is high, which partly limits the use of this module. In this work, we propose the lightweight self-attention network (LSANet) for semantic segmentation. Specifically, the Lightweight Self-Attentive Module (LSAM) captures information using a hand-designed compact feature representation, and weighted fusion of position information. In the decoder structure, an improved up-sampling module is proposed. Compared with the bilinear upsampling, this method achieves better results in restoring image details. The experimental results on PASCAL VOC 2012, and Cityscapes datasets show the effectiveness of our method, which simplifies operations and improves performance. Yan Zhou 0003, Haibin Zhou, Nanjun Li, Dongli Wang |
IJCNN | 5 |
| 2021 | Bilateral attention network for semantic segmentationabstractAbstract Enhancing network feature representation capabilities and reducing the loss of image details have become the focus of semantic segmentation task. This work proposes the bilateral attention network for semantic segmentation. The authors embed two attention modules in the encoder and decoder structures . Specifically, high‐level features of the encoder structure integrate all channel maps through dense channel relationships learned by the channel correlation coefficient attention module. The positively correlated channels promote each other, and the negatively correlated channels suppress each other. In the decoder structure, low‐level features selectively emphasize the edge detail information in the feature map through the position attention module. The feature expression of semantic segmentation is improved by feature fusion of the two attention modules to obtain more accurate segmentation results . Finally, to verify the effectiveness of the model, the authors conduct experiments on the PASCAL VOC 2012 and Cityscapes scene analysis benchmark data sets and achieve a mean intersection‐over‐union of 74.92% and 66.63%, respectively. Dongli Wang, Nanjun Li, Yan Zhou 0003, Jinzhen Mu |
IET Image Process. | 1 |
| 2021 | Self-attention feature fusion network for semantic segmentation
Yan Zhou 0003, Dongli Wang, Jinzhen Mu, Haibin Zhou |
Neurocomputing | 3 |
| 2019 | Human action recognition based on multi-mode spatial-temporal feature fusion
Dongli Wang, Yan Zhou 0003 |
FUSION | 1 |
| 2019 | LANet: A Ladder Attention Network for Image Semantic SegmentationabstractA Ladder Attention Network (LANet) is proposed to recalibrate feature maps by using channel information in image semantic segmentation. Unlike previous works which use global average pooling to obtain global information of channel map, we use sub-region average pooling (SAP) to extract more abundant channel information. Meanwhile, the attention of low-stage channel is related to that of high-stage channel. In order to make the attention information of low-stage channel better transmit to high-stage features, we add the attention expansion mechanism of low-stage channel. On the other hand, noting the influence of channel number on multi-scale features fusion of Atrous Spatial Pyramid Pooling (ASPP) model in DeepLabv3 network, we allocate different channels to different convolution or pooling operations without increasing the total number of channels. Finally, experimental results on MIT Scene Parsing Benchmark dataset (SceneParse150) are included to demonstrate the effectiveness of the proposed model. Dongli Wang, Yan Zhou 0003 |
IECON | 1 |
| 2019 | Moving Human Focus Inference Model for Action RecognitionabstractSpatio-temporal information is crucial to human action recognition. Inspired by the primary visual cortex of the human brain, a new end-to-end action recognition method based on deep learning was proposed to extract pure spatiotemporal information in videos. It is named moving human focus inference model that can capture long-term temporal information with low time costs and eliminates the interference of dynamic background. It obtains the temporal and spatial information in the video from the temporal pathway with the focus block and spatial pathway, respectively. Then the temporal and spatial information are merged as spatio-temporal information by fusion region. Finally, the softmax layer is used to recognize the actions, followed by the bayes inference block that is used to correct the recognition results through the methods of maximum posterior probability inference. The model is evaluated on two challenge datasets: UCF101 and HMDB51. The experimental results show that the approach has better accuracy and generalization ability compared to the state-of-the-art methods. Dongli Wang, Hexue Xiao, Fang Ou, Yan Zhou 0003 |
IECON | 1 |
| 2019 | A Hybrid Statistics-based Channel Pruning Method for Deep CNNsabstractIn recent years, depth and width Convolution Neural Networks (CNNs) have achieved significant results in various fields of machine vision. To be honest, deploying the depth models to devices directly is Incredible. Reducing the width of the framework is an effective strategy to make the model more slender. In this paper, we propose a channel pruning method to simultaneously accelerate and compress deep CNNs while maintaining their accuracy. First, a pre-trained CNN model is evaluated layer-by-layer according to the hybrid statistics-based criterion. The negative-score channel and the corresponding filters are discarded. For the performance damaged by channel pruning, knowledge distillation is adopted to fine-tune to achieve stronger generalization ability. Experiments on the ILSVRC-12 benchmark demonstrate the effectiveness of our method when applying to the popular CNN framework. We achieve $3.79\times$ speed-up and $19.1\times$ compression baselines on VGG-16, $3.6\times$ acceleration and $5.5\times$ compression on Darknet (fully convolutional network), both with a little accuracy decrease. Even for modern networks like ResNet, we achieve $2.1\times$ speed-up and $2.2\times$ compression, suffering only 0.4% accuracy loss. Yan Zhou 0003, Dongli Wang |
SMC | 3 |
| 2018 | Action Recognition Based on Multi-feature Depth Motion MapsabstractDepth motion maps (DMM), containing abundant information on appearance and motion, are captured from the absolute difference between two consecutive depth video sequences. In this paper, each depth frame is first projected onto three orthogonal planes (front, side, top). Then the DMMf, DMMs and DMMt are generated under the three projection view respectively. In order to describe DMM in local and global, histogram of oriented gradient (HOG), local binary patterns (LBP), a local Gist feature description based on a dense grid are computed respectively. Considering the advantages of features fusion and information entropy quantitative evaluation of the Principal Component Analysis (PCA), three descriptors are weighted and fused based on information entropy improved PCA to represent the depth video. A reconstruction error adaptively weighted combination collaborative classifier based on l1-norm and l2-norm is employed for action recognition, the adaptively weights are determined by Entropy Method. Experimental results on MSR Action3D dataset show that the present approach has strong robustness, discriminability and stability. Dongli Wang, Fang Ou, Yan Zhou 0003 |
IECON | 1 |
| 2018 | Network-Assisted Optimal Datalink Selection Scheme for Heterogeneous Aeronautical NetworkabstractCurrently, the existing datalink selection algorithms in heterogeneous aeronautical network unilaterally focus on the user side or network side and lack the two sides’ cooperation, which seriously affects the utilization of available datalinks. Firstly, to reflect the user‐side requirements and network‐side load, we propose a user‐centric datalink selection scheme combing the user‐side multiattribute utility and the network‐side access rate for single‐user selection. It can be established as a multiobjective optimization problem and solved by the selected optimal datalink. Secondly, for multiuser datalink selection, in order to guarantee the access requests of multiple users, the central control mode of Software Define Networking (SDN) is used to assist the user to perform datalink selection on the network side. By constructing multiuser multiobjective optimization problem, a priority distinction selection (PDS) algorithm is proposed to maximize the number of users accessing their optimal datalinks and minimize the modification of the users’ access request. Finally, the proposed user‐side datalink selection scheme and network‐side assisted selection algorithm are simulated to verify the feasibility and effectiveness. Dongli Wang, Guoce Huang, Shufu Dong, Yequn Wang, Weiting Gao |
Wirel. Commun. Mob. Comput. | 1 |
| 2015 | Distributed multi-object localisation by consensus on compressive sampling received signal strength fingerprintsabstractRecent growing interest for location‐based services has created a demand on object localisation approaches with low cost and high accuracy. In this study, the problem of distributed multi‐object localisation using fingerprints of received signal strength (RSS) is addressed by combining average consensus and compressed sensing. First, Bayesian compressed sensing is employed at each agent to recover the sparse index vector from RSS measurements, which are corrupted by noises. It relaxes the requirement on accurate prior position knowledge of beacon nodes and is applicable in non‐line‐of‐sight conditions. Then, average consensus is adopted to compel all agents to reach an agreement on the index vector, and in turn, on the location of objects. By using only one‐hop neighbours’ information, the proposed distributed localisation method is applicable to large‐scale networks. Moreover, the final location of each object is obtainable from each individual agent, which makes the proposed method flexible to the network administration. Experimental results are included to demonstrate the effectiveness of the proposed method. Dongli Wang, Yan Zhou 0003, Yanhua Wei, Tingrui Pei |
IET Commun. | 1 |
| 2014 | Consensus 3-D bearings-only tracking in switching senor networks
Yan Zhou 0003, Dongli Wang |
Signal Process. | 2 |
| 2013 | Cross-layer design of quantized-innovation-based target tracking in wireless sensor networks
Yan Zhou 0003, Dongli Wang, Tingrui Pei, Shujuan Tian |
FUSION | 2 |
| 2013 | Binding Polymorphism in the DNA Bound State of the Pdx1 HomeodomainabstractThe subtle effects of DNA-protein recognition are illustrated in the homeodomain fold. This is one of several small DNA binding motifs that, in spite of limited DNA binding specificity, adopts crucial, specific roles when incorporated in a transcription factor. The homeodomain is composed of a 3-helix domain and a mobile N-terminal arm. Helix 3 (the recognition helix) interacts with the DNA bases through the major groove, while the N-terminal arm becomes ordered upon binding a specific sequence through the minor groove. Although many structural studies have characterized the DNA binding properties of homeodomains, the factors behind the binding specificity are still difficult to elucidate. A crystal structure of the Pdx1 homeodomain bound to DNA (PDB 2H1K) obtained previously in our lab shows two complexes with differences in the conformation of the N-terminal arm, major groove contacts, and backbone contacts, raising new questions about the DNA recognition process by homeodomains. Here, we carry out fully atomistic Molecular Dynamics simulations both in crystal and aqueous environments in order to elucidate the nature of the difference in binding contacts. The crystal simulations reproduce the X-ray experimental structures well. In the absence of crystal packing constraints, the differences between the two complexes increase during the solution simulations. Thus, the conformational differences are not an artifact of crystal packing. In solution, the homeodomain with a disordered N-terminal arm repositions to a partially specific orientation. Both the crystal and aqueous simulations support the existence of different stable binding conformers identified in the original crystallographic data with different degrees of specificity. We propose that protein-protein and protein-DNA interactions favor a subset of the possible conformations. This flexibility in DNA binding may facilitate multiple functions for the same transcription factor. Volodymyr Babin, Dongli Wang, Robert B. Rose, Celeste Sagui |
PLoS Comput. Biol. | 2 |
| 2012 | Target tracking in wireless sensor networks using adaptive measurement quantization
Yan Zhou 0003, Dongli Wang |
Sci. China Inf. Sci. | 3 |
| 2011 | Binary tree of posterior probability support vector machinesabstractPosterior probability support vector machines (PPSVMs) prove robust against noises and outliers and need fewer storage support vectors (SVs). Gonen et al. (2008) extended PPSVMs to a multiclass case by both single-machine and multimachine approaches. However, these extensions suffer from low classification efficiency, high computational burden, and more importantly, unclassifiable regions. To achieve higher classification efficiency and accuracy with fewer SVs, a binary tree of PPSVMs for the multiclass classification problem is proposed in this letter. Moreover, a Fisher ratio separability measure is adopted to determine the tree structure. Several experiments on handwritten recognition datasets are included to illustrate the proposed approach. Specifically, the Fisher ratio separability accelerated binary tree of PPSVMs obtains overall test accuracy, if not higher than, at least comparable to those of other multiclass algorithms, while using significantly fewer SVs and much less test time. Dongli Wang, Jianguo Zheng, Yan Zhou 0003 |
J. Zhejiang Univ. Sci. C | 1 |
| 2010 | A scalable support vector machine for distributed classification in ad hoc sensor networks
Dongli Wang, Jianguo Zheng, Yan Zhou 0003 |
Neurocomputing | 1 |
| 2010 | Posterior CramÉr-Rao Lower Bounds for Target Tracking in Sensor Networks With Quantized Range-Only MeasurementsabstractWe consider the problem of target tracking in a wireless sensor network (WSN) that consists of randomly distributed range-only sensors. Quantized measurements are usually adopted in such a network to attack the problem of limited power supply and communication bandwidth. Assuming that local sensor noises are mutually independent, we derive the posterior Cramer-Rao lower bound (CRLB) on the mean squared error (MSE) of target tracking in WSNs with quantized range-only measurements. Recursion of posterior CRLB on tracking based on both constant velocity (CV) and constant acceleration (CA) model for target dynamics and a general range-only measuring model for local sensors are obtained. Due to the analytical difficulties, particle filter is applied to approximate the theoretical bounds. To illustrate the posterior CRLB, an example on tracking a target with noisy circular trajectories is given. Yan Zhou 0003, Dongli Wang |
IEEE Signal Process. Lett. | 3 |
| 2005 | Applying Neural Network to Reinforcement Learning in Continuous Spaces
Dongli Wang, Yang Gao 0001 |
ISNN (1) | 1 |