Bing Zhou 0003

dblp:90/3394-3 · DBLP profile ↗
← Back
52ranked-venue papers
0as first author
31since 2021 · last 2026
0000-0003-3446-3903ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ECG-Text multi-modal learning for zero-shot detection via time-frequency alignment and medical prompt learning
Ning Wang 0037, Haiyan Wang 0021, Panpan Feng, Shihua Li 0007, Zongmin Wang, Bing Zhou 0003
Expert Syst. Appl.7
2026 Dynamic Spatiotemporal Information Interaction for Multidrone Single Object Tracking
abstract
Multi-drone single object tracking is a key technology in Internet of Things-based aerial sensing systems. However, existing methods often neglect both temporal continuity of the object across frames and spatial complementarity from multi-drone viewpoints, limiting their robustness in large-scale and dynamic IoT environments. To address these issues, this paper proposes a dynamic spatiotemporal information interaction framework for multi-drone single object tracking (DSTII-MDOT), where spatiotemporal information refers to the object’s feature evolution over time (temporal) and its complementary multi-view representations (spatial). First, a dynamic temporal feature aggregation (DTFA) module is proposed, which captures the object’s motion and appearance variations by exploiting frame-to-frame differences, thereby ensuring feature continuity and stable predictions across frames in long-term tracking. Next, a spatial alignment method for cross-modal (text-image) semantic alignment (CMSA) is proposed. This method ensures the semantic alignment between visual features and textual descriptions by leveraging the CLIP model and employs a multi-head cross-attention mechanism to capture the top-k regions of interest within the search area. This enhances cross-drone collaboration and reduces redundant communications in IoT networks. Finally, a wavelet frequency domain feature refinement (WFDFR) module is proposed to enhance the texture features of the template image, effectively solving the problem of blurred texture features or missing details in complex scenes. Experimental results on the MDOT dataset demonstrate that the proposed DSTII-MDSOT tracker surpasses existing advanced methods in both success rate and precision, validating the effectiveness and superiority of the proposed method. The code and models are available at https://github.com/JerryBryant24/DSTII-MDOT.
Xiangqian Liu, Guangwei Zhang 0003, Bin Wang 0088, Lihong Zhong, Bing Zhou 0003
IEEE Internet Things J.5
2026 Global attention and multi-scale fusion dynamic extraction network for 12-lead ECG classification
Xiangqian Liu, Shihua Li 0007, Bing Zhou 0003
Multim. Syst.4
2026 Load-Balancing During Resource Allocation in Multi-Cell MIMO-NOMA
abstract
Load balancing during resource allocation (RA) in multicell multi-input multi-output non-orthogonal multiple access (MC-MIMO-NOMA) poses significant challenges for the deployment of massive MIMO-NOMA in next-generation networks. This paper introduces load-balancing algorithms to address RA challenges in massive MIMO-NOMA systems. First, we formulate the general RA problem as a load minimization problem, involving the joint optimization of resource allocation and user clustering. To minimize computation complexity, this problem is then reduced to a suboptimal RA problem within each cell. However, due to the non-convex nature of the problem, obtaining an optimal solution is not straightforward. To overcome this, we propose the power to targeted load optimization (PTLO) approach to solve the RA problem in each cell based on underload and overload scenarios. By leveraging the PTLO approach, we develop a user clustering strategy using branch-and-bound and weighted matching theory to optimize RA in each cell. Through simulations, we evaluate the performance of the proposed load-balancing algorithms (LB-BB-MIMO-NOMA and LB-M2O-MIMO-NOMA) against benchmark algorithms such as Correlated-MIMO-NOMA, Coalition-MIMO-NOMA, Gain-Diff-MIMO-NOMA, Greedy-MIMO-NOMA, and MIMO-OMA. The results demonstrate that the proposed algorithms consistently outperform the benchmarks. Furthermore, we provide a detailed analysis highlighting the efficiency of the proposed algorithms in the MC-MIMO-NOMA networks.
Aretor Samuel, Ruijie Zhu 0001, Bing Zhou 0003
IEEE Trans. Commun.4
2026 TSMMR: a mixed Mamba network with cross-drone template fusion and text-guided redetection for robust aerial tracking
Xiangqian Liu, Guangwei Zhang 0003, Yifan Zhu 0001, Bing Zhou 0003
Vis. Comput.4
2025 Rethinking Contrastive Learning for Electrocardiogram Anomaly Detection: A Time-Frequency Augmentations Perspective
Huihui Chang, Haoyi Fan, Mingzhe Han, Bing Zhou 0003, Zongmin Wang
PAKDD (1)5
2025 Workload-based adaptive decision-making for edge server layout with deep reinforcement learning
Shihua Li 0007, Yanjie Zhou, Bing Zhou 0003, Zongmin Wang
Eng. Appl. Artif. Intell.3
2025 Dynamic weight reinforcement learning method considering multiple factors in mobile edge computing system
Shihua Li 0007, Yanjie Zhou, Xiangqian Liu, Ning Wang 0037, Bing Zhou 0003, Zongmin Wang
Neurocomputing6
2025 Semi-supervised multi-label cardiovascular diseases detection via contrastive learning and label inference
Ning Wang 0037, Haiyan Wang 0021, Panpan Feng, Shihua Li 0007, Zongmin Wang, Bing Zhou 0003
Knowl. Based Syst.7
2025 An efficient feature aggregation network for small object detection in UAV aerial images
Xiangqian Liu, Guangwei Zhang 0003, Bing Zhou 0003
J. Supercomput.3
2025 TITFormer: Combining Textual Modality and Simulating Infrared Modality Based on Transformer for Image Enhancement
abstract
The image enhancement task requires a complex balance between extracting high-level contextual information and optimizing spatial details in the image to improve the visual quality. Most of existing methods have limited capability in capturing contextual features and optimizing spatial details when they only rely on a single modality. To address the above issues, this paper introduces a novel multi-modal image enhancement network based on Transformer, named as TITFormer, which combines textual and simulating infrared modalities firstly for this important task. TITFormer comprises a text channel attention fusion (TCF) network block and an infrared-guided spatial detail optimization (SDO) network block. The TCF extracts contextual features from the high-dimensional features compressed after spatial channel transformation of the textual feature and image feature. The SDO module uses simulating infrared images characterized by pixel intensity to guide the optimization of spatial details with contextual features adaptively. Experimental results demonstrate that TITFormer achieves state-of-the-art performance on two publicly available benchmark datasets.
Kaijiang Li, Haining Li, Miduo Cui, Junxin Li, Pei Lv, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Multim.6
2024 A Multi-Resolution Mutual Learning Network for Multi-Label ECG Classification
abstract
Electrocardiograms (ECG), essential for diagnosing cardiovascular diseases. In recent years, the application of deep learning techniques has significantly improved the performance of ECG signal classification. Multi-resolution feature analysis, which captures and processes information at different time scales, can extract subtle changes and overall trends in ECG signals, showing unique advantages. However, common multi-resolution analysis methods based on simple feature addition or concatenation may lead to the neglect of low-resolution features, affecting model performance. To address this issue, this paper proposes the Multi-Resolution Mutual Learning Network (MRMNet). MRM-Net includes a dual-resolution attention architecture and a feature complementary mechanism. The dual-resolution attention architecture processes high-resolution and low-resolution features in parallel. Through the attention mechanism, the high-resolution and low-resolution branches can focus on subtle waveform changes and overall rhythm patterns, enhancing the ability to capture critical features in ECG signals. Meanwhile, the feature complementary mechanism introduces mutual feature learning after each layer of the feature extractor. This allows features at different resolutions to reinforce each other, thereby reducing information loss and improving model performance and robustness. Experiments on the PTB-XL and CPSC2018 datasets demonstrate that MRM-Net significantly outperforms existing methods in multi-label ECG classification performance. The code for our framework will be publicly available at https://github.com/wxhdf/MRM.
Ning Wang 0037, Panpan Feng, Haiyan Wang 0021, Zongmin Wang, Bing Zhou 0003
BIBM6
2024 Adversarial Spatiotemporal Contrastive Learning for Electrocardiogram Signals
abstract
Extracting invariant representations in unlabeled electrocardiogram (ECG) signals is a challenge for deep neural networks (DNNs). Contrastive learning is a promising method for unsupervised learning. However, it should improve its robustness to noise and learn the spatiotemporal and semantic representations of categories, just like cardiologists. This article proposes a patient-level adversarial spatiotemporal contrastive learning (ASTCL) framework, which includes ECG augmentations, an adversarial module, and a spatiotemporal contrastive module. Based on the ECG noise attributes, two distinct but effective ECG augmentations, ECG noise enhancement, and ECG noise denoising, are introduced. These methods are beneficial for ASTCL to enhance the robustness of the DNN to noise. This article proposes a self-supervised task to increase the antiperturbation ability. This task is represented as a game between the discriminator and encoder in the adversarial module, which pulls the extracted representations into the shared distribution between the positive pairs to discard the perturbation representations and learn the invariant representations. The spatiotemporal contrastive module combines spatiotemporal prediction and patient discrimination to learn the spatiotemporal and semantic representations of categories. To learn category representations effectively, this article only uses patient-level positive pairs and alternately uses the predictor and the stop-gradient to avoid model collapse. To verify the effectiveness of the proposed method, various groups of experiments are conducted on four ECG benchmark datasets and one clinical dataset compared with the state-of-the-art methods. Experimental results showed that the proposed method outperforms the state-of-the-art methods.
Ning Wang 0037, Panpan Feng, Zhaoyang Ge, Yanjie Zhou, Bing Zhou 0003, Zongmin Wang
IEEE Trans. Neural Networks Learn. Syst.5
2023 ECG-MAKE: An ECG signal delineation approach based on medical attribute knowledge extraction
Zhaoyang Ge, Huiqing Cheng, Zhuang Tong, Ning Wang 0037, Adi Alhudhaif, Fayadh Alenezi, Haiyan Wang 0021, Bing Zhou 0003, Zongmin Wang
Inf. Sci.8
2023 Semantic-aware alignment and label propagation for cross-domain arrhythmia classification
Panpan Feng, Ning Wang 0037, Yanjie Zhou, Bing Zhou 0003, Zongmin Wang
Knowl. Based Syst.5
2023 Emotional Contagion-Aware Deep Reinforcement Learning for Antagonistic Crowd Simulation
abstract
The antagonistic behavior in the crowd usually exacerbates the seriousness of the situation in sudden riots, where the antagonistic emotional contagion and behavioral decision making play very important roles. However, the complex mechanism of antagonistic emotion influencing decision making, especially in the environment of sudden confrontation, has not yet been explored very clearly. In this paper, we propose an Emotional contagion-aware Deep reinforcement learning model for Antagonistic Crowd Simulation (ACSED). First, we build a group emotional contagion module based on the improved Susceptible Infected Susceptible (SIS) infection disease model, and estimate the emotional state of the group at each time step during the simulation. Then, the tendency of crowd antagonistic action is estimated based on Deep Q Network (DQN), where the agent learns the action autonomously, and leverages the mean field theory to quickly calculate the influence of other surrounding individuals on the central one. Finally, the rationality of the predicted actions by DQN is further analyzed in combination with group emotion, and the final action of the agent is determined. The proposed method in this paper is verified through several experiments with different settings. We can conclude antagonistic emotions play a critical role in the decision making of the crowd through influencing the individual behavior in the riot scenario, where individual behaviors are primarily driven by emotions and goals, rather than common rules. The experiment results also prove that the antagonistic emotion has a vital impact on the group combat, and positive emotional states are more conducive to combat. Moreover, by comparing the simulation results with real scenes, the feasibility of our method is further confirmed, which can provide good reference to formulate battle plans and improve the win rate of righteous groups in a variety of situations.
Pei Lv, Qingqing Yu, Boya Xu, Chaochao Li, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Affect. Comput.5
2023 User-Guided Personalized Image Aesthetic Assessment Based on Deep Reinforcement Learning
abstract
Personalized image aesthetic assessment (PIAA) has recently become a hot topic due to its wide applications, such as photography, film, television, e-commerce, fashion design, and so on. This task is more seriously affected by subjective factors and samples provided by users. In order to acquire precise personalized aesthetic distribution by small amount of samples, we propose a novel user-guided personalized image aesthetic assessment framework. This framework leverages user interactions to retouch and rank images for aesthetic assessment based on deep reinforcement learning (DRL), and generates personalized aesthetic distribution that is more in line with the aesthetic preferences of different users. It mainly consists of two stages. In the first stage, personalized aesthetic ranking is generated by interactive image enhancement and manual ranking, meanwhile, two policy networks will be trained. These two networks will be trained iteratively and alternatively to facilitate the final personalized aesthetic assessment. In the second stage, these modified images are labeled with aesthetic attributes by one style-specific classifier, and then the personalized aesthetic distribution is generated based on the multiple aesthetic attributes of these images, which conforms to the aesthetic preference of users better. Compared with other existing methods, our approach has achieved new state-of-the-art in the task of personalized image aesthetic assessment on the public AVA and FLICKR-AES datasets.
Pei Lv, Jianqi Fan, Xixi Nie, Weiming Dong, Xiaoheng Jiang, Bing Zhou 0003, Mingliang Xu 0001, Changsheng Xu
IEEE Trans. Multim.6
2022 Trajectory distributions: A new description of movement for trajectory prediction
abstract
Trajectory prediction is a fundamental and challenging task for numerous applications, such as autonomous driving and intelligent robots. Current works typically treat pedestrian trajectories as a series of 2D point coordinates. However, in real scenarios, the trajectory often exhibits randomness, and has its own probability distribution. Inspired by this observation and other movement characteristics of pedestrians, we propose a simple and intuitive movement description called a trajectory distribution, which maps the coordinates of the pedestrian trajectory to a 2D Gaussian distribution in space. Based on this novel description, we develop a new trajectory prediction method, which we call the social probability method . The method combines trajectory distributions and powerful convolutional recurrent neural networks. Both the input and output of our method are trajectory distributions, which provide the recurrent neural network with sufficient spatial and random information about moving pedestrians. Furthermore, the social probability method extracts spatio-temporal features directly from the new movement description to generate robust and accurate predictions. Experiments on public benchmark datasets show the effectiveness of the proposed method.
Pei Lv, Tianxin Gu, Xiaoheng Jiang, Bing Zhou 0003, Mingliang Xu 0001
Comput. Vis. Media6
2022 Transferring priors from virtual data for crowd counting in real world
Xiaoheng Jiang, Hao Liu 0057, Li Zhang 0072, Geyang Li, Mingliang Xu 0001, Pei Lv, Bing Zhou 0003
Frontiers Comput. Sci.7
2022 Unsupervised semantic-aware adaptive feature fusion network for arrhythmia detection
Panpan Feng, Zhaoyang Ge, Haiyan Wang 0021, Yanjie Zhou, Bing Zhou 0003, Zongmin Wang
Inf. Sci.6
2022 ACSEE: Antagonistic Crowd Simulation Model With Emotional Contagion and Evolutionary Game Theory
abstract
Antagonistic crowd behaviors are often observed in cases of serious conflict. Antagonistic emotions, which is the typical psychological state of agents in different roles (i.e., cops, activists, and civilians) in crowd violence scenes, and the way they spread through contagion in a crowd are important causes of crowd antagonistic behaviors. Moreover, games, which refers to the interaction between opposing groups adopting different strategies to obtain higher benefits and less casualties, determine the level of crowd violence. We present an antagonistic crowd simulation model (ACSEE), which is integrated with antagonistic emotional contagion and evolutionary game theories. Our approach models the antagonistic emotions between agents in different roles using two components: mental emotion and external emotion. We combine enhanced susceptible-infectious-susceptible (SIS) and game approaches to evaluate the role of antagonistic emotional contagion in crowd violence. Our evolutionary game theoretic approach incorporates antagonistic emotional contagion through deterrent force, which is modelled by a mixture of emotional forces and physical forces defeating the opponents. Antagonistic emotional contagion and evolutionary game theories influence each other to determine antagonistic crowd behaviors. We evaluate our approach on real-world scenarios consisting of different kinds of agents. We also compare the simulated crowd behaviors with real-world crowd videos and use our approach to predict the trends of crowd movements in violence incidents. We investigate the impact of various factors (number of agents, emotion, strategy, etc.) on the outcome of crowd violence. We present results from user studies suggesting that our model can simulate antagonistic crowd behaviors similar to those seen in real-world scenarios.
Chaochao Li, Pei Lv, Dinesh Manocha, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Affect. Comput.6
2022 Agent-Based Campus Novel Coronavirus Infection and Control Simulation
abstract
Corona Virus Disease 2019 (COVID-19), due to its extremely high infectivity, has been spreading rapidly around the world and bringing huge influence to socioeconomic development and people’s daily life. Taking for example the virus transmission that may occur after college students return to school, we analyze the quantitative influence of the key factors on the virus spread, including crowd density and self-protection. One Campus Virus Infection and Control Simulation (CVICS) model of the novel coronavirus is proposed in this article, fully considering the characteristics of repeated contact and strong mobility of crowd in the closed environment. Specifically, we build an agent-based infection model, introduce the mean field theory to calculate the probability of virus transmission, and microsimulate the daily prevalence of infection among individuals. The experimental results show that the proposed model in this article efficiently simulates how the virus spreads in the dense crowd in frequent contact under a closed environment. Furthermore, preventive and control measures, such as self-protection, crowd decentralization, and isolation during the epidemic, can effectively delay the arrival of infection peak, reduce the prevalence, and, finally, lower the risk of COVID-19 transmission after the students return to school.
Pei Lv, Boya Xu, Ran Feng, Chaochao Li, Junxiao Xue, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Comput. Soc. Syst.7
2022 Context-Aware Block Net for Small Object Detection
abstract
State-of-the-art object detectors usually progressively downsample the input image until it is represented by small feature maps, which loses the spatial information and compromises the representation of small objects. In this article, we propose a context-aware block net (CAB Net) to improve small object detection by building high-resolution and strong semantic feature maps. To internally enhance the representation capacity of feature maps with high spatial resolution, we delicately design the context-aware block (CAB). CAB exploits pyramidal dilated convolutions to incorporate multilevel contextual information without losing the original resolution of feature maps. Then, we assemble CAB to the end of the truncated backbone network (e.g., VGG16) with a relatively small downsampling factor (e.g., 8) and cast off all following layers. CAB Net can capture both basic visual patterns as well as semantical information of small objects, thus improving the performance of small object detection. Experiments conducted on the benchmark Tsinghua-Tencent 100K and the Airport dataset show that CAB Net outperforms other top-performing detectors by a large margin while keeping real-time speed, which demonstrates the effectiveness of CAB Net for small object detection.
Lisha Cui, Pei Lv, Xiaoheng Jiang, Zhimin Gao, Bing Zhou 0003, Ling Shao 0001, Mingliang Xu 0001
IEEE Trans. Cybern.5
2021 Interactive ECG annotation: An artificial intelligence method for smart ECG manipulation
Haiyan Wang 0021, Yanjie Zhou, Bing Zhou 0003, Xiangdong Niu, Zongmin Wang
Inf. Sci.3
2021 An effective feature extraction method based on GDS for atrial fibrillation detection
Haiyan Wang 0021, Honghua Dai 0001, Yanjie Zhou, Bing Zhou 0003, Peng Lu 0009, Hongpo Zhang, Zongmin Wang
J. Biomed. Informatics4
2021 Multi-label correlation guided feature fusion network for abnormal ECG diagnosis
Zhaoyang Ge, Xiaoheng Jiang, Zhuang Tong, Panpan Feng, Bing Zhou 0003, Mingliang Xu 0001, Zongmin Wang, Yanwei Pang
Knowl. Based Syst.5
2021 Bio-Inspired Deep Attribute Learning Towards Facial Aesthetic Prediction
abstract
Computational prediction of facial aesthetics has attracted ever-increasing research focus, which has wide range of prospects in multimedia applications. The key challenge lies in extracting discriminative and perception-aware features to characterize the facial beautifulness. To this end, the existing schemes simply adopt a direct feature mapping, which relies on handcraft-designed low-level features that cannot reflect human-level aesthetic perception. In this paper, we present a systematic framework towards designing biology-inspired, discriminative representation for facial aesthetic prediction. First, we design a group of biological experiments that adopt eye tracker to identify spatial regions of interest during the facial aesthetic judgments of subjects, which forms a Bio-inspired Facial Aesthetic Ontology (Bio-FAO) and is made public available. Second, we adopt the cutting-edge convolutional neural network to train a set of Bio-inspired Attribute features, termed Bio-AttriBank, which forms a mid-level interpretable representation corresponding to the aforementioned Bio-FAO. For a given image, the facial aesthetic prediction is then formulated as a classification problem over the Bio-AttriBank descriptor responses, which well bridges the affective gap, and provides explainable evidences on why/how a face is beautiful or not. We have carried out extensive experiments on both JAFFE and FaceWarehouse datasets, with comparisons to a set of state-of-the-art and alternative approaches. Superior performance gains in the experiments have demonstrated the merits of the proposed scheme.
Mingliang Xu 0001, Fuhai Chen, Pei Lv, Bing Zhou 0003, Rongrong Ji
IEEE Trans. Affect. Comput.6
2021 Emotion-Based Crowd Simulation Model Based on Physical Strength Consumption for Emergency Scenarios
abstract
Increasing attention is being given to the modeling and simulation of traffic flow and crowd movement, two phenomena that both deal with interactions between pedestrians and cars in many situations. In particular, crowd simulation is important for understanding mobility and transportation patterns. In this paper, we propose an emotion-based crowd simulation model integrating physical strength consumption. Inspired by the theory of “the devoted actor,” the movements of each individual in our model are determined by modeling the influence of physical strength consumption and the emotion of panic. In particular, human physical strength consumption is computed using a physics-based numerical method. Inspired by the James-Lange theory, panic levels are estimated by means of an enhanced emotional contagion model that leverages the inherent relationship between physical strength consumption and panic. To the best of our knowledge, our model is the first method integrating physical strength consumption into an emotion-based crowd simulation model by exploiting the relationship between physical strength consumption and emotion. We highlight the performance on different scenarios and compare the resulting behaviors with real-world video sequences. Our approach can reliably predict changes in physical strength consumption and panic levels of individuals in an emergency situation.
Mingliang Xu 0001, Chaochao Li, Pei Lv, Wei Chen 0001, Zhigang Deng 0001, Bing Zhou 0003, Dinesh Manocha
IEEE Trans. Intell. Transp. Syst.6
2021 Density-Aware Multi-Task Learning for Crowd Counting
abstract
In this paper, we present a method called density-aware convolutional neural network (DensityCNN) to perform the crowd counting task in various crowded scenes. The key idea of the DensityCNN is to utilize high-level semantic information to provide guidance and constraint when generating density maps. To this end, we implement the DensityCNN by adopting a multi-task CNN structure to jointly learn density-level classification and density map estimation. The density-level classification task learns multi-channel semantic features that are aware of the density distributions of the input image. This task is accomplished via our specially designed group-based convolutional structure in a supervised learning manner. In the density map estimation task, these semantic features are deployed together with high-dimension convolutional features to generate density maps with lower count errors. Extensive experiments on four challenging crowd datasets (ShanghaiTech, UCF_CC_50, UCF-QNCF, and WorldExpo'10) and one vehicle dataset TRANCOS demonstrate the effectiveness of the proposed method.
Xiaoheng Jiang, Li Zhang 0072, Tianzhu Zhang 0001, Pei Lv, Bing Zhou 0003, Yanwei Pang, Mingliang Xu 0001, Changsheng Xu
IEEE Trans. Multim.5
2021 ART-UP: A Novel Method for Generating Scanning-Robust Aesthetic QR Codes
abstract
Quick response (QR) codes are usually scanned in different environments, so they must be robust to variations in illumination, scale, coverage, and camera angles. Aesthetic QR codes improve the visual quality, but subtle changes in their appearance may cause scanning failure. In this article, a new method to generate scanning-robust aesthetic QR codes is proposed, which is based on a module-based scanning probability estimation model that can effectively balance the tradeoff between visual quality and scanning robustness. Our method locally adjusts the luminance of each module by estimating the probability of successful sampling. The approach adopts the hierarchical, coarse-to-fine strategy to enhance the visual quality of aesthetic QR codes, which sequentially generate the following three codes: a binary aesthetic QR code, a grayscale aesthetic QR code, and the final color aesthetic QR code. Our approach also can be used to create QR codes with different visual styles by adjusting some initialization parameters. User surveys and decoding experiments were adopted for evaluating our method compared with state-of-the-art algorithms, which indicates that the proposed approach has excellent performance in terms of both visual quality and scanning robustness.
Mingliang Xu 0001, Qingfeng Li 0004, Jianwei Niu 0002, Hao Su 0001, Xiting Liu, Weiwei Xu 0003, Pei Lv, Bing Zhou 0003, Yi Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.8
2021 Crowd Behavior Simulation With Emotional Contagion in Unexpected Multihazard Situations
abstract
Numerous research efforts have been conducted to simulate the crowd movements, while relatively few of them are specifically focused on multihazard situations. In this paper, we propose a novel crowd simulation method by modeling the generation and contagion of panic emotion under multihazard circumstances. In order to depict the effect from hazards and other agents to crowd movement, we first classify hazards into different types (transient and persistent, concurrent and nonconcurrent, and static and dynamic) based on their inherent characteristics. Second, we introduce the concept of perilous field for each hazard and further transform the critical level of the field to its invoked-panic emotion. After that, we propose an emotional contagion model to simulate the evolving process of panic emotion caused by multiple hazards. Finally, we introduce an emotional reciprocal velocity obstacles (RVOs) model to simulate the crowd behaviors by augmenting the traditional RVO model with emotional contagion, which for the first time combines the emotional impact and local avoidance together. Our experimental results demonstrate that the overall approach is robust, can better generate realistic crowds and the panic emotion dynamics in a crowd. Furthermore, it is recommended that our method can be applied to various complex multihazard environments.
Mingliang Xu 0001, Xiaozheng Xie, Pei Lv, Jianwei Niu 0002, Chaochao Li, Ruijie Zhu 0001, Zhigang Deng 0001, Bing Zhou 0003
IEEE Trans. Syst. Man Cybern. Syst.9
2020 Attention Scaling for Crowd Counting
abstract
Convolutional Neural Network (CNN) based methods generally take crowd counting as a regression task by outputting crowd densities. They learn the mapping between image contents and crowd density distributions. Though having achieved promising results, these data-driven counting networks are prone to overestimate or underestimate people counts of regions with different density patterns, which degrades the whole count accuracy. To overcome this problem, we propose an approach to alleviate the counting performance differences in different regions. Specifically, our approach consists of two networks named Density Attention Network (DANet) and Attention Scaling Network (ASNet). DANet provides ASNet with attention masks related to regions of different density levels. ASNet first generates density maps and scaling factors and then multiplies them by attention masks to output separate attention-based density maps. These density maps are summed to give the final density map. The attention scaling factors help attenuate the estimation errors in different regions. Furthermore, we present a novel Adaptive Pyramid Loss (APLoss) to hierarchically calculate the estimation losses of sub-regions, which alleviates the training bias. Extensive experiments on four challenging datasets (ShanghaiTech Part A, UCF_CC_50, UCF-QNRF, and WorldExpo'10) demonstrate the superiority of the proposed approach.
Xiaoheng Jiang, Li Zhang 0072, Mingliang Xu 0001, Tianzhu Zhang 0001, Pei Lv, Bing Zhou 0003, Xin Yang 0011, Yanwei Pang
CVPR6
2020 Semi-Dynamic Hypergraph Neural Network for 3D Pose Estimation
abstract
This paper proposes a novel Semi-Dynamic Hypergraph Neural Network (SD-HNN) to estimate 3D human pose from a single image. SD-HNN adopts hypergraph to represent the human body to effectively exploit the kinematic constrains among adjacent and non-adjacent joints. Specifically, a pose hypergraph in SD-HNN has two components. One is a static hypergraph constructed according to the conventional tree body structure. The other is the semi-dynamic hypergraph representing the dynamic kinematic constrains among different joints. These two hypergraphs are combined together to be trained in an end-to-end fashion. Unlike traditional Graph Convolutional Networks (GCNs) that are based on a fixed tree structure, the SD-HNN can deal with ambiguity in human pose estimation. Experimental results demonstrate that the proposed method achieves state-of-the-art performance both on the Human3.6M and MPI-INF-3DHP datasets.
Pei Lv, Junjin Cheng, Wanqing Li 0001, Bing Zhou 0003, Mingliang Xu 0001
IJCAI7
2020 MDSSD: multi-scale deconvolutional single shot detector for small objects
Lisha Cui, Pei Lv, Xiaoheng Jiang, Zhimin Gao, Bing Zhou 0003, Mingliang Xu 0001
Sci. China Inf. Sci.6
2020 Learning Multi-Level Density Maps for Crowd Counting
abstract
People in crowd scenes often exhibit the characteristic of imbalanced distribution. On the one hand, people size varies largely due to the camera perspective. People far away from the camera look smaller and are likely to occlude each other, whereas people near to the camera look larger and are relatively sparse. On the other hand, the number of people also varies greatly in the same or different scenes. This article aims to develop a novel model that can accurately estimate the crowd count from a given scene with imbalanced people distribution. To this end, we have proposed an effective multi-level convolutional neural network (MLCNN) architecture that first adaptively learns multi-level density maps and then fuses them to predict the final output. Density map of each level focuses on dealing with people of certain sizes. As a result, the fusion of multi-level density maps is able to tackle the large variation in people size. In addition, we introduce a new loss function named balanced loss (BL) to impose relatively BL feedback during training, which helps further improve the performance of the proposed network. Furthermore, we introduce a new data set including 1111 images with a total of 49 061 head annotations. MLCNN is easy to train with only one end-to-end training stage. Experimental results demonstrate that our MLCNN achieves state-of-the-art performance. In particular, our MLCNN reaches a mean absolute error (MAE) of 242.4 on the UCF_CC_50 data set, which is 37.2 lower than the second-best result.
Xiaoheng Jiang, Li Zhang 0072, Pei Lv, Yibo Guo, Ruijie Zhu 0001, Yanwei Pang, Xi Li 0001, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Neural Networks Learn. Syst.9
2019 Disassembling a 3D mechanism for efficient packing
abstract
Abstract This paper introduces a disassemble‐and‐pack algorithm to disassemble a mechanical 3D model in groups that can be efficiently packed within a box, with the objective of reassembling them easily after delivery. Its key feature is that, mostly, the mechanism can be disassembled at the joint and each part can be an adjusted motion structure based on its joint type. Our system consists of two steps: disassembling the mechanical object into a group set and packing them within a box efficiently. The first step consists in the creation of a hierarchy of possible group set of parts that can be tightly packed within their minimum bounding boxes. Use the breadth‐first search algorithm to traverse the hierarchy of possible group set in order to disconnect the joints and get the group set. In the second step, according to the reverse order of volume, each group in the set is inserted into the specified box. The fact that mechanism disassembly and shape packing are both an NP‐complete problem justifies finding approximated solutions according to efficacy and efficiency. Experimental results show that our approach can really efficiently pack a range of mechanisms from a simple model to complex objects.
Xiaoheng Jiang, Ning-Bo Gu, Weiwei Xu 0003, Junxiao Xue, Bing Zhou 0003, Mingliang Xu 0001
Comput. Animat. Virtual Worlds6
2019 Personalized training through Kinect-based games for physical education
Mingliang Xu 0001, Yafang Zhai, Yibo Guo, Pei Lv, Meng Wang 0001, Bing Zhou 0003
J. Vis. Commun. Image Represent.7
2019 Depth Information Guided Crowd Counting for complex crowd scenes
Mingliang Xu 0001, Zhaoyang Ge, Xiaoheng Jiang, Gaoge Cui, Pei Lv, Bing Zhou 0003, Changsheng Xu
Pattern Recognit. Lett.6
2019 D-STC: Deep learning with spatio-temporal constraints for train drivers detection from videos
Mingliang Xu 0001, Fang Hao, Pei Lv, Lisha Cui, Shuo Zhang 0014, Bing Zhou 0003
Pattern Recognit. Lett.6
2019 Crowd Behavior Evolution With Emotional Contagion in Political Rallies
abstract
In this paper, we present a novel crowd behavior evolution method with emotional contagion in political rallies. We first analyze the most representative political rally scenes in detail and model them into two kinds of abstract scenario. Furthermore, the “extroversion” and “empathy” factors from the OCEAN model are chosen to describe the most important individual personalities in such scenarios. Based on this, an improved emotional contagion model is proposed by combining the Susceptible-Infected-Recovered model and individual personality under different political viewpoints. Finally, the crowd in a political rally is driven to move according to the new potential moving direction generated by emotional contagion and the original direction of the individual together. The experiments show that our method can intuitively demonstrate the emotional changes of those individuals with different political perspectives and reasonably simulate the crowd movement under the political rally scenes.
Pei Lv, Zhujin Zhang, Chaochao Li, Yibo Guo, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Comput. Soc. Syst.5
2019 Traffic Simulation and Visual Verification in Smog
abstract
Smog causes low visibility on the road and it can impact the safety of traffic. Modeling traffic in smog will have a significant impact on realistic traffic simulations. Most existing traffic models assume that drivers have optimal vision in the simulations, making these simulations are not suitable for modeling smog weather conditions. In this article, we introduce the Smog Full Velocity Difference Model (SMOG-FVDM) for a realistic simulation of traffic in smog weather conditions. In this model, we present a stadia model for drivers in smog conditions. We introduce it into a car-following traffic model using both psychological force and body force concepts, and then we introduce the SMOG-FVDM. Considering that there are lots of parameters in the SMOG-FVDM, we design a visual verification system based on SMOG-FVDM to arrive at an adequate solution which can show visual simulation results under different road scenarios and different degrees of smog by reconciling the parameters. Experimental results show that our model can give a realistic and efficient traffic simulation of smog weather conditions.
Mingliang Xu 0001, Shili Chu, Yong Gan, Xiaoheng Jiang, Bing Zhou 0003
ACM Trans. Intell. Syst. Technol.7
2019 Stylized Aesthetic QR Code
abstract
With the continued proliferation of smart mobile devices, the Quick Response (QR) code has become one of the most-used types of two-dimensional code in the world. Aiming at beautifying the visual-unpleasant appearance of QR codes, existing works have developed a series of techniques. However, these works still leave much to be desired, such as personalization, artistry, and robustness. To address these issues, in this paper, we propose a novel type of aesthetic QR codes, Stylized aEsthEtic (SEE) QR code , and a three-stage approach to automatically produce such robust style-oriented codes. Specifically, in the first stage, we propose a method to generate an optimized baseline aesthetic QR code, which reduces the visual contrast between the noise-like black/white modules and the blended image. In the second stage, to obtain an art style QR code, we tailor an appropriate neural style transformation network to endow the baseline aesthetic QR code with artistic elements. In the third stage, we design a module-based robustness-optimization mechanism to ensure the performance robust by balancing two competing terms: visual quality and readability. Extensive experiments demonstrate that the SEE QR code has high quality in terms of both visual appearance and robustness and also offers a greater variety of personalized choices to users.
Mingliang Xu 0001, Hao Su 0001, Xi Li 0001, Jing Liao 0001, Jianwei Niu 0002, Pei Lv, Bing Zhou 0003
IEEE Trans. Multim.8
2018 USAR: An Interactive User-specific Aesthetic Ranking Framework for Images
abstract
When assessing whether an image is of high or low quality, it is indispensable to take personal preference into account. Existing aesthetic models lay emphasis on hand-crafted features or deep features commonly shared by high quality images, but with limited or no consideration for personal preference and user interaction. To that end, we propose a novel and user-friendly aesthetic ranking framework via powerful deep neural network and a small amount of user interaction, which can automatically estimate and rank the aesthetic characteristics of images in accordance with users' preference. Our framework takes as input a series of photos that users prefer, and produces as output a reliable, user-specific aesthetic ranking model matching with users' preference. Considering the subjectivity of personal preference and the uncertainty of user's single selection, a unique and exclusive dataset will be constructed interactively to describe the preference of one individual by retrieving the most similar images with regard to those specified by users. Based on this unique user-specific dataset and sufficient well-designed aesthetic attributes, a customized aesthetic distribution model can be learned, which concatenates both personalized preference and aesthetic rules. We conduct extensive experiments and user studies on two large-scale public datasets, and demonstrate that our framework outperforms those work based on conventional aesthetic assessment or ranking model.
Pei Lv, Meng Wang 0001, Yongbo Xu, Junyi Sun, Shi-Mei Su, Bing Zhou 0003, Mingliang Xu 0001
ACM Multimedia7
2018 Shadow traffic: A unified model for abnormal traffic behavior simulation
Mingliang Xu 0001, Fubao Zhu, Zhigang Deng 0001, Bing Zhou 0003
Comput. Graph.6
2018 An Efficient Method of Crowd Aggregation Computation in Public Areas
abstract
The crowd stampede and terrorist attacks in public areas have now become more serious and dangerous threats due to the rapid increase in the population and scale of cities. Therefore, the analysis of crowd aggregation behavior has been a new research focus in the field of intelligent video surveillance. However, such public area scenes not only contain moving crowd but also contain other types of objects. The sizes of these objects are usually small, which make their appearances quite similar. Moreover, the individuals in a crowd move randomly and often occlude each other. All the above factors make the analysis of crowd aggregation very difficult. In this paper, the authors attempt to solve this problem in three aspects. First, a novel global feature is used to represent the moving crowd. This feature can well describe the spatial and the temporal motion information of points-of-interest. Second, a strategy is adopted to cluster the feature points first and then calculate the collectiveness. This makes the collectiveness computation of individual groups more consistent and effective. Finally, more comprehensive collective crowd descriptors are proposed to provide a detailed description of the crowd status. Based on the proposed descriptor, the authors realize the evolution analysis of the group movement and the crowd abnormal detection. The experiment results show that the proposed method is able to efficiently compute the crowd collectiveness in various public areas and provide a reliable reference for the public safety management.
Mingliang Xu 0001, Chunxu Li, Pei Lv, Nie Lin, Rui Hou 0001, Bing Zhou 0003
IEEE Trans. Circuits Syst. Video Technol.6
2017 Mechanical Assembly Packing Problem Using Joint Constraints
Mingliang Xu 0001, Ning-Bo Gu, Weiwei Xu 0003, Junxiao Xue, Bing Zhou 0003
J. Comput. Sci. Technol.6
2017 Learning-Based Shadow Recognition and Removal From Monochromatic Natural Images
abstract
This paper addresses the problem of recognizing and removing shadows from monochromatic natural images from a learning-based perspective. Without chromatic information, shadow recognition and removal are extremely challenging in this paper, mainly due to the missing of invariant color cues. Natural scenes make this problem even harder due to the complex illumination condition and ambiguity from many near-black objects. In this paper, a learning-based shadow recognition and removal scheme is proposed to tackle the challenges above-mentioned. First, we propose to use both shadow-variant and invariant cues from illumination, texture, and odd order derivative characteristics to recognize shadows. Such features are used to train a classifier via boosting a decision tree and integrated into a conditional random field, which can enforce local consistency over pixel labels. Second, a Gaussian model is introduced to remove the recognized shadows from monochromatic natural scenes. The proposed scheme is evaluated using both qualitative and quantitative results based on a novel database of hand-labeled shadows, with comparisons to the existing state-of-the-art schemes. We show that the shadowed areas of a monochromatic image can be accurately identified using the proposed scheme, and high-quality shadow-free images can be precisely recovered after shadow removal.
Mingliang Xu 0001, Jiejie Zhu, Pei Lv, Bing Zhou 0003, Marshall F. Tappen, Rongrong Ji
IEEE Trans. Image Process.4
2016 Data-driven humanlike reaching behaviors synthesis
Pei Lv, Mingliang Xu 0001, Bailin Yang, Bing Zhou 0003
Neurocomputing5
2016 Medical image denoising by parallel non-local means
Mingliang Xu 0001, Pei Lv, Fang Hao, Hongling Zhao, Bing Zhou 0003, Yusong Lin, Li-Wei Zhou
Neurocomputing6
2015 Summarizing surveillance videos with local-patch-learning-based abnormality detection, blob sequence optimization, and type-based synopsis
Weiyao Lin, Jiwen Lu, Bing Zhou 0003, Jinjun Wang, Yu Zhou 0015
Neurocomputing4
2014 Facial expression cloning with elastic and muscle models
Weiyao Lin, Bing Zhou 0003, Zhenzhong Chen 0001, Bin Sheng 0001, Jianxin Wu 0001
J. Vis. Commun. Image Represent.3
2013 Improved image deblurring based on salient-region segmentation
Weiyao Lin, Wei Li 0209, Bing Zhou 0003, Jijia Li
Signal Process. Image Commun.4