Pei Lv

dblp:47/8698 · DBLP profile ↗
← Back
67ranked-venue papers
13as first author
45since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FDI-Net: Frequency Decomposition Guided Cross-Modal Interaction for Multimodal Object Detection
Yulu Xu, Kaijiang Li, Pei Lv
ICIC (17)4
2026 LiteWiFi: Ultra-low Power Wi-Fi Radio for Ubiquitous IoT Connection
Zeming Yang, Fengyuan Zhu 0001, Yibin Deng, Pei Lv, Xiaohua Tian
INFOCOM6
2026 DeckVis: visual analysis of carrier-based aircraft deck support operation scenarios
Xingyu Guo, Fangfei Liu, Zhipan Liu, Ke Wang 0064, Pei Lv, Mingliang Xu 0001
Frontiers Comput. Sci.7
2026 A general dual-view framework for instance weighted naive Bayes
Huan Zhang 0007, Kexin Meng, Pei Lv, Shuo He 0002, Mingliang Xu 0001
Pattern Recognit.3
2026 To the Best of Trust: Full-Stage Trusted Multi-Modal Clustering
abstract
Multi-modal clustering aims to integrate complementary information from different modalities to uncover latent consistent structures and improve clustering performance. However, existing methods mainly rely on predictive (result) uncertainty to improve robustness, while often neglecting aleatoric (data) uncertainty introduced by sample noise and epistemic (model) uncertainty induced by model parameters and structural variations. To this end, we propose a novel Full-Stage Trusted Multi-modal Clustering (FSTMC) method. To the best of trust, we jointly utilize aleatoric, epistemic, and predictive uncertainties to optimize the model and learn more reliable feature representations and clustering results. In the representation learning phase, probabilistic modeling is used to capture stable latent representations and estimate aleatoric uncertainty, while structured random perturbations are present to estimate epistemic uncertainty. In the clustering stage, instead of conventional feature-level fusion, we design an evidence-based fusion strategy, where soft labels from each modality are first mapped into categorical evidence while cluster distributions are parameterized via a Dirichlet model, with finally dynamic multi-modal fusion achieved by Dempster-Shafer theory. To mitigate overconfidence and modal conflicts, prior constraints guided by aleatoric and epistemic uncertainty are imposed, resulting in calibrated predictive uncertainty. Finally, we exploit predictive uncertainty to selectively incorporate pseudo labels for optimization. Benchmark experiments on a number of multi-modal datasets demonstrate that our approach significantly improves accuracy compared to state-of-the-art methods.
Shizhe Hu, Yucong Wu, Jinlan Wang, Xiaoheng Jiang, Pei Lv, Mingliang Xu 0001
IEEE Trans. Image Process.6
2026 Multidimensional Feature-Guided Cross-Population Human Activity Recognition and Prediction
abstract
With the rapid development of wearable devices and intelligent sensing technologies, the demand for human behavior recognition in rehabilitation medicine and human-machine collaboration has been increasing. To address the issue of high variability in gait features caused by individual differences in cross-population gait analysis, and to tackle the insufficient generalization ability of models due to the coupling of pathological features with normal gait, we propose a multi-dimensional spatiotemporal feature-guided SG-LSTM framework, based on a dual-branch architecture comprising symmetric LSTM (S-LSTM) and grouped LSTM (G-LSTM) networks, for cross-population lower-limb activity recognition and prediction. On the one hand, the S-LSTM module with a symmetric input structure is used to explicitly model the spatiotemporal symmetry of lower-limb joints in normal gait. On the other hand, the G-LSTM module with a joint functional grouping strategy and local motion decoupling is employed to explicitly model the abnormal motion coupling of lower-limb joints in pathological gait. Furthermore, a dynamically weighted multi-task loss function is designed to jointly optimize gait trajectory prediction and classification tasks, allowing the framework to simultaneously produce both outputs and enhance the adaptability of the model. Extensive experiments on our self-constructed gait dataset as well as the HuGaDB and WearGait-PD datasets demonstrate that the proposed method not only outperforms several existing approaches in cross-population human behavior prediction and gait recognition, but also holds potential clinical application value, achieving state-of-the-art (SOTA) performance.
Renbo Liu, Yangfei Zhao, Pei Lv, Ke Wang 0064, Zhaoyang Ge, Mingliang Xu 0001
IEEE J. Biomed. Health Informatics3
2025 Adaptive Non-disjoint Discretization for Tree Augmented Naive Bayes
Pei Lv, Huan Zhang 0007
ADMA (4)4
2025 SocialMP: Learning Social Aware Motion Patterns via Additive Fusion for Pedestrian Trajectory Prediction
abstract
Accurately capturing social interaction in complex scenarios is essential for pedestrian trajectory prediction task. The uncertainty in pedestrian interactions and the physical constraints imposed by the environment make this task challenging. To solve this problem, existing methods adopt dimensionality reduction algorithms to capture explainable human motions and behaviors. However, these approaches not only suffer from weak social awareness due to the inadequate feature extraction, but also overlook physical constraints, leading to predicted trajectories often cross unwalkable areas. To overcome these problems, we build an attention-based motion pattern representation, named SocialMP, which can effectively enhance the social awareness and environmental perception of motion patterns. Specifically, our method first characterizes the motion patterns through singular value decomposition and defines a visual field-based rule to model environmental social interaction. Then, an attention-based additive fusion mechanism is designed to enhance social awareness and environment perception of motion patterns. Therein, we integrate social interactions into motion patterns through cross-attention mechanism to generate latent motion patterns, and feed them into our devised additive fusion structure with backward connection for multiple iterations. Lastly, we design a map loss function by applying an additional penalty into average displacement error to prevent the pedestrians from passing through the unwalkable area. Extensive experiments on ETH-UCY and SDD datasets demonstrate that our SocialMP can not only improve prediction accuracy but also generate plausible trajectories.
Tianci Gao, Pei Lv
IJCAI4
2025 EMAWNB: Enhanced Multi-view Attribute Weighted Naive Bayes
Guanzhi Liu, Kexin Meng, Pei Lv, Huan Zhang 0007
PAKDD (3)3
2025 Ultra Fast Automatic Exposure Target Estimation for Smartphone Platform
Chong Tang 0004, Chenlu Wei, Pei Lv, Ruirong You
PRCV (18)5
2025 HSVNet: A dual-branch collaborative network for low-light image enhancement in the HSV color space
Mengjie Chen, Kaijiang Li, Pei Lv, Mingliang Xu 0001
Comput. Graph.4
2025 Virtual-physical digital twin testbed for heterogeneous crowd operations
Mingliang Xu 0001, Wencan Luo, Shuo He 0002, Chaochao Li, Yibo Guo, Pei Lv
Sci. China Inf. Sci.10
2025 FlexPlan: High-flexibility interactive floorplan design based on ArchiGraph
abstract
AI-aided floorplan design is a longstanding task in computer graphics. However, most of the existing methods focus on generating floorplans by limited architecture-level elements (e.g., room sizes, positions, and adjacencies), which ignore environmental factors and do not support customized designs. In this paper, we propose FlexPlan, an interactive approach for high-flexibility floorplan design. In FlexPlan, we propose a novel graph structure, named ArchiGraph, which enables flexible editing more comprehensive layout elements (e.g., architectures, environments, human needs) in a floorplan. First, we match similar floorplans according to the input architecture and environment features. Then, leveraging ArchiGraph, we interactively produce rooms’ attributes and quickly output the vectorized floorplans. For ArchiGraph, we design a RelationNet to predict room adjacencies, and propose a BoxNet to generate high-quality room boxes. Subjective and objective experiments show that our method is compatible with generating diverse complex floorplans (e.g., floorplans with irregular layout boundaries and room shapes). Compared with the state-of-the-art methods, our method can produce higher quality floorplans, and increase the speed of layout generation by nearly 20 times at most.
Zongpu Li, Hao Su 0001, Pei Lv, Mingliang Xu 0001
Graph. Model.5
2025 Local Cross-Patch Activation From Multi-Direction for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) learns to localize objects using only image-level labels. Recently, some studies apply transformers in WSOL to capture the long-range feature dependency and alleviate the partial activation issue of CNN-based methods. However, existing transformer-based methods still face two challenges. The first challenge is the over-activation of backgrounds. Specifically, the object boundaries and background are often semantically similar, and localization models may misidentify the background as a part of objects. The second challenge is the incomplete activation of occluded objects, since transformer architecture makes it difficult to capture local features across patches due to ignoring semantic and spatial coherence. To address these issues, in this paper, we propose LCA-MD, a novel transformer-based WSOL method using local cross-patch activation from multi-direction, which can capture more details of local features while inhibiting the background over-activation. In LCA-MD, first, combining contrastive learning with the transformer, we propose a token feature contrast module (TCM) that can maximize the difference between foregrounds and backgrounds and further separate them more accurately. Second, we propose a semantic-spatial fusion module (SFM), which leverages multi-directional perception to capture the local cross-patch features and diffuse activation across occlusions. Experiment results on the CUB-200-2011 and ILSVRC datasets demonstrate that our LCA-MD is significantly superior and has achieved state-of-the-art results in WSOL. The project code is available at https://github.com/rjy-fighting/LCA-MD.
Pei Lv, Junying Ren, Genwang Han, Jiwen Lu, Mingliang Xu 0001
IEEE Trans. Image Process.1
2025 CaliFree3DLane: Calibration Free Spatio-Temporal BEV Representation for Monocular 3D Lane Detection
abstract
Monocular 3D lane detection plays a crucial role in autonomous driving, assisting vehicles in safe navigation. Existing methods primarily utilize calibrated camera parameters in the dataset to conduct 3D lane detection from a single image. However, errors or sudden absence of camera parameters can pose significant challenges to safe driving. On one hand, this can lead to incorrect feature acquisition, which further affects the precision of lane detection. On the other hand, it renders methods relying on transformation matrices for temporal fusion ineffective. To address the above issue and achieve accurate 3D lane detection, we propose CaliFree3DLane, a calibration-free method for spatio-temporal 3D lane detection based on Transformer structure. Instead of using geometric projections to obtain static reference points on images, we propose a reference point refinement strategy that dynamically updates the reference points and finally generates appropriate sampling points for image feature extraction. To integrate multi-frame features, we generate sub-queries from the current scene query to focus on the image features of each frame independently. We then aggregate these sub-queries to form a more comprehensive scene query for 3D lane detection. Using these operations, CaliFree3DLane accurately transforms multi-frame image features into the current bird’s-eye view (BEV) space, enabling precise 3D lane detection. Experimental results show that our CaliFree3DLane achieves state-of-the-art 3D lane detection performance in various datasets. Compared to the Transformer-based methods of the same type, we have also improved$ {6.0\%}\sim {10.5\%}$at the F1 score. Code is available athttps://github.com/Ciisrlab/CaliFree3DLane.
Weizhi Guo, Chaochao Li, Kaijiang Li, Pei Lv, Mingliang Xu 0001
IEEE Trans. Intell. Transp. Syst.4
2025 Remember and Recall: Associative-Memory-Based Trajectory Prediction
abstract
Trajectory prediction is a key component of autonomous driving systems, enabling the application of accumulated movement experience to current scenarios. Although most existing methods concentrate on learning continuous representations to gain valuable experience, they often suffer from computational inefficiencies and struggle with unfamiliar situations. To address this issue, we propose the Fragmented-Memory-based Trajectory Prediction model (FMTP) inspired by the remarkable learning capabilities of humans, particularly their ability to leverage accumulated experience and recall relevant memories in unfamiliar situations. The FMTP model employs discrete representations to enhance computational efficiency by reducing information redundancy while maintaining the flexibility to utilize past experiences. Specifically, a learnable memory array is designed by consolidating continuous trajectory representations from the training set using vector quantization operations during the training phase. The array further eliminates redundant information while preserving essential features in discrete form. Additionally, an advanced reasoning engine based on language models is developed to deeply learn the associative rules among these discrete representations. Our method has been evaluated on various public datasets, including ETH-UCY, inD, SDD, nuScenes, Waymo, and VTL-TP. The extensive experimental results demonstrate that our approach achieves significant performance and extracts more valuable experience from past trajectories to inform the current state.
Tianci Gao, Junning Su, Pei Lv, Mingliang Xu 0001
IEEE Trans. Intell. Transp. Syst.5
2025 An Efficient Ungrouped Mask Method With two Learnable Parameters for 3D Object Detection
abstract
In 3D point cloud-based object detection, attention mechanism in Group-Free [1] learns direct relationships between proposals and all seed points, providing each proposal with a global context in the form of a cross-attention map. However, our analysis and experimental comparison show that the attention mechanism assigns inappropriately large attention weights to certain seed points far from a proposal, which is not conducive to detecting objects correctly. In this work, we alleviate the above problem by proposing a mask method. For an initial proposal, our method first calculates a spatial distance-based mask, which measures the spatial relationship between all seed points and the proposal. Then, we fuse the mask into cross-attention layers in stacked attention modules and get a refined cross-attention map. In essence, our mask gives each proposal a local context; after it is fused with the global context given by the attention mechanism, the refined cross-attention map could suppress the negative impact of some distant seed points on a proposal. We present two alternative strategies to compute the mask, a hard mask, and a soft mask. Experimental results demonstrate that the soft mask brings better performance. In the soft mask, for each initial proposal's 3D-box shape, we use a parametric approximate ellipsoid as the basis of the mask's calculation, which has only two learnable parameters. Experimental results show our work could outperform Group-Free 0.7 [email protected] at the cost of increasing inference time by less than 1%. The performance of our algorithm on the public dataset SUN RGB-D is 63.7 [email protected] and 45.5 [email protected], which is the best performance among algorithms that preserve the irregular of seed points.
Shuai Guo 0004, Lei Shi 0001, Xiaoheng Jiang, Pei Lv, Qidong Liu 0001, Yazhou Hu, Rongrong Ji, Mingliang Xu 0001
IEEE Trans. Multim.4
2025 TITFormer: Combining Textual Modality and Simulating Infrared Modality Based on Transformer for Image Enhancement
abstract
The image enhancement task requires a complex balance between extracting high-level contextual information and optimizing spatial details in the image to improve the visual quality. Most of existing methods have limited capability in capturing contextual features and optimizing spatial details when they only rely on a single modality. To address the above issues, this paper introduces a novel multi-modal image enhancement network based on Transformer, named as TITFormer, which combines textual and simulating infrared modalities firstly for this important task. TITFormer comprises a text channel attention fusion (TCF) network block and an infrared-guided spatial detail optimization (SDO) network block. The TCF extracts contextual features from the high-dimensional features compressed after spatial channel transformation of the textual feature and image feature. The SDO module uses simulating infrared images characterized by pixel intensity to guide the optimization of spatial details with contextual features adaptively. Experimental results demonstrate that TITFormer achieves state-of-the-art performance on two publicly available benchmark datasets.
Kaijiang Li, Haining Li, Miduo Cui, Junxin Li, Pei Lv, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Multim.5
2024 Foreground and Background Separate Adaptive Equilibrium Gradients Loss for Long-Tail Object Detection
Tianran Hao, Ying Tao, Meng Li 0039, Lisha Cui, Pei Lv, Mingliang Xu 0001
CVM (2)7
2024 GLAD: A Global-Attention-Based Diffusion Model for Infrared and Visible Image Fusion
Haozhe Guo, Mengjie Chen, Kaijiang Li, Pei Lv
ICIC (7)5
2024 Semantic-Aware Global and Local Fusion Model for Image Enhancement
Mengjie Chen, Pei Lv
PRCV (9)3
2024 S-CVAE: Stacked CVAE for Trajectory Prediction With Incremental Greedy Region
abstract
Predicting accurate future trajectories of agents is essential for autonomous navigation in complex scenarios. Although numerous work has made great progress on this goal, it is still challenging due to the uncertainty and continuity of behavioral intentions of agents, where uncertainty means the instantaneous multimodality of motion behavior, while the continuity refers to the consistency and stability of behavioral intention of an agent over a period of time constrained by its final destination. These factors easily affect the improvement of prediction accuracy. In this paper, we present a novel trajectory prediction method, Stacked Conditional VAE (S-CVAE) with Incremental Greedy Region (IGR). Specifically, the IGR is designed to enlarge the coverage of candidate waypoints/endpoints by reformulating the waypoints/endpoints prediction problem as candidate region generation, which can further encourage and model multimodality of behavioral intentions. Meanwhile, to exploit the inherent continuity between adjacent behavioral intentions of an agent, the S-CVAE architecture is constructed to transmit the behavioral intentions of one agent by inserting intermediate waypoints with IGR into the potential trajectories from the observed path to the final endpoint, and also enhances the reliability of the generated waypoints/endpoints in the next moment, further improve the accuracy of trajectory prediction. Our method is evaluated on several public datasets, including nuScenes, Apolloscape, SDD, INTERSECTION, Waymo, and VTPTL. The comprehensive experimental results demonstrate that our method achieves significant performance on these datasets. Especially in nuScenes and VTPTL, the accuracy is increased by at least 11.11% on average ADE and 2.40% on average FDE compared with state-of-the-arts.
Junning Su, Chaochao Li, Pei Lv, Mingliang Xu 0001
IEEE Trans. Intell. Transp. Syst.5
2024 SSAGCN: Social Soft Attention Graph Convolution Network for Pedestrian Trajectory Prediction
abstract
Pedestrian trajectory prediction is an important technique of autonomous driving. In order to accurately predict the reasonable future trajectory of pedestrians, it is inevitable to consider social interactions among pedestrians and the influence of surrounding scene simultaneously, which can fully represent the complex behavior information and ensure the rationality of predicted trajectories obeyed realistic rules. In this article, we propose one new prediction model named social soft attention graph convolution network (SSAGCN), which aims to simultaneously handle social interactions among pedestrians and scene interactions between pedestrians and environments. In detail, when modeling social interaction, we propose a new social soft attention function, which fully considers various interaction factors among pedestrians. Also, it can distinguish the influence of pedestrians around the agent based on different factors under various situations. For the scene interaction, we propose one new sequential scene sharing mechanism. The influence of the scene on one agent at each moment can be shared with other neighbors through social soft attention; therefore, the influence of the scene is expanded both in spatial and temporal dimensions. With the help of these improvements, we successfully obtain socially and physically acceptable predicted trajectories. The experiments on public available datasets prove the effectiveness of SSAGCN and have achieved state-of-the-art results. The project code is available at https://github.com/WW-Tong/ssagcn_for_path_prediction.
Pei Lv, Wentong Wang, Yunxin Wang, Mingliang Xu 0001, Changsheng Xu
IEEE Trans. Neural Networks Learn. Syst.1
2024 TraInterSim: Adaptive and Planning-Aware Hybrid-Driven Traffic Intersection Simulation
abstract
Traffic intersections are important scenes that can be seen almost everywhere in the traffic system. Currently, most simulation methods perform well at highways and urban traffic networks. In intersection scenarios, the challenge lies in the lack of clearly defined lanes, where agents with various motion plannings converge in the central area from different directions. Traditional model-based methods are difficult to drive agents to move realistically at intersections without enough predefined lanes, while data-driven methods often require a large amount of high-quality input data. Simultaneously, tedious parameter tuning is inevitable involved to obtain the desired simulation results. In this paper, we present a novel adaptive and planning-aware hybrid-driven method (TraInterSim) to simulate traffic intersection scenarios. Our hybrid-driven method combines an optimization-based data-driven scheme with a velocity continuity model. It guides the agent's movements using real-world data and can generate those behaviors not present in the input data. Our optimization method fully considers velocity continuity, desired speed, direction guidance, and planning-aware collision avoidance. Agents can perceive others' motion plannings and relative distances to avoid possible collisions. To preserve the individual flexibility of different agents, the parameters in our method are automatically adjusted during the simulation. TraInterSim can generate realistic behaviors of heterogeneous agents in different traffic intersection scenarios in interactive rates. Through extensive experiments as well as user studies, we validate the effectiveness and rationality of the proposed simulation method.
Pei Lv, Xinming Pei, Xinyu Ren, Chaochao Li, Mingliang Xu 0001
IEEE Trans. Vis. Comput. Graph.1
2023 TTA-GCN: Temporal Topology Aggregation for Skeleton-Based Action Recognition
Haoming Meng, Yangfei Zhao, Yijing Guo, Pei Lv
ICIG (2)4
2023 Multi-agent broad reinforcement learning for intelligent traffic light control
Ruijie Zhu 0001, Lulu Li 0010, Shuning Wu, Pei Lv, Mingliang Xu 0001
Inf. Sci.4
2023 Emotional Contagion-Aware Deep Reinforcement Learning for Antagonistic Crowd Simulation
abstract
The antagonistic behavior in the crowd usually exacerbates the seriousness of the situation in sudden riots, where the antagonistic emotional contagion and behavioral decision making play very important roles. However, the complex mechanism of antagonistic emotion influencing decision making, especially in the environment of sudden confrontation, has not yet been explored very clearly. In this paper, we propose an Emotional contagion-aware Deep reinforcement learning model for Antagonistic Crowd Simulation (ACSED). First, we build a group emotional contagion module based on the improved Susceptible Infected Susceptible (SIS) infection disease model, and estimate the emotional state of the group at each time step during the simulation. Then, the tendency of crowd antagonistic action is estimated based on Deep Q Network (DQN), where the agent learns the action autonomously, and leverages the mean field theory to quickly calculate the influence of other surrounding individuals on the central one. Finally, the rationality of the predicted actions by DQN is further analyzed in combination with group emotion, and the final action of the agent is determined. The proposed method in this paper is verified through several experiments with different settings. We can conclude antagonistic emotions play a critical role in the decision making of the crowd through influencing the individual behavior in the riot scenario, where individual behaviors are primarily driven by emotions and goals, rather than common rules. The experiment results also prove that the antagonistic emotion has a vital impact on the group combat, and positive emotional states are more conducive to combat. Moreover, by comparing the simulation results with real scenes, the feasibility of our method is further confirmed, which can provide good reference to formulate battle plans and improve the win rate of righteous groups in a variety of situations.
Pei Lv, Qingqing Yu, Boya Xu, Chaochao Li, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Affect. Comput.1
2023 BIP-Tree: Tree Variant With Behavioral Intention Perception for Heterogeneous Trajectory Prediction
abstract
An insightful understanding and relational reasoning of motion behavior are typical components for trajectory prediction to achieve safe planning when navigating in complex scenarios. Due to the differences in behavioral responses of heterogeneous agents and the existence of chain effect in message passing, an effective prediction method is desired to better acquire potential behavioral intention and model motion behavior. In this paper, we construct a trajectory prediction method to represent and encode the behavioral interactions among heterogeneous agents, called as Tree variant with Behavioral Intention Perception (BIP-Tree). Specifically, a dual-behavior interaction module is presented to deeply understand behavioral intention by simultaneously considering the behavioral perception and behavioral response in spatial interaction. The behavioral perception means that individual acquires behavioral features from interactive objects located in its perception range, while the behavioral response means that each agent makes distinctive reactions to different categories of agents (for example, due to different collision risks caused by pedestrian and vehicle, a pedestrian will respond differently to the interactive agents at the same distance). Meanwhile, we also introduce one new tree variant in message passing stage to enhance the acquisition of potential motion feature, denoting traffic agents as nodes and the interactions among them as tree trunks. The interaction message can be delivered along tree trunks from leaf nodes to root node, to further achieves the chain effect of high-order interactions beyond adjacent entities. Our method is evaluated on several public datasets, such as Apolloscape, nuScenes, Argoverse, SDD, INTERACTION, inD, and Waymo. The extensive experimental results demonstrate that our method can predict more plausible and realistic trajectories with multi-modality. Among them, the best performance is achieved on three datasets. More remarkably, compared with state-of-the-arts, our method achieves significant performance and decreases by at least 13.04% on average ADE and 19.42% on average FDE on inD dataset with four intersections. The dataset and code are available at: htpps://github.com/VTP-TL/BIP-Tree.
Weizhi Guo, Junning Su, Pei Lv, Mingliang Xu 0001
IEEE Trans. Intell. Transp. Syst.4
2023 User-Guided Personalized Image Aesthetic Assessment Based on Deep Reinforcement Learning
abstract
Personalized image aesthetic assessment (PIAA) has recently become a hot topic due to its wide applications, such as photography, film, television, e-commerce, fashion design, and so on. This task is more seriously affected by subjective factors and samples provided by users. In order to acquire precise personalized aesthetic distribution by small amount of samples, we propose a novel user-guided personalized image aesthetic assessment framework. This framework leverages user interactions to retouch and rank images for aesthetic assessment based on deep reinforcement learning (DRL), and generates personalized aesthetic distribution that is more in line with the aesthetic preferences of different users. It mainly consists of two stages. In the first stage, personalized aesthetic ranking is generated by interactive image enhancement and manual ranking, meanwhile, two policy networks will be trained. These two networks will be trained iteratively and alternatively to facilitate the final personalized aesthetic assessment. In the second stage, these modified images are labeled with aesthetic attributes by one style-specific classifier, and then the personalized aesthetic distribution is generated based on the multiple aesthetic attributes of these images, which conforms to the aesthetic preference of users better. Compared with other existing methods, our approach has achieved new state-of-the-art in the task of personalized image aesthetic assessment on the public AVA and FLICKR-AES datasets.
Pei Lv, Jianqi Fan, Xixi Nie, Weiming Dong, Xiaoheng Jiang, Bing Zhou 0003, Mingliang Xu 0001, Changsheng Xu
IEEE Trans. Multim.1
2022 Focal and Global Spatial-Temporal Transformer for Skeleton-Based Action Recognition
Zhimin Gao, Peitao Wang, Pei Lv, Xiaoheng Jiang, Qidong Liu 0001, Pichao Wang, Mingliang Xu 0001, Wanqing Li 0001
ACCV (4)3
2022 D2-TPred: Discontinuous Dependency for Trajectory Prediction Under Traffic Lights
Wentong Wang, Weizhi Guo, Pei Lv, Mingliang Xu 0001, Wei Chen 0001, Dinesh Manocha
ECCV (8)4
2022 Trajectory distributions: A new description of movement for trajectory prediction
abstract
Trajectory prediction is a fundamental and challenging task for numerous applications, such as autonomous driving and intelligent robots. Current works typically treat pedestrian trajectories as a series of 2D point coordinates. However, in real scenarios, the trajectory often exhibits randomness, and has its own probability distribution. Inspired by this observation and other movement characteristics of pedestrians, we propose a simple and intuitive movement description called a trajectory distribution, which maps the coordinates of the pedestrian trajectory to a 2D Gaussian distribution in space. Based on this novel description, we develop a new trajectory prediction method, which we call the social probability method . The method combines trajectory distributions and powerful convolutional recurrent neural networks. Both the input and output of our method are trajectory distributions, which provide the recurrent neural network with sufficient spatial and random information about moving pedestrians. Furthermore, the social probability method extracts spatio-temporal features directly from the new movement description to generate robust and accurate predictions. Experiments on public benchmark datasets show the effectiveness of the proposed method.
Pei Lv, Tianxin Gu, Xiaoheng Jiang, Bing Zhou 0003, Mingliang Xu 0001
Comput. Vis. Media1
2022 Transferring priors from virtual data for crowd counting in real world
Xiaoheng Jiang, Hao Liu 0057, Li Zhang 0072, Geyang Li, Mingliang Xu 0001, Pei Lv, Bing Zhou 0003
Frontiers Comput. Sci.6
2022 ACSEE: Antagonistic Crowd Simulation Model With Emotional Contagion and Evolutionary Game Theory
abstract
Antagonistic crowd behaviors are often observed in cases of serious conflict. Antagonistic emotions, which is the typical psychological state of agents in different roles (i.e., cops, activists, and civilians) in crowd violence scenes, and the way they spread through contagion in a crowd are important causes of crowd antagonistic behaviors. Moreover, games, which refers to the interaction between opposing groups adopting different strategies to obtain higher benefits and less casualties, determine the level of crowd violence. We present an antagonistic crowd simulation model (ACSEE), which is integrated with antagonistic emotional contagion and evolutionary game theories. Our approach models the antagonistic emotions between agents in different roles using two components: mental emotion and external emotion. We combine enhanced susceptible-infectious-susceptible (SIS) and game approaches to evaluate the role of antagonistic emotional contagion in crowd violence. Our evolutionary game theoretic approach incorporates antagonistic emotional contagion through deterrent force, which is modelled by a mixture of emotional forces and physical forces defeating the opponents. Antagonistic emotional contagion and evolutionary game theories influence each other to determine antagonistic crowd behaviors. We evaluate our approach on real-world scenarios consisting of different kinds of agents. We also compare the simulated crowd behaviors with real-world crowd videos and use our approach to predict the trends of crowd movements in violence incidents. We investigate the impact of various factors (number of agents, emotion, strategy, etc.) on the outcome of crowd violence. We present results from user studies suggesting that our model can simulate antagonistic crowd behaviors similar to those seen in real-world scenarios.
Chaochao Li, Pei Lv, Dinesh Manocha, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Affect. Comput.2
2022 Agent-Based Campus Novel Coronavirus Infection and Control Simulation
abstract
Corona Virus Disease 2019 (COVID-19), due to its extremely high infectivity, has been spreading rapidly around the world and bringing huge influence to socioeconomic development and people’s daily life. Taking for example the virus transmission that may occur after college students return to school, we analyze the quantitative influence of the key factors on the virus spread, including crowd density and self-protection. One Campus Virus Infection and Control Simulation (CVICS) model of the novel coronavirus is proposed in this article, fully considering the characteristics of repeated contact and strong mobility of crowd in the closed environment. Specifically, we build an agent-based infection model, introduce the mean field theory to calculate the probability of virus transmission, and microsimulate the daily prevalence of infection among individuals. The experimental results show that the proposed model in this article efficiently simulates how the virus spreads in the dense crowd in frequent contact under a closed environment. Furthermore, preventive and control measures, such as self-protection, crowd decentralization, and isolation during the epidemic, can effectively delay the arrival of infection peak, reduce the prevalence, and, finally, lower the risk of COVID-19 transmission after the students return to school.
Pei Lv, Boya Xu, Ran Feng, Chaochao Li, Junxiao Xue, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Comput. Soc. Syst.1
2022 Context-Aware Block Net for Small Object Detection
abstract
State-of-the-art object detectors usually progressively downsample the input image until it is represented by small feature maps, which loses the spatial information and compromises the representation of small objects. In this article, we propose a context-aware block net (CAB Net) to improve small object detection by building high-resolution and strong semantic feature maps. To internally enhance the representation capacity of feature maps with high spatial resolution, we delicately design the context-aware block (CAB). CAB exploits pyramidal dilated convolutions to incorporate multilevel contextual information without losing the original resolution of feature maps. Then, we assemble CAB to the end of the truncated backbone network (e.g., VGG16) with a relatively small downsampling factor (e.g., 8) and cast off all following layers. CAB Net can capture both basic visual patterns as well as semantical information of small objects, thus improving the performance of small object detection. Experiments conducted on the benchmark Tsinghua-Tencent 100K and the Airport dataset show that CAB Net outperforms other top-performing detectors by a large margin while keeping real-time speed, which demonstrates the effectiveness of CAB Net for small object detection.
Lisha Cui, Pei Lv, Xiaoheng Jiang, Zhimin Gao, Bing Zhou 0003, Ling Shao 0001, Mingliang Xu 0001
IEEE Trans. Cybern.2
2022 TriATNE: Tripartite Adversarial Training for Network Embeddings
abstract
Existing network embedding algorithms based on generative adversarial networks (GANs) improve the robustness of node embeddings by selecting high-quality negative samples with the generator to play against the discriminator. Since most of the negative samples can be easily discriminated from positive samples in graphs, their poor competitiveness weakens the function of the generator. Inspired by the sales skills in the market, in this article, we present tripartite adversarial training for network embeddings (TriATNE), a novel adversarial learning framework for learning stable and robust node embeddings. TriATNE consists of three players: 1) producer; 2) seller; and 3) customer. The producer strives to learn the representation of each sample (node pair), making it easy for the customer to differentiate between the positive and the negative, while the seller tries to confuse the customer by selecting realistic-looking samples. The customer, a biased evaluation metric, provides feedback for training the producer and the seller. To further enhance the robustness of node embedding, we model the customer as a two-layer neural network, where each unit in the hidden layer can be regarded as a customer with different preferences. TriATNE also plays against the producer by adjusting the weight of each customer. We test the performance of TriATNE on two common tasks: classification as well as link prediction. The experimental results on various publicly available datasets show that TriATNE can exploit the network structure well.
Qidong Liu 0001, Cheng Long 0001, Jie Zhang 0002, Mingliang Xu 0001, Pei Lv
IEEE Trans. Cybern.5
2022 Contrastive Proposal Extension With LSTM Network for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) has attracted more and more attention since it only uses image-level labels and can save huge annotation costs. Most of the WSOD methods use Multiple Instance Learning (MIL) as their basic framework, which regard it as an instance classification problem. However, these methods based on MIL tend to only converge on the most discriminative regions of different instances, rather than their corresponding complete regions, that is, insufficient integrity. Inspired by the human habit of observing things, we propose a new method by comparing the initial proposals and the extended ones to optimize those initial proposals. Specifically, we propose one new strategy for WSOD by involving contrastive proposal extension (CPE), which consists of multiple directional contrastive proposal extensions (D-CPEs), and each D-CPE contains LSTM-based encoders and dual-stream decoders. Firstly, the boundary of initial proposals in MIL is extended to different positions according to well-designed sequential order. Then, the CPE compares the extended proposal and the initial one by extracting the feature semantics of them using the encoders, and calculates the integrity of the initial proposal to optimize its score. These contrastive contextual semantics will guide the basic WSOD to suppress bad proposals and improve the scores of good ones. In addition, a simple dual-stream network is designed as the decoder to constrain the temporal coding of LSTM and improve the performance of WSOD further. Experiments on PASCAL VOC 2007, VOC 2012 and MS-COCO datasets show that our method has achieved the state-of-the-art results.
Pei Lv, Suqi Hu, Tianran Hao
IEEE Trans. Image Process.1
2021 Protected Resource Allocation in Space Division Multiplexing-Elastic Optical Networks with Fluctuating Traffic
Ruijie Zhu 0001, Aretor Samuel, Peisen Wang, Shihua Li 0007, Bounsou Kham Oun, Lulu Li 0010, Pei Lv, Mingliang Xu 0001, Shui Yu 0001
J. Netw. Comput. Appl.7
2021 Bio-Inspired Deep Attribute Learning Towards Facial Aesthetic Prediction
abstract
Computational prediction of facial aesthetics has attracted ever-increasing research focus, which has wide range of prospects in multimedia applications. The key challenge lies in extracting discriminative and perception-aware features to characterize the facial beautifulness. To this end, the existing schemes simply adopt a direct feature mapping, which relies on handcraft-designed low-level features that cannot reflect human-level aesthetic perception. In this paper, we present a systematic framework towards designing biology-inspired, discriminative representation for facial aesthetic prediction. First, we design a group of biological experiments that adopt eye tracker to identify spatial regions of interest during the facial aesthetic judgments of subjects, which forms a Bio-inspired Facial Aesthetic Ontology (Bio-FAO) and is made public available. Second, we adopt the cutting-edge convolutional neural network to train a set of Bio-inspired Attribute features, termed Bio-AttriBank, which forms a mid-level interpretable representation corresponding to the aforementioned Bio-FAO. For a given image, the facial aesthetic prediction is then formulated as a classification problem over the Bio-AttriBank descriptor responses, which well bridges the affective gap, and provides explainable evidences on why/how a face is beautiful or not. We have carried out extensive experiments on both JAFFE and FaceWarehouse datasets, with comparisons to a set of state-of-the-art and alternative approaches. Superior performance gains in the experiments have demonstrated the merits of the proposed scheme.
Mingliang Xu 0001, Fuhai Chen, Pei Lv, Bing Zhou 0003, Rongrong Ji
IEEE Trans. Affect. Comput.5
2021 Emotion-Based Crowd Simulation Model Based on Physical Strength Consumption for Emergency Scenarios
abstract
Increasing attention is being given to the modeling and simulation of traffic flow and crowd movement, two phenomena that both deal with interactions between pedestrians and cars in many situations. In particular, crowd simulation is important for understanding mobility and transportation patterns. In this paper, we propose an emotion-based crowd simulation model integrating physical strength consumption. Inspired by the theory of “the devoted actor,” the movements of each individual in our model are determined by modeling the influence of physical strength consumption and the emotion of panic. In particular, human physical strength consumption is computed using a physics-based numerical method. Inspired by the James-Lange theory, panic levels are estimated by means of an enhanced emotional contagion model that leverages the inherent relationship between physical strength consumption and panic. To the best of our knowledge, our model is the first method integrating physical strength consumption into an emotion-based crowd simulation model by exploiting the relationship between physical strength consumption and emotion. We highlight the performance on different scenarios and compare the resulting behaviors with real-world video sequences. Our approach can reliably predict changes in physical strength consumption and panic levels of individuals in an emergency situation.
Mingliang Xu 0001, Chaochao Li, Pei Lv, Wei Chen 0001, Zhigang Deng 0001, Bing Zhou 0003, Dinesh Manocha
IEEE Trans. Intell. Transp. Syst.3
2021 Top-$k$k Vehicle Matching in Social Ridesharing: A Price-Aware Approach
abstract
In the past few years ridesharing has largely reshaped the transportation marketplace. It is envisioned as a promising solution to transportation-related problems in metropolitan cities, such as traffic congestion and air pollution. In the current ridesharing research, social ridesharing, which makes use of social relations among drivers and riders to address safety issues, and dynamic pricing are two active directions with important business implications. Simultaneously optimizing social cohesion and revenue is vital to a commercial ridesharing platform's sustainable development, which, however, has not been previously studied. In this paper, we first present a new pricing scheme that better incentivizes drivers and riders to participate in ridesharing, and then propose a novel type of Price-aware Top-$k$Matching (PTkM) queries which retrieve the top-$k$vehicles for a rider's request by taking into account both social relations and revenue. We design an efficient algorithm with a set of powerful pruning techniques to tackle this problem. Moreover, we propose a novel index tailored to our problem to further speed up query processing. Extensive experimental results on real datasets show that our proposed algorithms achieve desirable performance for real-world deployment.
Ji Wan, Rui Chen 0012, Jianliang Xu, Xiaoyi Fu, Hongyan Gu, Pei Lv, Mingliang Xu 0001
IEEE Trans. Knowl. Data Eng.7
2021 Density-Aware Multi-Task Learning for Crowd Counting
abstract
In this paper, we present a method called density-aware convolutional neural network (DensityCNN) to perform the crowd counting task in various crowded scenes. The key idea of the DensityCNN is to utilize high-level semantic information to provide guidance and constraint when generating density maps. To this end, we implement the DensityCNN by adopting a multi-task CNN structure to jointly learn density-level classification and density map estimation. The density-level classification task learns multi-channel semantic features that are aware of the density distributions of the input image. This task is accomplished via our specially designed group-based convolutional structure in a supervised learning manner. In the density map estimation task, these semantic features are deployed together with high-dimension convolutional features to generate density maps with lower count errors. Extensive experiments on four challenging crowd datasets (ShanghaiTech, UCF_CC_50, UCF-QNCF, and WorldExpo'10) and one vehicle dataset TRANCOS demonstrate the effectiveness of the proposed method.
Xiaoheng Jiang, Li Zhang 0072, Tianzhu Zhang 0001, Pei Lv, Bing Zhou 0003, Yanwei Pang, Mingliang Xu 0001, Changsheng Xu
IEEE Trans. Multim.4
2021 ART-UP: A Novel Method for Generating Scanning-Robust Aesthetic QR Codes
abstract
Quick response (QR) codes are usually scanned in different environments, so they must be robust to variations in illumination, scale, coverage, and camera angles. Aesthetic QR codes improve the visual quality, but subtle changes in their appearance may cause scanning failure. In this article, a new method to generate scanning-robust aesthetic QR codes is proposed, which is based on a module-based scanning probability estimation model that can effectively balance the tradeoff between visual quality and scanning robustness. Our method locally adjusts the luminance of each module by estimating the probability of successful sampling. The approach adopts the hierarchical, coarse-to-fine strategy to enhance the visual quality of aesthetic QR codes, which sequentially generate the following three codes: a binary aesthetic QR code, a grayscale aesthetic QR code, and the final color aesthetic QR code. Our approach also can be used to create QR codes with different visual styles by adjusting some initialization parameters. User surveys and decoding experiments were adopted for evaluating our method compared with state-of-the-art algorithms, which indicates that the proposed approach has excellent performance in terms of both visual quality and scanning robustness.
Mingliang Xu 0001, Qingfeng Li 0004, Jianwei Niu 0002, Hao Su 0001, Xiting Liu, Weiwei Xu 0003, Pei Lv, Bing Zhou 0003, Yi Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2021 Crowd Behavior Simulation With Emotional Contagion in Unexpected Multihazard Situations
abstract
Numerous research efforts have been conducted to simulate the crowd movements, while relatively few of them are specifically focused on multihazard situations. In this paper, we propose a novel crowd simulation method by modeling the generation and contagion of panic emotion under multihazard circumstances. In order to depict the effect from hazards and other agents to crowd movement, we first classify hazards into different types (transient and persistent, concurrent and nonconcurrent, and static and dynamic) based on their inherent characteristics. Second, we introduce the concept of perilous field for each hazard and further transform the critical level of the field to its invoked-panic emotion. After that, we propose an emotional contagion model to simulate the evolving process of panic emotion caused by multiple hazards. Finally, we introduce an emotional reciprocal velocity obstacles (RVOs) model to simulate the crowd behaviors by augmenting the traditional RVO model with emotional contagion, which for the first time combines the emotional impact and local avoidance together. Our experimental results demonstrate that the overall approach is robust, can better generate realistic crowds and the panic emotion dynamics in a crowd. Furthermore, it is recommended that our method can be applied to various complex multihazard environments.
Mingliang Xu 0001, Xiaozheng Xie, Pei Lv, Jianwei Niu 0002, Chaochao Li, Ruijie Zhu 0001, Zhigang Deng 0001, Bing Zhou 0003
IEEE Trans. Syst. Man Cybern. Syst.3
2020 Attention Scaling for Crowd Counting
abstract
Convolutional Neural Network (CNN) based methods generally take crowd counting as a regression task by outputting crowd densities. They learn the mapping between image contents and crowd density distributions. Though having achieved promising results, these data-driven counting networks are prone to overestimate or underestimate people counts of regions with different density patterns, which degrades the whole count accuracy. To overcome this problem, we propose an approach to alleviate the counting performance differences in different regions. Specifically, our approach consists of two networks named Density Attention Network (DANet) and Attention Scaling Network (ASNet). DANet provides ASNet with attention masks related to regions of different density levels. ASNet first generates density maps and scaling factors and then multiplies them by attention masks to output separate attention-based density maps. These density maps are summed to give the final density map. The attention scaling factors help attenuate the estimation errors in different regions. Furthermore, we present a novel Adaptive Pyramid Loss (APLoss) to hierarchically calculate the estimation losses of sub-regions, which alleviates the training bias. Extensive experiments on four challenging datasets (ShanghaiTech Part A, UCF_CC_50, UCF-QNRF, and WorldExpo'10) demonstrate the superiority of the proposed approach.
Xiaoheng Jiang, Li Zhang 0072, Mingliang Xu 0001, Tianzhu Zhang 0001, Pei Lv, Bing Zhou 0003, Xin Yang 0011, Yanwei Pang
CVPR5
2020 Semi-Dynamic Hypergraph Neural Network for 3D Pose Estimation
abstract
This paper proposes a novel Semi-Dynamic Hypergraph Neural Network (SD-HNN) to estimate 3D human pose from a single image. SD-HNN adopts hypergraph to represent the human body to effectively exploit the kinematic constrains among adjacent and non-adjacent joints. Specifically, a pose hypergraph in SD-HNN has two components. One is a static hypergraph constructed according to the conventional tree body structure. The other is the semi-dynamic hypergraph representing the dynamic kinematic constrains among different joints. These two hypergraphs are combined together to be trained in an end-to-end fashion. Unlike traditional Graph Convolutional Networks (GCNs) that are based on a fixed tree structure, the SD-HNN can deal with ambiguity in human pose estimation. Experimental results demonstrate that the proposed method achieves state-of-the-art performance both on the Human3.6M and MPI-INF-3DHP datasets.
Pei Lv, Junjin Cheng, Wanqing Li 0001, Bing Zhou 0003, Mingliang Xu 0001
IJCAI2
2020 MDSSD: multi-scale deconvolutional single shot detector for small objects
Lisha Cui, Pei Lv, Xiaoheng Jiang, Zhimin Gao, Bing Zhou 0003, Mingliang Xu 0001
Sci. China Inf. Sci.3
2020 Learning Multi-Level Density Maps for Crowd Counting
abstract
People in crowd scenes often exhibit the characteristic of imbalanced distribution. On the one hand, people size varies largely due to the camera perspective. People far away from the camera look smaller and are likely to occlude each other, whereas people near to the camera look larger and are relatively sparse. On the other hand, the number of people also varies greatly in the same or different scenes. This article aims to develop a novel model that can accurately estimate the crowd count from a given scene with imbalanced people distribution. To this end, we have proposed an effective multi-level convolutional neural network (MLCNN) architecture that first adaptively learns multi-level density maps and then fuses them to predict the final output. Density map of each level focuses on dealing with people of certain sizes. As a result, the fusion of multi-level density maps is able to tackle the large variation in people size. In addition, we introduce a new loss function named balanced loss (BL) to impose relatively BL feedback during training, which helps further improve the performance of the proposed network. Furthermore, we introduce a new data set including 1111 images with a total of 49 061 head annotations. MLCNN is easy to train with only one end-to-end training stage. Experimental results demonstrate that our MLCNN achieves state-of-the-art performance. In particular, our MLCNN reaches a mean absolute error (MAE) of 242.4 on the UCF_CC_50 data set, which is 37.2 lower than the second-best result.
Xiaoheng Jiang, Li Zhang 0072, Pei Lv, Yibo Guo, Ruijie Zhu 0001, Yanwei Pang, Xi Li 0001, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2019 Crowd queuing simulation with an improved emotional contagion model
Junxiao Xue, Pei Lv, Mingliang Xu 0001
Sci. China Inf. Sci.3
2019 Personalized training through Kinect-based games for physical education
Mingliang Xu 0001, Yafang Zhai, Yibo Guo, Pei Lv, Meng Wang 0001, Bing Zhou 0003
J. Vis. Commun. Image Represent.4
2019 Depth Information Guided Crowd Counting for complex crowd scenes
Mingliang Xu 0001, Zhaoyang Ge, Xiaoheng Jiang, Gaoge Cui, Pei Lv, Bing Zhou 0003, Changsheng Xu
Pattern Recognit. Lett.5
2019 D-STC: Deep learning with spatio-temporal constraints for train drivers detection from videos
Mingliang Xu 0001, Fang Hao, Pei Lv, Lisha Cui, Shuo Zhang 0014, Bing Zhou 0003
Pattern Recognit. Lett.3
2019 Crowd Behavior Evolution With Emotional Contagion in Political Rallies
abstract
In this paper, we present a novel crowd behavior evolution method with emotional contagion in political rallies. We first analyze the most representative political rally scenes in detail and model them into two kinds of abstract scenario. Furthermore, the “extroversion” and “empathy” factors from the OCEAN model are chosen to describe the most important individual personalities in such scenarios. Based on this, an improved emotional contagion model is proposed by combining the Susceptible-Infected-Recovered model and individual personality under different political viewpoints. Finally, the crowd in a political rally is driven to move according to the new potential moving direction generated by emotional contagion and the original direction of the individual together. The experiments show that our method can intuitively demonstrate the emotional changes of those individuals with different political perspectives and reasonably simulate the crowd movement under the political rally scenes.
Pei Lv, Zhujin Zhang, Chaochao Li, Yibo Guo, Bing Zhou 0003, Mingliang Xu 0001
IEEE Trans. Comput. Soc. Syst.1
2019 Stylized Aesthetic QR Code
abstract
With the continued proliferation of smart mobile devices, the Quick Response (QR) code has become one of the most-used types of two-dimensional code in the world. Aiming at beautifying the visual-unpleasant appearance of QR codes, existing works have developed a series of techniques. However, these works still leave much to be desired, such as personalization, artistry, and robustness. To address these issues, in this paper, we propose a novel type of aesthetic QR codes, Stylized aEsthEtic (SEE) QR code , and a three-stage approach to automatically produce such robust style-oriented codes. Specifically, in the first stage, we propose a method to generate an optimized baseline aesthetic QR code, which reduces the visual contrast between the noise-like black/white modules and the blended image. In the second stage, to obtain an art style QR code, we tailor an appropriate neural style transformation network to endow the baseline aesthetic QR code with artistic elements. In the third stage, we design a module-based robustness-optimization mechanism to ensure the performance robust by balancing two competing terms: visual quality and readability. Extensive experiments demonstrate that the SEE QR code has high quality in terms of both visual appearance and robustness and also offers a greater variety of personalized choices to users.
Mingliang Xu 0001, Hao Su 0001, Xi Li 0001, Jing Liao 0001, Jianwei Niu 0002, Pei Lv, Bing Zhou 0003
IEEE Trans. Multim.7
2018 USAR: An Interactive User-specific Aesthetic Ranking Framework for Images
abstract
When assessing whether an image is of high or low quality, it is indispensable to take personal preference into account. Existing aesthetic models lay emphasis on hand-crafted features or deep features commonly shared by high quality images, but with limited or no consideration for personal preference and user interaction. To that end, we propose a novel and user-friendly aesthetic ranking framework via powerful deep neural network and a small amount of user interaction, which can automatically estimate and rank the aesthetic characteristics of images in accordance with users' preference. Our framework takes as input a series of photos that users prefer, and produces as output a reliable, user-specific aesthetic ranking model matching with users' preference. Considering the subjectivity of personal preference and the uncertainty of user's single selection, a unique and exclusive dataset will be constructed interactively to describe the preference of one individual by retrieving the most similar images with regard to those specified by users. Based on this unique user-specific dataset and sufficient well-designed aesthetic attributes, a customized aesthetic distribution model can be learned, which concatenates both personalized preference and aesthetic rules. We conduct extensive experiments and user studies on two large-scale public datasets, and demonstrate that our framework outperforms those work based on conventional aesthetic assessment or ranking model.
Pei Lv, Meng Wang 0001, Yongbo Xu, Junyi Sun, Shi-Mei Su, Bing Zhou 0003, Mingliang Xu 0001
ACM Multimedia1
2018 An Efficient Method of Crowd Aggregation Computation in Public Areas
abstract
The crowd stampede and terrorist attacks in public areas have now become more serious and dangerous threats due to the rapid increase in the population and scale of cities. Therefore, the analysis of crowd aggregation behavior has been a new research focus in the field of intelligent video surveillance. However, such public area scenes not only contain moving crowd but also contain other types of objects. The sizes of these objects are usually small, which make their appearances quite similar. Moreover, the individuals in a crowd move randomly and often occlude each other. All the above factors make the analysis of crowd aggregation very difficult. In this paper, the authors attempt to solve this problem in three aspects. First, a novel global feature is used to represent the moving crowd. This feature can well describe the spatial and the temporal motion information of points-of-interest. Second, a strategy is adopted to cluster the feature points first and then calculate the collectiveness. This makes the collectiveness computation of individual groups more consistent and effective. Finally, more comprehensive collective crowd descriptors are proposed to provide a detailed description of the crowd status. Based on the proposed descriptor, the authors realize the evolution analysis of the group movement and the crowd abnormal detection. The experiment results show that the proposed method is able to efficiently compute the crowd collectiveness in various public areas and provide a reliable reference for the public safety management.
Mingliang Xu 0001, Chunxu Li, Pei Lv, Nie Lin, Rui Hou 0001, Bing Zhou 0003
IEEE Trans. Circuits Syst. Video Technol.3
2017 Learning-Based Shadow Recognition and Removal From Monochromatic Natural Images
abstract
This paper addresses the problem of recognizing and removing shadows from monochromatic natural images from a learning-based perspective. Without chromatic information, shadow recognition and removal are extremely challenging in this paper, mainly due to the missing of invariant color cues. Natural scenes make this problem even harder due to the complex illumination condition and ambiguity from many near-black objects. In this paper, a learning-based shadow recognition and removal scheme is proposed to tackle the challenges above-mentioned. First, we propose to use both shadow-variant and invariant cues from illumination, texture, and odd order derivative characteristics to recognize shadows. Such features are used to train a classifier via boosting a decision tree and integrated into a conditional random field, which can enforce local consistency over pixel labels. Second, a Gaussian model is introduced to remove the recognized shadows from monochromatic natural scenes. The proposed scheme is evaluated using both qualitative and quantitative results based on a novel database of hand-labeled shadows, with comparisons to the existing state-of-the-art schemes. We show that the shadowed areas of a monochromatic image can be accurately identified using the proposed scheme, and high-quality shadow-free images can be precisely recovered after shadow removal.
Mingliang Xu 0001, Jiejie Zhu, Pei Lv, Bing Zhou 0003, Marshall F. Tappen, Rongrong Ji
IEEE Trans. Image Process.3
2016 Data-driven humanlike reaching behaviors synthesis
Pei Lv, Mingliang Xu 0001, Bailin Yang, Bing Zhou 0003
Neurocomputing1
2016 Medical image denoising by parallel non-local means
Mingliang Xu 0001, Pei Lv, Fang Hao, Hongling Zhao, Bing Zhou 0003, Yusong Lin, Li-Wei Zhou
Neurocomputing2
2016 Robust Lane Detection using Two-stage Feature Extraction with Curve Fitting
Jianwei Niu 0002, Jie Lu 0003, Mingliang Xu 0001, Pei Lv, Xiaoke Zhao
Pattern Recognit.4
2015 A Suggestive Interface for Sketch-based Character Posing
abstract
We present a user-friendly suggestive interface for sketch-based character posing. Our interface provides suggestive information on the sketching canvas in succession by combining image retrieval technique with 3D character posing, while the user is drawing. The system highlights the canvas region where the user should draw on and constrains the user's sketches in a reasonable solution space. This is based on an efficient image descriptor, which is used to measure the distance between the user's sketch and 2D views of 3D poses. In order to achieve faster query response, local sensitive hashing is involved in our system. In addition, sampling-based optimization algorithm is adopted to synthesize and optimize the retrieved 3D pose to match the user's sketches the best. Experiments show that our interface can provide smooth suggestive information to improve the reality of sketching poses and shorten the time required for 3D posing.
Pei Lv, Pengjie Wang 0001, Weiwei Xu 0003, Jinxiang Chai
Comput. Graph. Forum1
2015 miSFM: On combination of Mutual Information and Social Force Model towards simulating crowd evacuation
Yunpeng Wu, Pei Lv, Hao Jiang 0013, Mingxuan Luo, Yangdong Ye
Neurocomputing3
2012 Virtual Network Marathon with immersion, scientificalness, competitiveness, adaptability and learning
Mingliang Xu 0001, Lizhen Han, Yong Liu 0007, Pei Lv, Gaoqi He
Comput. Graph.5
2011 Biomechanics-based reaching optimization
Pei Lv, Huansen Li
Vis. Comput.1
2010 L4RW: Laziness-based Realistic Real-time Responsive Rebalance in Walking
abstract
Abstract We present a novel L4RW (Laziness‐based Realistic Real‐time Responsive Rebalance in Walking) technique to synthesize 4RW animations under unexpected external perturbations with minimal locomotion effort. We first devise a lazy dynamic rebalance model, which specifies the dynamic balance conditions, defines the rebalance effort, and selects the suitable rebalance strategy automatically using the laziness law after an unexpected perturbation. Based on the model, L4RW searches over a motion capture (mocap) database for an appropriate motion segment to follow, and the transition‐to motions is generated by interpolating the active response dynamic motion. A support vector machine (SVM) based training, classification, and predication algorithm is applied to reduce the search space, and it is trained offline only once. Our algorithm classifies the mocap database into many rebalance strategy‐specified subsets and then online predicts responsive motions in the subset according to the selected strategy. The rebalance effort, the ‘extrapolated center of mass’ (XCoM) and environment constraints are selected as feature attributes for the SVM feature vector. Furthermore, the subset's segments are sorted through the rebalance effort, then our algorithm searches for an acceptable segment starting from the least‐effort segment. Compared with previous methods, our search increases speed by over two orders of magnitude, and our algorithm creates more realistic and smooth 4RW animation.
Huansen Li, Pei Lv, Wenzhi Chen, Gengdai Liu
Comput. Graph. Forum3
2010 Moving-Target Pursuit Algorithm Using Improved Tracking Strategy
abstract
Pursuing a moving target in modern computer games presents several challenges to situated agents, including real-time response, large-scale search space, severely limited computation resources, incomplete environmental knowledge, adversarial escaping strategy, and outsmarting the opponent. In this paper, we propose a novel tracking automatic optimization moving-target pursuit (TAO-MTP) algorithm employing improved tracking strategy to effectively address all challenges above for the problem involving single hunter and single prey. TAO-MTP uses a queue to store prey's trajectory, and simultaneously runs real-time adaptive A* (RTAA*) repeatedly to approach the optimal position updated periodically in the trajectory within limited steps, which makes the overall pursuit cost smallest. In the process, the hunter speculatively moves to any position explored in the trajectory, not necessarily the optimal position, to speed up convergence, and then directly moves along the trajectory to pursue the prey. Moreover, automatic optimization methods, such as reducing trajectory storage and optimizing pursuit path, are used to further enhance its performance. As long as the hunter's moving speed is faster than that of the prey, and its sense scope is large enough, it will eventually capture the prey. Experiments using commercial game maps show that TAO-MTP is independent of adversarial escaping strategy, and outperforms all the classic and state-of-the-art moving-target pursuit algorithms such as extended moving-target search (eMTS), path refinement moving-target search (PR MTS), moving-target adaptive A* (MTAA*), and generalized adaptive A* (GAA*).
Hongxing Lu, Yangdong Ye, Pei Lv, Abdennour El Rhalibi
IEEE Trans. Comput. Intell. AI Games5