Tianlu Mao

dblp:60/1810 · DBLP profile ↗
← Back
45ranked-venue papers
7as first author
23since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 9 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2027 CGARF: a causality-guided framework for reliable automated program repair
Le Yuan, Shaohua Liu 0002, Yancheng Yao, Tianlu Mao
Empir. Softw. Eng.8
2026 Meta-enhanced code: leveraging structural and functional features for precise cross-modal code search
Le Yuan, Shaohua Liu 0002, Shangwei Zhu, Tianlu Mao, Songbo Shao
Empir. Softw. Eng.6
2026 DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo
abstract
Recently, patch deformation-based methods have demonstrated significant effectiveness in multi-view stereo due to their incorporation of deformable and expandable perception for reconstructing textureless areas. However, these methods generally focus on identifying reliable pixel correlations to mitigate matching ambiguity of patch deformation, while neglecting the deformation instability caused by edge-skipping and visibility occlusions, which may cause potential estimation deviations. To address these issues, we propose DVP-MVS++, an innovative approach that synergizes both depth-normal-edge aligned and harmonized cross-view priors for robust and visibility-aware patch deformation. Specifically, to avoid edge-skipping, we first apply DepthPro, Metric3Dv2 and Roberts operator to generate coarse depth maps, normal maps and edge maps, respectively. These maps are then aligned via an erosion-dilation strategy to produce fine-grained homogeneous boundaries for facilitating robust patch deformation. Moreover, we reformulate view selection weights as visibility maps, and then implement both an enhanced cross-view depth reprojection and an area-maximization strategy to help reliably restore visible areas and effectively balance deformed patch. Additionally, we obtain geometry consistency by adopting both aggregated normals via view selection and projection depth differences via epipolar lines, and then employ SHIQ for highlight correction to facilitate highlight perception capacity, thus improving reconstruction quality during propagation and refinement stage. Evaluations on ETH3D, Tanks & Temples and Strecha datasets exhibit the state-of-the-art performance and robust generalization capability of our proposed method.
Zhenlong Yuan, Chengxuan Qian, Jianing Chen 0007, Yinda Chen, Kehua Chen, Tianlu Mao, Zhaoxin Li, Hao Jiang 0013
IEEE Trans. Circuits Syst. Video Technol.8
2025 Dual-Level Precision Edges Guided Multi-View Stereo with Accurate Planarization
abstract
The reconstruction of low-textured areas is a prominent research focus in multi-view stereo (MVS). In recent years, traditional MVS methods have performed exceptionally well in reconstructing low-textured areas by constructing plane models. However, these methods often encounter issues such as crossing object boundaries and limited perception ranges, which undermine the robustness of plane model construction. Building on previous work (APD-MVS), we propose the DPE-MVS method. By introducing dual-level precision edge information, including fine and coarse edges, we enhance the robustness of plane model construction, thereby improving reconstruction accuracy in low-textured areas. Furthermore, by leveraging edge information, we refine the sampling strategy in conventional PatchMatch MVS and propose an adaptive patch size adjustment approach to optimize matching cost calculation in both stochastic and low-textured areas. This additional use of edge information allows for more precise and robust matching. Our method achieves state-of-the-art performance on the ETH3D and Tanks & Temples benchmarks. Notably, our method outperforms all published methods on the ETH3D benchmark.
Kehua Chen, Zhenlong Yuan, Tianlu Mao
AAAI3
2025 DVP-MVS: Synergize Depth-Edge and Visibility Prior for Multi-View Stereo
abstract
Patch deformation-based methods have recently exhibited substantial effectiveness in multi-view stereo, due to the incorporation of deformable and expandable perception to reconstruct textureless areas. However, such approaches typically focus on exploring correlative reliable pixels to alleviate match ambiguity during patch deformation, but ignore the deformation instability caused by mistaken edge-skipping and visibility occlusion, leading to potential estimation deviation. To remedy the above issues, we propose DVP-MVS, which innovatively synergizes depth-edge aligned and cross-view prior for robust and visibility-aware patch deformation. Specifically, to avoid unexpected edge-skipping, we first utilize Depth Anything V2 followed by the Roberts operator to initialize coarse depth and edge maps respectively, both of which are further aligned through an erosion-dilation strategy to generate fine-grained homogeneous boundaries for guiding patch deformation. In addition, we reform view selection weights as visibility maps and restore visible areas by cross-view depth reprojection, then regard them as cross-view prior to facilitate visibility-aware patch deformation. Finally, we improve propagation and refinement with multi-view geometry consistency by introducing aggregated visible hemispherical normals based on view selection and local projection depth differences based on epipolar lines, respectively. Extensive evaluations on ETH3D and Tanks & Temples benchmarks demonstrate that our method can achieve state-of-the-art performance with excellent robustness and generalization.
Zhenlong Yuan, Jinguo Luo, Fei Shen 0004, Zhaoxin Li, Tianlu Mao
AAAI6
2025 MSP-MVS: Multi-Granularity Segmentation Prior Guided Multi-View Stereo
abstract
Recently, patch deformation-based methods have demonstrated significant strength in multi-view stereo by adaptively expanding the reception field of patches to help reconstruct textureless areas. However, such methods mainly concentrate on searching for pixels without matching ambiguity (i.e., reliable pixels) when constructing deformed patches, while neglecting the deformation instability caused by unexpected edge-skipping, resulting in potential matching distortions. Addressing this, we propose MSP-MVS, a method introducing multi-granularity segmentation prior for edge-confined patch deformation. Specifically, to avoid unexpected edge-skipping, we first aggregate and further refine multi-granularity depth edges gained from Semantic-SAM as prior to guide patch deformation within depth-continuous (i.e., homogeneous) areas. Moreover, to address attention imbalance caused by edge-confined patch deformation, we implement adaptive equidistribution and disassemble-clustering of correlative reliable pixels (i.e., anchors), thereby promoting attention-consistent patch deformation. Finally, to prevent deformed patches from falling into local-minimum matching costs caused by the fixed sampling pattern, we introduce disparity-sampling synergistic 3D optimization to help identify global-minimum matching costs. Evaluations on ETH3D and Tanks & Temples benchmarks prove our method obtains state-of-the-art performance with remarkable generalization.
Zhenlong Yuan, Fei Shen 0004, Zhaoxin Li, Jinguo Luo, Tianlu Mao
AAAI6
2025 HAIF-GS: Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic Scene
abstract
Reconstructing dynamic 3D scenes from monocular videos remains a fundamental challenge in 3D vision. While 3D Gaussian Splatting (3DGS) achieves real-time rendering in static settings, extending it to dynamic scenes is challenging due to the difficulty of learning structured and temporally consistent motion representations. This challenge often manifests as three limitations in existing methods: redundant Gaussian updates, insufficient motion supervision, and weak modeling of complex non-rigid deformations. These issues collectively hinder coherent and efficient dynamic reconstruction. To address these limitations, we propose HAIF-GS, a unified framework that enables structured and consistent dynamic modeling through sparse anchor-driven deformation. It first identifies motion-relevant regions via an Anchor Filter to suppress redundant updates in static areas. A self-supervised Induced Flow-Guided Deformation module induces anchor motion using multi-frame feature aggregation, eliminating the need for explicit flow labels. To further handle fine-grained deformations, a Hierarchical Anchor Propagation mechanism increases anchor resolution based on motion complexity and propagates multi-level transformations. Extensive experiments on synthetic and real-world benchmarks validate that HAIF-GS significantly outperforms prior dynamic 3DGS methods in rendering quality, temporal coherence, and reconstruction efficiency.
Jianing Chen 0007, Yujun Cai, Hao Jiang 0013, Chengxuan Qian, Juyuan Kang, Shuqin Gao, Honglong Zhao, Tianlu Mao
NeurIPS9
2025 StyleGarNet: A Real-Time Garment Animation Generation Method With Diverse Styles
abstract
ABSTRACT Dressing animations have broad applications in film and animation, but current methods require retraining networks for new styles, which is resource‐intensive. While 2D pattern parameters are used for static 3D clothing modeling and editing, applying them to control style variations in dynamic clothing deformation is challenging due to the uncertainty introduced by multiple parameters. Ensuring consistent and stable style features across time‐varying deformation sequences is difficult. Therefore, we present StyleGarNet, a novel approach for style‐parameter‐controlled dressing animation generation and editing. We employ a conditional variational autoencoder architecture to build our network, with style parameters serving as constraint information. This allows us to learn the probabilistic distribution model of deformations under style constraints, crucial for enhancing the robust representation of styles during the deformation process. Simultaneously, considering the temporal nature of motion, we introduce a transformer layer to capture the temporal dependencies of both motion and clothing deformation, thereby enhancing the stability of clothing deformation. Ultimately, our approach enables flexible manipulation of dressing animation generation through inputting style and motion features. Evaluation results demonstrate the efficacy of our approach and show that StyleGarNet outperforms existing methods in terms of prediction speed, accuracy, and stability of deformation sequences.
Min Shi 0005, Xinru Zhuo, Yueyue Sun, Lin Gao 0004, Tianlu Mao, Dengming Zhu
Comput. Animat. Virtual Worlds6
2025 NeRF-based Polarimetric Multi-view Stereo
Jiakai Cao, Zhenlong Yuan, Tianlu Mao, Zhaoxin Li
Pattern Recognit.3
2025 Learning Multi-View Stereo With Geometry-Aware Prior
abstract
Multi-View Stereo (MVS) reconstructs detailed 3D structures from multi-view images by establishing spatial correspondences. While learning-based methods have significantly advanced the MVS task, challenges such as ambiguous matching caused by textureless surfaces and lighting variations persist. To address these issues, we propose GAP-MVSNet, a framework that leverages surface normals from a monocular normal foundation model as priors to enhance the geometric awareness of reconstruction targets. In this work, surface normal priors are seamlessly integrated into the MVS pipeline to improve depth prediction robustness and accuracy. Specifically, we introduce a structure-aware feature pyramid network that incorporates surface normal information and utilizes uncertainty-aware feature resampling to extract robust image features. Additionally, we present the spatial geometry enhanced regularization that combines sampled depth hypotheses with surface normals to generate a spatial geometric prior, guiding the cost regularization process and enforcing strong spatial coherence, particularly in textureless regions. Furthermore, we design a local consistency depth refinement module that utilizes surface normals to establish depth relationships as a local geometric prior, thereby refining classification-based depth predictions and aligning them with ground truth depth. Extensive experiments on the DTU and Tanks & Temples datasets demonstrate that our method achieves state-of-the-art performance.
Kehua Chen, Zhenlong Yuan, Haihong Xiao, Tianlu Mao
IEEE Trans. Circuits Syst. Video Technol.4
2024 GSMNet: Towards Long-Term Trajectory Prediction by Integrating Multi-scale Information
Shaohua Liu 0002, Yisu Wang, Yinglong Zhu, Pengfei Yao, Tianlu Mao
ACCV (2)5
2024 SpectrumNet: Spectrum-Based Trajectory Encode Neural Network for Pedestrian Trajectory Prediction
abstract
Extracting motion pattern implied in the history trajectory is important for the pedestrian trajectory prediction task. The motion pattern determines how a pedestrian moves, including but not limited to reaction of interaction, tendency of speed and direction change. Although the motion pattern is a comprehensive concept and can’t be described concretely, it is clear that it contains both long-term and short-term factors. Inspired by this, we introduce SpectrumNet which enables more effective encoding of historical motion patterns for trajectory prediction. Different from existing methods, which consider the history trajectory as a time sequence of position, SpectrumNet represents it in the frequency space by applying Fourier Transform (FT) to decompose the historical information on different time scales. SpectrumNet consists of two sub-networks, the Multi-Frequency Combination (MFC) encoder, which models the historical information by combining multiple frequency feature in the spectrum; and the Frequency Interaction (FI) encoder, which captures the interaction between pedestrians in the frequency domain. To validate the effect of SpectrumNet, we build a CVAE-based prediction system to predict stochastic future trajectory. Experiments conducted on ETH-UCY dataset show that our prediction system with SpectrumNet out-performs the previous state-of-the-art model and achieves a new record on ADE metric.
Shaohua Liu 0002, Yinglong Zhu, Pengfei Yao, Tianlu Mao
ICASSP4
2024 Modeling Scene-Agent Interaction for Pedestrian Trajectory Prediction
abstract
Modeling scene-agent interaction is required in reliable and feasible for pedestrian trajectory prediction. Most existing methods utilize segmentation result of scene images as semantic information and concatenate it with trajectory features to model scene-agent interaction by concatenating semantic features with trajectory features. Since the image and trajectory are heterogeneous data, directly concatenating features proves inadequate for modeling scene-agent interaction. We notice that agent trajectory is the result of the interaction between the agent movement decision and the scene. They could be regarded as examples that demonstrate the influence between the scene and pedestrians’ motion in different locations. Compared with semantic image, historical trajectories of neighboring agents do not cover the whole scene, but they are more direct and efficient information for scene-agent interaction modeling. We propose a novel scene-agent interaction modeling method which uses Fourier Transform to convert the temporal historical trajectories into the frequency domain, utilizes transformer to encode the scene-agent interaction feature(SAIF), and designs a spectral position embedding to represents the scene influence on the agents. We combine the spatio-temporal features of trajectories and inter-agent interaction features with SAIF for generating predicted trajectories. Experiments demonstrate that our method achieves the state-of-the-art on both ETH-UCY and SDD.
Pengfei Yao, Yinglong Zhu, Tianlu Mao
ICME3
2024 TrajCLIP: Pedestrian trajectory prediction method using contrastive learning and idempotent networks
abstract
The distribution of pedestrian trajectories is highly complex and influenced by the scene, nearby pedestrians, and subjective intentions. This complexity presents challenges for modeling and generalizing trajectory prediction. Previous methods modeled the feature space of future trajectories based on the high-dimensional feature space of historical trajectories, but this approach is suboptimal because it overlooks the similarity between historical and future trajectories. Our proposed method, TrajCLIP, utilizes contrastive learning and idempotent generative networks to address this issue. By pairing historical and future trajectories and applying contrastive learning on the encoded feature space, we enforce same-space consistency constraints. To manage complex distributions, we use idempotent loss and tightness loss to control over-expansion in the latent space. Additionally, we have developed a trajectory interpolation algorithm and synthetic trajectory data to enhance model capacity and improve generalization. Experimental results on public datasets demonstrate that TrajCLIP achieves state-of-the-art performance and excels in scene-to-scene transfer, few-shot transfer, and online learning tasks.
Pengfei Yao, Yinglong Zhu, Huikun Bi, Tianlu Mao
NeurIPS4
2024 VTSIM: Attention-Based Recurrent Neural Network for Intersection Vehicle Trajectory Simulation
abstract
ABSTRACT Simulating vehicle trajectories at intersections is one of the challenging tasks in traffic simulation. Existing methods are often ineffective due to the complexity and diversity of lane topologies at intersections, as well as the numerous interactions affecting vehicle motion. To address this issue, we propose a deep learning based vehicle trajectory simulation method. First, we employ a vectorized representation to uniformly extract features from traffic elements such as pedestrians, vehicles, and lanes. By fusing all factors that influence vehicle motion, this representation makes our method suitable for a variety of intersections. Second, we propose a deep learning model, which has an attention network to dynamically extract features from the surrounding environment of the vehicles. To address the issue of vehicles continuously entering and exiting the simulation scene, we employ an asynchronous recurrent neural network for the extraction of temporal features. Comparative evaluations against existing rule‐based and deep learning‐based methods demonstrate our model's superior simulation accuracy. Furthermore, experimental validation on public datasets demonstrates that our model can simulate vehicle trajectories among the urban intersections with different topologies including those not present in the training dataset.
Tianlu Mao
Comput. Animat. Virtual Worlds2
2023 CVTP3D: Cross-view Trajectory Prediction Using Shared 3D Queries for Autonomous Driving
abstract
Trajectory prediction with uncertainty is a critical and challenging task for autonomous driving. Nowadays, we can easily access sensor data represented in multiple views. However, cross-view consistency has not been evaluated by the existing models, which might lead to divergences between the multimodal predictions from different views. It is not practical and effective when the network does not comprehend the 3D scene, which could cause the downstream module in a dilemma. Instead, we predicts multimodal trajectories while maintaining cross-view consistency. We presented a cross-view trajectory prediction method using shared 3D Queries (XVTP3D). We employ a set of 3D queries shared across views to generate multi-goals that are cross-view consistent. We also proposed a random mask method and coarse-to-fine cross-attention to capture robust cross-view features. As far as we know, this is the first work that introduces the outstanding top-down paradigm in BEV detection field to a trajectory prediction problem. The results of experiments on two publicly available datasets show that XVTP3D achieved state-of-the-art performance with consistent cross-view predictions.
Zijian Song 0002, Huikun Bi, Ruisi Zhang, Tianlu Mao
IJCAI4
2023 Motion-Inspired Real-Time Garment Synthesis with Temporal-Consistency
Yukun Wei, Min Shi 0005, Wenke Feng, Dengming Zhu, Tianlu Mao
J. Comput. Sci. Technol.5
2023 Data-driven based double-layer bicycle simulation model
abstract
Abstract Bicycle motion simulation is fundamental to urban transportation planning, virtual reality and other areas. This article proposes a data‐driven based double‐layer bicycle simulation model to consider the cyclist's decision‐making process and the bicycle's kinematic structure. This proposed model consists of two layers, the decision‐making layer and the motion layer. First, the decision‐making layer using machine learning algorithms models the decision‐making process as a regression problem to output the cyclist's decision. Then, the motion layer applies a bicycle kinematics model to output bicycle motion under physical constraints. In addition, a solution to calculate bicycles' dynamic information is proposed for the data‐driven method. Quantitative and qualitative experiments have been conducted, and results show that the double‐layer model and the parameter calculation solution can generate realistic bicycle motion simulations.
Tianlu Mao, Zhong Fang, Qinyuan Yan, Ruoyu Meng, Shaohua Liu 0002
Comput. Animat. Virtual Worlds1
2022 A Fusion Crowd Simulation Method: Integrating Data with Dynamics, Personality with Common
abstract
Abstract This paper proposes a novel crowd simulation method which integrates not only modelling ideas but also advantages from both data‐driven methods and crowd dynamics methods. To seamlessly integrate these two different modelling ideas, first, a fusion crowd motion model is developed. In this model the motion of crowd are driven dynamically by different forces. Part of the forces are modeled under a universal interaction mechanism, which describe the common parts of crowd dynamics. Others are modeled by examples from real data, which describe the personality parts of the agent motion. Second, a construction method for example dataset is proposed to support the fusion model. In the dataset, crowd trajectories captured in the real world are decomposed and re‐described under the structure of the fusion model. Thus, personality parts hidden in the real data could be locked and extracted, making the data understandable and migratable for our fusion model. A comprehensive crowd motion generation workflow using the fusion model and example dataset is also proposed. Quantitative and qualitative experiments and user studies are conducted. Results show that the proposed fusion crowd simulation method can generate crowd motion with the great motion fidelity, which not only match the macro characteristics of real data, but also has lots of micro personality showing the diversity of crowd motion.
Tianlu Mao, Ruoyu Meng, Qinyuan Yan, Shaohua Liu 0002
Comput. Graph. Forum1
2022 An efficient Spatial-Temporal model based on gated linear units for trajectory prediction
Shaohua Liu 0002, Yisu Wang, Jingkai Sun, Tianlu Mao
Neurocomputing4
2022 DeepORCA: Realistic crowd simulation for varying scenes
abstract
Abstract Crowd simulation is a challenging problem, aiming to generate realistic pedestrians motions in virtual environment. Nowadays, ORCA is a widely used simulation algorithm in practice because of its stable and efficient performance. However, this algorithm cannot regenerate continuity and diversity of pedestrian motions in real data, leading to defects in motion fidelity. Otherwise, trajectory prediction methods based on deep learning have progressed in real pedestrians movement patterns mining. However, they are rarely applied in simulation due to the lack of ability to avoid collision and adapt to manufactured scenarios. Our work proposes a simulation method DeepORCA that integrates ORCA with a CVAE‐based velocity probability generator, which can model motion continuity, variable intentions, and scene semantics. Moreover, DeepORCA converts the velocity optimization into quadratic programming, which accelerates the calculation while maintaining the collision‐avoidance ability of ORCA. In the experiments of real and artificial scenes, our method produces more realistic crowd simulation results than ORCA quantitatively and qualitatively, while keeps the computational efficiency at the same order of magnitude.
Yaqiang Li, Tianlu Mao, Ruoyu Meng, Qinyuan Yan
Comput. Animat. Virtual Worlds2
2022 R-CTM: A data-driven macroscopic simulation model for heterogeneous traffic
abstract
Abstract There is a well‐known trade‐off between computational efficiency and computational accuracy in the field of traffic simulation. In this article, we propose a novel recurrent neural network based model with an integrated attention mechanism, called R‐CTM, to simulate heterogeneous traffic flow with multiple types of vehicles. It can effectively extract the traffic flow patterns of spatial and temporal changes from training traffic data, which can be real‐world traffic data or synthetic traffic data via microscopic simulation models. Through experiments and comparisons, we show that it can significantly outperform the state of the art methods in terms of simulation accuracy. Besides accuracy, we also demonstrate its scalability: its runtime consumption does not linearly increase with respect to the spatial extent.
Zhigang Deng 0001, Tianlu Mao
Comput. Animat. Virtual Worlds3
2021 Learning a shared deformation space for efficient design-preserving garment transfer
Min Shi 0005, Yukun Wei, Dengming Zhu, Tianlu Mao
Graph. Model.5
2020 How Can I See My Future? FvTraj: Using First-Person View for Pedestrian Trajectory Prediction
Huikun Bi, Ruisi Zhang, Tianlu Mao, Zhigang Deng 0001
ECCV (7)3
2020 Multi-Scale Residual Pyramid Attention Network for Monocular Depth Estimation
abstract
Monocular depth estimation is a challenging problem in computer vision and is crucial for understanding 3D scene geometry. Recently, deep convolutional neural networks (DCNNs) based methods have improved the estimation accuracy significantly. However, existing methods fail to consider complex textures and geometries in scenes, thereby resulting in loss of local details, distorted object boundaries, and blurry reconstruction. In this paper, we proposed an end-to-end multi-scale residual pyramid attention network (MRPAN) to mitigate these problems. First, we propose a multi-scale attention context aggregation (MACA) module, which consists of spatial attention module (SAM) and global attention module (GAM). By considering the position and scale correlation of pixels from spatial and global perspectives, the proposed module can adaptively learn the similarity between pixels so as to obtain more global context information of the image and recover complex structures in the scene. Then we proposed an improved residual refinement module (RRM) to further refine the scene structure, giving rise to deeper semantic information and retain more local details. Experimental results show that our method achieves more promising performance in object boundaries and local details compared with other state-of-the-art methods.
Jing Liu 0004, Xiaona Zhang, Zhaoxin Li, Tianlu Mao
ICPR4
2020 A Survey on Visual Traffic Simulation: Models, Evaluations, and Applications in Autonomous Driving
abstract
Abstract Virtualized traffic via various simulation models and real‐world traffic data are promising approaches to reconstruct detailed traffic flows. A variety of applications can benefit from the virtual traffic, including, but not limited to, video games, virtual reality, traffic engineering and autonomous driving. In this survey, we provide a comprehensive review on the state‐of‐the‐art techniques for traffic simulation and animation. We start with a discussion on three classes of traffic simulation models applied at different levels of detail. Then, we introduce various data‐driven animation techniques, including existing data collection methods, and the validation and evaluation of simulated traffic flows. Next, we discuss how traffic simulations can benefit the training and testing of autonomous vehicles. Finally, we discuss the current states of traffic simulation and animation and suggest future research directions.
Qianwen Chao, Huikun Bi, Weizi Li, Tianlu Mao, Ming C. Lin, Zhigang Deng 0001
Comput. Graph. Forum4
2020 A Deep Learning-Based Framework for Intersectional Traffic Simulation and Editing
abstract
Most of existing traffic simulation methods have been focused on simulating vehicles on freeways or city-scale urban networks. However, relatively little research has been done to simulate intersectional traffic to date despite its broad potential applications. In this paper, we propose a novel deep learning-based framework to simulate and edit intersectional traffic. Specifically, based on an in-house collected intersectional traffic dataset, we employ the combination of convolution network (CNN) and recurrent network (RNN) to learn the patterns of vehicle trajectories in intersectional traffic. Besides simulating novel intersectional traffic, our method can be used to edit existing intersectional traffic. Through many experiments as well as comparative user studies, we demonstrate that the results by our method are visually indistinguishable from ground truth, and our method can outperform existing methods.
Huikun Bi, Tianlu Mao, Zhigang Deng 0001
IEEE Trans. Vis. Comput. Graph.2
2019 Joint Prediction for Kinematic Trajectories in Vehicle-Pedestrian-Mixed Scenes
abstract
Trajectory prediction for objects is challenging and critical for various applications (e.g., autonomous driving, and anomaly detection). Most of the existing methods focus on homogeneous pedestrian trajectories prediction, where pedestrians are treated as particles without size. However, they fall short of handling crowded vehicle-pedestrian-mixed scenes directly since vehicles, limited with kinematics in reality, should be treated as rigid, non-particle objects ideally. In this paper, we tackle this problem using separate LSTMs for heterogeneous vehicles and pedestrians. Specifically, we use an oriented bounding box to represent each vehicle, calculated based on its position and orientation, to denote its kinematic trajectories. We then propose a framework called VP-LSTM to predict the kinematic trajectories of both vehicles and pedestrians simultaneously. In order to evaluate our model, a large dataset containing the trajectories of both vehicles and pedestrians in vehicle-pedestrian-mixed scenes is specially built. Through comparisons between our method with state-of-the-art approaches, we show the effectiveness and advantages of our method on kinematic trajectories prediction in vehicle-pedestrian-mixed scenes.
Huikun Bi, Zhong Fang, Tianlu Mao, Zhigang Deng 0001
ICCV3
2019 STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory Prediction
abstract
Human trajectory prediction is challenging and critical in various applications (e.g., autonomous vehicles and social robots). Because of the continuity and foresight of the pedestrian movements, the moving pedestrians in crowded spaces will consider both spatial and temporal interactions to avoid future collisions. However, most of the existing methods ignore the temporal correlations of interactions with other pedestrians involved in a scene. In this work, we propose a Spatial-Temporal Graph Attention network (STGAT), based on a sequence-to-sequence architecture to predict future trajectories of pedestrians. Besides the spatial interactions captured by the graph attention mechanism at each time-step, we adopt an extra LSTM to encode the temporal correlations of interactions. Through comparisons with state-of-the-art methods, our model achieves superior performance on two publicly available crowd datasets (ETH and UCY) and produces more "socially" plausible trajectories for pedestrians.
Yingfan Huang, Huikun Bi, Zhaoxin Li, Tianlu Mao
ICCV4
2019 Vehicle tracking by detection in UAV aerial video
Shaohua Liu 0002, Suqin Wang, Zhaoxin Li, Tianlu Mao
Sci. China Inf. Sci.6
2018 An emotion evolution based model for collective behavior simulation
abstract
Current crowd simulation progresses still fall short of simulating many real-world collective behaviors. Arguably, one of the main reasons is that some essential qualities of human beings such as emotion have not been effectively modeled and incorporated into crowd simulation algorithms. In this paper, we propose a novel computational model for emotion evolution and demonstrate its applications for crowd simulation. Specifically, our approach is designed to tackle three major issues in the emotion evolution process: (i) how to perceive and evaluate emotion when individuals face emergency or external events, (ii) how to evolve the emotion during induction, and (iii) how specific actions of individuals in a crowd are impacted by emotion. Through many experiments, we demonstrate that our method can effectively simulate emergent dynamic collective patterns observed in real-world crowd footages.
Hao Jiang 0013, Zhigang Deng 0001, Xiangjun He, Tianlu Mao
I3D5
2018 Behavioral Simulation of Passengers in a Waiting Hall
abstract
In this paper, we introduced a behavioral decision and execution method to simulate crowded passengers in a waiting hall. The method, as well as its simulation framework, is designed under the special purpose of passenger safety investigation. It supports the simulation of both regular crowded passenger behaviors and emergency passenger behavior. Situations under different time tables and density control measure could easily be conducted and simulated for safety purposes.
Shaohua Liu 0002, Xiyuan Song, Hao Jiang 0013, Min Shi 0005, Tianlu Mao
VR5
2016 Groupnect: Integrating group interaction into large display system
abstract
Large display systems have been successfully applied in virtual reality domains because they can provide full sense of immersion through large visual space and high display resolution. However, only a few users can interact with these systems by using pen-like or marker-based devices. In addition, user experience and application mode are constrained in many areas. In this paper, we propose a novel application framework called “Groupnect”, which gives users unique experience of group interaction in a large display system. By using optical tracking and 3D gesture recognition technologies, our approach can automatically recognize gesture-based control signals for 12 users simultaneously, and the backend system can trigger corresponding actions in real time. We conduct a user study and compare the results with a standard interaction mode. The results demonstrate that our approach greatly increases recorded objective activities and subjective efforts. Moreover, the physical and mental participation of users can be promoted by Groupnect. It indicates great potential to design novel applications in entertainment, education and training areas.
Hao Jiang 0013, Tianlu Mao
VR3
2015 An efficient lane model for complex traffic simulation
abstract
Abstract Traffic simulation heavily relies on lane model. This paper presents a novel method to model lanes based on the road axis under the Frenet frame. The road axis is generated from the geographic information system data after curve approximation, discretization, and compression. This lane model couples mileage information with three‐dimensional geometric information, so it offers an easy and fast position transformation from mileage to the Cartesian coordinate. It also keeps strictly consistent for mileage among neighboring lanes so that it facilitates lane‐change processing. Compared with existing methods that depict lanes as simple polylines or curves, the proposed lane model is more functional and more efficient, especially for complex traffic simulation with a large number of lane‐changes. Copyright © 2015 John Wiley & Sons, Ltd.
Tianlu Mao, Zhigang Deng 0001
Comput. Animat. Virtual Worlds1
2014 Modeling interactions in continuum traffic
abstract
It is a big challenge to generate the traffic scenarios with frequent lane changes in flow-based continuum traffic simulations. In this paper, we present a novel macroscopic method, named interactable cooperative driving lattice hydrodynamic model (Interactable CDL-H model). We describe traffic flow along lanes and flow interactions between lanes in a uniformly continuum frame. We further consider various constraints for a detailed lane-changing simulation. The model owns the efficiency of traditional macroscopic traffic models and can describe lane-changing behaviors effectively. It physically describes where/when/how traffic flow goes into (out) of a lane, which make it possible to simulate and display lane-changing behaviors in large-scale virtual environments. The validity and efficiency of the interactable CDLH model are demonstrated by comparing simulation results with real traffic data and one-dimensional CDLH model.
Tianlu Mao
VR2
2014 Optimization-based group performance deducing
abstract
ABSTRACT Large‐scale group performance animation has been an important research topic because of its diverse range of applications including virtual rehearsal and film production. Animating hundreds of virtual actors as what the director wishes is a tough task. In this paper, we address this challenge by introducing an optimization method that generates large‐scale group performance by deducing a small‐scale one with fewer actors. We introduced group motion bigraph technique and transformed the motion‐deducing problem into a constrained optimization problem. A solving process is then presented to automatically obtain the motion of the large group with velocity constraints. Moreover, an interactive system of constructing the group motion bigraph has been implemented, which provides flexible edit and control on deducing group motion. The animation results show that our method is competent for deducing large‐scale group performance from only several motion clips performed by small groups. Copyright © 2013 John Wiley & Sons, Ltd.
Lei Lv, Tianlu Mao, Xuecheng Liu
Comput. Animat. Virtual Worlds2
2014 An all-in-one efficient lane-changing model for virtual traffic
abstract
ABSTRACT Modeling lane changes realistically play an important role in traffic animations. Existing models for traffic simulations mostly focus on lane‐changing decision‐making. They cannot describe how lane‐changing processes go on. Though some methods in motion planning can be further used to describe the processes, they are time‐consuming. In this paper, we present an all‐in‐one model for lane‐changing animations. We transform all conditions for lane‐changing decision‐making into constraints on lane‐changing trajectories. We use two polynomials designed in moving reference frames to describe the trajectories. The whole lane‐changing process, whether it can occur and how it goes on, is determined by whether there is a valid trajectory. It only takes O(1) time for the determination. Experiment results show that our model can describe lane‐changing processes realistically and efficiently. Copyright © 2014 John Wiley & Sons, Ltd.
Tianlu Mao, Xingchen Kang
Comput. Animat. Virtual Worlds2
2010 Parallelizing continuum crowds
abstract
In this paper, we present a novel parallelizing method for crowd simulators constructed with a continuum model rather than an agent-based model. The basic idea is to partition a crowded virtual environment into some districts, each of which keeps its own dynamic continuum fields and has several transitional blocks to make individuals keep continuum motion from one district to another. Our method makes continuum models to be parallelizable while preserving their existing superiority of generating smooth motion. Moreover, for most of large-scale applications, our partitioning method effectively simplifies the complexity of simulation. Experiments show that our method has achieved super-linear speedup and could employ more than one hundred worker processors to simulate 1 million people in an area of 672,400m2.
Tianlu Mao, Hao Jiang 0013, Jian Li 0057, Shihong Xia
VRST1
2010 Continuum crowd simulation in complex environments
Hao Jiang 0013, Tianlu Mao, Chunpeng Li, Shihong Xia
Comput. Graph.3
2009 Crowds flow in complex environment
abstract
This paper presents a hybrid approach based on the continuum model proposed by Treuille et al.. Compared to the original method, our solution is well suited for complex environment. We first present an environment structure and a corresponding discretization scheme that help us to organize and simulate crowds in large-scale scenarios. Second, additional discomforts around obstacles are auto-generated for keeping a certain distance between pedestrians and obstacles which is psychologically plausible, and it could obtain smoother trajectory when people move around many obstacles. Thirdly, we propose a technique for density conversion; the density field is dynamically affected by each individual so that it could be adapted to different grid resolution. The experiment results demonstrate that our hybrid solution can perform plausible crowds flow in complex dynamic environments.
Hao Jiang 0013, Tianlu Mao, Chunpeng Li, Shihong Xia
CAD/Graphics3
2009 Evaluating simplified air force models for cloth simulation
abstract
Any cloth simulation system needs the aerodynamics model to describe the dynamic behavior of the cloth interacting with the air. Different air force model has different simplification treatment with different approximation degree. In this paper, we present a quantitative evaluation method for simplified air force models, and experimentally investigate 5 different air force models which are all commonly used in cloth simulations. The results show that the lift component, which was usually neglected in air force models, actually plays an important role in the simulation. Moreover, when the air field is simplified as the linear global wind field, the air force model containing the linear drag component and simple upward lift component matches the real cloth motion better than the complicated nonlinear ones.
Tianlu Mao, Shihong Xia
CAD/Graphics1
2009 A semantic environment model for crowd simulation in multilayered complex environment
abstract
Simulating crowds in complex environment is fascinating and challenging, however, modeling of the environment is always neglected in the past, which is one of the essential problems in crowd simulation especially for multilayered complex environment. This paper presents a semantic model for representing the complex environment, where the semantic information is described with a three-tier framework: a geometric level, a semantic level and an application level. Each level contains different maps for different purposes and our approach greatly facilitates the interactions between individuals and virtual environment. And then a modified continuum crowd method is designed to fit the proposed virtual environment model so that realistic behaviors of large dense crowds could be simulated in multilayered complex environments such as buildings and subway stations. Finally, we implement this method and test it in two complex synthetic urban spaces. The experiment results demonstrate that the semantic environment model can provide sufficient and accurate information for crowd simulation in multilayered complex environment.
Hao Jiang 0013, Tianlu Mao, Chunpeng Li, Shihong Xia
VRST3
2008 Facial animation by optimized blendshapes from motion capture data
abstract
Abstract This paper presents a labor‐saving method to construct optimal facial animation blendshapes from given blendshape sketches and facial motion capture data. At first, a mapping function is established between target “Marker Face” and performer's face by RBF interpolating selected feature points. Sketched blendshapes are transferred to performer's “Marker Face” by using motion vector adjustment technique. Then, the blendshapes of performer's “Marker Face” are optimized according to the facial motion capture data. At last, the optimized blendshapes are inversely transferred to target facial model. Apart from that, this paper also proposes a method of computing blendshape weights from facial motion capture data more accurately. Experiments show that expressive facial animation can be acquired. Copyright © 2008 John Wiley & Sons, Ltd.
Xuecheng Liu, Tianlu Mao, Shihong Xia
Comput. Animat. Virtual Worlds2
2007 CrowdViewer: from simple script to large-scale virtual crowds
abstract
Visualization of large-scale virtual crowds is ubiquitous in many applications of computer graphics. For reasons of efficiency in modeling, animating and rendering, it is difficult to populate scenes with a large number of individually animated virtual characters in real-time applications. In this paper, we present an effective and readily usable solution to this problem. It accepts simple script which includes motion state and position information of each individual at each time step. Supported by material database and motion database, various human models are generated from model templates, and then driven by an agile on-line animation approach. A developed point-based rendering approach is presented to accelerate rendering. We test our system with script including 30,000 people evacuating from a sports arena. The results demonstrate that our approach provides a very effective way to visualize large-scale crowds with high visual realism in real-time.
Tianlu Mao, Bo Shu, Shihong Xia
VRST1
2004 A Faster Method for Modeling Virtual Colony
Tianlu Mao
VR1