VLDB 2026 Research / reviewers in the wild / expert
Hao Jiang 0013
dblp:38/6049-13
· DBLP profile ↗
18ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-8406-6845ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View StereoabstractRecently, patch deformation-based methods have demonstrated significant effectiveness in multi-view stereo due to their incorporation of deformable and expandable perception for reconstructing textureless areas. However, these methods generally focus on identifying reliable pixel correlations to mitigate matching ambiguity of patch deformation, while neglecting the deformation instability caused by edge-skipping and visibility occlusions, which may cause potential estimation deviations. To address these issues, we propose DVP-MVS++, an innovative approach that synergizes both depth-normal-edge aligned and harmonized cross-view priors for robust and visibility-aware patch deformation. Specifically, to avoid edge-skipping, we first apply DepthPro, Metric3Dv2 and Roberts operator to generate coarse depth maps, normal maps and edge maps, respectively. These maps are then aligned via an erosion-dilation strategy to produce fine-grained homogeneous boundaries for facilitating robust patch deformation. Moreover, we reformulate view selection weights as visibility maps, and then implement both an enhanced cross-view depth reprojection and an area-maximization strategy to help reliably restore visible areas and effectively balance deformed patch. Additionally, we obtain geometry consistency by adopting both aggregated normals via view selection and projection depth differences via epipolar lines, and then employ SHIQ for highlight correction to facilitate highlight perception capacity, thus improving reconstruction quality during propagation and refinement stage. Evaluations on ETH3D, Tanks & Temples and Strecha datasets exhibit the state-of-the-art performance and robust generalization capability of our proposed method. Zhenlong Yuan, Chengxuan Qian, Jianing Chen 0007, Yinda Chen, Kehua Chen, Tianlu Mao, Zhaoxin Li, Hao Jiang 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2025 | All-in-One: Transferring Vision Foundation Models into Stereo MatchingabstractAs a fundamental vision task, stereo matching has made remarkable progress. While recent iterative optimization-based methods have achieved promising performance, their feature extraction capabilities still have room for improvement. Inspired by the ability of vision foundation models (VFMs) to extract general representations, in this work, we propose AIO-Stereo which can flexibly select and transfer knowledge from multiple heterogeneous VFMs to a single stereo matching model. To better reconcile features between heterogeneous VFMs and the stereo matching model and fully exploit prior knowledge from VFMs, we proposed a dual-level feature utilization mechanism that aligns heterogeneous features and transfers multi-level knowledge. Based on the mechanism, a dual-level selective knowledge transfer module is designed to selectively transfer knowledge and integrate the advantages of multiple VFMs. Experimental results show that AIO-Stereo achieves start-of-the-art performance on multiple datasets and ranks 1st on the Middlebury dataset and outperforms all the published work on the ETH3D benchmark. Jiakang Yuan, Peng Ye 0006, Tao Chen 0003, Hao Jiang 0013, Meiya Chen |
AAAI | 6 |
| 2025 | Learning to Predict the Future from Monocular Vision for Efficient Human-Aware NavigationabstractHuman-aware navigation (HAN) aims to build autonomous agents that robustly and naturally navigate in human-centered environments. Due to the complex and dynamic nature of this task, existing approaches typically rely on sophisticated pipelines that separately process perception and decision-making to solve it. In this work, we propose an Obstruction Distance Vector based End-to-End Model (ODVEEM), using monocular vision for navigation around humans. The Obstruction Distance Vector (ODV) is an intermediate representation in our model, leveraged to describe the Obstruction Distance to the first future collision in all possible directions in the horizontal field of view. As ODV cannot be calculated directly in the real world, we design a neural network for ODV estimation, formulating it as a classification problem with auxiliary proxy tasks, which play a key role in effectively predicting the implicit future motion of nearby humans. Taking advantage of ODV, ODVEEM supervised by human behavioral heuristics is employed to guide the agent to reach a goal efficiently and avoid potential collisions. Several challenging experiments show our method's substantial improvement over a number of baseline methods, attaining solid performance with zero-shot transfer to unseen simulated and real-world environments. Yushuang Huang, Hao Jiang 0013, Wanli Ouyang |
ICRA | 2 |
| 2025 | HAIF-GS: Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic SceneabstractReconstructing dynamic 3D scenes from monocular videos remains a fundamental challenge in 3D vision. While 3D Gaussian Splatting (3DGS) achieves real-time rendering in static settings, extending it to dynamic scenes is challenging due to the difficulty of learning structured and temporally consistent motion representations. This challenge often manifests as three limitations in existing methods: redundant Gaussian updates, insufficient motion supervision, and weak modeling of complex non-rigid deformations. These issues collectively hinder coherent and efficient dynamic reconstruction. To address these limitations, we propose HAIF-GS, a unified framework that enables structured and consistent dynamic modeling through sparse anchor-driven deformation. It first identifies motion-relevant regions via an Anchor Filter to suppress redundant updates in static areas. A self-supervised Induced Flow-Guided Deformation module induces anchor motion using multi-frame feature aggregation, eliminating the need for explicit flow labels. To further handle fine-grained deformations, a Hierarchical Anchor Propagation mechanism increases anchor resolution based on motion complexity and propagates multi-level transformations. Extensive experiments on synthetic and real-world benchmarks validate that HAIF-GS significantly outperforms prior dynamic 3DGS methods in rendering quality, temporal coherence, and reconstruction efficiency. Jianing Chen 0007, Yujun Cai, Hao Jiang 0013, Chengxuan Qian, Juyuan Kang, Shuqin Gao, Honglong Zhao, Tianlu Mao |
NeurIPS | 4 |
| 2025 | SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo With Depth Restoration and Occlusion ConstraintabstractRecently, patch-deformation methods have exhibited significant effectiveness in multi-view stereo owing to the deformable and expandable patches in reconstructing textureless areas. However, existing approaches neglect to address the problem of deformation instability caused by easily overlooked edge-skipping, potentially leading to matching distortions, thus leaving room for further improvement. To fill this gap, we propose SED-MVS, which adopts panoptic segmentation and multi-trajectory diffusion strategy for segmentation-driven and edge-aligned patch deformation. Specifically, to prevent unanticipated edge-skipping, we first employ SAM2 for panoptic segmentation as depth-edge guidance to guide patch deformation, followed by multi-trajectory diffusion strategy to ensure patches are comprehensively aligned with depth edges. Moreover, to avoid potential inaccuracy of random initialization, we combine both sparse points from LoFTR and monocular depth map from DepthAnything V2 to restore reliable and realistic depth map for initialization and supervised guidance. Finally, we integrate the segmentation image with the monocular depth map to exploit inter-instance occlusion relationship, then further regard them as occlusion map to implement two distinct edge constraint, thereby facilitating occlusion-aware patch deformation. Extensive results on ETH3D, Tanks & Temples, BlendedMVS, Strecha and DL3DV-10K datasets validate the state-of-the-art performance and robust generalization capability of our proposed method. Zhenlong Yuan, Zhidong Yang, Yujun Cai, Kuangxin Wu, Mufan Liu, Hao Jiang 0013, Zhaoxin Li |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | An Embarrassingly Simple Approach to Enhance Transformer Performance in Genomic Selection for Crop Breeding
Renqi Chen, Wenwei Han, Haohao Zhang, Haoyang Su 0001, Zhefan Wang 0002, Hao Jiang 0013, Wanli Ouyang, Nanqing Dong |
IJCAI | 7 |
| 2024 | Large Scale Farm Scene Modeling from Remote Sensing ImageryabstractIn this paper we propose a scalable framework for large-scale farm scene modeling that utilizes remote sensing data, specifically satellite images. Our approach begins by accurately extracting and categorizing the distributions of various scene elements from satellite images into four distinct layers: fields, trees, roads, and grasslands. For each layer, we introduce a set of controllable Parametric Layout Models (PLMs). These models are capable of learning layout parameters from satellite images, enabling them to generate complex, large-scale farm scenes that closely reproduce reality across multiple scales. Additionally, our framework provides intuitive control for users to adjust layout parameters to simulate different stages of crop growth and planting patterns. This adaptability makes our model an excellent tool for graphics and virtual reality applications. Experimental results demonstrate that our approach can rapidly generate a variety of realistic and highly detailed farm scenes with minimal inputs. Zhiqi Xiao, Hao Jiang 0013, Zhigang Deng 0001, Wenwei Han |
ACM Trans. Graph. | 2 |
| 2021 | High accuracy and geometry-consistent confidence prediction network for multi-view stereo
Zhaoxin Li, Xiaoge Zhang 0003, Kangkan Wang, Hao Jiang 0013 |
Comput. Graph. | 4 |
| 2018 | An emotion evolution based model for collective behavior simulationabstractCurrent crowd simulation progresses still fall short of simulating many real-world collective behaviors. Arguably, one of the main reasons is that some essential qualities of human beings such as emotion have not been effectively modeled and incorporated into crowd simulation algorithms. In this paper, we propose a novel computational model for emotion evolution and demonstrate its applications for crowd simulation. Specifically, our approach is designed to tackle three major issues in the emotion evolution process: (i) how to perceive and evaluate emotion when individuals face emergency or external events, (ii) how to evolve the emotion during induction, and (iii) how specific actions of individuals in a crowd are impacted by emotion. Through many experiments, we demonstrate that our method can effectively simulate emergent dynamic collective patterns observed in real-world crowd footages. Hao Jiang 0013, Zhigang Deng 0001, Xiangjun He, Tianlu Mao |
I3D | 1 |
| 2018 | Behavioral Simulation of Passengers in a Waiting HallabstractIn this paper, we introduced a behavioral decision and execution method to simulate crowded passengers in a waiting hall. The method, as well as its simulation framework, is designed under the special purpose of passenger safety investigation. It supports the simulation of both regular crowded passenger behaviors and emergency passenger behavior. Situations under different time tables and density control measure could easily be conducted and simulated for safety purposes. Shaohua Liu 0002, Xiyuan Song, Hao Jiang 0013, Min Shi 0005, Tianlu Mao |
VR | 3 |
| 2016 | Groupnect: Integrating group interaction into large display systemabstractLarge display systems have been successfully applied in virtual reality domains because they can provide full sense of immersion through large visual space and high display resolution. However, only a few users can interact with these systems by using pen-like or marker-based devices. In addition, user experience and application mode are constrained in many areas. In this paper, we propose a novel application framework called “Groupnect”, which gives users unique experience of group interaction in a large display system. By using optical tracking and 3D gesture recognition technologies, our approach can automatically recognize gesture-based control signals for 12 users simultaneously, and the backend system can trigger corresponding actions in real time. We conduct a user study and compare the results with a standard interaction mode. The results demonstrate that our approach greatly increases recorded objective activities and subjective efforts. Moreover, the physical and mental participation of users can be promoted by Groupnect. It indicates great potential to design novel applications in entertainment, education and training areas. Hao Jiang 0013, Tianlu Mao |
VR | 1 |
| 2015 | Collective Crowd Formation Transform with Mutual Information-Based Runtime FeedbackabstractAbstract This paper introduces a new crowd formation transform approach to achieve visually pleasing group formation transition and control. Its core idea is to transform crowd formation shapes with a least effort pair assignment using the Kuhn–Munkres algorithm, discover clusters of agent subgroups using affinity propagation and Delaunay triangulation algorithms and apply subgroup‐based social force model (SFM) to the agent subgroups to achieve alignment, cohesion and collision avoidance. Meanwhile, mutual information of the dynamic crowd is used to guide agents' movement at runtime. This approach combines both macroscopic (involving least effort position assignment and clustering) and microscopic (involving SFM) controls of the crowd transformation to maximally maintain subgroups' local stability and dynamic collective behaviour, while minimizing the overall effort (i.e. travelling distance) of the agents during the transformation. Through simulation experiments and comparisons, we demonstrate that this approach is efficient and effective to generate visually pleasing and smooth transformations and outperform several existing crowd simulation approaches including reciprocal velocity avoidances, optimal reciprocal collision avoidance and OpenSteer. Yunpeng Wu, Yangdong Ye, Illés J. Farkas, Hao Jiang 0013, Zhigang Deng 0001 |
Comput. Graph. Forum | 5 |
| 2015 | miSFM: On combination of Mutual Information and Social Force Model towards simulating crowd evacuation
Yunpeng Wu, Pei Lv, Hao Jiang 0013, Mingxuan Luo, Yangdong Ye |
Neurocomputing | 4 |
| 2014 | Crowd Simulation and Its Applications: Recent Advances
Mingliang Xu 0001, Hao Jiang 0013, Xiaogang Jin 0001, Zhigang Deng 0001 |
J. Comput. Sci. Technol. | 2 |
| 2010 | Parallelizing continuum crowdsabstractIn this paper, we present a novel parallelizing method for crowd simulators constructed with a continuum model rather than an agent-based model. The basic idea is to partition a crowded virtual environment into some districts, each of which keeps its own dynamic continuum fields and has several transitional blocks to make individuals keep continuum motion from one district to another. Our method makes continuum models to be parallelizable while preserving their existing superiority of generating smooth motion. Moreover, for most of large-scale applications, our partitioning method effectively simplifies the complexity of simulation. Experiments show that our method has achieved super-linear speedup and could employ more than one hundred worker processors to simulate 1 million people in an area of 672,400m2. Tianlu Mao, Hao Jiang 0013, Jian Li 0057, Shihong Xia |
VRST | 2 |
| 2010 | Continuum crowd simulation in complex environments
Hao Jiang 0013, Tianlu Mao, Chunpeng Li, Shihong Xia |
Comput. Graph. | 1 |
| 2009 | Crowds flow in complex environmentabstractThis paper presents a hybrid approach based on the continuum model proposed by Treuille et al.. Compared to the original method, our solution is well suited for complex environment. We first present an environment structure and a corresponding discretization scheme that help us to organize and simulate crowds in large-scale scenarios. Second, additional discomforts around obstacles are auto-generated for keeping a certain distance between pedestrians and obstacles which is psychologically plausible, and it could obtain smoother trajectory when people move around many obstacles. Thirdly, we propose a technique for density conversion; the density field is dynamically affected by each individual so that it could be adapted to different grid resolution. The experiment results demonstrate that our hybrid solution can perform plausible crowds flow in complex dynamic environments. Hao Jiang 0013, Tianlu Mao, Chunpeng Li, Shihong Xia |
CAD/Graphics | 1 |
| 2009 | A semantic environment model for crowd simulation in multilayered complex environmentabstractSimulating crowds in complex environment is fascinating and challenging, however, modeling of the environment is always neglected in the past, which is one of the essential problems in crowd simulation especially for multilayered complex environment. This paper presents a semantic model for representing the complex environment, where the semantic information is described with a three-tier framework: a geometric level, a semantic level and an application level. Each level contains different maps for different purposes and our approach greatly facilitates the interactions between individuals and virtual environment. And then a modified continuum crowd method is designed to fit the proposed virtual environment model so that realistic behaviors of large dense crowds could be simulated in multilayered complex environments such as buildings and subway stations. Finally, we implement this method and test it in two complex synthetic urban spaces. The experiment results demonstrate that the semantic environment model can provide sufficient and accurate information for crowd simulation in multilayered complex environment. Hao Jiang 0013, Tianlu Mao, Chunpeng Li, Shihong Xia |
VRST | 1 |