VLDB 2026 Research / reviewers in the wild / expert
Hao Liu 0125
dblp:09/3214-125
· DBLP profile ↗
16ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-8121-2333ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NLOS-MT: A Hybrid Mamba and Windowed Attention Transformer for Non-Line-of-Sight Imaging
Shaohui Jin, Xiu Ye, Mengge Liu, Yang Lu 0016, Hao Liu 0125, Mingliang Xu 0001 |
ICPR (5) | 6 |
| 2026 | Hybrid attention triple branch transformer net for underwater image enhancement
Shaohui Jin, Guangpeng Li, Ziqin Xu, Zhengguang Qin, Hao Liu 0125, Mingliang Xu 0001 |
Pattern Recognit. Lett. | 6 |
| 2025 | Enhancing Non-line-of-Sight Imaging Through Contrastive Multiscale Context Aggregation
Shaohui Jin, Zhenjie Yu, Hao Liu 0125 |
ICIC (6) | 5 |
| 2025 | Multispectral non-line-of-sight imaging via deep fusion photography
Hao Liu 0125 |
Sci. China Inf. Sci. | 1 |
| 2025 | Hyperspectral passive non-line-of-sight imaging with band selection
Shaohui Jin, Mengge Liu, Ziqin Xu, Hao Liu 0125, Mingliang Xu 0001 |
Expert Syst. Appl. | 5 |
| 2025 | SSRA: Semantic Segmentation-Guided Region-Attention Colorization MethodabstractInfrared image colorization has witnessed notable advances in recent years; however, existing methods still suffer from color distortions, edge blurring, and artifacts, particularly in complex scenes. To mitigate these issues, we propose a Semantic Segmentation-Guided Region-Attention (SSRA) framework, which enhances colorization fidelity with a stronger emphasis on semantically salient regions. Specifically, we employ the Segment Anything Model (SAM2) to produce high-quality semantic masks, enabling regional self-attention to operate within consistent object boundaries. This design effectively suppresses background interference and facilitates more precise color reconstruction in foreground regions. Furthermore, a focal region loss is introduced to adaptively weight reconstruction errors based on regional importance, thereby directing the model’s representational capacity towards critical areas such as vehicles, pedestrians, and road infrastructure. Extensive experiments on two drone-acquired datasets validate the efficacy of our approach, demonstrating superior performance in both quantitative metrics and qualitative assessments, with notable improvements in structural detail preservation and semantic-aware color consistency. Shaohui Jin, Hao Liu 0125, Mingliang Xu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Enhanced passive non-line-of-sight imaging via multi-scale polarization-guided diffusion model
Shaohui Jin, Guangpeng Li, Hao Liu 0125 |
Vis. Comput. | 7 |
| 2024 | Non-Line-of-Sight Long-Wave Infrared Imaging based on NestformerabstractDue to its portability, low cost, and unnoticed detection mode, passive non-line-of-sight (NLOS) imaging has garnered widespread attention in recent years. This technology reconstructs hidden objects beyond the direct line of sight by analyzing the diffuse reflection on a relay surface. However, issues such as scene complexity and interference from ambient light lead to suboptimal reconstruction results. Long-wave infrared (LWIR) imaging is more significantly affected by ambient temperature but less by illumination, resulting superior resistance to environmental light interference. Therefore, we conduct NLOS imaging experiments using an LWIR camera and propose a novel NLOS image reconstruction network called Nestformer. The feature extraction component of this network incorporates a Transformer-CNN dual-branch mechanism. This dual-branch mechanism includes a Base Transformer Module for capturing global features and a Spatial Channel Attention Module that focuses on extracting local information. Additionally, to enhance the robustness of our model, a combination of multiple loss functions is employed to optimize its feedback capability. Experimental results on our self-collected NLOS-LR dataset indicate that Nestformer outperforms current passive NLOS imaging methods in terms of reconstruction performance. Shaohui Jin, Yayong Zhao, Hao Liu 0125, Zhenjie Yu, Mingliang Xu 0001 |
CSCWD | 3 |
| 2024 | Long-Wave Infrared Non-Line-of-Sight Imaging with Visible Conversion
Shaohui Jin, Hao Liu 0125, Mingliang Xu 0001 |
ICPR (20) | 3 |
| 2024 | TaDFusion: Infrared and Visible Image Fusion Network Based on The Target Detection Task-driven MethodabstractThe combination of visible and infrared images is intended to facilitate complex vision tasks by combining target information and rich texture. By focusing solely on visual perception enhancement, current fusion algorithms do not take into account performance on high-level vision tasks. As a solution to these problems, this research develops a high-level vision task-driven image fusion network (TaDFusion) that combines image fusion and target identification tasks. Through cascading of the image fusion and target detection modules we can significantly improve the performance of advanced vision tasks by using detection loss to guide the information back to the image fusion module. Our algorithm provides better texture preservation and pixel intensity distribution than existing methods based on extensive comparisons and generalization experiments. Besides our framework demonstrates the greatest advantages in facilitating advanced vision tasks by not only generating visually appealing fused images but also detecting higher mAPs than state-of-the-art methods, according to a comparison of the performance of various fusion algorithms in target detection tasks. Shaohui Jin, Qixu Liu, Hao Liu 0125, Mingliang Xu 0001 |
IJCNN | 3 |
| 2024 | Corner Detection: Passive Non-Lin-of-Sight Pedestrian Detection
Shaohui Jin, Xiaoheng Jiang, Jiyue Wang, Hao Liu 0125, Mingliang Xu 0001 |
PRCV (9) | 6 |
| 2024 | Adaptive Dual Attention Fusion Network for RGB-D Surface Defect Detection
Xiaoheng Jiang, Jingqi Liu, Yang Lu 0016, Shaohui Jin, Hao Liu 0125, Mingliang Xu 0001 |
PRCV (9) | 6 |
| 2023 | DFAR-Net: Dual-Input Three-Branch Attention Fusion Reconstruction Network for Polarized Non-Line-of-Sight Imaging
Hao Liu 0125, Ke Wang 0064, Shaohu Jin, Pengyun Chen, Xiaoheng Jiang, Mingliang Xu 0001 |
PRCV (6) | 1 |
| 2022 | Naturalistic Driving Scenario Recognition with Multimodal DataabstractDriving Scenario recognition is one fundamental technology of automated driving systems or advanced driver assistance systems. A common practice of driving scenario recog-nition is to conduct classification tasks with the data collected by in-vehicle data acquisition system or driving simulator. In most existing works, visual data were used since the relevant information of driving scenarios is usually inferable from their visual appearance. However, the non-visual information, e.g. physiological state of the driver, also provide complementary information for the scenario recognition task especially when the visual appearance is insufficient to differentiate similar naturalistic scenarios in some cases. In this paper, we propose a hybrid driving scenario recognition model with multimodal input. The model consists of a convolutional neural network based visual data sub-model, a stacked autoencoder based physiological data sub-model, and a fusion sub-model that combines the extracted features from both visual and physiological sub-models. Besides, a post-processing is adopted to correct the recognition results of some ambiguous scenarios. Experimental results on the trip data collected in the naturalistic driving context demonstrated the effectiveness of the proposed method. Junxiao Xue, Hao Liu 0125 |
MDM | 6 |
| 2019 | Panoptic Studio: A Massively Multiview System for Social Interaction CaptureabstractWe present an approach to capture the 3D motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent; (2) subtle motion needs to be measured over a space large enough to host a social group; (3) human appearance and configuration variation is immense; and (4) attaching markers to the body may prime the nature of interactions. The Panoptic Studio is a system organized around the thesis that social interactions should be measured through the integration of perceptual analyses over a large variety of view points. We present a modularized system designed around this principle, consisting of integrated structural, hardware, and software innovations. The system takes, as input, 480 synchronized video streams of multiple people engaged in social activities, and produces, as output, the labeled time-varying 3D structure of anatomical landmarks on individuals in the space. Our algorithm is designed to fuse the "weak" perceptual processes in the large number of views by progressively generating skeletal proposals from low-level appearance cues, and a framework for temporal refinement is also presented by associating body parts to reconstructed dense 3D trajectory stream. Our system and method are the first in reconstructing full body motion of more than five people engaged in social interactions without using markers. We also empirically demonstrate the impact of the number of views in achieving this goal. Hanbyul Joo, Tomas Simon, Xulong Li 0001, Hao Liu 0125, Sean Banerjee, Timothy Godisart, Bart C. Nabbe, Iain A. Matthews, Takeo Kanade, Shohei Nobuhara, Yaser Sheikh |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Panoptic Studio: A Massively Multiview System for Social Motion CaptureabstractWe present an approach to capture the 3D structure and motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent, (2) subtle motion needs to be measured over a space large enough to host a social group, and (3) human appearance and configuration variation is immense. The Panoptic Studio is a system organized around the thesis that social interactions should be measured through the perceptual integration of a large variety of view points. We present a modularized system designed around this principle, consisting of integrated structural, hardware, and software innovations. The system takes, as input, 480 synchronized video streams of multiple people engaged in social activities, and produces, as output, the labeled time-varying 3D structure of anatomical landmarks on individuals in the space. The algorithmic contributions include a hierarchical approach for generating skeletal trajectory proposals, and an optimization framework for skeletal reconstruction with trajectory re-association. Hanbyul Joo, Hao Liu 0125, Bart C. Nabbe, Iain A. Matthews, Takeo Kanade, Shohei Nobuhara, Yaser Sheikh |
ICCV | 2 |