Junhao Zhao

dblp:223/1501 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GAR-EvoRL: Teaching LLMs to Ask Under Incomplete Information
Menghan Lu, Junhao Zhao, Chaopeng Guo
ICIC (22)3
2026 CARE: Contrastive Alignment for ADL Recognition from Event-Triggered Sensor Streams
abstract
The recognition of Activities of Daily Living (ADLs) from event-triggered ambient sensors is an essential task in Ambient Assisted Living, yet existing methods remain constrained by representation-level limitations. Sequence-based approaches preserve temporal order of sensor activations but are sensitive to noise and lack spatial awareness, while image-based approaches capture global patterns and implicit spatial correlations but compress fine-grained temporal dynamics and distort sensor layouts. Naïve fusion (e.g., feature concatenation) fails to enforce alignment between sequence- and image-based representation views, underutilizing their complementary strengths. We propose Contrastive Alignment for ADL Recognition from Event-Triggered Sensor Streams (CARE), an end-to-end framework that jointly optimizes representation learning via Sequence–Image Contrastive Alignment (SICA) and classification via cross-entropy, ensuring both cross-representation alignment and task-specific discriminability. CARE integrates (i) time-aware, noise-resilient sequence encoding with (ii) spatially-informed and frequency-sensitive image representations, and employs (iii) a joint contrastive-classification objective for end-to-end learning of aligned and discriminative embeddings. Evaluated on three CASAS datasets, CARE achieves state-of-the-art performance (89.8% on Milan, 88.9% on Cairo, and 73.3% on Kyoto7) and demonstrates robustness to sensor malfunctions and layout variability, highlighting its potential for reliable ADL recognition in smart homes. We release our code at https://github.com/Jhziiiig/CARE.
Junhao Zhao, Zishuai Liu, Ruili Fang, Jin Lu 0001, Linghan Zhang, Fei Dou
PerCom1
2026 Vib-ner: a model for out-of-vocabulary recognition in cybersecurity threat intelligence based on variational bottleneck and mutual information
abstract
Abstract The cybersecurity defense strategy of “proactive defense, traceability, and responsiveness” has gained increasing attention. This strategy relies on the collection and application of massive cybersecurity threat intelligence (CTI). However, in the named entity recognition task for threat intelligence processing, traditional models suffer from severe out-of-vocabulary (OOV) issues due to their over-reliance on explicit entity mention information. To address this technical bottleneck, this study designs a novel named entity recognition model–VIB-NER. The model leverages the variational information bottleneck to compress redundant features and strengthen key OOV characteristics, complemented by a mutual information dynamic balance loss. Experimental results show that this method achieves F1 score, recall, and precision of 79%, 77%, and 80% in entity extraction tasks for cybersecurity threat intelligence data, representing a 4–8% improvement over mainstream models such as E-NER. Meanwhile, in terms of training efficiency, the time consumption per training batch is reduced by 50% compared to existing models.
YueDong Wang, Changzheng Liu, Junhao Zhao, XuQing Wang
Cybersecur.5
2026 Enhancing neural topic modeling for social media text via semantic bag of word clusters and log-domain Sinkhorn transport
abstract
Topic modeling has been widely applied to analyze text data from social media platforms. Under this scenario, traditional Neural Topic Models (NTMs) encounter three primary challenges: (1) initial text representation; (2) the long-tail nature of topic distributions in social network texts; (3) approximation of Optimal Transport. Motivated by these challenges, we propose an end-to-end solution spanning from text representation to topic modeling. First, we propose SBoWC, a novel text representation method that performs dimensionality reduction while absorbing semantic information through base terms, achieved by combining word embeddings with clustering statistics. Subsequently, we propose GSWTM, a Wasserstein-based autoencoder topic model that fits the long-tail topic distribution in social network texts via Gamma priors and innovatively employs log-domain Sinkhorn to approximate Optimal Transport. Ablation studies demonstrate the transferability and effectiveness of SBoWC in text representation. GSWTM demonstrates significantly better performance than baselines in TU, C V , and the comprehensive metrics TQ across four real social network datasets of varying sizes. The log-domain Sinkhorn approximation exhibits excellent stability, allowing the regularization parameter ϵ to be reduced to 0.1–0.01, thereby approaching the original Optimal Transport.
Junhao Zhao, Changzheng Liu, Limengzi Yuan
Inf. Process. Manag.2
2025 Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment
abstract
Test-Time adaptation (TTA) aims to improve the performance of source-domain pre-trained models on previously unseen, shifted target domains. Traditional TTA methods primarily adapt model weights based on target data streams, making model performance sensitive to the amount and order of target data. The recently proposed diffusion-driven TTA methods mitigate this by adapting model inputs instead of weights, where an unconditional diffusion model, trained on the source domain, transforms target-domain data into a synthetic domain that is expected to approximate the source domain. However, in this paper, we reveal that although the synthetic data in diffusion-driven TTA seems indistinguishable from the source data, it is unaligned with, or even markedly different from the latter for deep networks. To address this issue, we propose a Synthetic-Domain Alignment (SDA) framework. Our key insight is to fine-tune the source model with synthetic data to ensure better alignment. Specifically, we first employ a conditional diffusion model to generate labeled samples, creating a synthetic dataset. Subsequently, we use the aforementioned unconditional diffusion model to add noise to and denoise each sample before fine-tuning. This Mix of Diffusion (MoD) process mitigates the potential domain misalignment between the conditional and unconditional models. Extensive experiments across classifiers, segmenters, and multimodal large language models (MLLMs, e.g., LLaVA) demonstrate that SDA achieves superior domain alignment and consistently outperforms existing diffusion-driven TTA methods. Our code is available at https://github.com/SHI-Labs/Diffusion-Driven-Test-Time-Adaptation-Via-Synthetic-Domain-Alignment.
Junhao Zhao, Chaoqun Du, Yulin Wang 0002, Chunjiang Ge, Zanlin Ni, Shiji Song, Humphrey Shi, Gao Huang 0001
CVPR2
2025 Next-Generation Access Network Based on Coherent Optics With Hybrid Transceivers, Multi Formats, and Flexible Rates
abstract
Coherent detection has emerged as a key technology for advancing passive optical networks (PON) beyond 100 Gbps per wavelength, due to its advantages over the intensity and direct-detection (IM/DD) approach, which was previously dominant in PON systems. In coherent PON architectures, the cost of the optical network unit (ONU) is substantial due to the complexity of its transceivers. Consequently, numerous studies have focused on simplifying the transmitters and receivers on the ONU side. Yet, simplifying the ONU transceiver introduces issues like reduced dynamic range and lower data rates. How to make a balance between the performance and the cost, particularly in terms of component complexity, represents a critical challenge for the advancement of coherent PON. Therefore, we propose and demonstrate a hybrid, multi-format, and flexible-rate bidirectional coherent PON system, supporting a compatible OLT and ONUs with different types of transceivers. ONUs are categorized into different tiers based on their performance and cost requirements, with transceivers of varying complexity allocated accordingly. As a demonstration of concept, we have successfully conducted experimental transmissions of 25-GBaud 4/16/64- quadrature-amplitude-modulation signals across 20-km fiber in a bidirectional setup, achieving data rates ranging from 50-Gbps to 300-Gbps. The low-end, middle-end, and high-end ONUs attain power budgets of 39/31/21 dB, 43/36/26 dB, and 40/33/23 dB, respectively. This architecture serves as an effective bridge from the current 50G IM/DD PON to the anticipated 200G coherent PON, meeting the varied requirements of users at different service levels.
Aolong Sun, Sizhe Xing, Guoqiang Li 0010, Wangwei Shen, Yongzhu Hu, Junhao Zhao, Ouhan Huang, Jifan Cai, Jianyang Shi, Nan Chi, Junwen Zhang 0001
IEEE J. Sel. Areas Commun.7
2025 Optimizing Superdiffusion of Multiplex Networks Based on Spectral Graph Theory
abstract
Superdiffusion refers to the faster diffusion process in a multiplex network compared to that in an individual network. In this work, we study how interlayer connectivity affects the diffusion performance of a multiplex network. Based on spectral graph theory, we explore the principles of superdiffusion in multiplex networks. We prove that in a duplex network with identical structures, superdiffusion cannot occur under one-to-one interlayer connections. In addition, we prove that the dissimilarity of the Fiedler vector significantly enhances the network superdiffusion performance, which can lead to superdiffusion when selecting nodes with differential eigenvector components in the Fiedler vector for interlayer connections. We also prove that the upper bound of network diffusion with interlayer crossing-connections is limited by the maximum difference of the eigenvector components in the Fiedler vector. Finally, we verify the effectiveness of the theoretical results by numerical analysis.
Hui Liu 0004, Shiqi Dai, Junhao Zhao, Xiaoqun Wu, Shaolin Tan, Guanrong Chen, Zhigang Zeng, Jinhu Lü 0001
IEEE Trans. Syst. Man Cybern. Syst.3
2024 A Novel High-Precision and Low-Latency Abandoned Object Detection Method Under the Hybrid Cloud-Fog Computing Architecture
abstract
Abandoned object Detection (Aod) is of critical importance in the field of public safety. However, the demand on detection accuracy and latency hinders the development of ubiquitous Aod in safety protection, especially for some surveillance devices with relatively low-computational capacity. To this end, a novel high-precision and low-latency Aod method under the hybrid cloud-fog computing architecture is proposed in this article. To be specific, a YOLO-various hidden (YOLO-VH) Aod network model, which is integrated with an efficient dynamic convolution-based ghost module and a Haar wavelet-based downsampling convolution module, is presented to improve the detection accuracy of Aod. In addition, a flexible task offloading strategy is proposed to offload some of the Aod tasks based on the expectation cursor, which is designed to determine the local optimal offloading amount at different times. Finally, extensive experiments are conducted to verify the performance of our proposal through simulations. Our proposal exhibits a reduction of approximately 5.31 million parameters and 30.4 GFLOPs in computation compared with YOLOv9, while demonstrating performance improvements of 25.0% and 38.8% relative to cloud and fog computing, respectively. Furthermore, the total latency for image acquisition, task offloading, and task processing has been observed to be approximately 60% and 15% lower than cloud and fog computing, respectively.
Deyu Lin, Junhao Zhao, Fuxin Yu, Weidong Min, Yong Liang Guan 0001
IEEE Internet Things J.2
2024 Pinning Control of Multiplex Dynamical Networks Using Spectral Graph Theory
abstract
Pinning control has been attracting wide attention for the study of various complex networks for decades. This article explores grounded theory on the pinning synchronization of the emerging multiplex dynamical networks. The multiplex dynamical networks under study can describe many real-world scenarios, in which different layers have distinct individual dynamics of node. In this work, we build the bridge between multiplex structures and network dynamics by using the Lyapunov stability theory and the spectral graph theory. Furthermore, by analyzing spectral properties of the grounded super-Laplacian matrices, we set up several graph-based synchronization criteria for multiplex networks via pinning control. In addition, we overcome the difficulties induced by distinct node dynamics in different layers, and find that interlayer coupling strengths promote intralayer synchronization of multiplex networks. Finally, a collection of numerical simulations verifies the effectiveness of theoretical results.
Hui Liu 0004, Jie Li 0084, Junhao Zhao, Xiaoqun Wu, Zhigang Zeng, Jinhu Lü 0001
IEEE Trans. Cybern.3
2023 FLAME-Based Multi-view 3D Face Reconstruction
Wenzhuo Zheng, Junhao Zhao, Xiaohong Liu 0001, Yongyang Pan, Zhenghao Gan, Haozhe Han
CGI (4)2
2023 SemanticFlow: Semantic Segmentation of Sequential LiDAR Point Clouds From Sparse Frame Annotations
abstract
Sequential point clouds acquired by light detection and ranging (LiDAR) technology provide accurate spatial information for environmental sensing. However, semantic segmentation of point cloud sequences relies on many manual point-wise annotations, which are error-prone and expensive. Existing mainstream weakly supervised methods tackle this by reducing the percentage of labeled points, but they are mostly designed for static indoor scenes and are hard to apply practically. From the viewpoint of realistic annotation procedures and the nature of point cloud sequences, this paper proposes a novel semantic segmentation method, SemanticFlow, for LiDAR point cloud sequences using sparse frames with annotations. The proposed method achieves competitive performance compared with fully supervised methods. Specifically, we designed a bidirectional cross-frame pseudo label propagation module that uses scene flow to learn the correlation and propagate pseudo labels across neighboring frames. In addition, a label refinement mechanism is proposed to select reliable pseudo labels for learning. Extensive experiments on SemanticKITTI, SemanticPOSS, and Synthia 4D datasets demonstrate that our sparse frame annotation method is compatible with some fully supervised counterparts.
Junhao Zhao, Chenglu Wen, Bo Yang 0027, Yulan Guo, Cheng Wang 0003
IEEE Trans. Geosci. Remote. Sens.1
2022 Feature Encoding by Location-Enhanced Word2Vec Embedding for Human Activity Recognition in Smart Homes
Junhao Zhao, Basem Suleiman, Muhammad Johan Alibasa
MobiQuitous1
2021 Simplicity Driven Edge Refinement and Color Reconstruction in Image Vectorization
Junhao Zhao, Shi-Qing Xin, Shuang-Min Chen, Yuanfeng Zhou, Changhe Tu, Wenping Wang 0001
CGI2
2021 GemNN: Gating-enhanced Multi-task Neural Networks with Feature Interaction Learning for CTR Prediction
abstract
Deep neural network (DNN) models have been widely used for click-through rate (CTR) prediction in online advertising. The training framework typically consists of embedding layers and multi-layer perceptions (MLP). At Baidu Search Ads (a.k.a. Phoenix Nest), the new generation of CTR training platform has become PaddleBox, a GPU-based parameter server system. In this paper, we present Baidu's recently updated CTR training framework, called Gating-enhanced Multi-task Neural Networks (GemNN). In particular, we develop a neural network based multi-task learning model to predict CTR in a coarse-to-fine manner, which gradually reduces ad candidates and allows parameter sharing from upstream tasks to downstream tasks to improve the training efficiency. Also, we introduce a gating mechanism between embedding layers and MLP to learn feature interactions and control the information flow fed to MLP layers. We have launched our solution in Baidu PaddleBox platform and observed considerable improvements in both offline and online evaluations. It is now part of the current production~system.
Hongliang Fei, Xingxuan Zhou, Junhao Zhao, Xinyang Qi, Ping Li 0001
SIGIR4
2020 Optimization of RDMA-Based HDFS Data Distribution Mechanism
Junhao Zhao, Cong Dong
NPC3
2020 Computing Smooth Quasi-geodesic Distance Field (QGDF) with Quadratic Programming
Luming Cao, Junhao Zhao, Jian Xu 0023, Shuang-Min Chen, Guozhu Liu, Shi-Qing Xin, Yuanfeng Zhou, Ying He 0001
Comput. Aided Des.2
2020 Automatically modeling piecewise planar furniture shapes from unorganized point cloud
Junhao Zhao, Chen Zong, Luming Cao, Shuang-Min Chen, Guozhu Liu, Jian Xu 0023, Shi-Qing Xin
Comput. Graph.1
2020 A novel tree-based dynamic heterogeneous ensemble method for credit scoring
Yufei Xia, Junhao Zhao, Lingyun He, Yinguo Li, Mengyi Niu
Expert Syst. Appl.2
2019 Object Fusion Tracking Based on Visible and Infrared Images Using Fully Convolutional Siamese Networks
Junhao Zhao, Shengyun Peng, Gang Xiao 0002
FUSION4
2018 A New Image and Video Fusion Method Based on Cross Bilateral Filter
abstract
Image fusion is quite common in applications such as digital photography, medical imaging, surveillance and remote sensing. Designing a common fusion algorithm for all image fusion applications is a challenging task. The cross bilateral filter based image fusion (CBFF) is one such general-purpose method which can be applied to both mono and multimodal image fusion applications. However, CBFF has some drawbacks. 1) It introduces artifacts or extra information into the fused image. 2) The runtime of the CBFF is high. To solve these problems and improve the performance further, we propose a new image fusion algorithm based on the cross bilateral filter by designing a simple and an efficient image fusion rule. Experiments are conducted on both images and videos. Results are analyzed using recent fusion metrics in addition to the qualitative and run time analysis. Results demonstrated that the proposed algorithm can be used as an alternative method to the CBFF.
Durga Prasad Bavirisetti, Gang Xiao 0002, Junhao Zhao
FUSION3