EDBT 2026 Demo / reviewers in the wild / expert
Junhong Zhao
dblp:71/3420
· DBLP profile ↗
26ranked-venue papers
11as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NCPM: A lightweight node-aware channel personalization mechanism for graph neural networks
Feng Hu 0001, Xinfeng Liu, Zuqiang Su, Junhong Zhao, Hong Yu 0007 |
Neurocomputing | 6 |
| 2025 | LLM-Based Simulation Tool for Clinician-Patient Communication Training: A Dual-Mode AI Approach
Magezi Julius, Junhong Zhao, Xiaoying Gao, Jon Herries, Melita MacDonald, Brad Peckler |
PRICAI | 2 |
| 2025 | Local Context-Aware Buoyancy Prediction for Mussel Farm Floats
Carl McMillan, Junhong Zhao, Bing Xue 0001, Ross Vennell, Mengjie Zhang 0001 |
PRICAI | 2 |
| 2025 | Cluster output synchronization analysis of coupled fractional-order uncertain neural networks
Junhong Zhao, Yunliu Li, Peng Liu 0038, Junwei Sun 0002 |
Inf. Sci. | 1 |
| 2025 | Anisotropic Spherical Gaussians Lighting Priors for Indoor Environment Map EstimationabstractHigh Dynamic Range (HDR) environment lighting is essential for augmented reality and visual editing applications, enabling realistic object relighting and seamless scene composition. However, the acquisition of accurate HDR environment maps remains resource-intensive, often requiring specialized devices such as light probes or 360° capture systems, and necessitating stitching during postprocessing. Existing deep learning-based methods attempt to estimate global illumination from partial-view images but often struggle with complex lighting conditions, particularly in indoor environments with diverse lighting variations. To address this challenge, we propose a novel method for estimating indoor HDR environment maps from single standard images, leveraging Anisotropic Spherical Gaussians (ASG) to model intricate lighting distributions as priors. Unlike traditional Spherical Gaussian (SG) representations, ASG can better capture anisotropic lighting properties, including complex shape, rotation, and spatial extent. Our approach introduces a transformer-based network with a two-stage training scheme to predict ASG parameters effectively. To leverage these predicted lighting priors for environment map generation, we introduce a novel generative projector that synthesizes environment maps with high-frequency textures. To train the generative projector, we propose a parameter-efficient adaptation method that transfers knowledge from SG-based guidance to ASG, enabling the model to preserve the generalizability of SG (e.g., spatial distribution and dominance of light sources) while enhancing its capacity to capture fine-grained anisotropic lighting characteristics. Experimental results demonstrate that our method yields environment maps with more precise lighting conditions and environment textures, facilitating the realistic rendering of lighting effects. The implementation code for ASG extraction can be found at https://github.com/junhong-jennifer-zhao/ASG-lighting. Junhong Zhao, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Full-Body Human De-lighting with Semi-supervised Learning
Joshua Weir, Junhong Zhao, Andrew Chalmers, Taehyun Rhee |
ACCV (1) | 2 |
| 2024 | Neural Radiance Fields for Dynamic View Synthesis Using Local Temporal Priors
Rongsen Chen, Junhong Zhao, Andrew Chalmers, Taehyun Rhee |
CVM (1) | 2 |
| 2024 | SGformer: Boosting transformers for indoor lighting estimation from a single imageabstractEstimating lighting from standard images can effectively circumvent the need for resource-intensive high-dynamic-range (HDR) lighting acquisition. However, this task is often ill-posed and challenging, particularly for indoor scenes, due to the intricacy and ambiguity inherent in various indoor illumination sources. We propose an innovative transformer-based method called SGformer for lighting estimation through modeling spherical Gaussian (SG) distributions—a compact yet expressive lighting representation. Diverging from previous approaches, we explore underlying local and global dependencies in lighting features, which are crucial for reliable lighting estimation. Additionally, we investigate the structural relationships spanning various resolutions of SG distributions, ranging from sparse to dense, aiming to enhance structural consistency and curtail potential stochastic noise stemming from independent SG component regressions. By harnessing the synergy of local-global lighting representation learning and incorporating consistency constraints from various SG resolutions, the proposed method yields more accurate lighting estimates, allowing for more realistic lighting effects in object relighting and composition. Our code and model implementing our work can be found at https://github.com/junhong-jennifer-zhao/SGformer . Junhong Zhao, Bing Xue 0001, Mengjie Zhang 0001 |
Comput. Vis. Media | 1 |
| 2024 | Prescribed-time cluster synchronization of coupled inertial neural networks: a lifting dimension approach
Peng Liu 0038, Jian Yong, Junwei Sun 0002, Yanfeng Wang 0002, Junhong Zhao |
Neural Comput. Appl. | 5 |
| 2024 | SALENet: Structure-Aware Lighting Estimations From a Single Image for Indoor EnvironmentsabstractHigh Dynamic Range (HDR) lighting plays a pivotal role in modern augmented and mixed-reality (AR/MR) applications, facilitating immersive experiences through realistic object insertion and dynamic relighting. However, the acquisition of precise HDR environment maps remains cost-prohibitive and impractical when using standard devices. To bridge this gap, this paper introduces SALENet, a novel deep network for estimating global lighting conditions from a single image, to effectively mitigate the need for resource-intensive acquisition methods. In contrast to earlier studies, we focus on exploring the inherent structural relationships within the lighting distribution. We design a hierarchical transformer-based neural network architecture with a proposed cross-attention mechanism between different resolution lighting source representations, optimizing the spatial distribution of lighting sources simultaneously for enhanced consistency. To further improve accuracy, a structure-based contrastive learning method is proposed to select positive-negative pairs based on lighting distribution similarity. By harnessing the synergy of hierarchical transformers and structure-based contrastive learning, our framework yields a significant enhancement in lighting prediction accuracy, enabling high-fidelity augmented and mixed reality to achieve cost-effectively immersive and realistic lighting effects. Junhong Zhao, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | 360$^\circ$∘ Stereo Image Composition With Depth Adaptionabstract360° images and videos have become an economic and popular way to provide VR experiences using real-world content. However, the manipulation of the stereo panoramic content remains less explored. In this paper, we focus on the 360° image composition problem, and develop a solution that can take an object from a stereo image pair and insert it at a given 3D position in a target stereo panorama, with well-preserved geometry information. Our method uses recovered 3D point clouds to guide the composited image generation. More specifically, we observe that using only a one-off operation to insert objects into equirectangular images will never produce satisfactory depth perception and generate ghost artifacts when users are watching the result from different view directions. Therefore, we propose a novel per-view projection method that segments the object in 3D spherical space with the stereo camera pair facing in that direction. A deep depth densification network is proposed to generate depth guidance for the stereo image generation of each view segment according to the desired position and pose of the inserted object. We finally combine the synthesized view segments and blend the objects into the target stereo 360° scene. A user study demonstrates that our method can provide good depth perception and removes ghost artifacts. The per-view solution is a potential paradigm for other content manipulation methods for 360° images and videos. Junhong Zhao, Neil A. Dodgson |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Deep Learning-based Simulator Sickness Estimation from 3D MotionabstractThis paper presents a novel solution for estimating simulator sickness in HMDs using machine learning and 3D motion data, informed by user-labeled simulator sickness data and user analysis. We conducted a novel VR user study, which decomposed motion data and used an instant dial-based sickness scoring mechanism. We were able to emulate typical VR usage and collect user simulator sickness scores. Our user analysis shows that translation and rotation differently impact user simulator sickness in HMDs. In addition, users’ demographic information and self-assessed simulator sickness susceptibility data are collected and show some indication of potential simulator sickness. Guided by the findings from the user study, we developed a novel deep learning-based solution to better estimate simulator sickness with decomposed 3D motion features and user profile information. The model was trained and tested using the 3D motion dataset with user-labeled simulator sickness and profiles collected from the user study. The results show higher estimation accuracy when using the 3D motion data compared with methods based on optical flow extracted from the recorded video, as well as improved accuracy when decomposing the motion data and incorporating user profile information. Junhong Zhao, Kien T. P. Tran, Andrew Chalmers, Weng Khuan Hoh, Richard Yao, Arindam Dey 0001, James Wilmott, Mark Billinghurst, Robert W. Lindeman, Taehyun Rhee |
ISMAR | 1 |
| 2023 | A novel parallel merge neural network with streams of spiking neural network and artificial neural network
Jie Yang 0007, Junhong Zhao |
Inf. Sci. | 2 |
| 2023 | A Survey on 360° Images and Videos in Mixed Reality: Algorithms and Applications
Junhong Zhao, Stefanie Zollmann |
J. Comput. Sci. Technol. | 2 |
| 2022 | Deep Portrait Delighting
Joshua Weir, Junhong Zhao, Andrew Chalmers, Taehyun Rhee |
ECCV (16) | 2 |
| 2022 | Spiking Neural Network Regularization With Fixed and Adaptive Drop-Keep ProbabilitiesabstractDropout and DropConnect are two techniques to facilitate the regularization of neural network models, having achieved the state-of-the-art results in several benchmarks. In this paper, to improve the generalization capability of spiking neural networks (SNNs), the two drop techniques are first applied to the state-of-the-art SpikeProp learning algorithm resulting in two improved learning algorithms called SPDO (SpikeProp with Dropout) and SPDC (SpikeProp with DropConnect). In view that a higher membrane potential of a biological neuron implies a higher probability of neural activation, three adaptive drop algorithms, SpikeProp with Adaptive Dropout (SPADO), SpikeProp with Adaptive DropConnect (SPADC), and SpikeProp with Group Adaptive Drop (SPGAD), are proposed by adaptively adjusting the keep probability for training SNNs. A convergence theorem for SPDC is proven under the assumptions of the bounded norm of connection weights and a finite number of equilibria. In addition, the five proposed algorithms are carried out in a collaborative neurodynamic optimization framework to improve the learning performance of SNNs. The experimental results on the four benchmark data sets demonstrate that the three adaptive algorithms converge faster than SpikeProp, SPDO, and SPDC, and the generalization errors of the five proposed algorithms are significantly smaller than that of SpikeProp. Furthermore, the experimental results also show that the five algorithms based on collaborative neurodynamic optimization can be improved in terms of several measures. Junhong Zhao, Jie Yang 0007, Jun Wang 0002, Wei Wu 0010 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Learning Discriminative Speaker Embedding by Improving Aggregation Strategy and Loss Function for Speaker VerificationabstractThe embedding-based speaker verification (SV) technology has witnessed significant progress due to the advances of deep convolutional neural networks (DCNN). However, how to improve the discrimination of speaker embedding in the open world SV task is still the focus of current research in the community. In this paper, we improve the discriminative power of speaker embedding from three-fold: (1) NeXtVLAD is introduced to aggregate frame-level features, which decomposes the high-dimensional frame-level features into a group of low-dimensional vectors before applying VLAD aggregation. (2) A multi-scale aggregation strategy (MSA) assembled with NeXtVLAD is designed with the purpose of fully extract speaker information from the frame-level feature in different hidden layers of DCNN. (3) A mutually complementary assembling loss function is proposed to train the model, which consists of a prototypical loss and a marginal-based softmax loss. Extensive experiments have been conducted on the VoxCeleb-1 dataset, and the experimental results show that our proposed system can obtain significant performance improvements compared with the baseline, and obtains new state-of-the-art results. The source code of this paper is available at https://github.com/LCF2764/Discriminative-Speaker-Embedding. Chengfang Luo, Aiwen Deng, Junhong Zhao, Wenxiong Kang |
IJCB | 5 |
| 2021 | Reconstructing Reflection Maps Using a Stacked-CNN for Mixed Reality RenderingabstractCorresponding lighting and reflectance between real and virtual objects is important for spatial presence in augmented and mixed reality (AR and MR) applications. We present a method to reconstruct real-world environmental lighting, encoded as a reflection map (RM), from a conventional photograph. To achieve this, we propose a stacked convolutional neural network (SCNN) that predicts high dynamic range (HDR) 360° RMs with varying roughness from a limited field of view, low dynamic range photograph. The SCNN is progressively trained from high to low roughness to predict RMs at varying roughness levels, where each roughness level corresponds to a virtual object's roughness (from diffuse to glossy) for rendering. The predicted RM provides high-fidelity rendering of virtual objects to match with the background photograph. We illustrate the use of our method with indoor and outdoor scenes trained on separate indoor/outdoor SCNNs showing plausible rendering and composition of virtual objects in AR/MR. We show that our method has improved quality over previous methods with a comparative user study and error metrics. Andrew Chalmers, Junhong Zhao, Daniel Medeiros 0001, Taehyun Rhee |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Adaptive Light Estimation using Dynamic Filtering for Diverse Lighting ConditionsabstractHigh dynamic range (HDR) panoramic environment maps are widely used to illuminate virtual objects to blend with real-world scenes. However, in common applications for augmented and mixed-reality (AR/MR), capturing 360° surroundings to obtain an HDR environment map is often not possible using consumer-level devices. We present a novel light estimation method to predict 360° HDR environment maps from a single photograph with a limited field-of-view (FOV). We introduce the Dynamic Lighting network (DLNet), a convolutional neural network that dynamically generates the convolution filters based on the input photograph sample to adaptively learn the lighting cues within each photograph. We propose novel Spherical Multi-Scale Dynamic (SMD) convolutional modules to dynamically generate sample-specific kernels for decoding features in the spherical domain to predict 360° environment maps. Using DLNet and data augmentations with respect to FOV, an exposure multiplier, and color temperature, our model shows the capability of estimating lighting under diverse input variations. Compared with prior work that fixes the network filters once trained, our method maintains lighting consistency across different exposure multipliers and color temperature, and maintains robust light estimation accuracy as FOV increases. The surrounding lighting information estimated by our method ensures coherent illumination of 3D objects blended with the input photograph, enabling high fidelity augmented and mixed reality supporting a wide range of environmental lighting conditions and device sensors. Junhong Zhao, Andrew Chalmers, Taehyun Rhee |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Coherent video generation for multiple hand-held cameras with dynamic foregroundabstractFor many social events such as public performances, multiple hand-held cameras may capture the same event. This footage is often collected by amateur cinematographers who typically have little control over the scene and may not pay close attention to the camera. For these reasons, each individually captured video may fail to cover the whole time of the event, or may lose track of interesting foreground content such as a performer. We introduce a new algorithm that can synthesize a single smooth video sequence of moving foreground objects captured by multiple hand-held cameras. This allows later viewers to gain a cohesive narrative experience that can transition between different cameras, even though the input footage may be less than ideal. We first introduce a graph-based method for selecting a good transition route. This allows us to automatically select good cut points for the hand-held videos, so that smooth transitions can be created between the resulting video shots. We also propose a method to synthesize a smooth photorealistic transition video between each pair of hand-held cameras, which preserves dynamic foreground content during this transition. Our experiments demonstrate that our method outperforms previous state-of-the-art methods, which struggle to preserve dynamic foreground content. Connelly Barnes, Hao-Tian Zhang, Junhong Zhao, Gabriel Salas |
Comput. Vis. Media | 4 |
| 2020 | Learning imbalanced datasets based on SMOTE and Gaussian distribution
Tingting Pan, Junhong Zhao, Wei Wu 0010, Jie Yang 0007 |
Inf. Sci. | 2 |
| 2018 | FV-Net: learning a finger-vein feature representation based on a CNNabstractFinger vein pattern has been proven to be an effective biometric for personal identification in recent years. Nevertheless, there remain challenges that need to be solved, such as finger-vein features that lack robustness and expressiveness. In this paper, we propose a deep convolutional neural network (CNN) model, named the Finger-vein Network (FV-Net), to learn the features representative of a finger vein that is more discriminative and robust than handcrafted features. Next, to address the issue of translation and rotation in vein imaging, we propose a template-like matching strategy while designing the top architecture of the FV-net to extract features with spatial information. Finally, the extensive experimental results show that our proposed method can achieve excellent performance on several public datasets. Wenxiong Kang, Yuxun Fang, Junhong Zhao, Feiqi Deng |
ICPR | 6 |
| 2018 | The convergence analysis of SpikeProp algorithm with smoothing L1∕2 regularization
Junhong Zhao, Jacek M. Zurada, Jie Yang 0007, Wei Wu 0010 |
Neural Networks | 1 |
| 2013 | Audiovisual synthesis of exaggerated speech for corrective feedback in computer-assisted pronunciation trainingabstractIn second language learning, unawareness of the differences between correct and incorrect pronunciations is one of the largest obstacles for mispronunciation correction. In order to make the feedback more discriminatively perceptible, this paper presents a novel method for corrective feedback generation, namely, exaggerated feedback, for language learning. To produce exaggeration effect, the neutral audio and visual speech are both exaggerated and then re-synthesized synchronously based on the audiovisual synthesis technology. The audio speech exaggeration is realized by adjusting the acoustic features related to duration, pitch and energy of the speech according to different phonemes conditions. The visual speech exaggeration is realized by increasing the range of articulatory movement and slowing down the movement around the key actions. The results show that our methods can effectively generate bimodal exaggeration effect for feedback provision and make them more distinctive to be perceived. Junhong Zhao, Wai-Kim Leung, Helen M. Meng, Jia Liu 0001, Shanhong Xia |
ICASSP | 1 |
| 2013 | Exploiting articulatory features for pitch accent detectionabstractArticulatory features describe how articulators are involved in making sounds. Speakers often use a more exaggerated way to pronounce accented phonemes, so articulatory features can be helpful in pitch accent detection. Instead of using the actual articulatory features obtained by direct measurement of articulators, we use the posterior probabilities produced by multi-layer perceptrons (MLPs) as articulatory features. The inputs of MLPs are frame-level acoustic features pre-processed using the split temporal context-2 (STC-2) approach. The outputs are the posterior probabilities of a set of articulatory attributes. These posterior probabilities are averaged piecewise within the range of syllables and eventually act as syllable-level articulatory features. This work is the first to introduce articulatory features into pitch accent detection. Using the articulatory features extracted in this way, together with other traditional acoustic features, can improve the accuracy of pitch accent detection by about 2%. Junhong Zhao, Weiqiang Zhang 0001, Jia Liu 0001, Shanhong Xia |
J. Zhejiang Univ. Sci. C | 1 |
| 2007 | CMOS Current-controlled OscillatorsabstractThe work presented in this paper is about the design of current-controlled oscillators (ICO). Two ICOs are proposed. Aiming at reducing the duration of the short-circuit currents caused by slowly-changing voltages in the circuits, signal conversion blocks are introduced to generate sharp pulses. In this way, the power efficiency of the circuits is improved, which leads to an extensive performance improvement of the circuits. Both ICOs can operate over a frequency range from 100 KHz to 900 MHz. The quality of the output waveforms before buffers is good and over the entire frequency range the rise/fall time is consistently short. The power dissipation of the ICOs is very low, the same as that of a 5-stage current-starving ring oscillator. Moreover, the scheme of the ICOs allows an easy adjustment of the duty cycle of the output pulse signals. A simple digital control structure of the duty cycle has also been proposed. Junhong Zhao, Chunyan Wang 0004 |
ISCAS | 1 |