Maozheng Zhao

dblp:164/2232 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-2628-2523ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Tap&Say: Touch Location-Informed Large Language Model for Multimodal Text Correction on Smartphones
abstract
layer that integrates the tap location into the LLM's attention mechanism, enabling it to utilize the tap location for text correction. We fine-tuned the touch location-informed LLM on synthetic touch locations and correction commands, achieving significantly higher correction accuracy than the state-of-the-art method VT [45]. A 16-person user study demonstrated that Tap&Say outperforms VT [45] with 16.4% shorter task completion time and 47.5% fewer keyboard clicks and is preferred by users.
Maozheng Zhao, Michael Xuelin Huang, Nathan G. Huang, Shanqing Cai, Henry Huang, Michael G. Huang, Shumin Zhai, I. V. Ramakrishnan, Xiaojun Bi 0001
CHI1
2023 Gaze Speedup: Eye Gaze Assisted Gesture Typing in Virtual Reality
abstract
Mid-air text input in augmented or virtual reality (AR/VR) is an open problem. One proposed solution is gesture typing where the user performs a gesture trace over the keyboard. However, this requires the user to move their hands precisely and continuously, potentially causing arm fatigue. With eye tracking available on AR/VR devices, multiple works have proposed gaze-driven gesture typing techniques. However, such techniques require the explicit use of gaze which are prone to Midas touch problems, conflicting with other gaze activities in the same moment. In this work, the user is not made aware that their gaze is being used to improve the interaction, making the use of gaze completely implicit. We observed that a user’s implicit gaze fixation location during gesture typing is usually the gesture cursor’s target location if the gesture cursor is moving toward it. Based on this observation, we propose the Speedup method in which we speed up the gesture cursor toward the user’s gaze fixation location, the speedup rate depends on how well the gesture cursor’s moving direction aligns with the gaze fixation. To reduce the overshooting near the target in the Speedup method, we further proposed the Gaussian Speedup method in which the speedup rate is dynamically reduced with a Gaussian function when the gesture cursor gets nearer to the gaze fixation. Using a wrist IMU as input, a 12-person study demonstrated that the Speedup method and Gaussian Speedup method reduced users’ hand movement by and respectively without any loss of typing speed or accuracy.
Maozheng Zhao, Alec M. Pierce, Ran Tan, Ting Zhang 0013, Tianyi Wang 0004, Tanya R. Jonker, Hrvoje Benko, Aakar Gupta
IUI1
2022 Select or Suggest? Reinforcement Learning-based Method for High-Accuracy Target Selection on Touchscreens
abstract
Suggesting multiple target candidates based on touch input is a possible option for high-accuracy target selection on small touchscreen devices. But it can become overwhelming if suggestions are triggered too often. To address this, we propose SATS, a Suggestion-based Accurate Target Selection method, where target selection is formulated as a sequential decision problem. The objective is to maximize the utility: the negative time cost for the entire target selection procedure. The SATS decision process is dictated by a policy generated using reinforcement learning. It automatically decides when to provide suggestions and when to directly select the target. Our user studies show that SATS reduced error rate and selection time over Shift [51], a magnification-based method, and MUCS, a suggestion-based alternative that optimizes the utility for the current selection. SATS also significantly reduced error rate over BayesianCommand [58], which directly selects targets based on posteriors, with only a minor increase in selection time.
Zhi Li 0052, Maozheng Zhao, Hang Zhao 0005, Yan Ma 0006, Wanyu Liu 0001, Michel Beaudouin-Lafon, Fusheng Wang 0001, I. V. Ramakrishnan, Xiaojun Bi 0001
CHI2
2022 EyeSayCorrect: Eye Gaze and Voice Based Hands-free Text Correction for Mobile Devices
abstract
Text correction on mobile devices usually requires precise and repetitive manual control. In this paper, we present EyeSayCorrect, an eye gaze and voice based hands-free text correction method for mobile devices. To correct text with EyeSayCorrect, the user first utilizes the gaze location on the screen to select a word, then speaks the new phrase. EyeSayCorrect would then infer the user’s correction intention based on the inputs and the text context. We used a Bayesian approach for determining the selected word given an eye-gaze trajectory. Given each sampling point in an eye-gaze trajectory, the posterior probability of selecting a word is calculated and accumulated. The target word would be selected when its accumulated interest is larger than a threshold. The misspelt words have higher priors. Our user studies showed that using priors for misspelt words reduced the task completion time up to 23.79% and the text selection time up to 40.35%, and EyeSayCorrect is a feasible hands-free text correction method on mobile devices.
Maozheng Zhao, Henry Huang, Zhi Li 0052, Wenzhe Cui, Kajal Toshniwal, Ananya Goel, Sina Rashidian, Furqan Baig, Khiem Phi, Shumin Zhai, I. V. Ramakrishnan, Fusheng Wang 0001, Xiaojun Bi 0001
IUI1
2021 BayesGaze: A Bayesian Approach to Eye-Gaze Based Target Selection
abstract
Selecting targets accurately and quickly with eye-gaze input remains an open research question. In this paper, we introduce BayesGaze, a Bayesian approach of determining the selected target given an eye-gaze trajectory. This approach views each sampling point in an eye-gaze trajectory as a signal for selecting a target. It then uses the Bayes' theorem to calculate the posterior probability of selecting a target given a sampling point, and accumulates the posterior probabilities weighted by sampling interval to determine the selected target. The selection results are fed back to update the prior distribution of targets, which is modeled by a categorical distribution. Our investigation shows that BayesGaze improves target selection accuracy and speed over a dwell-based selection method, and the Center of Gravity Mapping (CM) method. Our research shows that both accumulating posterior and incorporating the prior are effective in improving the performance of eye-gaze based target selection.
Zhi Li 0052, Maozheng Zhao, Sina Rashidian, Furqan Baig, Wanyu Liu 0001, Michel Beaudouin-Lafon, Brooke Ellison, Fusheng Wang 0001, I. V. Ramakrishnan, Xiaojun Bi 0001
Graphics Interface2
2021 Voice and Touch Based Error-tolerant Multimodal Text Editing and Correction for Smartphones
abstract
Editing operations such as cut, copy, paste, and correcting errors in typed text are often tedious and challenging to perform on smartphones. In this paper, we present VT, a voice and touch-based multi-modal text editing and correction method for smartphones. To edit text with VT, the user glides over a text fragment with a finger and dictates a command, such as "bold" to change the format of the fragment, or the user can tap inside a text area and speak a command such as "highlight this paragraph" to edit the text. For text correcting, the user taps approximately at the area of erroneous text fragment and dictates the new content for substitution or insertion. VT combines touch and voice inputs with language context such as language model and phrase similarity to infer a user's editing intention, which can handle ambiguities and noisy input signals. It is a great advantage over the existing error correction methods (e.g., iOS's Voice Control) which require precise cursor control or text selection. Our evaluation shows that VT significantly improves the efficiency of text editing and text correcting on smartphones over the touch-only method and the iOS's Voice Control method. Our user studies showed that VT reduced the text editing time by 30.80%, and text correcting time by 29.97% over the touch-only method. VT reduced the text editing time by 30.81%, and text correcting time by 47.96% over the iOS's Voice Control method.
Maozheng Zhao, Wenzhe Cui, I. V. Ramakrishnan, Shumin Zhai, Xiaojun Bi 0001
UIST1
2017 Shadow Detection with Conditional Generative Adversarial Networks
abstract
We introduce scGAN, a novel extension of conditional Generative Adversarial Networks (GAN) tailored for the challenging problem of shadow detection in images. Previous methods for shadow detection focus on learning the local appearance of shadow regions, while using limited local context reasoning in the form of pairwise potentials in a Conditional Random Field. In contrast, the proposed adversarial approach is able to model higher level relationships and global scene characteristics. We train a shadow detector that corresponds to the generator of a conditional GAN, and augment its shadow accuracy by combining the typical GAN loss with a data loss term. Due to the unbalanced distribution of the shadow labels, we use weighted cross entropy. With the standard GAN architecture, properly setting the weight for the cross entropy would require training multiple GANs, a computationally expensive grid procedure. In scGAN, we introduce an additional sensitivity parameter w to the generator. The proposed approach effectively parameterizes the loss of the trained detector. The resulting shadow detector is a single network that can generate shadow maps corresponding to different sensitivity levels, obviating the need for multiple models and a costly training procedure. We evaluate our method on the large-scale SBU and UCF shadow datasets, and observe up to 17% error reduction with respect to the previous state-of-the-art method.
Vu Nguyen 0004, Tomás F. Yago Vicente, Maozheng Zhao, Minh Hoai, Dimitris Samaras
ICCV3
2015 No-reference image quality assessment based on phase congruency and spectral entropies
abstract
We develop an efficient general-purpose blind/no-reference image quality assessment (IQA) algorithm that utilizes curvelet domain features of phase congruency values and local spectral entropy values in distorted images. A 2-stage framework of distortion classification followed by quality assessment is used for mapping feature vectors to prediction scores. We utilize a support vector machine (SVM) to train an image distortion and quality prediction model. The resulting algorithm which we name Phase Congruency and Spectral Entropy based Quality (PCSEQ) index is capable of assessing the quality of distorted images across multiple distortion categories. We explain the advantages of phase congruency features and spectral entropy features. We also thoroughly test the algorithm on the LIVE IQA databse and find that PCSEQ correlates well with human judgments of quality. It is superior to the full-reference (FR) IQA algorithm SSIM and several top-performance no-reference (NR) IQA methods such as DIIVINE and SSEQ. We also tested PCSEQ on the TID2008 database to ascertain whether it has performance that is database independent.
Maozheng Zhao, Qin Tu, Yongyu Chang, Bo Yang 0007, Aidong Men
PCS1
2015 Graph based spatiotemporal saliency detection incorporating low and high level features
abstract
In this paper, we propose a novel graph based spatiotemporal saliency detection method which models eye movements using random walk with restart. The method is executed in superpixel domain and unify low level features and high level features into the framework of random walk with restart. The boundary prior, as a high level feature, is employed to obtain a boundary prior based restarting distribution. The temporal saliency map, which is achieved utilizing the low level motion features, is regarded as another restarting distribution. Then the spatiotemporal saliency map is implemented by incorporating two restarting distributions and spatial transition matrix into the random walk with restart framework. Experiment results tested on two public databases show that the proposed method outperforms the existing saliency detection methods.
Qin Tu, Cuiwei Li, Maozheng Zhao, Guangtao Fu
VCIP4
2015 Gradient magnitude similarity for tone-mapped image quality assessment
abstract
Recently, high dynamic range (HDR) image which can accurately reflect the real scene prompts increasing interest. For visualizing the HDR images on standard display devices, it is needed to convert it to low dynamic range (LDR) images. This process is named as tone-mapping. Due to the reduction of the dynamic range, evaluating the tone-mapped image becomes important. Gradient magnitude measure has been used in many state-of-the-art image quality assessment methods. In our research we find that the gradient magnitude measure is also effective for the quality assessment of tone-mapped images. The more different between the gradient magnitude maps of the HDR image and the corresponding LDR image, the worse quality of the LDR image. This paper using the gradient magnitude feature develops a gradient magnitude based method for tone-mapped image quality assessment. The similarity of gradient magnitude maps between the HDR image and the corresponding LDR image is computed. A tone-mapped image should look natural. The naturalness measure used in tone-mapped image quality index (TMQI) is as the supplement measure. In the experiment on the subject-rated tone-mapped image database provided by the authors of TMQI, our proposed method gets better performance than state-of-the-art tone-mapped image assessment methods.
Qin Tu, Maozheng Zhao, Aidong Men, Bo Yang 0007
VCIP3