Zixin Guo

dblp:141/9965 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-7088-2331ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SeekUI: Predicting Visual Search Behavior on Graphical User Interfaces with a Reward-Augmented Vision Language Model
abstract
Visual search is key to understanding and improving interaction with graphical user interfaces (GUIs), yet predicting scanpaths on real GUIs remains an open challenge. Unlike free-viewing, visual search is goal-driven and shaped by both linguistic and visual features of the GUI. State-of-the-art models of visual search, trained on natural images, fail with GUIs because they cannot capture the effects of grouping and semantics on search strategies. We present SeekUI, a reward-augmented Vision Language Model (VLM) that predicts scanpaths directly from a GUI screenshot and a text cue describing the desired target. Our model extends the capability of VLMs to reproduce human-like visual search behavior on GUIs and outperforms baseline models across different types of GUIs. Importantly, it reproduces key empirical phenomena established in eye-tracking studies of visual search, including the Guess–Scan–Confirm strategy. In sum, SeekUI provides a foundation for predicting visual search behavior and has potential for informing GUI evaluation and optimization.
Zixin Guo, Yue Jiang 0002, Luis A. Leiva, Antti Oulasvirta
CHI1
2025 Valor32k-AVQA v2.0: Open-Ended Audio-Visual Question Answering Dataset and Benchmark
abstract
Despite growing interest in Audio-Visual Question Answering (AVQA), existing datasets often suffer from limited diversity, rigid formats, and insufficient integration of audio and visual modalities. To address these limitations, we introduce Valor32k-AVQA v2.0, a large-scale dataset containing 28,863 real-world videos and over 225,000 QA pairs, designed to support diverse and realistic multimodal understanding. The dataset features both open-ended and multiple-choice questions, each annotated with the required modality ( visual, audio, or audio-visual ) and question category ( description, action, count, temporal, location, or relative position ). All annotations-including questions, answers, and metadata-are generated through a fully automated prompting pipeline using GPT-4o, with human validation performed on a representative sample to ensure quality. We benchmark a few state-of-the-art models, with additional evaluations available on the project page, and observe that incorporating audio consistently improves performance during fine-tuning without compromising visual reasoning capabilities. These findings highlight that the audio signals in our dataset are not only well integrated, but also informative and complementary, establishing Valor32k-AVQA v2.0 as a valuable resource for developing and evaluating robust audio-visual question answering systems.
Ines Riahi, Abduljalil Radman, Zixin Guo, Rachid Hedjam, Jorma Laaksonen
ACM Multimedia3
2025 FastTalker: An unified framework for generating speech and conversational gestures from text
Zixin Guo, Minggui He, Osamu Yoshie
Neurocomputing2
2025 Prompt-based Weakly-supervised Vision-language Pre-training
abstract
Weakly-supervised Vision-Language Pre-training (W-VLP) explores methods leveraging weak cross-modal supervision, typically relying on object tags generated by a pre-trained object detector (OD) from images. However, training such an OD necessitates dense cross-modal information, including images paired with numerous object-level annotations. To alleviate that requirement, this paper addresses W-VLP in two stages: (1) creating data with weaker cross-modal supervision and (2) pre-training a vision-language (VL) model with the created data. The data creation process involves collecting knowledge from large language models (LLMs) to describe images. Given a category label of an image, its descriptions generated by an LLM are used as the language counterpart. This knowledge supplements what can be obtained using an OD, such as spatial relationships among objects most likely appearing in a scene. To mitigate the noise in the LLM-generated descriptions that destabilizes the training process and may lead to overfitting, we incorporate knowledge distillation and external retrieval-augmented knowledge during pre-training. Furthermore, we present an effective VL model pre-trained with the created data. Empirically, despite its weaker cross-modal supervision, our pre-trained VL model notably outperforms other W-VLP works in image and text retrieval tasks, e.g., VLMixer by 17.7% on MSCOCO and RELIT by 11.25% on Flickr30K relatively in Recall@1 in text-to-image retrieval task. It also shows superior performance on other VL downstream tasks, making a big stride towards matching the performances of strongly supervised VLP models. The results reveal the effectiveness of the proposed W-VLP methodology.
Zixin Guo, Julius Wang, Selen Pehlivan, Abduljalil Radman, Min Cao 0005, Jorma Laaksonen
Pattern Recognit. Lett.1
2024 Diffusion-Based Multimodal Video Captioning
Jaakko Kainulainen, Zixin Guo, Jorma Laaksonen
ACCV (3)2
2024 TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling
Dong Huo, Zixin Guo, Xinxin Zuo, Zhihao Shi, Juwei Lu, Peng Dai 0002, Songcen Xu, Li Cheng 0001, Yee-Hong Yang
ECCV (38)2
2024 EyeFormer: Predicting Personalized Scanpaths with Transformer-Guided Reinforcement Learning
abstract
From a visual-perception perspective, modern graphical user interfaces (GUIs) comprise a complex graphics-rich two-dimensional visuospatial arrangement of text, images, and interactive objects such as buttons and menus. While existing models can accurately predict regions and objects that are likely to attract attention “on average”, no scanpath model has been capable of predicting scanpaths for an individual. To close this gap, we introduce EyeFormer, which utilizes a Transformer architecture as a policy network to guide a deep reinforcement learning algorithm that predicts gaze locations. Our model offers the unique capability of producing personalized predictions when given a few user scanpath samples. It can predict full scanpath information, including fixation positions and durations, across individuals and various stimulus types. Additionally, we demonstrate applications in GUI layout optimization driven by our model.
Yue Jiang 0002, Zixin Guo, Hamed Rezazadegan Tavakoli, Luis A. Leiva, Antti Oulasvirta
UIST2
2024 Impact of Design Decisions in Scanpath Modeling
abstract
Modeling visual saliency in graphical user interfaces (GUIs) allows to understand how people perceive GUI designs and what elements attract their attention. One aspect that is often overlooked is the fact that computational models depend on a series of design parameters that are not straightforward to decide. We systematically analyze how different design parameters affect scanpath evaluation metrics using a state-of-the-art computational model (DeepGaze++). We particularly focus on three design parameters: input image size, inhibition-of-return decay, and masking radius. We show that even small variations of these design parameters have a noticeable impact on standard evaluation metrics such as DTW or Eyenalysis. These effects also occur in other scanpath models, such as UMSS and ScanGAN, and in other datasets such as MASSVIS. Taken together, our results put forward the impact of design decisions for predicting users' viewing behavior on GUIs.
Parvin Emami, Yue Jiang 0002, Zixin Guo, Luis A. Leiva
Proc. ACM Hum. Comput. Interact.3
2023 PiTL: Cross-modal Retrieval with Weakly-supervised Vision-language Pre-training via Prompting
abstract
Vision-language (VL) Pre-training (VLP) has shown to well generalize VL models over a wide range of VL downstream tasks, especially for cross-modal retrieval. However, it hinges on a huge amount of image-text pairs, which requires tedious and costly curation. On the contrary,weakly-supervised VLP (W-VLP) explores means with object tags generated by a pre-trained object detector (OD) from images. Yet, they still require paired information, i.e. images and object-level annotations, as supervision to train an OD.
Zixin Guo, Julius Wang, Selen Pehlivan, Abduljalil Radman, Jorma Laaksonen
SIGIR1
2022 Post-Attention Modulator for Dense Video Captioning
abstract
Dense video captioning (VC) aims at generating a paragraph-long description for events in video segments. Borrowing from the success in language modeling, Transformer-based models for VC have been shown effective also in modeling cross-domain video-text representations with cross-attention (Xatt). Despite Xatt’s effectiveness, the queries and outputs of attention, which are from different domains, tend to be weakly related. In this paper, we argue that the weak relatedness, or domain discrepancy, could impede a model from learning meaningful cross-domain representations. Hence, we propose a simple yet effective Post-Attention Modulator (PAM) that post-processes Xatt’s outputs to narrow the discrepancy. Specifically, PAM modulates and enhances the average similarity over Xatt’s queries and outputs. The modulated similarities are then utilized as a weighting basis to interpolate PAM’s outputs. In our experiments, PAM was applied to two strong VC baselines, VTransformer and MART, with two different video features on the well-known VC benchmark datasets ActivityNet Captions and YouCookII. According to the results, the proposed PAM brings consistent improvements in, e.g., CIDEr-D at most to 14.5%, as well as other metrics, BLEU and METEOR, considered.
Zixin Guo, Julius Wang, Jorma Laaksonen
ICPR1
2021 Global Fusion Attention for Vision and Language Understanding (Student Abstract)
abstract
We extend the popular transformer architecture to a multi-modal model, processing both visual and textual inputs. We propose a new attention mechanism on Transformer-based architecture for the joint vision and language understanding tasks. Our model fuses multi-level comprehension between images and texts in a weighted manner, which could better curve the internal relationships. Experiments on benchmark VQA dataset CLEVR demonstrate the effectiveness of the proposed attention mechanism. We also observe the improvements in sample efficiency of reinforcement learning through the experiments on grounded language understanding tasks of BabyAI platform.
Zixin Guo, Ziyu Wan
AAAI1
2021 EdgeKeeper: a trusted edge computing framework for ubiquitous power Internet of Things
abstract
Ubiquitous power Internet of Things (IoT) is a smart service system oriented to all aspects of the power system, and has the characteristics of universal interconnection, human-computer interaction, comprehensive state perception, efficient information processing, and other convenient and flexible applications. It has become a hot topic in the field of IoT. We summarize some existing research work on the IoT and edge computing framework. Because it is difficult to meet the requirements of ubiquitous power IoT for edge computing in terms of real time, security, reliability, and business function adaptation using the general edge computing framework software, we propose a trusted edge computing framework, named “EdgeKeeper,” adapting to the ubiquitous power IoT. Several key technologies such as security and trust, quality of service guarantee, application management, and cloud-edge collaboration are desired to meet the needs of the edge computing framework. Experiments comprehensively evaluate EdgeKeeper from the aspects of function, performance, and security. Comparison results show that EdgeKeeper is the most suitable edge computing framework for the electricity IoT. Finally, future directions for research are proposed.
Weiyong Yang, Xingshen Wei, Zixin Guo, Kangle Yang, Longyun Qi
Frontiers Inf. Technol. Electron. Eng.4
2013 Graph-based multiple instance learning for action recognition
abstract
This paper presents a novel framework for recognizing realistic actions captured from unconstrained environments. We describe an action as a collection of space-time activity parts, which are adaptively extracted by clustering foreground trajectories. Each video part is associated with a Bag-of-Features (BoF) histogram, yielding our bag-of-histograms representation for video. We formulate our action classification problem within the graph-based Multiple Instance Learning (MIL) framework, in which each activity part is cast as an instance and a graphical model is incorporated with MIL to leverage the interaction information among the instances. We evaluate our method on two challenging action datasets and demonstrate significant improvements over the state-of-the-art BoF baseline algorithm.
Zixin Guo, Yang Yi 0003
ICIP1