Sihui Zhang

dblp:178/5996 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Face, body and person analysis · 57% Transfer learning and domain adaptation · 27% Vision and language · 8%
Human-computer interaction and pervasive computing
1 paper
Accessibility and assistive technology · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
gaze estimation
1.622025
'Disengage AND Integrate': Personalized Causal Network for Gaze Estimation · IEEE Trans. Image Process. 2025
Domain-Consistent and Uncertainty-Aware Network for Generalizable Gaze Estimation · IEEE Trans. Multim. 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.812024
Domain-Consistent and Uncertainty-Aware Network for Generalizable Gaze Estimation · IEEE Trans. Multim. 2024
Accessibility and assistive technology › autism
autism screening
0.812024
Visual Question Answering Driven Eye Tracking Paradigm for Identifying Children with Autism Spectrum Disorder · ACM Multimedia 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
0.212024
Domain-Consistent and Uncertainty-Aware Network for Generalizable Gaze Estimation · IEEE Trans. Multim. 2024
Computer vision › Vision and language
visual question answering
0.212024
Visual Question Answering Driven Eye Tracking Paradigm for Identifying Children with Autism Spectrum Disorder · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

cooperative network · 1.5prototype-based subject identification · 0.9episodic training · 0.9causal intervention · 0.9uncertainty perception · 0.8scanpath analysis · 0.8scan path analysis · 0.8domain consistency constraint · 0.8
YearPublicationVenuePosition
2026 Novel dynamic event-triggered fuzzy control for interval type-2 fuzzy singular semi-Markovian jump systems
Sihui Zhang
Fuzzy Sets Syst.2
2025 'Disengage AND Integrate': Personalized Causal Network for Gaze Estimation
abstract
Gaze estimation task aims to predict a 3D gaze direction or a 2D gaze point given a face or eye image. To improve generalization of gaze estimation models to unseen new users, existing methods either disentangle personalized information of all subjects from their gaze features, or integrate unrefined personalized information into blended embeddings. Their methodologies are not rigorous whose performance is still unsatisfactory. In this paper, we put forward a comprehensive perspective named 'Disengage AND Integrate' to deal with personalized information, which elaborates that for specified users, their irrelevant personalized information should be discarded while relevant one should be considered. Accordingly, a novel Personalized Causal Network (PCNet) for generalizable gaze estimation has been proposed. The PCNet adopts a two-branch framework, which consists of a subject-deconfounded appearance sub-network (SdeANet) and a prototypical personalization sub-network (ProPNet). The SdeANet aims to explore causalities among facial images, gazes, and personalized information and extract a subject-invariant appearance-aware feature of each image by means of causal intervention. The ProPNet aims to characterize customized personalization-aware features of arbitrary users with the help of a prototype-based subject identification task. Furthermore, our whole PCNet is optimized in a hybrid episodic training paradigm, which further improve its adaptability to new users. Experiments on three challenging datasets over within-domain and cross-domain gaze estimation tasks demonstrate the effectiveness of our method.
Xiyun Wang, Sihui Zhang, Wanru Xu, Yi Jin 0001
IEEE Trans. Image Process.3
2024 Visual Question Answering Driven Eye Tracking Paradigm for Identifying Children with Autism Spectrum Disorder
abstract
As a non-contact method, eye-tracking data can be used to diagnose people with Autism Spectrum Disorder (ASD) by comparing the differences of eye movements between ASD and healthy people. However, existing works mainly employ a simple free-viewing paradigm or visual search paradigm with restricted or unnatural stimuli to collect the gaze patterns of adults or children with an average age of 6-to-8 years, hindering the early diagnosis and intervention of preschool children with ASD. In this paper, we propose a novel method for identifying children with ASD in three unique features: First, we design a novel eye-tracking paradigm that records Visual Question Answering (VQA) driven gaze patterns in complex natural scenes as a powerful guide for differentiating children with ASD. Second, we contribute a carefully designed dataset, named VQA4ASD, for collecting VQA-driven eye-tracking data from 2-to-6-year-old ASD and healthy children. To the best of our knowledge, this is the first dataset focusing on the early diagnosis of preschool children, which could facilitate the community to understand and explore the visual behaviors of ASD children. Third, we further develop a VQA-guided cooperative ASD screening network (VQA-CASN), in which both task-agnostic and task-specific visual scanpaths are explored simultaneously for ASD screening. Extensive experiments demonstrate that the proposed VQA-CASN achieves competitive performance with the proposed VQA-driven eye-tracking paradigm. The code and dataset is available at: https://github.com/qijiansong/VQA4ASD.
Jiansong Qi, Ying Zhang 0092, Sihui Zhang, Lin Guan 0004, Tianyi Chang
ACM Multimedia4
2024 Domain-Consistent and Uncertainty-Aware Network for Generalizable Gaze Estimation
abstract
Unsupervised domain adaptive (UDA) gaze estimation aims to predict gaze directions of unlabeled target face or eye images given a set of annotated source images, which has been widely applied in practical applications. However, existing methods still perform poorly due to two major challenges. 1) There exists large personalized differences and style discrepancies between source and target samples, which leads the learned source model easily collapsing to biased results; 2) Data uncertainties inherent in reference samples will affect the generalization ability of their models. To tackle the above challenges, in this paper, we propose a novel Domain-Consistent and Uncertainty-Aware (DCUA) network for generalizable gaze estimation. Our DCUA network employs a two-phase framework where a primary training sub-network (PTNet) and a refined adaptation sub-network (RANet) are trained on the source and target domain, respectively. Firstly, to obtain robust and pure gaze-related features, we propose twain domain consistent constraints, that is, the intra-domain consistent constraint and the inter-domain consistent constraint. These two constraints could eliminate the impact of gaze-irrelevant factors by maintaining consistency between label and feature space. Secondly, to further improve the adaptability of our model, we propose dual uncertainty perception modules, which include an intrinsic uncertainty module and an extrinsic uncertainty module. These modules help DCUA network distinguish inferior reference samples and avoid overfitting to them. Experiments on four cross-domain gaze estimation tasks demonstrate the effectiveness of our method.
Sihui Zhang
IEEE Trans. Multim.1
2023 Uncertainty Inspired Autism Spectrum Disorder Screening
Jiansong Qi, Sihui Zhang
MICCAI (5)4
2023 Dual-Uncertainty Guided Cycle-Consistent Network for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to identify novel categories via transferring shared semantic knowledge from seen classes to unseen ones. Since labeled samples of novel categories are unavailable in training phase, visual and semantic spaces are difficult to align precisely. Besides, the uncertainties inherent in fixed visual features and predefined semantic prototypes are always neglected, which also play important roles in modeling unbiased visual-semantic embeddings. In this paper, we propose a Dual-uncertainty Guided Cycle-consistent Network (DGCNet) for ZSL, which aims to learn a robust semantic-to-visual mapping to generate visual centers based on semantic prototypes. Firstly, we propose a cycle-consistent embedding framework, which consists of visual generation sub-network and semantic preservation sub-network. The former generates a primary visual center for each category, while the latter remaps obtained centers back to semantic space to further ensure the consistency between reconstructed semantic embeddings and original prototypes. These two sub-networks explore the intrinsic bidirectional relationships between visual and semantic features complementarily, thus effectively mitigating the alignment shift problem. Furthermore, we develop dual uncertainty perception modules, namely visual uncertainty module and semantic uncertainty module, on the basis of the above two sub-networks. These modules are designed to measure visual and semantic uncertainties of sample features and class prototypes, respectively, which avoid our model overfitting to noisy data and unreliable prototypes. Substantially, the dual uncertainty perception modules contribute to improving the discriminability and adaptability of our DGCNet. Extensive experiments on various datasets demonstrate the effectiveness of our proposed method.
Sihui Zhang
IEEE Trans. Circuits Syst. Video Technol.3