Zhenyue Qin

dblp:222/6232 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-3471-6280ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DermEVAL: A Dermatologist-Reviewed Benchmark for Multimodal Large Language Models
abstract
Clinical photographs play a vital role in conversational computer-aided diagnosis, particularly in dermatology. However, existing skin disease benchmarks contain limitations like insufficient dataset size, the sole presence of categorical labels, the lack of expert inspections, and limited diversity in annotations. To address these shortcomings, we introduce DermEVAL, a large-scale benchmark specifically designed to evaluate the performance of Multimodal Large Language Models (MLLMs) in dermatology. Our benchmark includes image-text pairs depicting 16 distinct skin diseases, featuring a total of 11,347 representative images drawn from various dermatological datasets, carefully selected and annotated with the guidance of dermatologists. DermEVAL enables two primary tasks: visual question answering (VQA) and medical report generation (MRG), designed to simulate real-world medical diagnostics. We evaluate the performance of MLLMs in dermatology using multiple metrics, including traditional metrics and GPT-4V-based assessments. Our results indicate that accurately diagnosing skin diseases remains challenging for state-of-the-art MLLMs. We also demonstrate that fine-tuning MLLMs using DermEVAL significantly improves their performance on dermatology-related image–text tasks.
Hongjin Zhao, Zhenyue Qin, Ge-Peng Ji, Tom Gedeon, Nick Barnes
WACV3
2026 LMOD\(\boldsymbol{+}\): A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
abstract
The rising prevalence of vision-threatening eye diseases poses a major global health and economic burden, yet timely diagnosis remains limited by workforce shortages, diagnostic delays, and restricted access to specialized care. Artificial intelligence (AI) offers potential solutions. In particular, recent progress in foundation models and large language models-especially multimodal large language models (MLLMs)-has shown promise in medical image interpretation and automated clinical documentation. However, advancing MLLMs for ophthalmology is hindered by the lack of unified, comprehensive benchmark datasets for development and evaluation. Most existing benchmarks were designed for earlier models, which focused on narrow tasks or specific disease conditions. These benchmarks typically provide outputs in the form of disease labels rather than free-text responses. As a result, they are less suitable for assessing emerging generative models. In this work, we present LMOD+, a large-scale multimodal ophthalmology benchmark dataset comprising 32,633 instances with multi-granular annotations across 12 common ophthalmic conditions and 5 imaging modalities. The dataset integrates imaging, anatomical structures, demographics, and free-text annotations. It supports primary ophthalmic applications such as anatomical structure recognition, disease screening, disease staging, and demographic prediction for potential performance bias evaluation. Alongside the dataset, we introduce a systematic and unified data curation pipeline that repurposes existing or new datasets for MLLM development. LMOD+ extends our preliminary LMOD benchmark-the first multimodal ophthalmology benchmark for MLLMs-with three major enhancements. First, we expanded the dataset by nearly 50% (from 21,933 to 32,633 instances). The color fundus photography (CFP) modality, the most accessible imaging modality in ophthalmology, was significantly enlarged to cover a broader range of pathological conditions. Second, we broadened task coverage to include (a) 12 binary disease diagnosis tasks for prevalent conditions such as diabetic retinopathy, age-related macular degeneration, and retinal vein occlusion; (b) multi-class ophthalmic disease diagnosis; (c) disease severity classification, including a diabetic retinopathy staging task, which uses two internationally adopted grading standards: the international clinical diabetic retinopathy classification and the Scottish diabetic retinopathy grading scheme classification; and (d) demographic prediction (age and sex) to assess potential model bias. Third, we systematically evaluated 24 state-of-the-art MLLMs, including recent models from the InternVL, Qwen, and DeepSeek families. Our evaluations highlight both the promise and limitations of current MLLMs in ophthalmology. For example, Qwen-7B and InternVL achieved accuracies of 58.26% and 57.83% in disease screening under a zero-shot setting with a single model-a considerably more challenging paradigm than traditional fine-tuning, where separate models are trained for each specific task. InternVL also demonstrated potential in anatomical recognition. Nonetheless, overall performance remained suboptimal and often close to random baselines for challenging tasks such as disease staging, underscoring the substantial gap between general-domain MLLMs and the specialized requirements of ophthalmology. We publicly release the dataset, curation pipeline, and leaderboard to encourage community-wide development and evaluation of MLLMs, with the goal of advancing ophthalmic applications and ultimately reducing the global burden of vision-threatening diseases through AI. The dataset website, benchmark leaderboard, and download link are available at https://kfzyqin.github.io/lmod_plus.
Zhenyue Qin, Yang Liu 0249, Jinyu Ding, Anran Li 0001, Dylan Campbell, Xuansheng Wu, Ke Zou, Tiarnan D. Keenan, Emily Y. Chew, Zhiyong Lu, Ninghao Liu 0001, Xiuzhen Zhang 0001, Qingyu Chen 0001
ACM Trans. Comput. Heal.1
2026 Representation-centric survey of supervised skeletal action recognition and the new benchmark
abstract
3D skeletal action recognition has emerged as a powerful alternative to traditional RGB and depth-based approaches, offering robustness to environmental variations, computational efficiency, and enhanced privacy. Despite remarkable progress, current research remains fragmented across diverse input representations and lacks evaluation under scenarios that reflect real-world challenges. This paper presents a representation-centric review of supervised skeletal action recognition, systematically categorizing state-of-the-art methods by their input feature types: joint coordinates, bone vectors, motion flows, and extended representations, and analyzing how these choices influence spatiotemporal modeling strategies. Building on the insights from this review, we introduce ANUBIS, a large-scale, challenging dataset designed to address critical gaps in existing benchmarks. ANUBIS incorporates multi-view recordings with back-view perspectives, complex multi-person interactions, fine-grained and violent actions, and contemporary social behaviors. We benchmark a diverse set of state-of-the-art models on ANUBIS and conduct an in-depth analysis of how different feature types affect recognition performance across 102 action categories. Our results show strong action-feature dependencies, highlight the limitations of naïve multi-representational fusion, and point toward the need for task-aware, semantically aligned integration strategies. This work offers both a comprehensive foundation and a practical benchmarking resource, aiming to guide the next generation of robust, generalizable skeleton-based action recognition systems for complex real-world scenarios. The dataset, benchmarking framework, and code are available at https://yliu1082.github.io/ANUBIS/ .
Yang Liu 0249, Jiyao Yang, Madhawa Perera, Pan Ji, Dongwoo Kim 0002, Min Xu 0009, Tianyang Wang 0004, Saeed Anwar, Tom Gedeon, Lei Wang 0108, Zhenyue Qin
Pattern Recognit.11
2025 HandCraft: Anatomically Correct Restoration of Malformed Hands in Diffusion Generated Images
abstract
Generative text-to-image models, such as Stable Diffusion, have demonstrated a remarkable ability to generate diverse, high-quality images. However, they are surprisingly inept when it comes to rendering human hands, which are often anatomically incorrect or reside in the “uncanny valley”. In this paper, we propose a method HandCraft for restoring such malformed hands. This is achieved by auto-matically constructing masks and depth images for hands as conditioning signals using a parametric model, allowing a diffusion-based image editor to fix the hand's anatomy and adjust its pose while seamlessly integrating the changes into the original image, preserving pose, color, and style. Our plug-and-play hand restoration solution is compatible with existing pretrained diffusion models, and the restoration process facilitates adoption by eschewing any fine-tuning or training requirements for the diffusion models. We also contribute MalHand datasets that contain generated images with a wide variety of malformed hands in several styles for hand detector training and hand restoration benchmarking, and demonstrate through qualitative and quantitative eval-uation that HandCraft not only restores anatomical correctness but also maintains the integrity of the overall image.
Zhenyue Qin, Yang Liu 0249, Dylan Campbell
WACV1
2025 Position-Sensing Graph Neural Networks: Proactively Learning Nodes Relative Positions
abstract
Most existing graph neural networks (GNNs) learn node embeddings using the framework of message passing and aggregation. Such GNNs are incapable of learning relative positions between graph nodes within a graph. To empower GNNs with the awareness of node positions, some nodes are set as anchors. Then, using the distances from a node to the anchors, GNNs can infer relative positions between nodes. However, position-aware GNNs (P-GNNs) arbitrarily select anchors, leading to compromising position awareness and feature extraction. To eliminate this compromise, we demonstrate that selecting evenly distributed and asymmetric anchors is essential. On the other hand, we show that choosing anchors that can aggregate embeddings of all the nodes within a graph is NP-complete. Therefore, devising efficient optimal algorithms in a deterministic approach is practically not feasible. To ensure position awareness and bypass NP-completeness, we propose position-sensing GNNs (PSGNNs), learning how to choose anchors in a backpropagatable fashion. Experiments verify the effectiveness of PSGNNs against state-of-the-art GNNs, substantially improving performance on various synthetic and real-world graph datasets while enjoying stable scalability. Specifically, PSGNNs on average boost area under the curve (AUC) more than 14% for pairwise node classification and 18% for link prediction over the existing state-of-the-art position-aware methods. Our source code is publicly available at: https://github.com/ZhenyueQin/PSGNN.
Zhenyue Qin, Saeed Anwar, Dongwoo Kim 0002, Yang Liu 0249, Pan Ji, Tom Gedeon
IEEE Trans. Neural Networks Learn. Syst.2
2024 Visual Prompting in LLMs for Enhancing Emotion Recognition
abstract
Vision Large Language Models (VLLMs) are transforming the intersection of computer vision and natural language processing.Nonetheless, the potential of using visual prompts for emotion recognition in these models remains largely unexplored and untapped.Traditional methods in VLLMs struggle with spatial localization and often discard valuable global context.To address this problem, we propose a Set-of-Vision prompting (SoV) approach that enhances zero-shot emotion recognition by using spatial information, such as bounding boxes and facial landmarks, to mark targets precisely.SoV improves accuracy in face count and emotion categorization while preserving the enriched image context.Through a battery of experimentation and analysis of recent commercial or open-source VLLMs, we evaluate the SoV model's ability to comprehend facial expressions in natural environments.Our findings demonstrate the effectiveness of integrating spatial visual prompts into VLLMs for improving emotion recognition performance. Question:How many visible faces are there in the image?Analyze each face's emotion and group the face's emotions. Answer:In the image, there are 22 visible faces.Emotions can be grouped into three main categories. Neutral Emotion: Approximately 16Faces.This neutral expression is typical in settings requiring attention to a speaker or event, reflecting an engaged yet reserved emotional state. Mildly Positive Emotion: Approximately 5Faces.This could be a response to something interesting or mildly amusing happening at the event, but their reactions are subdued, in keeping with the formal setting.Happy: 1 Face.This expression stands out as more joyful or enthusiastic about the proceedings.
Qixuan Zhang, Zhifeng Wang 0004, Dylan Zhang, Wenjia Niu, Sabrina B. Caldwell, Tom Gedeon, Yang Liu 0003, Zhenyue Qin
EMNLP8
2024 Fusing Higher-Order Features in Graph Neural Networks for Skeleton-Based Action Recognition
abstract
Skeleton sequences are lightweight and compact and thus are ideal candidates for action recognition on edge devices. Recent skeleton-based action recognition methods extract features from 3-D joint coordinates as spatial-temporal cues, using these representations in a graph neural network for feature fusion to boost recognition performance. The use of first- and second-order features, that is, joint and bone representations, has led to high accuracy. Nonetheless, many models are still confused by actions that have similar motion trajectories. To address these issues, we propose fusing higher-order features in the form of angular encoding (AGE) into modern architectures to robustly capture the relationships between joints and body parts. This simple fusion with popular spatial-temporal graph neural networks achieves new state-of-the-art accuracy in two large benchmarks, including NTU60 and NTU120, while employing fewer parameters and reduced run time. Our source code is publicly available at: https://github.com/ZhenyueQin/Angular-Skeleton-Encoding.
Zhenyue Qin, Yang Liu 0249, Pan Ji, Dongwoo Kim 0002, Lei Wang 0108, Robert I. McKay, Saeed Anwar, Tom Gedeon
IEEE Trans. Neural Networks Learn. Syst.1
2023 Anonymization for Skeleton Action Recognition
abstract
Skeleton-based action recognition attracts practitioners and researchers due to the lightweight, compact nature of datasets. Compared with RGB-video-based action recognition, skeleton-based action recognition is a safer way to protect the privacy of subjects while having competitive recognition performance. However, due to improvements in skeleton recognition algorithms as well as motion and depth sensors, more details of motion characteristics can be preserved in the skeleton dataset, leading to potential privacy leakage. We first train classifiers to categorize private information from skeleton trajectories to investigate the potential privacy leakage from skeleton datasets. Our preliminary experiments show that the gender classifier achieves 87% accuracy on average, and the re-identification classifier achieves 80% accuracy on average with three baseline models: Shift-GCN, MS-G3D, and 2s-AGCN. We propose an anonymization framework based on adversarial learning to protect potential privacy leakage from the skeleton dataset. Experimental results show that an anonymized dataset can reduce the risk of privacy leakage while having marginal effects on action recognition performance even with simple anonymizer architectures. The code used in our experiments is available at https://github.com/ml-postech/Skeleton-anonymization/
Saemi Moon, Myeonghyeon Kim, Zhenyue Qin, Yang Liu 0249, Dongwoo Kim 0002
AAAI3
2022 Resolving Anomalies in the Behaviour of a Modularity-Inducing Problem Domain with Distributional Fitness Evaluation
abstract
Discrete gene regulatory networks (GRNs) play a vital role in the study of robustness and modularity. A common method of evaluating the robustness of GRNs is to measure their ability to regulate a set of perturbed gene activation patterns back to their unperturbed forms. Usually, perturbations are obtained by collecting random samples produced by a predefined distribution of gene activation patterns. This sampling method introduces stochasticity, in turn inducing dynamicity. This dynamicity is imposed on top of an already complex fitness landscape. So where sampling is used, it is important to understand which effects arise from the structure of the fitness landscape, and which arise from the dynamicity imposed on it. Stochasticity of the fitness function also causes difficulties in reproducibility and in post-experimental analyses. We develop a deterministic distributional fitness evaluation by considering the complete distribution of gene activity patterns, so as to avoid stochasticity in fitness assessment. This fitness evaluation facilitates repeatability. Its determinism permits us to ascertain theoretical bounds on the fitness, and thus to identify whether the algorithm has reached a global optimum. It enables us to differentiate the effects of the problem domain from those of the noisy fitness evaluation, and thus to resolve two remaining anomalies in the behaviour of the problem domain of Espinosa-Soto and A. Wagner (2010). We also reveal some properties of solution GRNs that lead them to be robust and modular, leading to a deeper understanding of the nature of the problem domain. We conclude by discussing potential directions toward simulating and understanding the emergence of modularity in larger, more complex domains, which is key both to generating more useful modular solutions, and to understanding the ubiquity of modularity in biological systems.
Zhenyue Qin, Tom Gedeon, Robert I. McKay
Artif. Life1
2021 Invertible Denoising Network: A Light Solution for Real Noise Removal
abstract
Invertible networks have various benefits for image de-noising since they are lightweight, information-lossless, and memory-saving during back-propagation. However, applying invertible models to remove noise is challenging because the input is noisy, and the reversed output is clean, following two different distributions. We propose an invertible denoising network, InvDN, to address this challenge. InvDN transforms the noisy input into a low-resolution clean image and a latent representation containing noise. To discard noise and restore the clean image, InvDN replaces the noisy latent representation with another one sampled from a prior distribution during reversion. The de-noising performance of InvDN is better than all the existing competitive models, achieving a new state-of-the-art result for the SIDD dataset while enjoying less run time. Moreover, the size of InvDN is far smaller, only having 4.2% of the number of parameters compared to the most recently proposed DANet. Further, via manipulating the noisy latent representation, InvDN is also able to generate noise more similar to the original one. Our code is available at: https://github.com/Yang-Liu1082/InvDN.git.
Yang Liu 0249, Zhenyue Qin, Saeed Anwar, Pan Ji, Dongwoo Kim 0002, Sabrina B. Caldwell, Tom Gedeon
CVPR2
2021 Skeletons on the Stairs: Are They Deceptive?
Yang Liu 0249, Zhenyue Qin, Xuanying Zhu, Sabrina B. Caldwell, Tom Gedeon
ICONIP (6)3
2021 Classification Models for Medical Data with Interpretative Rules
Zhenyue Qin, Yang Liu 0249
ICONIP (1)3
2020 Are Deep Neural Architectures Losing Information? Invertibility is Indispensable
Yang Liu 0249, Zhenyue Qin, Saeed Anwar, Sabrina B. Caldwell, Tom Gedeon
ICONIP (3)2
2019 Your Eyes Say You're Lying: An Eye Movement Pattern Analysis for Face Familiarity and Deceptive Cognition
abstract
Eye movement patterns reflect human latent internal cognitive activities. We aim to discover eye movement patterns during face recognition under different conditions of information concealment. These conditions include the degrees of face familiarity and deception or not, namely telling the truth when observing familiar and unfamiliar faces, and deceiving in front of familiar and unfamiliar faces. We apply Hidden Markov models with Gaussian emission to generalise regions and trajectories of eye fixation points under the above four conditions. Our results show that both eye movement patterns and eye gaze regions become significantly different during deception compared with truth-telling. We show the feasibility of detecting deception and further cognitive activity classification using eye movement patterns.
Jiaxu Zuo, Tom Gedeon, Zhenyue Qin
IJCNN3
2018 Neural Networks Assist Crowd Predictions in Discerning the Veracity of Emotional Expressions
Zhenyue Qin, Tom Gedeon, Sabrina B. Caldwell
ICONIP (6)1
2018 Artificial Neural Networks Can Distinguish Genuine and Acted Anger by Synthesizing Pupillary Dilation Signals from Different Participants
Zhenyue Qin, Tom Gedeon, Xuanying Zhu
ICONIP (5)1
2018 Detecting the Doubt Effect and Subjective Beliefs Using Neural Networks and Observers' Pupillary Responses
Xuanying Zhu, Zhenyue Qin, Tom Gedeon, Richard Jones 0002, Sabrina B. Caldwell
ICONIP (4)2