VLDB 2026 Research / reviewers in the wild / expert
Zhi Zhang 0004
dblp:36/5594-4
· DBLP profile ↗
17ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0001-7017-2375ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMsabstractLarge Language Models (LLMs) have demonstrated remarkable proficiency in diverse tasks. This success raises a fundamental question in machine composition: Can symbolic music be considered a special form of language that can be jointly modeled with natural language for composition tasks? Recent studies validate that symbolic music can be modeled as a human language, yet composing structured music from partial symbolic inputs through natural language interaction remains underexplored. Even LLMs struggle to generate structurally coherent compositions in such hybrid input-output scenarios, highlighting a fundamental gap that calls for a domain-specific learning paradigm. To this end, we propose Inspiration-to-Structure (IoS), a cognitively inspired framework that enables LLMs to generate structured musical sections from melodic ideas. IoS employs a three-phase process—semantic, structural, and collaborative cognition—and is supported by two key components: (1) a new dataset and construction protocol called Structured Triplet Data (STD), and (2) a training method, Dual-Instance Structural Contrastive Optimization (DiSCO), designed to enhance structural awareness. Experiments show that IoS improves structural coherence by 47.8% and artistic creativity by 21.8% compared to conventional language modeling paradigm, supervised fine-tuning, and even enables smaller LLMs to surpass larger LLMs. These results suggest that symbolic music, while language-like, demands specialized modeling beyond standard language modeling paradigms. IoS enables LLMs to transform music theory knowledge into structured composition, empowering users to compose music interactively via language and advancing toward general creative AI. Zhejing Hu, Yan Liu 0004, Zhi Zhang 0004, Aiwei Zhang, Shenghua Zhong, Bruce X. B. Yu, Gong Chen 0006 |
AAAI | 3 |
| 2026 | From Pixels to Logic: A Perception-Reasoning Decomposition Framework for Open-World Referring Expression ComprehensionabstractRecent advances in Referring Expression Comprehension (REC) have been largely driven by supervised learning on curated datasets, where each expression is assumed to refer to exactly one known object. However, such assumptions rarely hold in real-world scenarios, where expressions can refer to multiple objects, fail to refer to any, or involve novel categories and complex semantics. These challenges define the task of open-world REC, which demands robust generalization and structured reasoning beyond the scope of traditional REC methods. In this work, we introduce a novel, training-free framework that decouples visual perception from linguistic reasoning to address open-world REC. Our method first transforms the visual scene into a rich textual representation using an open-vocabulary multimodal perception module. It then employs a reasoning language model to interpret the referring expression and perform explicit logical inference over the perceived scene, enabling transparent decision-making and strong generalization in open-world scenarios. Experiments on three standard REC benchmarks as well as two more challenging ones, gRefCOCO and D³, demonstrate that our framework achieves highly competitive zero-shot performance, often surpassing supervised baselines. Shenghua Zhong, Zhi Zhang 0004, Yan Liu 0004 |
AAAI | 3 |
| 2026 | MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AIabstractAdolescence is marked by strong creative impulses but limited strategies for structured expression, often leading to frustration or disengagement. While generative AI lowers technical barriers and delivers efficient outputs, its role in fostering adolescents’ expressive growth has been overlooked. We propose MusicScaffold, an adolescent-centered framework that transforms classical AI roles from broad conceptualizations into stage-specific, actionable developmental scaffolds designed to make expressive strategies transparent and learnable and to support adolescents in mastering creative expression. In a four-week study with middle school students (ages 12–14), MusicScaffold enhanced cognitive specificity, behavioral regulation, and affective autonomy in music creation. By reframing generative AI as a scaffold rather than a generator, this work bridges the machine efficiency of generative systems with human growth in adolescent creativity education. Zhejing Hu, Yan Liu 0004, Zhi Zhang 0004, Gong Chen 0006, Bruce X. B. Yu, Jiannong Cao 0001 |
CHI | 3 |
| 2026 | SEmotion: Knowledge-Guided Sigmoid-Constrained Network for EEG-Based Emotion RecognitionabstractElectroencephalography (EEG)-based affective computing has achieved significant progress due to the rapidly developing learning model. However, compared to computer vision or natural language processing-related tasks, EEG data presents challenges such as a low signal-to-noise ratio, small sample size, and nonstationary properties. These factors may affect the model’s performance on cross-domain tasks, suggesting that further improvements are necessary. Existing models aim to minimize intraclass distance and maximize interclass distance to achieve an optimal solution. This can potentially have a negative impact on the model’s generalization ability because they ignore the existence of low-quality training samples that may not be suitable for strict optimization. This article presents a method for determining sample quality based on the guidance of knowledge from emotional neuroscience. The method differentiates between high- and low-quality training samples and designs a corresponding loss function to impose intraclass and interclass constraints on hyperspherical manifolds based on the quality of the samples. Therefore, it can be argued that our proposed method achieves a better balance between reducing the intraclass distance of high-quality samples and preventing the overfitting of low-quality ones. This could potentially help build a more robust EEG emotion recognition model. Numerous experiments are conducted on the SJTU Emotion EEG Dataset (SEED) and SJTU Emotion EEG Dataset (IV) (SEED-IV) datasets under cross-subject and cross-session scenarios, which show the superior performance of our proposed method and its higher recognition accuracy against adversarial attacks. Wenjie Rao, Shenghua Zhong, Zhi Zhang 0004, Yan Liu 0004 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2026 | Mutual Generation for Cross-Domain Challenge in Stroke Patients' Motor Imagery Classification and Functional Recovery PredictionabstractThe accumulating body of research indicates that Motor Imagery (MI)-BCIs have the potential to enhance the quality of life for individuals with disabilities and to advance our understanding of brain function and rehabilitation strategies. Among these diseases, stroke is the leading cause of long-term motor disability across the globe, thereby underscoring the need for innovative rehabilitation strategies, such as MI-BCI technologies. In contrast with these expectations, the majority of existing research is built upon data obtained from healthy subjects. The construction of effective classification models for Motor Imagery tasks in patients with brain diseases, particularly stroke, remains a significant challenge. The lateralization of the left and right hemispheres is more pronounced in patients who have suffered a stroke than in healthy individuals. Moreover, the specific locations of lesions and the regions of influence result in significant variations in the electroencephalogram (EEG) data of patients with different hemiplegic sides. This paper explores the potential of generative models in addressing the issue of domain differences arising from different hemiplegic sides EEG data. Furthermore, this paper circumvents the potential adverse effects of rigorous optimization of low-quality samples on model performance through the utilization of label softening algorithm. Two MI-EEG datasets of stroke patients performing Motor Imagery tasks are used to validate our method. In comparison to both classical machine learning methods and those state-of-the-art models for MI classification, the classification model in this paper achieves a noticeable performance improvement in different data partitioning strategies, including subject-dependent and subject-independent scenarios. Each sub-module, and each designed loss function, contributes to the final performance growth. In addition, this paper also investigates the potential of the proposed framework for predicting a patient's level of functional recovery. Our findings indicate that the addition of a prediction layer to the proposed model enables the accurate prediction of functional recovery level in stroke patients. Rongrong Lu, Wenchang Deng, Tianhao Gao, Songhua Huang, Zhi Zhang 0004, Yan Liu 0004, Shenghua Zhong |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Mixture of Knowledge Minigraph Agents for Literature Review GenerationabstractLiterature reviews play a crucial role in scientific research for understanding the current state of research, identifying gaps, and guiding future studies on specific topics. However, the process of conducting a comprehensive literature review is yet time-consuming. This paper proposes a novel framework, collaborative knowledge minigraph agents (CKMAs), to automate scholarly literature reviews. A novel prompt-based algorithm, the knowledge minigraph construction agent (KMCA), is designed to identify relations between concepts from academic literature and automatically constructs knowledge minigraphs. By leveraging the capabilities of large language models on constructed knowledge minigraphs, the multiple path summarization agent (MPSA) efficiently organizes concepts and relations from different viewpoints to generate literature review paragraphs. We evaluate CKMAs on three benchmark datasets. Experimental results show the effectiveness of the proposed method, further revealing promising applications of LLMs in scientific research. Zhi Zhang 0004, Yan Liu 0004, Shenghua Zhong, Gong Chen 0006, Yu Yang 0012, Jiannong Cao 0001 |
AAAI | 1 |
| 2025 | Attention-guided universal adversarial perturbations for EEG-based brain-computer interfaces
Shenghua Zhong, Sijia Zhao, Zhijiao Xiao, Zhi Zhang 0004, Yan Liu 0004 |
Expert Syst. Appl. | 4 |
| 2025 | Intellectual Property Protection for Deep Models: Pioneering Cross-Domain Fingerprinting SolutionsabstractThe high cost of developing high-performance deep models highlights their value as intellectual property for creators. However, it is important to consider the potential risks of theft. Although various techniques have been developed to protect the intellectual property of deep models, there is still room for improvement in terms of efficiency, comprehensiveness, and generalization. Compared with the intrusiveness of watermarking methods, fingerprinting methods do not affect the training process of the source model. Consequently, this paper proposes a fingerprinting method to address the paucity of attempts in fingerprinting methods for model protection. Our method consists of two efficient algorithms for generating fingerprinting samples, where the first one possesses the advantage of efficiency, while the second one is better in terms of robustness. The first algorithm takes a comprehensive approach to modeling the fingerprint of the deep model. The generated samples are distributed within the stable region and near the decision boundary of the model, taking into account both the duality and the conviction factors. Then, a heuristic sample perturbation algorithm is introduced, which generates a fingerprint with solid stability and generalization across multiple domains. The two algorithms proposed in this paper have been shown to be capable of withstanding attacks on intellectual property removal, detection, and evasion. They also show some advantages in terms of efficiency. In addition, the proposed method is the first to apply fingerprinting techniques in a cross-domain context. Shenghua Zhong, Zhi Zhang 0004, Yan Liu 0004 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Beyond Mimicking Under-Represented Emotions: Deep Data Augmentation with Emotional Subspace Constraints for EEG-Based Emotion RecognitionabstractIn recent years, using Electroencephalography (EEG) to recognize emotions has garnered considerable attention. Despite advancements, limited EEG data restricts its potential. Thus, Generative Adversarial Networks (GANs) are proposed to mimic the observed distributions and generate EEG data. However, for imbalanced datasets, GANs struggle to produce reliable augmentations for under-represented minority emotions by merely mimicking them. Thus, we introduce Emotional Subspace Constrained Generative Adversarial Networks (ESC-GAN) as an alternative to existing frameworks. We first propose the EEG editing paradigm, editing reference EEG signals from well-represented to under-represented emotional subspaces. Then, we introduce diversity-aware and boundary-aware losses to constrain the augmented subspace. Here, the diversity-aware loss encourages a diverse emotional subspace by enlarging the sample difference, while boundary-aware loss constrains the augmented subspace near the decision boundary where recognition models can be vulnerable. Experiments show ESC-GAN boosts emotion recognition performance on benchmark datasets, DEAP, AMIGOS, and SEED, while protecting against potential adversarial attacks. Finally, the proposed method opens new avenues for editing EEG signals under emotional subspace constraints, facilitating unbiased and secure EEG data augmentation. Zhi Zhang 0004, Shenghua Zhong, Yan Liu 0004 |
AAAI | 1 |
| 2024 | TorchEEGEMO: A deep learning toolbox towards EEG-based emotion recognition
Zhi Zhang 0004, Shenghua Zhong, Yan Liu 0004 |
Expert Syst. Appl. | 1 |
| 2023 | Knowledge-guided Network Pruning for EEG-based Emotion RecognitionabstractWith the development of deep learning in EEG-related tasks, the complexity of learning models has gradually increased. These complex models often result in long inference times, high energy consumption, and an increased risk of overfitting. Therefore, model compression has become an important consideration. Although some EEG models have used lightweight techniques, such as separable convolution, no existing work has directly attempted to compress EEG models to reduce their complexity. In this paper, we integrate neuroscience knowledge into EEG model pruning recovery, and innovatively propose two loss functions in the learning process, the knowledge-guided region-wise loss that enforces the classification evidence consistent with the importance of the prefrontal lobe, and the knowledge-guided sample-wise loss that constrains the learning process by distinguishing the importance of different samples. Wenjie Rao, Shenghua Zhong, Zhi Zhang 0004, Yan Liu 0004 |
BIBM | 3 |
| 2023 | GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular Videoabstract3D human pose estimation has been researched for decades with promising fruits. 3D human pose lifting is one of the promising research directions toward the task where both estimated pose and ground truth pose data are used for training. Existing pose lifting works mainly focus on improving the performance of estimated pose, but they usually underperform when testing on the ground truth pose data. We observe that the performance of the estimated pose can be easily improved by preparing good quality 2D pose, such as fine-tuning the 2D pose or using advanced 2D pose detectors. As such, we concentrate on improving the 3D human pose lifting via ground truth data for the future improvement of more quality estimated pose data. Towards this goal, a simple yet effective model called Global-local Adaptive Graph Convolutional Network (GLA-GCN) is proposed in this work. Our GLA-GCN globally models the spatiotemporal structure via a graph representation and backtraces local joint features for 3D human pose estimation via individually connected layers. To validate our model design, we conduct extensive experiments on three benchmark datasets: Human3.6M, HumanEva-I, and MPI-INF-3DHP. Experimental results show that our GLA-GCN1implemented with ground truth 2D poses significantly outperforms state-of-the-art methods (e.g., up to 3%, 17%, and 14% error reductions on Human3.6M, HumanEva-I, and MPI-INF-3DHP, respectively). Bruce X. B. Yu, Zhi Zhang 0004, Yongxu Liu 0003, Shenghua Zhong, Yan Liu 0004, Chang Wen Chen |
ICCV | 2 |
| 2023 | GANSER: A Self-Supervised Data Augmentation Framework for EEG-Based Emotion RecognitionabstractElectroencephalography (EEG)-based affective computing has a scarcity problem. As a result, it is difficult to build effective, highly accurate and stable models using machine learning algorithms, especially deep learning models. Data augmentation has recently shown performance improvements in deep learning models with increased accuracy, stability and reduced overfitting. In this paper, we propose a novel data augmentation framework, named the generative adversarial network-based self-supervised data augmentation (GANSER). As the first to combine adversarial training with self-supervised learning for EEG-based emotion recognition, the proposed framework generates high-quality and high-diversity simulated EEG samples. In particular, we utilize adversarial training to learn an EEG generator and force the generated EEG signals to approximate the distribution of real samples, ensuring the quality of the augmented samples. A transformation operation is employed to mask parts of the EEG signals and force the generator to synthesize potential EEG signals based on the unmasked parts to produce a wide variety of samples. A masking possibility during transformation is introduced as prior knowledge to generalize the classifier for the augmented sample space. Finally, numerous experiments demonstrate that our proposed method can improve emotion recognition with an increase in performance and achieve state-of-the-art results. Zhi Zhang 0004, Yan Liu 0004, Shenghua Zhong |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Detecting abnormality with separated foreground and background: Mutual Generative Adversarial Networks for video abnormal event detection
Zhi Zhang 0004, Shenghua Zhong, Ahmed Fares, Yan Liu 0004 |
Comput. Vis. Image Underst. | 1 |
| 2021 | Video Abnormal Event Detection Via Context Cueing Generative Adversarial NetworkabstractVideo abnormal event detection is a significant research subject in intelligent surveillance, referring to identifying events that do not conform to expected behaviors. It is a challenging task due to the diversity of video content and the various definition of anomaly depending on contexts. How to grasp the nuances of anomalies from diverse video content under associated contexts? In this paper, we propose a context cueing generative adversarial network (CC-GAN) to scop attention to specific regions of interest (ROIs). Different from existing work, the designed framework takes advantage of the context cues to preserve the associations between ROIs and contexts. By intrinsic learning of static spatial layouts and apparent velocities, context cues can help to understand ROIs. Optimized in a crop-inpainting fashion, CC-GAN learns to generate ROIs matching given spatio-temporal context and discriminate the distorted or incongruous ones, corresponding to abnormal scenes. Compared with the state-of-the-art method and other abnormal event detection approaches, the proposed framework demonstrates its effectiveness, reliability, and good generalization ability across various scenes and events. Zhi Zhang 0004, Shenghua Zhong, Yan Liu 0004 |
ICME | 1 |
| 2021 | Steganographer detection via a similarity accumulation graph convolutional network
Zhi Zhang 0004, Mingjie Zheng 0002, Shenghua Zhong, Yan Liu 0004 |
Neural Networks | 1 |
| 2020 | Steganographer Detection Via Enhancement-Aware Graph Convolutional NetworkabstractSteganographer detection aims to find guilty users who hide secret information in images or other multimedia data in the social network. In existing work, the distances between users are calculated based on the distributions of all images shared by the corresponding users, then users lying an abnormal distance from others are detected as guilty users. This flattened method is difficult to grasp the nuances of the guilty and innocent users. In this paper, we are the first to propose a graph-based deep learning framework for steganographer detection. The proposed Enhancement-aware Graph Convolutional Network (EGCN) represents each user as a weighted complete graph and learns to highlight the differences between guilty users and innocent users based on the structured graph. Compared with the state-of-the-art method and other representative graph-based models, the proposed framework demonstrates its effectiveness across image domains, and even under the context of large-scale social media scenario. Zhi Zhang 0004, Mingjie Zheng 0002, Shenghua Zhong, Yan Liu 0004 |
ICME | 1 |