VLDB 2026 Research / reviewers in the wild / expert
Yongjie Zhu
dblp:173/4671
· DBLP profile ↗
21ranked-venue papers
10as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARC: Adaptive Resource Coordination for Write Stall Mitigation in LSM-Tree
Guangxun Zhao, Yongjie Zhu, Suhwan Shin, See-hwan Yoo, Jongmoo Choi |
CCGrid | 2 |
| 2025 | MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion UnderstandingabstractMultimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by generalizable attention architecture. Advanced methods predominantly focus on language-centric tuning while less exploring multimodal tokens mixed through attention, posing challenges in high-level tasks that require fine-grained cognition and emotion understanding. In this work, we identify the attention deficit disorder problem in multimodal learning, caused by inconsistent cross-modal attention and layer-by-layer decayed attention activation. To address this, we propose a novel attention mechanism, termed MOdular Duplex Attention (MODA), simultaneously conducting the inner-modal refinement and inter-modal interaction. MODA employs a correct-after-align strategy to effectively decouple modality alignment from cross-layer token mixing. In the alignment phase, tokens are mapped to duplex modality spaces based on the basis vectors, enabling the interaction between visual and language modality. Further, the correctness of attention scores is ensured through adaptive masked attention, which enhances the model's flexibility by allowing customizable masking patterns for different modalities. Extensive experiments on 21 benchmark datasets verify the effectiveness of MODA in perception, cognition, and emotion tasks. Wuyou Xia, Chenxi Zhao 0002, Zhou Yan, Yongjie Zhu, Wenyu Qin, Pengfei Wan 0001, Di Zhang 0026, Jufeng Yang |
ICML | 6 |
| 2025 | VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsabstractUnderstanding and predicting emotions from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis, the intrinsic nature of emotions—characterized by their open-set, dynamic, and context-dependent properties—poses challenge in understanding complex and evolving emotional states with reasonable rationale. To tackle these challenges, we propose a novel affective cues-guided reasoning framework that unifies fundamental attribute perception, expression analysis, and high-level emotional understanding in a stage-wise manner. At the core of our approach is a family of video emotion foundation models (VidEmo), specifically designed for emotion reasoning and instruction-following. These models undergo a two-stage tuning process: first, curriculum emotion learning for injecting emotion knowledge, followed by affective-tree reinforcement learning for emotion reasoning. Moreover, we establish a foundational data infrastructure and introduce a emotion-centric fine-grained dataset (Emo-CFG) consisting of 2.1M diverse instruction-based samples. Emo-CFG includes explainable emotional question-answering, fine-grained captions, and associated rationales, providing essential resources for advancing emotion understanding tasks. Experimental results demonstrate that our approach achieves competitive performance, setting a new milestone across 15 face perception tasks. Yongjie Zhu, Wenyu Qin, Pengfei Wan 0001, Di Zhang 0026, Jufeng Yang |
NeurIPS | 3 |
| 2025 | Auto-Visualization Driven Visual Analytics for Elderly Care InstitutionsabstractTo tackle unstructured, scattered information and user-need gaps in elderly care info management, this study created SeniorAdvisor. It uses BERT-based classification, chart decision trees, and EDA for automated visualization. By blending user needs with EDA insights via iterative visualization and custom coding, it recommends fitting institutions, boosting efficiency and satisfaction. Case studies confirm SeniorAdvisor enhances info management and meets elderly care needs, presenting innovative approaches for elderly care institutions. Yongjie Zhu, Laipan Zuo |
SMC | 1 |
| 2024 | SPLiT: Single Portrait Lighting Estimation via a Tetrad of Face IntrinsicsabstractThis paper proposes a novel pipeline to estimate a non-parametric environment map with high dynamic range from a single human face image. Lighting-independent and -dependent intrinsic images of the face are first estimated separately in a cascaded network. The influence of face geometry on the two lighting-dependent intrinsics, diffuse shading and specular reflection, are further eliminated by distributing the intrinsics pixel-wise onto spherical representations using the surface normal as indices. This results in two representations simulating images of a diffuse sphere and a glossy sphere under the input scene lighting. Taking into account the distinctive nature of light sources and ambient terms, we further introduce a two-stage lighting estimator to predict both accurate and realistic lighting from these two representations. Our model is trained supervisedly on a large-scale and high-quality synthetic face image dataset. We demonstrate that our method allows accurate and detailed lighting estimation and intrinsic decomposition, outperforming state-of-the-art methods both qualitatively and quantitatively on real face images. Yean Cheng, Yongjie Zhu, Si Li 0001, Gang Pan 0001, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Complementary Intrinsics from Neural Radiance Fields and CNNs for Outdoor Scene RelightingabstractRelighting an outdoor scene is challenging due to the diverse illuminations and salient cast shadows. Intrinsic image decomposition on outdoor photo collections could partly solve this problem by weakly supervised labels with albedo and normal consistency from multiview stereo. With neural radiance fields (NeRF), editing the appearance code could produce more realistic results without interpreting the outdoor scene image formation explicitly. This paper proposes to complement the intrinsic estimation from volume rendering using NeRF and from inversing the photometric image formation model using convolutional neural networks (CNNs). The former produces richer and more reliable pseudo labels (cast shadows and sky appearances in addition to albedo and normal) for training the latter to predict interpretable and editable lighting parameters via a single-image prediction pipeline. We demonstrate the advantages of our method for both intrinsic image decomposition and relighting for various real outdoor scenes. Xuanning Cui, Yongjie Zhu, Jiajun Tang 0001, Si Li 0001, Zhaofei Yu, Boxin Shi |
CVPR | 3 |
| 2023 | PFED-AGG: A Personalized Private Federated Learning Aggregation AlgorithmabstractFederated learning is a special kind of distributed machine learning, in which multiple clients work together to solve a machine learning problem with the collaboration of a central server, and the clients only need to upload parameters for server aggregation instead of uploading raw data, so the privacy of the clients can be protected. However, existing research shows that an attacker who obtains the parameters uploaded by the client can reverse the privacy information of the client, and Federated Learning applies a local differential privacy approach to protect the information of the parameters uploaded by the client from being leaked. However, this privacy protection approach assumes the same level of privacy protection for all clients. To the best of our knowledge, there needs to be work that satisfies the personalized privacy needs of clients. To address this problem, we propose a personalized local differential privacy-based federation framework that satisfies the personalized privacy needs of clients and better protects the privacy of clients by making the specific privacy needs of clients inaccessible to the server. We have conducted extensive experiments on six benchmark datasets, and our approach works better and achieves personalized privacy protection compared to the same privacy-preserving method. Yongjie Zhu, Yukun Yan, Qilong Han |
IJCNN | 1 |
| 2023 | BCMask: a finer leaf instance segmentation with bilayer convolution mask
Xingjian Gu, Yongjie Zhu, Shougang Ren, Xiangbo Shu |
Multim. Syst. | 2 |
| 2022 | Estimating Spatially-Varying Lighting in Urban Scenes with Disentangled Representation
Jiajun Tang 0001, Yongjie Zhu, Jun Hoong Chan, Si Li 0001, Boxin Shi |
ECCV (6) | 2 |
| 2022 | AdsCVLR: Commercial Visual-Linguistic Representation Modeling in Sponsored SearchabstractSponsored search advertisements (ads) appear next to search results when consumers look for products and services on search engines. As the fundamental basis of search ads, relevance modeling has attracted increasing attention due to the significant research challenges and tremendous practical value. In this paper, we address the problem of multi-modal modeling in sponsored search, which models the relevance between user query and commercial ads with multi-modal structured information. To solve this problem, we propose a transformer architecture with Ads data on Commercial Visual-Linguistic Representation (AdsCVLR) with contrastive learning that naturally extends the transformer encoder with the complementary multi-modal inputs, serving as a strong aggregator of image-text features. We also make a public advertising dataset, which includes 480K labeled query-ad pairwise data with structured information of image, title, seller, description, and so on. Empirically, we evaluate the AdsCVLR model over the large industry dataset, and the experimental results of online/offline tests show the superiority of our method. Yongjie Zhu, Chunhui Han, Yuefeng Zhan, Bochen Pang, Zhaoju Li, Hao Sun 0015, Si Li 0001, Boxin Shi, Nan Duan 0001, Ruofei Zhang, Liangjie Zhang, Qi Zhang 0066 |
ACM Multimedia | 1 |
| 2022 | Hybrid Face Reflectance, Illumination, and Shape From a Single ImageabstractWe propose HyFRIS-Net to jointly estimate the hybrid reflectance and illumination models, as well as the refined face shape from a single unconstrained face image in a pre-defined texture space. The proposed hybrid reflectance and illumination representation ensure photometric face appearance modeling in both parametric and non-parametric spaces for efficient learning. While forcing the reflectance consistency constraint for the same person and face identity constraint for different persons, our approach recovers an occlusion-free face albedo with disambiguated color from the illumination color. Our network is trained in a self-evolving manner to achieve general applicability on real-world data. We conduct comprehensive qualitative and quantitative evaluations with state-of-the-art methods to demonstrate the advantages of HyFRIS-Net in modeling photo-realistic face albedo, illumination, and shape. Yongjie Zhu, Chen Li 0031, Si Li 0001, Boxin Shi, Yu-Wing Tai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Spatially-Varying Outdoor Lighting Estimation From IntrinsicsabstractWe present SOLID-Net, a neural network for spatially- varying outdoor lighting estimation from a single outdoor image for any 2D pixel location. Previous work has used a unified sky environment map to represent outdoor lighting. Instead, we generate spatially-varying local lighting environment maps by combining global sky environment map with warped image information according to geometric information estimated from intrinsics. As no outdoor dataset with image and local lighting ground truth is readily available, we introduce the SOLID-Img dataset with physically- based rendered images and their corresponding intrinsic and lighting information. We train a deep neural network to regress intrinsic cues with physically-based constraints and use them to conduct global and local lightings estimation. Experiments on both synthetic and real datasets show that SOLID-Net significantly outperforms previous methods. Yongjie Zhu, Yinda Zhang 0001, Si Li 0001, Boxin Shi |
CVPR | 1 |
| 2021 | DeRenderNet: Intrinsic Image Decomposition of Urban Scenes with Shape-(In)dependent Shading RenderingabstractWe propose DeRenderNet, a deep neural network to decompose the albedo and latent lighting, and render shape-(in)dependent shadings, given a single image of an outdoor urban scene, trained in a self-supervised manner. To achieve this goal, we propose to use the albedo maps extracted from scenes in videogames as direct supervision and pre-compute the normal and shadow prior maps based on the depth maps provided as indirect supervision. Compared with state-of-the-art intrinsic image decomposition methods, DeRenderNet produces shadow-free albedo maps with clean details and an accurate prediction of shadows in the shape-independent shading, which is shown to be effective in re-rendering and improving the accuracy of high-level vision tasks for urban scenes. Yongjie Zhu, Jiajun Tang 0001, Si Li 0001, Boxin Shi |
ICCP | 1 |
| 2021 | Disentangling Identifiable Features from Noisy Data with Structured Nonlinear ICAabstractWe introduce a new general identifiable framework for principled disentanglement referred to as Structured Nonlinear Independent Component Analysis (SNICA). Our contribution is to extend the identifiability theory of deep generative models for a very broad class of structured models. While previous works have shown identifiability for specific classes of time-series models, our theorems extend this to more general temporal structures as well as to models with more complex structures such as spatial dependencies. In particular, we establish the major result that identifiability for this framework holds even in the presence of noise of unknown distribution. Finally, as an example of our framework's flexibility, we introduce the first nonlinear ICA model for time-series that combines the following very useful properties: it accounts for both nonstationarity and autocorrelation in a fully unsupervised setting; performs dimensionality reduction; models hidden states; and enables principled estimation and inference by variational maximum-likelihood. Hermanni Hälvä, Sylvain Le Corff, Luc Lehéricy, Jonathan So, Yongjie Zhu, Elisabeth Gassiat, Aapo Hyvärinen |
NeurIPS | 5 |
| 2021 | Altered EEG Oscillatory Brain Networks During Music-Listening in Major DepressionabstractTo examine the electrophysiological underpinnings of the functional networks involved in music listening, previous approaches based on spatial independent component analysis (ICA) have recently been used to ongoing electroencephalography (EEG) and magnetoencephalography (MEG). However, those studies focused on healthy subjects, and failed to examine the group-level comparisons during music listening. Here, we combined group-level spatial Fourier ICA with acoustic feature extraction, to enable group comparisons in frequency-specific brain networks of musical feature processing. It was then applied to healthy subjects and subjects with major depressive disorder (MDD). The music-induced oscillatory brain patterns were determined by permutation correlation analysis between individual time courses of Fourier-ICA components and musical features. We found that (1) three components, including a beta sensorimotor network, a beta auditory network and an alpha medial visual network, were involved in music processing among most healthy subjects; and that (2) one alpha lateral component located in the left angular gyrus was engaged in music perception in most individuals with MDD. The proposed method allowed the statistical group comparison, and we found that: (1) the alpha lateral component was activated more strongly in healthy subjects than in the MDD individuals, and that (2) the derived frequency-dependent networks of musical feature processing seemed to be altered in MDD participants compared to healthy subjects. The proposed pipeline appears to be valuable for studying disrupted brain oscillations in psychiatric disorders during naturalistic paradigms. Yongjie Zhu, Klaus Mathiak, Petri Toiviainen, Tapani Ristaniemi, Fengyu Cong |
Int. J. Neural Syst. | 1 |
| 2021 | Response to Discussion on Y. Zhu, X. Wang, K. Mathiak, P. Toiviainen, T. Ristaniemi, J. Xu, Y. Chang and F. Cong, Altered EEG Oscillatory Brain Networks During Music-Listening in Major Depression, International Journal of Neural Systems, Vol. 31 No. 3 (2021)
Yongjie Zhu, Klaus Mathiak, Petri Toiviainen, Tapani Ristaniemi, Fengyu Cong |
Int. J. Neural Syst. | 1 |
| 2021 | Data fusion of atmospheric ozone remote sensing Lidar according to deep learning
Ru Qiao, Yongjie Zhu, Guibao Wang |
J. Supercomput. | 3 |
| 2020 | Distinct Patterns of Functional Connectivity During the Comprehension of Natural, Narrative SpeechabstractRecent continuous task studies, such as narrative speech comprehension, show that fluctuations in brain functional connectivity (FC) are altered and enhanced compared to the resting state. Here, we characterized the fluctuations in FC during comprehension of speech and time-reversed speech conditions. The correlations of Hilbert envelope of source-level EEG data were used to quantify FC between spatially separate brain regions. A symmetric multivariate leakage correction was applied to address the signal leakage issue before calculating FC. The dynamic FC was estimated based on a sliding time window. Then, principal component analysis (PCA) was performed on individually concatenated and temporally concatenated FC matrices to identify FC patterns. We observed that the mode of FC induced by speech comprehension can be characterized with a single principal component. The condition-specific FC demonstrated decreased correlations between frontal and parietal brain regions and increased correlations between frontal and temporal brain regions. The fluctuations of the condition-specific FC characterized by a shorter time demonstrated that dynamic FC also exhibited condition specificity over time. The FC is dynamically reorganized and FC dynamic pattern varies along a single mode of variation during speech comprehension. The proposed analysis framework seems valuable for studying the reorganization of brain networks during continuous task experiments. Yongjie Zhu, Jia Liu 0050, Tapani Ristaniemi, Fengyu Cong |
Int. J. Neural Syst. | 1 |
| 2019 | Measuring the Task Induced Oscillatory Brain Activity Using Tensor DecompositionabstractThe characterization of dynamic electrophysiological brain activity, which form and dissolve in order to support ongoing cognitive function, is one of the most important goals in neuroscience. Here, we introduce a method with tensor decomposition for measuring the task-induced oscillations in the human brain using electroencephalography (EEG). The time frequency representation of source-reconstructed single-trail EEG data constructed a third-order tensor with three factors of time · trails, frequency and source points. We then used a non-negative Canonical Polyadic decomposition (NCPD) to identify the temporal, spectral and spatial changes in electrophysiological brain activity. We validate this method using both simulation EEG data and real EEG data recorded during a task of irony comprehension. The results demonstrated that proposed method can track dynamics of the temporal-spectral modes of the rhythm in the brain on a timescale commensurate to the task they are undertaking. Yongjie Zhu, Xueqiao Li, Tapani Ristaniemi, Fengyu Cong |
ICASSP | 1 |
| 2018 | Increasing Stability of EEG Components Extraction Using Sparsity Regularized Tensor Decomposition
Deqing Wang 0003, Yongjie Zhu, Petri Toiviainen, Minna Huotilainen, Tapani Ristaniemi, Fengyu Cong |
ISNN | 3 |
| 2016 | Cyclic correntropy and its spectrum in frequency estimation in the presence of impulsive noise
Shengyang Luan, Tianshuang Qiu, Yongjie Zhu |
Signal Process. | 3 |