VLDB 2026 Research / reviewers in the wild / expert
Lu Zhang 0062
dblp:82/10609-62
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2024
0000-0002-2225-9772ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-Modal Traumatic Brain Injury Prognosis via Structure-Aware Field-Wise LearningabstractTraumatic brain injury (TBI) remains a growing significant public health problem and prognosis of outcome is difficult due to the multitude of factors that underlie the heterogeneity of TBI. Prognosis aims to forecast the likely development of the disease and significantly affects patient's recovery and healthcare. Traditionally, TBI prognosis relies on the physician's insights and their empirical knowledge which makes it infeasible for large-scale implementation. Existing methods utilize a single modality (i.e., either clinical data or Computed Tomography scan images) for TBI prognosis, leaving crucial information from multi-modal data largely underexplored. To address this concern, we explore a Multi-modal Structure-aware Field-wise learning (MSF) method that is capable of mining complex correlations between multi-modal data and TBI outcomes for prognosis on a real-world dataset. Specifically, we develop a High-Level Structure-Aware (HSA) module to capture the structure information of the multilayered clinical data. Experimental results on the publicly available TRACK-TBI dataset demonstrate the viability and effectiveness of our proposed method, by achieving the top-3 accuracy of 96.07% and 98.13% for 3-month and 6-month predictions after injury, respectively. Lu Zhang 0062, Zhibin Li 0002, Shekhar Chandra, Fatima A. Nasrallah |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Modeling Multiple Aesthetic Views for Series Photo SelectionabstractNumerous photos are taken in daily life, and sorting them is laborious and time consuming. The large number of similar images exacerbates the difficulty of album management, under this scenario, serial photo selection (SPS) emerges. As an important branch of image aesthetic quality assessment, it focuses on identifying the best image among a series of almost identical photos. Currently, most existing SPS methods focus only on extracting features from the original image, while neglecting the fact that multiple views of the image can provide much more detailed aesthetic information. In this article, we propose a Siamese network structure called SPSNet to enhance the representation learning of multi-view features by acquiring the depth, generic, and handcrafted features of images. In specific, we implement a parallel structure to extract deep and shallow features, fusing local and global representations at different resolutions interactively. The aggregation of multiple views of image via a self-attentive module with adaptive weights enables the model to discriminate the importance of each view. Moreover, we employ a graph neural network to construct the relationships among the multi-view features. Our proposed method, which is trained by a Siamese network, can effectively distinguish the nuances of similar images, and thus, select the best one from a series of almost identical photos. Extensive experiments conducted on the aesthetic dataset demonstrate that our method outperforms other state-of-the-art SPS methods, which achieves the 75.36% accuracy on the Phototriage dataset. Besides, our model is up to 3.04% better than the baseline methods in terms of the average accuracy. Yongshun Gong, Lu Zhang 0062, Jian Zhang 0002, Liqiang Nie, Yilong Yin |
IEEE Trans. Multim. | 3 |
| 2023 | Exploiting Field Dependencies for Learning on Categorical DataabstractTraditional approaches for learning on categorical data underexploit the dependencies between columns (a.k.a. fields) in a dataset because they rely on the embedding of data points driven alone by the classification/regression loss. In contrast, we propose a novel method for learning on categorical data with the goal of exploiting dependencies between fields. Instead of modelling statistics of features globally (i.e., by the covariance matrix of features), we learn a global field dependency matrix that captures dependencies between fields and then we refine the global field dependency matrix at the instance-wise level with different weights (so-called local dependency modelling) w.r.t. each field to improve the modelling of the field dependencies. Our algorithm exploits the meta-learning paradigm, i.e., the dependency matrices are refined in the inner loop of the meta-learning algorithm without the use of labels, whereas the outer loop intertwines the updates of the embedding matrix (the matrix performing projection) and global dependency matrix in a supervised fashion (with the use of labels). Our method is simple yet it outperforms several state-of-the-art methods on six popular dataset benchmarks. Detailed ablation studies provide additional insights into our method. Zhibin Li 0002, Piotr Koniusz, Lu Zhang 0062, Daniel Edward Pagendam, Peyman Moghadam |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Series Photo Selection via Multi-View Graph LearningabstractSeries photo selection (SPS) is an important branch of the image aesthetics quality assessment, which focuses on finding the best one from a series of nearly identical photos. While a great progress has been observed, most of the existing SPS approaches concentrate solely on extracting features from the original image, neglecting that multiple views, e.g, saturation level, color histogram and depth of field of the image, will be of benefit to successfully reflecting the subtle aesthetic changes. Taken multi-view into consideration, we leverage a graph neural network to construct the relationships between multi-view features. Besides, multiple views are aggregated with an adaptive-weight self-attention module to verify the significance of each view. Finally, a siamese network is proposed to select the best one from a series of nearly identical photos. Experimental results demonstrate that our model accomplish the highest success rates compared with competitive methods. Lu Zhang 0062, Yongshun Gong, Jian Zhang 0002, Xiushan Nie, Yilong Yin |
ICME | 2 |
| 2022 | Multimodal Marketing Intent Analysis for Effective Targeted AdvertisingabstractPeople’s daily information sharing and acquisition through the Internet has become more and more popular. The comprehensive multimodal marketing advertorial generated by ‘We Media’ accounts besides the normal social news is gaining its importance on social media platforms. In order to achieve effective advertising, the marketing intent understanding is a key step towards generating targeted advertising strategies (push advertorials to specific people at a specific time). However, advertorials in real are usually designed to pretend as normal social news with a wide range of contents. This poses big challenges to the platforms on accurately recognizing and analyzing the marketing intents behind the advertorials. As a pioneering study, we address this new problem of multimodal-based marketing intent analysis and answer three core questions: (1) does a piece of social news contain marketing intent? (2) what is the topic of marketing intent? (3) what is the extent of marketing intent? Towards this end, we propose a novel Multimodal-based Marketing Intent Analysis scheme (MMIA) to estimate the marketing intent embedded in the multimodal contents. Specifically, a novel supervised neural autoregressive model (SmiDocNADE) is proposed to enhance the discriminative capacity of the learned hidden features so that a single system is capable of solving the three questions. In order to effectively model inter-correlations between images and text in advertorials, we fuse multimodal data and extract features by Graph Convolution Networks as an enhancement to SmiDocNADE. The extensive evaluations demonstrate the advantages of our proposed system in multimodal-based marketing intent analysis from multiple aspects. Lu Zhang 0062, Jialie Shen 0001, Jian Zhang 0002, Jingsong Xu, Zhibin Li 0002, Yazhou Yao, Litao Yu |
IEEE Trans. Multim. | 1 |
| 2022 | Unsupervised Image and Text Fusion for Travel Information EnhancementabstractWith the explosive growth of the shared information on social media platforms, people are increasingly interested in sharing and making their travel plans by referring to others’ travel experiences. However, different social media sources render the heterogeneity of these valuable data, bringing difficulties for data collection and fusion. Thus, facing massive information online, one of the biggest challenges to enhance travel information is how to integrate and match these multi-source data without clear labels. In this paper, we propose an unsupervised method to fuse and match images and travelogues. We first use the three textual components (title, tag, and description) of the descriptive texts of images as three criteria to embed travelogues and the descriptive texts of images, and further introduce images into our method by joint embedding texts and images. Finally, a multiple kernel clustering approach is adopted for matching travelogues and images. Extensive experiments conducted on the real dataset crawled from two websites (Flickr and TripAdvisor) demonstrate the effectiveness and robustness of our proposed method. Lu Zhang 0062, Jingsong Xu, Yongshun Gong, Litao Yu, Jian Zhang 0002, Jialie Shen 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Incorporating Multimodal Cues for Advertorial DiscoveryabstractCommercial advertorials shared on websites are usually designed to pretend as normal social news for commercial benefits. The analysis of the commercial intents embedded in advertorials can greatly help media platforms personalize content. However, commercial intents are not only concealed in news texts but also conveyed by news images explicitly or implicitly. Consequently, how to effectively extract and incorporate the crucial cues of multiple modalities has been emerging as an important but challenging problem. Motivated by this observation, we propose a framework Multimodal Advertorial Discovery Model (MADM) to estimate the commercial intents embedded in the multimodal social news. Specifically, a novel Cross-graph Fusion (CGF) strategy is developed to achieve a soft assignment to incorporate images and text and generate comprehensive multimodal representations. The extensive evaluations demonstrate the superiority of our proposed system in multimodal-based advertorial detection and analysis. Lu Zhang 0062, Jian Zhang 0002, Jialie Shen 0001, Jingsong Xu, Zhibin Li 0002, Litao Yu |
ICME | 1 |
| 2020 | Towards Better Graph Representation: Two-Branch Collaborative Graph Neural Networks For Multimodal Marketing Intention DetectionabstractInspired by the fact that spreading and collecting information through the Internet becomes the norm, more and more people choose to post for-profit contents (images and texts) in social networks. Due to the difficulty of network censors, malicious marketing may be capable of harming the society. Therefore, it is meaningful to detect marketing intentions online automatically. However, gaps between multimodal data make it difficult to fuse images and texts for content marketing detection. To this end, this paper proposes Two-Branch Collaborative Graph Neural Networks to collaboratively represent multimodal data by Graph Convolution Networks (GCNs) in an end-to-end fashion. We first separately embed groups of images and texts by GCNs layers from two views and further adopt the proposed multimodal fusion strategy to learn the graph representation collaboratively. Experimental results demonstrate that our proposed method achieves superior graph classification performance for marketing intention detection. Lu Zhang 0062, Jian Zhang 0002, Zhibin Li 0002, Jingsong Xu |
ICME | 1 |