VLDB 2026 Research / reviewers in the wild / expert
Zehang Lin
dblp:206/7306
· DBLP profile ↗
27ranked-venue papers
7as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Robust Cross-Platform News Event Detection via Self-Supervised Modality ComplementationabstractMultimodal news event detection aims to identify and categorize significant events across media platforms using multimodal data. Previous work was limited to a single platform and assumed complete multimodal data. In this paper, we explore a novel task of cross-platform multimodal news event detection to enhance model generalization for cross-platform scenarios. We propose a Self-Supervised Modality Complementation (SSMC) method to tackle the challenges of incomplete modalities and platform heterogeneity presented in this task. Specifically, a Missing Data Complementation (MDC) module is designed to overcome the limitations caused by incomplete modalities. It employs a separation mechanism that distinguishes between modality-specific and modality-shared features across all modalities, allowing for the augmentation of missing modalities with information extracted from common features. Meanwhile, a Multimodal Self-Learning (MSL) module addresses platform heterogeneity by extracting pseudo labels from the target platform's multimodal views and incorporating a self-penalization mechanism to reduce reliance on low-confidence labels. Additionally, we collect a comprehensive cross-platform news event detection (CNED) dataset encompassing 37,711 multimodal samples from Twitter, Flickr, and online news media, covering 40 public news events verified by Wikipedia. Extensive experiments on the CNED dataset demonstrate the superior performance of our proposed method. The dataset is available athttps://github.com/RetrainIt/CNED. Zehang Lin, Zhenguo Yang, Qing Li 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | MixFormer: An End-To-End UAV Image-Matching Network Based on Mixed Attention MechanismabstractAchieving autonomous navigation and positioning for UAV through computer vision technology has emerged as a key focus and hot topic in current research. The precise matching of UAV images is crucial for realizing this functionality. However, the intrinsic difference between satellite images and UAV images resulted in the sub-optimal performance of traditional matching algorithms. Although the depth-local feature matching of linear Transformers has demonstrated superior hierarchical matching capabilities in UAV image matching, the lack of fine local interactions between pixel tokens restricts their ability to extract highly accurate and localized correspondences. To address the limitations of existing linear Transformer matching algorithms, this paper proposes a novel method called MixFormer, a Mixed Attention Matching Transformer for UAV image feature matching. The algorithm comprises two main aspects: firstly, it introduces a hierarchical feature matching attention framework that operates attention at different scales, integrating global attention and local attention to achieve global context awareness and fine-grained matching, thereby enabling efficient and compact implementation. Secondly, it introduces a feature transformation module for transitional Convolutional Neural Network (CNN) to extract features with full receptive fields. Experimental results on benchmark tests confirm the effectiveness and efficiency of the proposed technique. Haitao Jia, Shiyi Xu, Zehang Lin, Wenbo Xu 0004 |
IGARSS | 4 |
| 2024 | Generalized News Event Discovery via Dynamic Augmentation and Entropy OptimizationabstractNews event discovery refers to the identification and detection of news events using multimodal data on social media. Currently, most works assume that the test set consists of known events. However, in real life, the emergence of new events is more frequent, which invalidates this assumption. In this paper, we propose a Dynamic Augmentation and Entropy Optimization (DAEO) model to address the scenario of generalized news event discovery, which requires the model to not only identify known events but also distinguish various new events. Specifically, we first introduce a multimodal augmentation module, which utilizes adversarial learning to enhance the multimodal representation capability. Secondly, we design an adaptive entropy optimization strategy combined with a self-distillation method, which uses multi-view pseudo-label consistency to improve the model's performance on both known and new events. In addition, we collect a multimodal news event discovery (MNED) dataset of 161,350 samples annotated with 66 real-world events. Extensive experimental results on the MNED dataset demonstrate the effectiveness of our proposed method. Our dataset is available on https://github.com/RetrainIt/MNED. Zehang Lin, Jiayuan Xie, Zhenguo Yang, Yi Yu 0001, Qing Li 0001 |
ACM Multimedia | 1 |
| 2024 | Multi-modal news event detection with external knowledge
Zehang Lin, Jiayuan Xie, Qing Li 0001 |
Inf. Process. Manag. | 1 |
| 2024 | Context-Aware Dynamic Word Embeddings for Aspect Term ExtractionabstractThe aspect term extraction (ATE) task aims to extract aspect terms describing a part or an attribute of a product from review sentences. Most existing works rely on either general or domain embedding to address this problem. Despite the promising results, the importance of general and domain embeddings is still ignored by most methods, resulting in degraded performances. Besides, word embedding is also related to downstream tasks, and how to regularize word embeddings to capture context-aware information is an unresolved problem. To solve these issues, we first propose context-aware dynamic word embedding (CDWE), which could simultaneously consider general meanings, domain-specific meanings, and the context information of words. Based on CDWE, we propose an attention-based convolution neural network, called ADWE-CNN for ATE, which could adaptively capture the previous meanings of words by utilizing an attention mechanism to assign different importance to the respective embeddings. The experimental results show that ADWE-CNN achieves a comparable performance with the state-of-the-art approaches. Various ablation studies have been conducted to explore the benefit of each component. Our code is publicly available athttp://github.com/xiejiajia2018/ADWE-CNN. Jiayuan Xie, Yi Cai 0001, Zehang Lin, Ho-fung Leung, Qing Li 0001, Tat-Seng Chua |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Towards Temporal Event Detection: A Dataset, Benchmarks and ChallengesabstractThe availability of datasets annotated with verified events by the public is a necessary prerequisite for unleashing the potential of multimodal deep learning for news event detection. Publicly available datasets are either incompletely annotated due to expensive cost, or ignore the verifiability of event labels, which are susceptible to bias and errors introduced by a limited number of annotators. In this article, we provide a YouTube dataset labelled by real-world news events that can be verified by Wikipedia-like crowd sourcing platforms, with the target of advancing temporal event detection. The events in our dataset cover a wide range of event topics including public security, natural disasters, elections, sports, and entertainment events, etc. In the dataset, each sample is labelled with real-world event that is verifiable by the public. We extensively evaluate the performance of 13 state-of-the-art algorithms on our dataset in a temporal manner, involving the multiple relationships between training and testing event labels, and provide a thorough analysis of the findings. Zhenguo Yang, Zhuopan Yang, Zhiwei Guo 0001, Zehang Lin, Haizhong Zhu, Qing Li 0001, Wenyin Liu |
IEEE Trans. Multim. | 4 |
| 2023 | Graph convolutional network for difficulty-controllable visual question generation
Jiayuan Xie, Yi Cai 0001, Zehang Lin, Qing Li 0001, Tao Wang 0036 |
World Wide Web (WWW) | 4 |
| 2022 | Towards Temporal Event Detection: A Dataset and BenchmarksabstractIn this paper, we present a new dataset with the target of advancing temporal real-world event detection from video-sharing social media. The collected temporal event dataset (TED) has certain characteristics. 1) Certainty of event labels that can be recognized as real-world happening. 2) Wide range of event categories covering public security, natural disasters, elections, sports and entertainment events, etc. 3) High overlap among event topics ensuring the difficulties in distinguishing the labels. 4) Multiple data modalities involving textual, acoustic, and visual information, etc. More specifically, two scenarios are defined based on the close or open domains, i.e., temporal event detection without/with new events, denoted as TED-W and TED-N, respectively. For comparisons, a few benchmarks are investigated on the two scenarios with the dataset. Zhenguo Yang, Haizhong Zhu, Zhiwei Guo 0001, Zehang Lin, Qing Li 0001, Wenyin Liu |
ICME | 5 |
| 2022 | Intra-class low-rank regularization for supervised and semi-supervised cross-modal retrieval
Peipei Kang, Zehang Lin, Zhenguo Yang, Xiaozhao Fang, Alexander M. Bronstein, Qing Li 0001, Wenyin Liu |
Appl. Intell. | 2 |
| 2022 | Deep fused two-step cross-modal hashing with multiple semantic supervision
Peipei Kang, Zehang Lin, Zhenguo Yang, Alexander M. Bronstein, Qing Li 0001, Wenyin Liu |
Multim. Tools Appl. | 2 |
| 2022 | Learning semantic alignment from image for text-guided image inpainting
Yucheng Xie, Zehang Lin, Zhenguo Yang, Xingcai Wu, Xudong Mao, Qing Li 0001, Wenyin Liu |
Vis. Comput. | 2 |
| 2021 | A Double Phases Generation Network for Yes or No Question Generation (Student Abstract)abstractThis paper aims to solve the task of generating yes or no questions, which generates yes/no questions based on given passages. These questions can be used for evaluation automatically. We propose a double phases generation network that can identify specific phrases related to facts from the input passage and use them as auxiliary information for generation. Specifically, the 1st-phase prediction uses the extracted phrases as assistance to generate an initial question. Then, the 2nd-phase prediction utilizes an attention network to focus on the relevant phrases related to the initial question in the passage to generate questions that are more relevant to the specific facts contained in the initial question. Extensive experiments we performed on BoolQ dataset demonstrate the effectiveness of our framework. Jiayuan Xie, Yi Cai 0001, Zehang Lin |
AAAI | 4 |
| 2021 | Temporal Graph Convolutional Network for Multimodal Sentiment AnalysisabstractIn this paper, we propose a temporal graph convolutional network (TGCN) to recognize the sentiments from language (textual), acoustic, and visual (facial expressions) modalities. TGCN constructs a modality-specific graph whose nodes are the aligned segments in the multimodal utterances and edges are weighted according to the distances between their features, in order to learn node embeddings with sequential semantics underlying the utterances. In particular, we use positional encoding by interleaving sine and cosine embedding to encode the positions of the segments in the utterances into their features. Given the modality-specific embeddings of the segments in utterances, we create an attention mechanism corresponding to the segments to capture the sentiment-related ones and obtain the unified embeddings of utterances. Furthermore, we fuse the attended embeddings of the multimodel utterances and conduct the attention to capture their interaction. Finally, the fused embeddings together with their raw features are concatenated together for sentiment predictions. Extensive experiments on three publicly available datasets show that TGCN outperforms the state-of-the-art methods. Zehang Lin, Zhenguo Yang, Wenyin Liu |
ICMI | 2 |
| 2021 | DepthGrasp: Depth Completion of Transparent Objects Using Self-Attentive Adversarial Network with Spectral Residual for GraspingabstractTransparent objects with unique visual properties often make depth cameras fail to scan their reflective and refractive surfaces. Recent studies on depth completion of transparent objects have leveraged a linear system based on the geometric constraints to predict the missing depth, which is hard to be employed in an end-to-end framework and achieve joint optimization. In this paper, we propose DepthGrasp - a deep learning approach for depth completion of transparent objects from a raw RGB-D image. More specifically, we use a generative adversarial network, which utilizes the generator to complete the depth maps by predicting the missing or inaccurate depth values, and use discriminator to guide the completed depth maps against the groundtruth. In the generator, we devise spectral residual blocks (SRB) with spectral normalization for network stability, and residual block to pass the attention map in order to capture the structure information and distinguish the geometric shape of transparent objects. In the discriminator, we use a patch-based convolutional network to adapt the data distributions of the predicted depth maps according to groundtruth. Extensive experiments conducted on ClearGrasp dataset show the effectiveness and generalization of the DepthGrasp for depth completion, and the deployed robotic picking system makes significant improvement on the performance of grasping on transparent objects. Yingjie Tang, Zhenguo Yang, Zehang Lin, Qing Li 0001, Wenyin Liu |
IROS | 4 |
| 2021 | Event Cube for Suicidal Event Analysis: A Case Study
Qing Li 0001, Zhihan Yan, Jun Li 0130, Zhenguo Yang, Zehang Lin, Hong Va Leong, Lei Chen 0002, Nancy Xiaonan Yu |
WISE (1) | 5 |
| 2020 | MMED: A multi-domain and Multi-modality event dataset
Zhenguo Yang, Zehang Lin, Lingni Guo, Qing Li 0001, Wenyin Liu |
Inf. Process. Manag. | 2 |
| 2020 | Learning Shared Semantic Space with Correlation Alignment for Cross-Modal Event RetrievalabstractIn this article, we propose to learn shared semantic space with correlation alignment ( S 3 CA ) for multimodal data representations, which aligns nonlinear correlations of multimodal data distributions in deep neural networks designed for heterogeneous data. In the context of cross-modal (event) retrieval, we design a neural network with convolutional layers and fully connected layers to extract features for images, including images on Flickr-like social media. Simultaneously, we exploit a fully connected neural network to extract semantic features for text documents, including news articles from news media. In particular, nonlinear correlations of layer activations in the two neural networks are aligned with correlation alignment during the joint training of the networks. Furthermore, we project the multimodal data into a shared semantic space for cross-modal (event) retrieval, where the distances between heterogeneous data samples can be measured directly. In addition, we contribute a Wiki-Flickr Event dataset, where the multimodal data samples are not describing each other in pairs like the existing paired datasets, but all of them are describing semantic events. Extensive experiments conducted on both paired and unpaired datasets manifest the effectiveness of S 3 CA , outperforming the state-of-the-art methods. Zhenguo Yang, Zehang Lin, Peipei Kang, Jianming Lv, Qing Li 0001, Wenyin Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Deep Semantic Space with Intra-class Low-rank Constraint for Cross-modal RetrievalabstractIn this paper, a novel Deep Semantic Space learning model with Intra-class Low-rank constraint (DSSIL) is proposed for cross-modal retrieval, which is composed of two subnetworks for modality-specific representation learning, followed by projection layers for common space mapping. In particular, DSSIL takes into account semantic consistency to fuse the cross-modal data in a high-level common space, and constrains the common representation matrix within the same class to be low-rank, in order to induce the intra-class representations more relevant. More formally, two regularization terms are devised for the two aspects, which have been incorporated into the objective of DSSIL. To optimize the modality-specific subnetworks and the projection layers simultaneously by exploiting the gradient decent directly, we approximate the nonconvex low-rank constraint by minimizing a few smallest singular values of the intra-class matrix with theoretical analysis. Extensive experiments conducted on three public datasets demonstrate the competitive superiority of DSSIL for cross-modal retrieval compared with the state-of-the-art methods. Peipei Kang, Zehang Lin, Zhenguo Yang, Xiaozhao Fang, Qing Li 0001, Wenyin Liu |
ICMR | 2 |
| 2019 | Catboost-based Framework with Additional User Information for Social Media Popularity PredictionabstractIn this paper, a Catboost-based framework is proposed to predict social media popularity. The framework is constituted by two components: feature representation and Catboost training. In the component of feature representation, numerical features are directly used, while categorical features are converted into numerical features by a method of order target statistics in Catboost. Besides, some additional user information is also tracked to enrich the feature space. In the other component, Catboost is adopted as the regression model which is trained by using post-related, user-related and additional user information. Moreover, to make full use of the dataset for model training, a dataset augmentation strategy based on pseudo labels is proposed. This strategy involves in two-stage training. In the first stage, it trains a first-stage model that is used to label the test set as pseudo labeled. In the next stage, a final model is trained based on the new training set that includes original validation set and the pseudo labeled test set. The proposed method achieves the 2nd place in the leader board of the Grand Challenge of Social Media Prediction. Peipei Kang, Zehang Lin, Shaohua Teng, Guipeng Zhang, Lingni Guo, Wei Zhang 0005 |
ACM Multimedia | 2 |
| 2019 | Cross-domain Beauty Item Retrieval via Unsupervised Embedding LearningabstractCross-domain image retrieval is always encountering insufficient labelled data in real world. In this paper, we propose unsupervised embedding learning (UEL) for cross-domain beauty and personal care product retrieval to finetune the convolutional neural network (CNN). More specifically, UEL utilizes the non-parametric softmax to train the CNN model as instance-level classification, which reduces the influence of some inevitable problems (e.g., shape variations). In order to obtain better performance, we integrate a few existing retrieval methods trained on different datasets. Furthermore, a query expansion strategy (i.e., diffusion) is adopted to improve the performance. Extensive experiments conducted on a dataset including half million images of beauty and personal product items (Perfect-500K) manifest the effectiveness of our proposed method. Our approach achieves the 2nd place in the leader board of the Grand Challenge of AI Meets Beauty in ACM Multimedia 2019. Our code is available at: https://github.com/RetrainIt/Perfect-Half-Million-Beauty-Product-Image-Recognition-Challenge-2019. Zehang Lin, Haoran Xie 0001, Peipei Kang, Zhenguo Yang, Wenyin Liu, Qing Li 0001 |
ACM Multimedia | 1 |
| 2019 | Single-Stage Detector with Semantic Attention for Occluded Pedestrian Detection
Fang Wen 0003, Zehang Lin, Zhenguo Yang, Wenyin Liu |
MMM (2) | 2 |
| 2019 | A layer-wise deep stacking model for social image popularity prediction
Zehang Lin, Feitao Huang, Zhenguo Yang, Wenyin Liu |
World Wide Web | 1 |
| 2018 | Random Forest Exploiting Post-related and User-related Features for Social Media Popularity PredictionabstractSocial media headline prediction (SMHP) is a thriving application scenario, which aims to predict the popularity of the post data shared on social media. In this paper, we propose to use multi-aspect features combined with the random forest (RF) model for popularity predictions. Firstly, we extract features by combining both metadata of the posts and users' features. More specifically, we adopt the binary coding strategy for dimensionality reduction and deal with the missing values by using some strategies, i.e., estimating the missing geographic information according to the information of users, and filling the missing features with median. Furthermore, regression models can be used directly to make predictions. In particular, a random forest (RF) model is adopted since it does not require much effort in tuning hyper-parameters and performs effectively. Extensive experiments conducted on the SMHP dataset consisting of 340K image posts shared by 80K users manifest the effectiveness of our method. Our approach achieves the 4nd place in the leader board of the Grand Challenge of SMHP in ACM Multimedia 2018. Feitao Huang, Zehang Lin, Peipei Kang, Zhenguo Yang |
ACM Multimedia | 3 |
| 2018 | Regional Maximum Activations of Convolutions with Attention for Cross-domain Beauty and Personal Care Product RetrievalabstractCross-domain beauty and personal care product image retrieval is a challenging problem due to data variations (e.g., brightness, viewpoint, and scale), and the rich types of items. In this paper, we present a regional maximum activations of convolutions with attention (RA-MAC) descriptor to extract image features for retrieval. RA-MAC improves the regional maximum activations of convolutions (R-MAC) descriptor considering the influence of background in cross-domain images (i.e., shopper domain and seller domain). More specifically, RA-MAC utilizes the characteristics of the convolutional layer to find the attention of an image, and reduces the influence of the unimportant regions in an unsupervised manner. Furthermore, a few strategies have been exploited to improve the performance, such as multiple features fusion, query expansion, and database augmentation. Extensive experiments conducted on a dataset consisting of half a million images of beauty care products (Perfect-500K) manifest the effectiveness of RA-MAC. Our approach achieves the 2nd place in the leader board of the Grand Challenge of AI Meets Beauty in ACM Multimedia 2018. Our code is available at: https://github.com/RetrainIt/Perfect-Half-Million-Beauty-Product-Image-Recognition-Challenge. Zehang Lin, Zhenguo Yang, Feitao Huang |
ACM Multimedia | 1 |
| 2018 | Improving Maximum Classifier Discrepancy by Considering Joint Distribution for Domain Adaptation
Zehang Lin, Zhenguo Yang, Runwei Situ, Feitao Huang, Jianming Lv, Qing Li 0001, Wenyin Liu |
WISE (2) | 1 |
| 2017 | Fire and Smoke Dynamic Textures Characterization by Applying Periodicity Index Based on Motion Features
Kanoksak Wattanachote, Zehang Lin, Mingchao Jiang, Liuwu Li, Gongliang Wang, Wenyin Liu |
ICVS | 2 |
| 2017 | Cross-Domain and Cross-Modality Transfer Learning for Multi-domain and Multi-modality Event Detection
Zhenguo Yang, Min Cheng 0003, Qing Li 0001, Zehang Lin, Wenyin Liu |
WISE (1) | 5 |