EDBT 2026 Demo / reviewers in the wild / expert
Tianze Luo
dblp:297/4000
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video UnderstandingabstractMultimodal large language models have become a popular topic in deep visual understanding due to many promising real-world applications. However, hour-long video understanding, spanning over one hour and containing tens of thousands of visual frames, remains under-explored because of 1) challenging long-term video analyses, 2) inefficient large-model approaches, and 3) lack of large-scale benchmark datasets. Among them, in this paper, we focus on building a large-scale hour-long long video benchmark, HLV-1K1, designed to evaluate long video understanding models. HLV-1K comprises 1009 hour-long videos with 14,847 high-quality question answering (QA) and multi-choice question asnwering (MCQA) pairs with time-aware query and diverse annotations, covering frame-level, within-event-level, cross-event-level, and long-term reasoning tasks. We evaluate our benchmark using existing state-of-the-art methods and demonstrate its value for testing deep long video understanding capabilities at different levels and for various tasks. This includes promoting future long video understanding tasks at a granular level, such as deep understanding of long live videos, meeting recordings, and movies. Heqing Zou, Tianze Luo, Guiyang Xie, Victor Xiao Jie Zhang, Fengmao Lv, Guangcong Wang, Junyang Chen 0001, Zhuochen Wang, Hansheng Zhang, Huaijian Zhang |
ICME | 2 |
| 2025 | WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow MatchingabstractTianze Luo, Xingchen Miao, Wenbo Duan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tianze Luo, Xingchen Miao, Wenbo Duan |
NAACL (Long Papers) | 1 |
| 2025 | MSTI-Plus: Introducing Non-Sarcasm Reference Materials to Enhance Multimodal Sarcasm Target IdentificationabstractSarcasm is a subtle expression that indicates the incongruity between literal meanings and factual opinions. For multimodal posts in social medias which consist of both images and texts, sarcasm expressions are even more widespread. Recent works have paid attentions to Multimodal Sarcasm Target Identification (MSTI), which focuses on detecting aspect terms of mockery or ridicule as sarcasm targets. However, the current MSTI benchmark only contains annotations on fine-grained sarcasm targets within sarcastic samples. In practice, it will be featured by two major limitations. First, there lack annotations on non-sarcasm aspects to inform deep models to perceive the semantic difference between sarcasm targets and non-sarcasm aspects. As a result, deep models will tend to incorrectly recognize non-sarcasm aspects as sarcasm targets. Second, there lack non-sarcasm samples to inform deep models to perceive the inherent semantics of sarcasm intentions. Due to the subtle characteristic of sarcasm expressions, models trained with only fine-grained supervision signals cannot thoroughly understand the sarcasm semantics, making the fine-grained task of sarcasm target identification restricted. Motivated by these limitations, this work reconstructs a more comprehensive MSTI benchmark by introducing both fine-grained non-sarcasm aspect annotations for existing sarcasm samples and non-sarcastic samples as non-sarcasm references to enable deep models to clearly perceive the mentioned information during training. Based on the multi-granularity (i.e., both aspect-level and sample-level) non-sarcasm information introduced into this new benchmark, this work further proposes a pluggable Semantics-aware Sarcasm Target Identification mechanism to enhance sarcasm target identification by modeling the overall semantics of sarcasm intentions via an auxiliary sample-level sarcasm recognition task. By modeling the overall semantics of sarcasm intention, deep models can obtain a more comprehensive understanding on sarcasm semantics, leading to improved performance on fine-grained sarcasm target identification. Extensive experiments are conducted to validate our contribution. Both the dataset and code are available at https://github.com/tiggers23/MSTI-Plus. Fengmao Lv, Mengting Xiong, Junlin Fang, Tianze Luo, Weichao Liang, Tianrui Li 0001 |
WWW | 5 |
| 2024 | Progressive Multimodal Pivot Learning: Towards Semantic Discordance Understanding as HumansabstractMultimodal recognition can achieve enhanced performance by leveraging the complementary information from different modali- ties. However, in real-world scenarios, multimodal samples often express discordant semantic meanings across modalities, lacking evident complementary information. Unlike humans who can easily understand the intrinsic semantic information of these semantically discordant samples, existing multimodal recognition models show poor performance on them. With the motivation of improving the robustness of multimodal recognition models in practical scenar- ios, this work poses a new challenge in multimodal recognition, which is coined as Semantic Discordance Understanding. Unlike ex- isting works only focusing on detecting semantically discordant samples as noisy data, this new challenge requires deep models to follow humans’ ability in understanding the inherent seman- tic meanings of semantically discordant samples. To address this challenge, we further propose the Progressive Multimodal Pivot Learning (PMPL) approach by introducing a learnable pivot mem- ory to explore the inherent semantics meaning hidden under dis- cordant modalities. To this end, our approach inserts Pivot Memory Learning (PML) modules into multiple layers of unimodal foun- dation models to progressively trade-off the conflict information across modalities. By introducing the multimodal pivot learning paradigm for multimodal recognition, the proposed PMPL approach can alleviate the negative effect of semantic discordance caused by the cross-modal information exchange mechanism of existingmultimodal recognition models. Experiments on different bench- marks validate the superiority of our approach. Code is available at https://github.com/tiggers23/PMPL. Junlin Fang, Wenya Wang 0001, Tianze Luo, Yanyong Huang, Fengmao Lv |
CIKM | 3 |
| 2024 | Learning Adaptive Multiresolution Transforms via Meta-Framelet-based Graph Convolutional NetworkabstractGraph Neural Networks are popular tools in graph representation learning that capture the graph structural properties. However, most GNNs employ single-resolution graph feature extraction, thereby failing to capture micro-level local patterns (high resolution) and macro-level graph cluster and community patterns (low resolution) simultaneously. Many multiresolution methods have been developed to capture graph patterns at multiple scales, but most of them depend on predefined and handcrafted multiresolution transforms that remain fixed throughout the training process once formulated. Due to variations in graph instances and distributions, fixed handcrafted transforms can not effectively tailor multiresolution representations to each graph instance. To acquire multiresolution representation suited to different graph instances and distributions, we introduce the Multiresolution Meta-Framelet-based Graph Convolutional Network (MM-FGCN), facilitating comprehensive and adaptive multiresolution analysis across diverse graphs. Extensive experiments demonstrate that our MM-FGCN achieves SOTA performance on various graph learning tasks. Tianze Luo, Zhanfeng Mo, Sinno Jialin Pan |
ICLR | 1 |
| 2024 | Graph Principal Flow Network for Conditional Graph Generation
Zhanfeng Mo, Tianze Luo, Sinno Jialin Pan |
WWW | 2 |
| 2024 | Fast Graph Generation via Spectral DiffusionabstractGenerating graph-structured data is a challenging problem, which requires learning the underlying distribution of graphs. Various models such as graph VAE, graph GANs, and graph diffusion models have been proposed to generate meaningful and reliable graphs, among which the diffusion models have achieved state-of-the-art performance. In this paper, we argue that running full-rank diffusion SDEs on the whole graph adjacency matrix space hinders diffusion models from learning graph topology generation, and hence significantly deteriorates the quality of generated graph data. To address this limitation, we propose an efficient yet effective Graph Spectral Diffusion Model (GSDM), which is driven by low-rank diffusion SDEs on the graph spectrum space. Our spectral diffusion model is further proven to enjoy a substantially stronger theoretical guarantee than standard diffusion models. Extensive experiments across various datasets demonstrate that our proposed GSDM turns out to be the SOTA model, by exhibiting both significantly higher generation quality and much less computational consumption than the baselines. Tianze Luo, Zhanfeng Mo, Sinno Jialin Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Collaborative Sequential Recommendations via Multi-view GNN-transformersabstractSequential recommendation systems aim to exploit users’ sequential behavior patterns to capture their interaction intentions and improve recommendation accuracy. Existing sequential recommendation methods mainly focus on modeling the items’ chronological relationships in each individual user behavior sequence, which may not be effective in making accurate and robust recommendations. On the one hand, the performance of existing sequential recommendation methods is usually sensitive to the length of a user’s behavior sequence (i.e., the list of a user’s historically interacted items). On the other hand, besides the context information in each individual user behavior sequence, the collaborative information among different users’ behavior sequences is also crucial to make accurate recommendations. However, this kind of information is usually ignored by existing sequential recommendation methods. In this work, we propose a new sequential recommendation framework, which encodes the context information in each individual user behavior sequence as well as the collaborative information among the behavior sequences of different users, through building a local dependency graph for each item. We conduct extensive experiments to compare the proposed model with state-of-the-art sequential recommendation methods on five benchmark datasets. The experimental results demonstrate that the proposed model is able to achieve better recommendation performance than existing methods, by incorporating collaborative information. Tianze Luo, Yong Liu 0020, Sinno Jialin Pan |
ACM Trans. Inf. Syst. | 1 |
| 2024 | STARNet: An Efficient Spatiotemporal Feature Sharing Reconstructing Network for Automatic Modulation ClassificationabstractAutomatic Modulation Classification (AMC) is a crucial task in the field of wireless communication, allowing for the identification of the modulation scheme of a received radio signal without prior knowledge of the communication system. Recently, AMC approaches based on Deep Learning (DL) have achieved outstanding results. However, the majority of current DL-based AMC methods face challenges in achieving high recognition accuracy while remaining computationally efficient. Some researchers have designed autoencoder-based models to generate low-dimensional temporal feature embeddings of the radio signal, thereby reducing the number of model parameters while maintaining high performance in recognizing modulation formats. However, when further improving AMC performance via learning low-dimensional spatial-temporal feature representations, traditional autoencoder models require both a convolutional decoder and an LSTM decoder to reconstruct temporal and spatial features separately, which unavoidably raises the model parameters. In this paper, we propose a spatiotemporal feature sharing reconstructing network (STARNet) to simultaneously extract low-dimensional spatial and temporal feature representations of radio signals using a single autoencoder structure, thereby reducing the number of model parameters and improving AMC performance. Additionally, we construct a Hybrid Attentive Ghost (HA-Ghost) to automatically extract discriminative radio signal spatial information according to signal reconstruction performance. Extensive experiments on benchmark datasets demonstrate that the proposed STARNet achieves an average modulation classification accuracy of 63.64%, outperforming previous state-of-the-art models. Despite extracting more types of features, STARNet has only 14,860 parameters, which is smaller than existing spatiotemporal autoencoder-based methods. Xiangli Zhang, Zishuo Wang, Tianze Luo, Yong Xiao 0001, Dapeng Luo |
IEEE Trans. Wirel. Commun. | 4 |
| 2022 | Domain Confused Contrastive Learning for Unsupervised Domain AdaptationabstractIn this work, we study Unsupervised Domain Adaptation (UDA) in a challenging selfsupervised approach.One of the difficulties is how to learn task discrimination in the absence of target labels.Unlike previous literature which directly aligns cross-domain distributions or leverages reverse gradient, we propose Domain Confused Contrastive Learning (DCCL) to bridge the source and the target domains via domain puzzles, and retain discriminative representations after adaptation.Technically, DCCL searches for a most domainchallenging direction and exquisitely crafts domain confused augmentations as positive pairs, then it contrastively encourages the model to pull representations towards the other domain, thus learning more stable and effective domain invariances.We also investigate whether contrastive learning necessarily helps with UDA when performing other data augmentations.Extensive experiments demonstrate that DCCL significantly outperforms baselines. Quanyu Long, Tianze Luo, Wenya Wang 0001, Sinno Jialin Pan |
NAACL-HLT | 2 |
| 2021 | Mitigating Performance Saturation in Neural Marked Point Processes: Architectures and Loss FunctionsabstractAttributed event sequences are commonly encountered in practice. A recent research line focuses on incorporating neural networks with the statistical model--marked point processes, which is the conventional tool for dealing with attributed event sequences. Neural marked point processes possess good interpretability of probabilistic models as well as the representational power of neural networks. However, we find that performance of neural marked point processes is not always increasing as the network architecture becomes more complicated and larger, which is what we call the performance saturation phenomenon. This is due to the fact that the generalization error of neural marked point processes is determined by both the network representational ability and the model specification at the same time. Therefore we can draw two major conclusions: first, simple network structures can perform no worse than complicated ones for some cases; second, using a proper probabilistic assumption is as equally, if not more, important as improving the complexity of the network. Based on this observation, we propose a simple graph-based network structure called GCHP, which utilizes only graph convolutional layers, thus it can be easily accelerated by the parallel mechanism. We directly consider the distribution of interarrival times instead of imposing a specific assumption on the conditional intensity function, and propose to use a likelihood ratio loss with a moment matching mechanism for optimization and model selection. Experimental results show that GCHP can significantly reduce training time and the likelihood ratio loss with interarrival time probability assumptions can greatly improve the model performance. Tianbo Li, Tianze Luo, Yiping Ke, Sinno Jialin Pan |
KDD | 2 |