EDBT 2026 Demo / reviewers in the wild / expert
Liangjian Wen
dblp:231/7379
· DBLP profile ↗
23ranked-venue papers
7as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E-ViC: Reasoning Beyond Text via Embodied Visual Chain for Spatial IntelligenceabstractJunbo Qi, Yi Zhang, Hanchu Ni, Che Liu, Zhimin Yao, Ruilin Yang, Xiancong Ren, Liangjian Wen, Wei Ge, Yuya Ieiri, Osamu Yoshie, Yong Dai, Xiaozhu Ju. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junbo Qi, Yi Zhang 0001, Hanchu Ni, Che Liu 0002, Zhimin Yao, Ruilin Yang, Xiancong Ren, Liangjian Wen, Yuya Ieiri, Osamu Yoshie, Yong Dai 0001, Xiaozhu Ju |
ACL (1) | 8 |
| 2026 | GazeCLIP: Enhancing gaze estimation through text-guided multimodal learning
Jun Wang 0089, Hao Ruan, Liangjian Wen, Yong Dai 0001, Mingjie Wang 0002 |
Neurocomputing | 3 |
| 2026 | MRDNet: Multivariable Relational Decomposition Network for Multivariate Time Series Forecasting
Ao Hu, Liangjian Wen, Yong Dai 0001, Dongkai Wang, Jun Wang 0089, Jiang Duan |
Knowl. Based Syst. | 2 |
| 2026 | TimeCNN: Refining inscross-variable interaction on time point for time series forecasting
Ao Hu, Liangjian Wen, Yong Dai 0001, Shiyi Qi, Jun Wang 0089, Xun Zhou 0001, Dongkai Wang, Zenglin Xu, Jiang Duan |
Neural Networks | 2 |
| 2026 | FDNet: High-frequency disentanglement network with information-theoretic guidance for multivariate time series forecasting
Ao Hu, Liangjian Wen, Jiang Duan, Yong Dai 0001, Dongkai Wang, Shudong Huang, Jun Wang 0089, Zenglin Xu |
Pattern Recognit. | 2 |
| 2025 | Generalizable Object Keypoint Localization from Generative PriorsabstractGeneralizable object keypoint localization is a fundamental computer vision task in understanding the object structure. It is challenging for existing keypoint localization methods because their limited training data cannot provide generalizable shape and semantic cues, leading to inferior performance and generalization capability. Instead of relying on large scale training data, this work tackles this challenge by exploiting the rich priors from large generative models. We propose a data-efficient generalizable localization method named GenLoc. GenLoc extracts the generative priors from a pre-trained image generation model by calculating the correlation map between image latent feature and condition embedding. Those priors are hence optimized with our proposed heatmap expectation loss to perform object keypoint localization. Benefited by the rich knowledge of generative priors in understanding of object semantics and structures, GenLoc achieves superior performance on various object keypoint localization benchmarks. It shows more substantial performance enhancements in cross-domain, few-shot and zero-shot evaluation settings, e.g., getting 20%+ AP enhancement over CLAMP [43] in various zero-shot settings. Dongkai Wang, Jiang Duan, Liangjian Wen, Shiyu Xuan, Hao Chen 0061, Shiliang Zhang |
CVPR | 3 |
| 2025 | Disentangling Homophily and Heterophily in Multimodal Graph ClusteringabstractMultimodal graphs, which integrate unstructured heterogeneous data with structured interconnections, offer substantial real-world utility but remain insufficiently explored in unsupervised learning. In this work, we initiate the study of multimodal graph clustering, aiming to bridge this critical gap. Through empirical analysis, we observe that real-world multimodal graphs often exhibit hybrid neighborhood patterns, combining both homophilic and heterophilic relationships. To address this challenge, we propose a novel framework---Disentangled Multimodal Graph Clustering (DMGC) ---which decomposes the original hybrid graph into two complementary views: (1) a homophily-enhanced graph that captures cross-modal class consistency, and (2) heterophily-aware graphs that preserve modality-specific inter-class distinctions. We introduce a Multimodal Dual-frequency Fusion mechanism that jointly filters these disentangled graphs through a dual-pass strategy, enabling effective multimodal integration while mitigating category confusion. Our self-supervised alignment objectives further guide the learning process without requiring labels. Extensive experiments on both multimodal and multi-relational graph datasets demonstrate that DMGC achieves state-of-the-art performance, highlighting its effectiveness and generalizability across diverse settings. Our code is available at https://github.com/Uncnbb/DMGC. Zhaochen Guo, Zhixiang Shen, Xuanting Xie, Liangjian Wen, Zhao Kang 0001 |
ACM Multimedia | 4 |
| 2025 | InfMasking: Unleashing Synergistic Information by Contrastive Multimodal InteractionsabstractIn multimodal representation learning, synergistic interactions between modalities not only provide complementary information but also create unique outcomes through specific interaction patterns that no single modality could achieve alone. Existing methods may struggle to effectively capture the full spectrum of synergistic information, leading to suboptimal performance in tasks where such interactions are critical. This is particularly problematic because synergistic information constitutes the fundamental value proposition of multimodal representation. To address this challenge, we introduce InfMasking, a contrastive synergistic information extraction method designed to enhance synergistic information through an Infinite Masking strategy. InfMasking stochastically occludes most features from each modality during fusion, preserving only partial information to create representations with varied synergistic patterns. Unmasked fused representations are then aligned with masked ones through mutual information maximization to encode comprehensive synergistic information. This infinite masking strategy enables capturing richer interactions by exposing the model to diverse partial modality combinations during training. As computing mutual information estimates with infinite masking is computationally prohibitive, we derive an InfMasking loss to approximate this calculation. Through controlled experiments, we demonstrate that InfMasking effectively enhances synergistic information between modalities. In evaluations on large-scale real-world datasets, InfMasking achieves state-of-the-art performance across seven benchmarks. Code is released at https://github.com/brightest66/InfMasking. Liangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng, Yong Dai 0001, Dongkai Wang, Zhao Kang 0001, Jun Wang 0089, Zenglin Xu, Jiang Duan |
NeurIPS | 1 |
| 2025 | Adaptive Downscaling on Inputs Improves Time Series Classification
Xuanxuan Li, Zenglin Xu, Liangjian Wen, Xun Zhou 0001 |
PRCV (4) | 3 |
| 2024 | MVEB: Self-Supervised Learning With Multi-View Entropy BottleneckabstractSelf-supervised learning aims to learn representation that can be effectively generalized to downstream tasks. Many self-supervised approaches regard two views of an image as both the input and the self-supervised signals, assuming that either view contains the same task-relevant information and the shared information is (approximately) sufficient for predicting downstream tasks. Recent studies show that discarding superfluous information not shared between the views can improve generalization. Hence, the ideal representation is sufficient for downstream tasks and contains minimal superfluous information, termed minimal sufficient representation. One can learn this representation by maximizing the mutual information between the representation and the supervised view while eliminating superfluous information. Nevertheless, the computation of mutual information is notoriously intractable. In this work, we propose an objective termed multi-view entropy bottleneck (MVEB) to learn minimal sufficient representation effectively. MVEB simplifies the minimal sufficient learning to maximizing both the agreement between the embeddings of two views and the differential entropy of the embedding distribution. Our experiments confirm that MVEB significantly improves performance. For example, it achieves top-1 accuracy of 76.9% on ImageNet with a vanilla ResNet-50 backbone on linear evaluation. To the best of our knowledge, this is the new state-of-the-art result with ResNet-50. Liangjian Wen, Xiasi Wang, Jianzhuang Liu, Zenglin Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Cross-Scale Attention for Long-Term Time Series ForecastingabstractTransformer-based models, especially PatchTST, have demonstrated remarkable success in time series forecasting tasks. However, the unique nature of time series data, which often contains jitter, and noise, and has inherently lower information density compared to images and texts, poses significant challenges. Specifically, the ViT-inspired patching design is suboptimal for time series data due to the sparse semantic relationships in such data. Moreover, modelling these sparse semantic relationships requires more resources and longer processing times. To overcome these limitations, we introduce a novel approach that leverages cross-scale attention interaction via a multi-scale patching technique. Our method initially treats the entire sequence as a single patch, then progressively divides it into increasing patches, doubling each time. This strategy improves the information density within each patch and reduces the total number of patches needed to model the time series to typically just seven effectively. We design a single layer of attention to model the cross-scale relationships among patches. These qualities significantly enhance computational efficiency. Extensive experiments have shown that our method surpasses or closely approaches existing methods in time-series forecasting benchmarks. Additionally, it achieves a speed that is 12x times faster than PatchTST on the large dataset. Liangjian Wen, Quan Hu, Ao Hu |
IEEE Signal Process. Lett. | 1 |
| 2023 | Generative Oversampling for Imbalanced Data via Majority-Guided VAEabstractLearning with imbalanced data is a challenging problem in deep learning. Over-sampling is a widely used technique to re-balance the sampling distribution of training data. However, most existing over-sampling methods only use intra-class information of minority classes to augment the data but ignore the inter-class relationships with the majority ones, which is prone to overfitting, especially when the imbalance ratio is large. To address this issue, we propose a novel over-sampling model, called Majority-Guided VAE(MGVAE), which generates new minority samples under the guidance of a majority-based prior. In this way, the newly generated minority samples can inherit the diversity and richness of the majority ones, thus mitigating overfitting in downstream tasks. Furthermore, to prevent model collapse under limited data, we first pre-train MGVAE on sufficient majority samples and then fine-tune based on minority samples with Elastic Weight Consolidation(EWC) regularization. Experimental results on benchmark image datasets and real-world tabular data show that MGVAE achieves competitive improvements over other over-sampling methods in downstream classification tasks, demonstrating the effectiveness of our method. Qingzhong Ai, Pengyun Wang, Lirong He, Liangjian Wen, Lujia Pan, Zenglin Xu |
AISTATS | 4 |
| 2022 | Self-Supervision Can Be a Good Few-Shot Learner
Yuning Lu, Liangjian Wen, Jianzhuang Liu, Xinmei Tian 0001 |
ECCV (19) | 2 |
| 2022 | LayouTransformer: Generating Layout Patterns with Transformer via Sequential Pattern ModelingabstractGenerating legal and diverse layout patterns to establish large pattern libraries is fundamental for many lithography design applications. Existing pattern generation models typically regard the pattern generation problem as image generation of layout maps and learn to model the patterns via capturing pixel-level coherence, which is insufficient to achieve polygon-level modeling, e.g., shape and layout of patterns, thus leading to poor generation quality. In this paper, we regard the pattern generation problem as an unsupervised sequence generation problem, in order to learn the pattern design rules by explicitly modeling the shapes of polygons and the layouts among polygons. Specifically, we first propose a sequential pattern representation scheme that fully describes the geometric information of polygons by encoding the 2D layout patterns as sequences of tokens, i.e., vertexes and edges. Then we train a sequential generative model to capture the long-term dependency among tokens and thus learn the design rules from training examples. To generate a new pattern in sequence, each token is generated conditioned on the previously generated tokens that are from the same polygon or different polygons in the same layout map. Our framework, termed LayouTransformer, is based on the Transformer architecture due to its remarkable ability in sequence modeling. Comprehensive experiments show that our LayouTransformer not only generates a large amount of legal patterns but also maintains high generation diversity, demonstrating its superiority over existing pattern generative models. Liangjian Wen, Yi Zhu 0004, Guojin Chen, Bei Yu 0001, Jianzhuang Liu, Chunjing Xu |
ICCAD | 1 |
| 2022 | Structure-Preserving Graph Representation LearningabstractThough graph representation learning (GRL) has made significant progress, it is still a challenge to extract and embed the rich topological structure and feature information in an adequate way. Most existing methods focus on local structure and fail to fully incorporate the global topological structure. To this end, we propose a novel Structure-Preserving Graph Representation Learning (SPGRL) method, to fully capture the structure information of graphs. Specifically, to reduce the uncertainty and misinformation of the original graph, we construct a feature graph as a complementary view via k-Nearest Neighbor method. The feature graph can be used to contrast at node-level to capture the local relation. Besides, we retain the global topological structure information by maximizing the mutual information (MI) of the whole graph and feature embeddings, which is theoretically reduced to exchanging the feature embeddings of the feature and the original graphs to reconstruct themselves. Extensive experiments show that our method has quite superior performance on semi-supervised node classification task and excellent robustness under noise perturbation on graph structure or node features. The source code is available at https://github.com/uestc-lese/SPGRL. Ruiyi Fang, Liangjian Wen, Zhao Kang 0001, Jianzhuang Liu |
ICDM | 2 |
| 2022 | AFINet: Attentive Feature Integration Networks for image classification
Xinglin Pan, Yu Pan 0005, Liangjian Wen, Wenxiang Lin, Hongguang Fu, Zenglin Xu |
Neural Networks | 4 |
| 2021 | Boosting Few-Shot Classification with View-Learnable Contrastive LearningabstractThe goal of few-shot classification is to classify new categories with few labeled examples within each class. Nowadays, the excellent performance in handling few-shot classification problems is shown by metric-based meta-learning methods. However, it is very hard for previous methods to discriminate the fine-grained sub-categories in the embedding space without fine-grained labels. This may lead to unsatisfactory generalization to fine-grained sub-categories, and thus affects model interpretation. To tackle this problem, we introduce the contrastive loss into few-shot classification for learning latent fine-grained structure in the embedding space. Furthermore, to overcome the drawbacks of random image transformation used in current contrastive learning in producing noisy and inaccurate image pairs (i.e., views), we develop a learning-to-learn algorithm to automatically generate different views of the same image. Extensive experiments on standard few-shot learning benchmarks demonstrate the superiority of our method. Xu Luo 0003, Liangjian Wen, Lili Pan 0001, Zenglin Xu |
ICME | 3 |
| 2021 | Self-supervised Consensus Representation Learning for Attributed GraphabstractAttempting to fully exploit the rich information of topological structure and node features for attributed graph, we introduce self-supervised learning mechanism to graph representation learning and propose a novel Self-supervised Consensus Representation Learning (SCRL) framework. In contrast to most existing works that only explore one graph, our proposed SCRL method treats graph from two perspectives: topology graph and feature graph. We argue that their embeddings should share some common information, which could serve as a supervisory signal. Specifically, we construct the feature graph of node features via k-nearest neighbour algorithm. Then graph convolutional network (GCN) encoders extract features from two graphs respectively. Self-supervised loss is designed to maximize the agreement of the embeddings of the same node in the topology graph and the feature graph. Extensive experiments on real citation networks and social networks demonstrate the superiority of our proposed SCRL over the state-of-the-art methods on semi-supervised node classification task. Meanwhile, compared with its main competitors, SCRL is rather efficient. Changshu Liu, Liangjian Wen, Zhao Kang 0001, Guangchun Luo, Ling Tian |
ACM Multimedia | 2 |
| 2021 | Rectifying the Shortcut Learning of Background for Few-Shot LearningabstractThe category gap between training and evaluation has been characterised as one of the main obstacles to the success of Few-Shot Learning (FSL). In this paper, we for the first time empirically identify image background, common in realistic images, as a shortcut knowledge helpful for in-class classification but ungeneralizable beyond training categories in FSL. A novel framework, COSOC, is designed to tackle this problem by extracting foreground objects in images at both training and evaluation without any extra supervision. Extensive experiments carried on inductive FSL tasks demonstrate the effectiveness of our approaches. Xu Luo 0003, Longhui Wei, Liangjian Wen, Lingxi Xie, Zenglin Xu, Qi Tian 0001 |
NeurIPS | 3 |
| 2021 | Gradient estimation of information measures in deep learning
Liangjian Wen, Haoli Bai, Lirong He, Yiji Zhou, Mingyuan Zhou, Zenglin Xu |
Knowl. Based Syst. | 1 |
| 2020 | Mutual Information Gradient Estimation for Representation Learning
Liangjian Wen, Yiji Zhou, Lirong He, Mingyuan Zhou, Zenglin Xu |
ICLR | 1 |
| 2020 | Structured pruning of recurrent neural networks through neuron selection
Liangjian Wen, Xuanyang Zhang, Haoli Bai, Zenglin Xu |
Neural Networks | 1 |
| 2019 | Low-rank kernel learning for graph-based clustering
Zhao Kang 0001, Liangjian Wen, Wenyu Chen 0001, Zenglin Xu |
Knowl. Based Syst. | 2 |