EDBT 2026 Demo / reviewers in the wild / expert
Jianglin Lu
dblp:252/0023
· DBLP profile ↗
21ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-4191-7734ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revealing the Seen, Imagining the Beyond: A Survey of Image-Grounded Chain-of-Thought Reasoning in Multimodal LLMsabstractQihua Dong, Yitian Zhang, Huimin Zeng, Yizhou Wang, Jianglin Lu, Kuo Yang, Yun Fu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qihua Dong, Yizhou Wang 0006, Jianglin Lu, Yun Fu 0001 |
ACL (1) | 5 |
| 2026 | ACBQ: Adaptive Cross-Block Quantization of Large Language ModelsabstractPost-training quantization (PTQ) has emerged as a promising approach for reducing the memory footprint and computational cost of large language models (LLMs), enabling efficient deployment without full model retraining.However, existing PTQ methods struggle to simultaneously support weight-activation joint quantization and extreme low-bit weight quantization.This limitation primarily arises from the depth of LLMs and their strong cross-layer dependencies, which cause quantization errors to propagate and accumulate across layers, ultimately leading to significant performance degradation.In this paper, we present ACBQ, a simple yet effective framework that simultaneously addresses weight-activation joint quantization and extreme weight quantization.We first propose a granular quantization strategy that treats self-attention and FFN as separate quantization units with module-specific optimization objectives.To mitigate the propagation and accumulation of quantization errors across layers, we introduce an adaptive cross-block quantization strategy that explicitly accounts for crosslayer dependencies by encouraging consistency across blocks.Extensive experiments across diverse LLMs, including OPT and the LLaMA family, demonstrate that ACBQ achieves superior performance under both W4A4 and highly aggressive W2 settings, while incurring negligible additional computational overhead. Jianglin Lu, Yun Fu 0001 |
ACL (1) | 2 |
| 2026 | Fuzzy Multi-Subspace ClusteringabstractProjection-based subspace clustering methods typically learn a single projection subspace shared across all clusters. This strategy prioritizes consistency across clusters at the expense of their distinctiveness—potentially misaligning the projection objective with the clustering goal. To overcome this limitation, we propose Fuzzy Multi-Subspace Clustering (FMSC), which simultaneously learns multiple cluster-specific subspaces and membership assignments in a mutually reinforcing manner. The refinement of subspaces helps improve membership accuracy, while more accurate membership assignments guide more precise subspace learning. Moreover, incorporating fuzzy learning enables FMSC to perform one-step clustering and handle scenarios in which a sample simultaneously belongs to multiple subspaces. We also propose a novel subspace regularization method for FMSC that maintains cross-subspace consistency without sacrificing the fidelity of cluster-specific subspace reconstruction. The time complexity of FMSC scales linearly with the sample size, ensuring its scalability. Extensive experiments on synthetic datasets with millions of samples and widely used benchmarks demonstrate that FMSC outperforms several state-of-the-art methods. Yangbo Wang, Jie Zhou 0009, Mingli Song, Jianglin Lu |
IEEE Trans. Fuzzy Syst. | 5 |
| 2025 | Representation Potentials of Foundation Models for Multimodal Alignment: A SurveyabstractFoundation models learn highly transferable representations through large-scale pretraining on diverse data.An increasing body of research indicates that these representations exhibit a remarkable degree of similarity across architectures and modalities.In this survey, we investigate the representation potentials of foundation models, defined as the latent capacity of their learned representations to capture task-specific information within a single modality while also providing a transferable basis for alignment and unification across modalities.We begin by reviewing representative foundation models and the key metrics that make alignment measurable.We then synthesize empirical evidence of representation potentials from studies in vision, language, speech, multimodality, and neuroscience.The evidence suggests that foundation models often exhibit structural regularities and semantic consistencies in their representation spaces, positioning them as strong candidates for cross-modal transfer and alignment.We further analyze the key factors that foster representation potentials, discuss open questions, and highlight potential challenges. Jianglin Lu, Yi Xu 0005, Yizhou Wang 0006, Yun Fu 0001 |
EMNLP | 1 |
| 2025 | Outlier-Aware Post-Training Quantization for Image Super-ResolutionabstractQuantization techniques, including quantization-aware training (QAT) and post-training quantization (PTQ), have become essential for inference acceleration of image super-resolution (SR) networks. Compared to QAT, PTQ has garnered significant attention as it eliminates the need for ground truth and model retraining. However, existing PTQ methods for SR often fail to achieve satisfactory performance as they overlook the impact of outliers in activation. Our empirical analysis reveals that these prevalent activation outliers are strongly correlated with image color information, and directly removing them leads to significant performance degradation. Motivated by this, we propose a dual-region quantization strategy that partitions activations into an outlier region and a dense region, applying uniform quantization to each region independently to better balance bit-width allocation. Furthermore, we observe that different network layers exhibit varying sensitivities to quantization, leading to different levels of performance degradation. To address this, we introduce sensitivity-aware finetuning that encourages the model to focus more on highly sensitive layers, further enhancing quantization performance. Extensive experiments demonstrate that our method outperforms existing PTQ approaches across various SR networks and datasets, while achieving performance comparable to QAT methods in most scenarios with at least a 75 speedup. Jianglin Lu, Yun Fu 0001 |
ICCV | 2 |
| 2025 | Scale-Free Graph-Language ModelsabstractGraph-language models (GLMs) have demonstrated great potential in graph-based semi-supervised learning. A typical GLM consists of two key stages: graph generation and text embedding, which are usually implemented by inferring a latent graph and finetuning a language model (LM), respectively. However, the former often relies on artificial assumptions about the underlying edge distribution, while the latter requires extensive data annotations. To tackle these challenges, this paper introduces a novel GLM that integrates graph generation and text embedding within a unified framework. Specifically, for graph generation, we leverage an inherent characteristic of real edge distribution—the scale-free property—as a structural prior. We unexpectedly find that this natural property can be effectively approximated by a simple k-nearest neighbor (KNN) graph. For text embedding, we develop a graph-based pseudo-labeler that utilizes scale-free graphs to provide complementary supervision for improved LM finetuning. Extensive experiments on representative datasets validate our findings on the scale-free structural approximation of KNN graphs and demonstrate the effectiveness of integrating graph generation and text embedding with a real structural prior. Our code is available at https://github.com/Jianglin954/SFGL. Jianglin Lu, Yun Fu 0001 |
ICLR | 1 |
| 2025 | The Indra Representation Hypothesis for Multimodal AlignmentabstractRecent studies have uncovered an interesting phenomenon: unimodal foundation models tend to learn convergent representations, regardless of differences in architecture, training objectives, or data modalities. However, these representations are essentially internal abstractions of samples that characterize samples independently, leading to limited expressiveness. In this paper, we propose The Indra Representation Hypothesis, inspired by the philosophical metaphor of Indra’s Net. We argue that representations from unimodal foundation models are converging to implicitly reflect a shared relational structure underlying reality, akin to the relational ontology of Indra’s Net. We formalize this hypothesis using the V-enriched Yoneda embedding from category theory, defining the Indra representation as a relational profile of each sample with respect to others. This formulation is shown to be unique, complete, and structure-preserving under a given cost function. We instantiate the Indra representation using angular distance and evaluate it in cross-model and cross-modal scenarios involving vision, language, and audio. Extensive experiments demonstrate that Indra representations consistently enhance robustness and alignment across architectures and modalities, providing a theoretically grounded and practical framework for training-free alignment of unimodal foundation models. Our code is available at https://github.com/Jianglin954/Indra. Jianglin Lu, Simon Jenni, Yun Fu 0001 |
NeurIPS | 1 |
| 2024 | Robust Self-expression Learning with Adaptive Noise PerceptionabstractSelf-expression learning methods often obtain a coefficient matrix to measure the similarity between pairs of samples. However, directly using the raw data to represent each sample under the self-expression framework may not be ideal, as noise points are inevitably involved in the process of representing clean samples. To address this issue, this work proposes a novel self-expression model called robust Self-Expression learning with adaptive Noise Perception (SENP). SENP decomposes each sample into a clean part and a noisy part, and samples with large self-expression losses can be recognized as the noise points. A reliable coefficient matrix can then be learned by using only the clean points to reconstruct the clean part of each sample. By simultaneously detecting the noisy part of each sample and noise points, and adaptively mitigating their negative impacts, the representative ability of the generated coefficient matrix is improved. Moreover, inspired by the solution of non-negative matrix factorization (NMF), an effective algorithm is formed to optimize SENP. Extensive experiments on well-known benchmark datasets demonstrate the superiority of SENP compared to several state-of-the-art methods. Yangbo Wang, Jie Zhou 0009, Jianglin Lu, Jun Wan 0005, Can Gao, Qingshui Lin |
Pattern Recognit. | 3 |
| 2024 | Asymmetric Transfer Hashing With Adaptive Bipartite Graph LearningabstractThanks to the efficient retrieval speed and low storage consumption, learning to hash has been widely used in visual retrieval tasks. However, the known hashing methods assume that the query and retrieval samples lie in homogeneous feature space within the same domain. As a result, they cannot be directly applied to heterogeneous cross-domain retrieval. In this article, we propose a generalized image transfer retrieval (GITR) problem, which encounters two crucial bottlenecks: 1) the query and retrieval samples may come from different domains, leading to an inevitable domain distribution gap and 2) the features of the two domains may be heterogeneous or misaligned, bringing up an additional feature gap. To address the GITR problem, we propose an asymmetric transfer hashing (ATH) framework with its unsupervised/semisupervised/supervised realizations. Specifically, ATH characterizes the domain distribution gap by the discrepancy between two asymmetric hash functions, and minimizes the feature gap with the help of a novel adaptive bipartite graph constructed on cross-domain data. By jointly optimizing asymmetric hash functions and the bipartite graph, not only can knowledge transfer be achieved but information loss caused by feature alignment can also be avoided. Meanwhile, to alleviate negative transfer, the intrinsic geometrical structure of single-domain data is preserved by involving a domain affinity graph. Extensive experiments on both single-domain and cross-domain benchmarks under different GITR subtasks indicate the superiority of our ATH method in comparison with the state-of-the-art hashing methods. Jianglin Lu, Jie Zhou 0009, Yudong Chen 0002, Witold Pedrycz, Kwok-Wai Hung |
IEEE Trans. Cybern. | 1 |
| 2023 | Latent Graph Inference with Limited SupervisionabstractLatent graph inference (LGI) aims to jointly learn the underlying graph structure and node representations from data features. However, existing LGI methods commonly suffer from the issue of supervision starvation, where massive edge weights are learned without semantic supervision and do not contribute to the training loss. Consequently, these supervision-starved weights, which determine the predictions of testing samples, cannot be semantically optimal, resulting in poor generalization. In this paper, we observe that this issue is actually caused by the graph sparsification operation, which severely destroys the important connections established between pivotal nodes and labeled ones. To address this, we propose to restore the corrupted affinities and replenish the missed supervision for better LGI. The key challenge then lies in identifying the critical nodes and recovering the corrupted affinities. We begin by defining the pivotal nodes as k-hop starved nodes, which can be identified based on a given adjacency matrix. Considering the high computational burden, we further present a more efficient alternative inspired by CUR matrix decomposition. Subsequently, we eliminate the starved nodes by reconstructing the destroyed connections. Extensive experiments on representative benchmarks demonstrate that reducing the starved nodes consistently improves the performance of state-of-the-art LGI methods, especially under extremely limited supervision (6.12% improvement on Pubmed with a labeling rate of only 0.3%). Jianglin Lu, Yi Xu 0005, Huan Wang 0014, Yun Fu 0001 |
NeurIPS | 1 |
| 2022 | Uncertainty-Guided Pixel Contrastive Learning for Semi-Supervised Medical Image SegmentationabstractRecently, contrastive learning has shown great potential in medical image segmentation. Due to the lack of expert annotations, however, it is challenging to apply contrastive learning in semi-supervised scenes. To solve this problem, we propose a novel uncertainty-guided pixel contrastive learning method for semi-supervised medical image segmentation. Specifically, we construct an uncertainty map for each unlabeled image and then remove the uncertainty region in the uncertainty map to reduce the possibility of noise sampling. The uncertainty map is determined by a well-designed consistency learning mechanism, which generates comprehensive predictions for unlabeled data by encouraging consistent network outputs from two different decoders. In addition, we suggest that the effective global representations learned by an image encoder should be equivariant to different geometric transformations. To this end, we construct an equivariant contrastive loss to strengthen global representation learning ability of the encoder. Extensive experiments conducted on popular medical image benchmarks demonstrate that the proposed method achieves better segmentation performance than the state-of-the-art methods. Jianglin Lu, Zhihui Lai 0001, Jiajun Wen 0001, Heng Kong |
IJCAI | 2 |
| 2022 | Deep asymmetric hashing with dual semantic regression and class structure quantization
Jianglin Lu, Jie Zhou 0009, Mengfan Yan, Jiajun Wen 0001 |
Inf. Sci. | 1 |
| 2022 | Generalized Embedding Regression: A Framework for Supervised Feature ExtractionabstractSparse discriminative projection learning has attracted much attention due to its good performance in recognition tasks. In this article, a framework called generalized embedding regression (GER) is proposed, which can simultaneously perform low-dimensional embedding and sparse projection learning in a joint objective function with a generalized orthogonal constraint. Moreover, the label information is integrated into the model to preserve the global structure of data, and a rank constraint is imposed on the regression matrix to explore the underlying correlation structure of classes. Theoretical analysis shows that GER can obtain the same or approximate solution as some related methods with special settings. By utilizing this framework as a general platform, we design a novel supervised feature extraction approach called jointly sparse embedding regression (JSER). In JSER, we construct an intrinsic graph to characterize the intraclass similarity and a penalty graph to indicate the interclass separability. Then, the penalty graph Laplacian is used as the constraint matrix in the generalized orthogonal constraint to deal with interclass marginal points. Moreover, the$L_{2,1}$-norm is imposed on the regression terms for robustness to outliers and data’s variations and the regularization term for jointly sparse projection learning, leading to interesting semantic interpretability. An effective iterative algorithm is elaborately designed to solve the optimization problem of JSER. Theoretically, we prove that the subproblem of JSER is essentially an unbalanced Procrustes problem and can be solved iteratively. The convergence of the designed algorithm is also proved. Experimental results on six well-known data sets indicate the competitive performance and latent properties of JSER. Jianglin Lu, Zhihui Lai 0001, Yudong Chen 0002, Jie Zhou 0009, LinLin Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Deep Semantic and Attentive Network for Unsupervised Video SummarizationabstractWith the rapid growth of video data, video summarization is a promising approach to shorten a lengthy video into a compact version. Although supervised summarization approaches have achieved state-of-the-art performance, they require frame-level annotated labels. Such an annotation process is time-consuming and tedious. In this article, we propose a novel deep summarization framework named Deep Semantic and Attentive Network for Video Summarization (DSAVS) that can select the most semantically representative summary by minimizing the distance between video representation and text representation without any frame-level labels. Another challenge associated with video summarization tasks mainly originates from the difficulty of considering temporal information over a long time. Long Short-Term Memory (LSTM) performs well for temporal dependencies modeling but does not work well with long video clips. Therefore, we introduce a self-attention mechanism into our summarization framework to capture the long-range temporal dependencies among the frames. Extensive experiments on two popular benchmark datasets, i.e., SumMe and TVSum, show that our proposed framework outperforms other state-of-the-art unsupervised approaches and even most supervised methods. Shenghua Zhong, Jingxu Lin, Jianglin Lu, Ahmed Fares, Tongwei Ren |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Progressive Distribution Alignment Based on Label Correction for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to transfer knowledge between different domains. Most of the existing UDA methods try to align the conditional distribution between the source and target domains by utilizing the information of pseudo labels induced from the target domain. To tackle the negative transfer caused by inaccurate pseudo labels, we propose a novel UDA method named progressive distribution alignment based on label correction (PDALC). Specifically, PDALC uses the class discriminative information to perform subspace learning to obtain the domain invariance subspace. Furthermore, a new mechanism of pseudo label correction is introduced to measure the reliability of pseudo labels and utmostly correct the inaccurate pseudo labels. By combining subspace learning with label correction, the performance of PDALC can be continuously improved, which in turn reduces the generation of inaccurate pseudo labels. The experimental results show that the proposed method outperforms the state-of-the-art UDA methods. Yuwu Lu, Can Gao, Jianglin Lu |
ICME | 6 |
| 2021 | Local Graph Convolutional Networks for Cross-Modal HashingabstractCross-modal hashing aims to map the data of different modalities into a common binary space to accelerate the retrieval speed. Recently, deep cross-modal hashing methods have shown promising performance by applying deep neural networks to facilitate feature learning. However, the known supervised deep methods mainly rely on the labeled information of datasets, which is insufficient to characterize the latent structures that exist among different modalities. To mitigate this problem, in this paper, we propose to use Graph Convolutional Networks (GCNs) to exploit the local structure information of datasets for cross-modal hash learning. Specifically, a local graph is constructed according to the neighborhood relationships between samples in deep feature spaces and fed into GCNs to generate graph embeddings. Then, a within-modality loss is designed to measure the inner products between deep features and graph embeddings so that hashing networks and GCNs can be jointly optimized. By taking advantage of GCNs to assist model's training, the performance of hashing networks can be improved. Extensive experiments on benchmarks verify the effectiveness of the proposed method. Yudong Chen 0002, Sen Wang 0001, Jianglin Lu, Zhi Chen 0010, Zheng Zhang 0006, Zi Huang |
ACM Multimedia | 3 |
| 2021 | Target redirected regression with dynamic neighborhood structure
Jianglin Lu, Jingxu Lin, Zhihui Lai 0001, Jie Zhou 0009 |
Inf. Sci. | 1 |
| 2021 | Low-rank adaptive graph embedding for unsupervised feature extraction
Jianglin Lu, Jie Zhou 0009, Yudong Chen 0002, Zhihui Lai 0001, Qinghua Hu |
Pattern Recognit. | 1 |
| 2020 | Deep Superpixel Cut for Unsupervised Image SegmentationabstractImage segmentation, one of the most critical vision tasks, has been studied for many years. Most of the early algorithms are unsupervised methods, which use hand-crafted features to divide the image into many regions. Recently, owing to the great success of deep learning technology, CNNs based methods show superior performance in image segmentation. However, these methods rely on a large number of human annotations, which are expensive to collect. In this paper, we propose a deep unsupervised method for image segmentation, which contains the following two stages. First, a Superpixelwise Autoencoder (SuperAE) is designed to learn the deep embedding and reconstruct a smoothed image, then the smoothed image is passed to generate superpixels. Second, we present a novel clustering algorithm called Deep Superpixel Cut (DSC), which measures the deep similarity between superpixels and formulates image segmentation as a soft partitioning problem. Via backpropagation, DSC adaptively partitions the superpixels into perceptual regions. Experimental results on the BSDS500 dataset demonstrate the effectiveness of the proposed method. Qinghong Lin, Weichan Zhong, Jianglin Lu |
ICPR | 3 |
| 2020 | Label Self-Adaption Hashing for Image RetrievalabstractHashing has attracted widespread attention in image retrieval because of its fast retrieval speed and low storage cost. Compared with supervised methods, unsupervised hashing methods are more reasonable and suitable for large-scale image retrieval since it is always difficult and expensive to collect true labels of the massive data. Without label information, however, unsupervised hashing methods can not guarantee the quality of learned binary codes. To resolve this dilemma, this paper proposes a novel unsupervised hashing method called Label Self-Adaption Hashing (LSAH), which contains effective hashing function learning part and self-adaption label generation part. In the first part, we utilize anchor graph to keep the local structure of the data and introduce joint sparsity into the model to extract effective features for high-quality binary code learning. In the second part, a self-adaptive cluster label matrix is learned from the data under the assumption that the nearest neighbor points should have a large probability to be in the same cluster. Therefore, the proposed LSAH can make full use of the potential discriminative information of data to guide the learning of binary codes. It is worth noting that LSAH can learn effective binary codes, hashing function and cluster labels simultaneously in a unified optimization framework. To solve the resulting optimization problem, an Augmented Lagrange Multiplier based iterative algorithm is elaborately designed. Extensive experiments on three large-scale data sets indicate the promising performance of the proposed LSAH. Jianglin Lu, Zhihui Lai 0001, Jingxu Lin, Qinghong Lin, Jie Zhou 0009 |
ICPR | 1 |
| 2019 | Robust Embedding Regression for Face Recognition
Jiaqi Bao, Jianglin Lu, Zhihui Lai 0001, Yuwu Lu |
PRCV (2) | 2 |