VLDB 2026 Research / reviewers in the wild / expert
Lei Li 0051
dblp:13/7007-51
· DBLP profile ↗
16ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-4514-3617ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accurate is Not Necessarily the Best: Edge-Assisted Bitrate Re-Adaptation for Video StreamingabstractThe increasing volume of video traffic presents significant challenges to network transmission, while edge computing accelerates video delivery by leveraging caching and computation to optimize content forwarding. However, as edge computing is generally deployed by service providers in a transparent manner, clients cannot perceive edge states, e.g., cache availability, potentially resulting in suboptimal bitrate decisions. This issue persists even with intelligent bitrate selection approaches on the client side, as the inaccurate estimation of network delivery capacity due to edge cache transparency remains unresolved. Meanwhile, single-edge servers or nodes, with limited cache space and computational capacity for a small number of users, can be more effective by aggregating into clusters to better serve users and optimize resource utilization. Therefore, we propose an edge-assisted bitrate re-adaptation scheme (e-BitRead) for adaptive streaming, utilizing neighbor edges to accelerate video deliveries.e-BitReadintroduces three key innovations: (i) it employs a bitrate re-adaptation mechanism that intelligently selects alternative bitrates from edge servers instead of strictly responding with the requested bitrate, (ii) it utilizes collaborative caching across multiple edge servers to expand available bitrate options through coordinated resource sharing, and (iii) it enhances the learning efficiency through joint optimization of network architecture and reward design, which leverages actor-critic structure to fit into multi-edge bitrate adaptation. In experiments with an intelligent client ABR,e-BitReaddemonstrates its superiority by achieving a 1.43x higher hit ratio compared to the baseline, while improving QoE by 1.93x over the scheme without smart bitrate matching and delivering a 33% gain over the single-edge re-adaptation approach. Wanxin Shi, Weijia Lang, Qing Li 0006, Gengbiao Shen, Lei Li 0051, Yang Xu 0010, Yong Jiang 0001, Gabriel-Miro Muntean |
IEEE Trans. Netw. | 6 |
| 2025 | Visual-Semantic Dual Calibration Network for Zero-Shot Learning
Qingyang Hao, Lei Li 0051, Yu Li 0006, Chun Yuan 0003 |
ICIC (9) | 2 |
| 2025 | Overlooked Factors in Continual Zero-Shot Learning: Inflexible Semantic Prototypes, Simplistic Loss Functions, and SGD NoiseabstractZero-Shot Learning (ZSL) enables the recognition of unseen classes by transferring semantic knowledge from seen classes, typically through shared attributes. Continual Zero-Shot Learning (CZSL) advances the concept by enabling the model to continuously recognize new classes while retaining the ability to recognize previously learned ones. A common approach in CZSL is generative replay, where pseudo-data is synthesized to retain recognition of old classes while adapting to new ones. Despite this, existing methods rely on manually defined semantic prototypes, which may not align well with diverse visual representations. In this paper, we propose a method that refines semantic prototypes in the continual setting, treating them similarly to visual feature generation for better alignment with category semantics. We also introduce a contrastive loss on generated embeddings, which outperforms simplistic loss functions used in prior work. Furthermore, we investigate the impact of SGD (Stochastic Gradient Descent) noise, a factor often overlooked in CZSL, highlighting its significant role in model convergence and generalization. Our contributions improve CZSL performance and enhance the understanding of SGD noise in this context. Qingyang Hao, Lei Li 0051, Chun Yuan 0003 |
ICIP | 2 |
| 2025 | Boosting Long-Tailed Recognition With Label Descriptor and BeyondabstractLong-Tailed Recognition (LTR) poses significant challenges due to the heavily imbalanced nature of real-world data, which severely skews data-driven deep neural networks. Despite the rapid progress of Vision-Language Models (VLMs), they still face challenges in effectively learning from long-tailed visual data. In this paper, we present a comprehensive analysis of the reasons behind the underperformance of VLMs and propose a hierarchical inference framework to address this issue. Specifically, we prompt the large language models to generatesentence-leveldescriptors for class labels and conduct the open vocabulary classification by computing the average similarity between the image and each descriptor. Areweightingmechanism is further proposed to filter out uninformative descriptors. To mitigate model bias incurred by the long-tail distribution, we propose a feature adapter with the logit adjustment technique and fine-tune the CLIP model via visual prompt tokens. We introduce the Shared Feature space Mixup (SFM) to enhance the interaction between modalities to address tail visual feature insufficiency. Finally, we propose a hierarchical inference manner to combine the aforementioned proposals. Extensive evaluations demonstrate that our approach achieves state-of-the-art performance by fine-tuning only a few parameters on the Places-LT, ImageNet-LT, and iNaturalist 2018 benchmarks. Zhengzhuo Xu, Ruikang Liu, Zenghao Chai, Yiyan Qi, Lei Li 0051, Haiqin Yang, Chun Yuan 0003 |
IEEE Trans. Multim. | 5 |
| 2024 | Sparse Model Inversion: Efficient Inversion of Vision Transformers for Data-Free ApplicationsabstractModel inversion, which aims to reconstruct the original training data from pre-trained discriminative models, is especially useful when the original training data is unavailable due to privacy, usage rights, or size constraints. However, existing dense inversion methods attempt to reconstruct the entire image area, making them extremely inefficient when inverting high-resolution images from large-scale Vision Transformers (ViTs). We further identify two underlying causes of this inefficiency: the redundant inversion of noisy backgrounds and the unintended inversion of spurious correlations—a phenomenon we term “hallucination” in model inversion. To address these limitations, we propose a novel sparse model inversion strategy, as a plug-and-play extension to speed up existing dense inversion methods with no need for modifying their original loss functions. Specifically, we selectively invert semantic foregrounds while stopping the inversion of noisy backgrounds and potential spurious correlations. Through both theoretical and empirical studies, we validate the efficacy of our approach in achieving significant inversion acceleration (up to $\times$3.79) while maintaining comparable or even enhanced downstream performance in data-free model quantization and data-free knowledge transfer. Code is available at https://github.com/Egg-Hu/SMI. Yongxian Wei, Li Shen 0008, Zhenyi Wang 0001, Lei Li 0051, Chun Yuan 0003, Dacheng Tao |
ICML | 5 |
| 2024 | Semantic Distillation from Neighborhood for Composed Image RetrievalabstractThe challenging task composed image retrieval targets at identifying the matched image from the multi-modal query with a reference image and a textual modifier. Most existing methods are devoted to composing the unified query representations from the query images and texts, yet the distribution gaps between the hybrid-modal query representations and visual target representations are neglected. However, directly incorporating target features on the query may cause ambiguous rankings and poor robustness due to the insufficient exploration of the distinguishments and overfitting issues. To address the above concerns, we propose a novel framework termed SemAntic Distillation from Neighborhood (SADN) for composed image retrieval. For mitigating the distribution divergences, we construct neighborhood sampling from the target domain for each query and aggregate neighborhood features with adaptive weights to restructure the query representations. Specifically, the adaptive weights are determined by the collaboration of two individual modules, as correspondence-induced adaption and divergence-based correction. Correspondence-induced adaption accounts for capturing the correlation alignments from neighbor features under the guidance of the positive representations, and the divergence-based correction regulates the weights based on the embedding distances between hard negatives and the query in the latent space. Extensive results and ablation studies on CIRR and FashionIQ validate that the proposed semantic distillation from neighborhood significantly outperforms baseline methods. Yifan Wang 0027, Wuliang Huang, Lei Li 0051, Chun Yuan 0003 |
ACM Multimedia | 3 |
| 2024 | Progress and opportunities of foundation models in bioinformaticsabstractBioinformatics has undergone a paradigm shift in artificial intelligence (AI), particularly through foundation models (FMs), which address longstanding challenges in bioinformatics such as limited annotated data and data noise. These AI techniques have demonstrated remarkable efficacy across various downstream validation tasks, effectively representing diverse biological entities and heralding a new era in computational biology. The primary goal of this survey is to conduct a general investigation and summary of FMs in bioinformatics, tracing their evolutionary trajectory, current research landscape, and methodological frameworks. Our primary focus is on elucidating the application of FMs to specific biological problems, offering insights to guide the research community in choosing appropriate FMs for tasks like sequence analysis, structure prediction, and function annotation. Each section delves into the intricacies of the targeted challenges, contrasting the architectures and advancements of FMs with conventional methods and showcasing their utility across different biological domains. Further, this review scrutinizes the hurdles and constraints encountered by FMs in biology, including issues of data noise, model interpretability, and potential biases. This analysis provides a theoretical groundwork for understanding the circumstances under which certain FMs may exhibit suboptimal performance. Lastly, we outline prospective pathways and methodologies for the future development of FMs in biological research, facilitating ongoing innovation in the field. This comprehensive examination not only serves as an academic reference but also as a roadmap for forthcoming explorations and applications of FMs in biology. Qing Li 0075, Zhihang Hu, Lei Li 0051, Yimin Fan, Irwin King, Gengjie Jia, Yu Li 0006 |
Briefings Bioinform. | 4 |
| 2024 | Learning meaningful representation of single-neuron morphology via large-scale pre-trainingabstractSUMMARY: Single-neuron morphology, the study of the structure, form, and shape of a group of specialized cells in the nervous system, is of vital importance to define the type of neurons, assess changes in neuronal development and aging and determine the effects of brain disorders and treatments. Despite the recent surge in the amount of available neuron morphology reconstructions due to advancements in microscopy imaging, existing computational and deep learning methods for modeling neuron morphology have been limited in both scale and accuracy. In this paper, we propose MorphRep, a model for learning meaningful representation of neuron morphology pre-trained with over 250 000 existing neuron morphology data. By encoding the neuron morphology into graph-structured data, using graph transformers for feature encoding and enforcing the consistency between multiple augmented views of neuron morphology, MorphRep achieves the state of the art performance on widely used benchmarking datasets. Meanwhile, MorphRep can accurately characterize the neuron morphology space across neuron morphometrics, fine-grained cell types, brain regions and ages. Furthermore, MorphRep can be applied to distinguish neurons under a wide range of conditions, including genetic perturbation, drug injection, environment change and disease. In summary, MorphRep provides an effective strategy to embed and represent neuron morphology and can be a valuable tool in integrating cell morphology into single-cell multiomics analysis. AVAILABILITY AND IMPLEMENTATION: The codebase has been deposited in https://github.com/YaxuanLi-cn/MorphRep. Yimin Fan, Yaxuan Li 0002, Yunhua Zhong, Lei Li 0051, Yu Li 0006 |
Bioinform. | 5 |
| 2024 | Meta-Learning Without Data via Unconditional Diffusion ModelsabstractAlthough few-shot learning aims to address data scarcity, it still requires large, annotated datasets for training, which are often unavailable due to cost and privacy concerns. Previous studies have utilized pre-trained diffusion models, either to synthesize auxiliary data besides limited labeled samples, or to employ diffusion models as zero-shot classifiers. However, they are limited to conditional diffusion models needing class prior information (e.g., carefully crafted text prompts) about unseen tasks. To overcome this, we leverage unconditional diffusion models without needs for class information to train a meta-model capable of generalizing to unseen tasks. The framework contains(1)a meta-learning without data approach that uses synthetic data during training; and(2)a diffusion model-based data augmentation to calibrate the distribution shift during testing. During meta-training, we implement aself-taughtclass-learner to gradually capture class concepts, guiding unconditional diffusion models to generate alabeledpseudo dataset. This pseudo dataset is then used to jointly train the class-learner and the meta-model, allowing for iterative refinement and clear differentiation between classes. During meta-testing, we introduce a data augmentation that employs the diffusion models used in meta-training, to narrow the gap between meta-training and meta-testing task distribution. This enables the meta-model trained onsyntheticimages to effectively classifyrealimages in unseen tasks. Comprehensive experiments showcase the superiority and adaptability of our approach in four real-world scenarios. Code available athttps://github.com/WalkerWorldPeace/MLWDUDM. Yongxian Wei, Li Shen 0008, Zhenyi Wang 0001, Lei Li 0051, Yu Li 0006, Chun Yuan 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | StrokeNet: Stroke Assisted and Hierarchical Graph Reasoning NetworksabstractScene text detection is still a challenging task, as there may be extremely small or low-resolution strokes and close or arbitrary-shaped texts. In this paper, StrokeNet proposes to effectively detect the texts by capturing the fine-grained strokes and inferring structural relations between the hierarchical representations of each text area in the graph-based network. Different from existing approaches that represent the text area by a series of points or rectangular boxes, we directly localize the strokes of each text instance. We introduce Stroke Assisted Prediction Network (SAPN), which performs hierarchical representation learning of text areas, effectively capturing extremely small or low-resolution texts. We extract a series of text- and stroke-level rectangular boxes on the predicted text areas, which are treated as graph nodes and grouped to form the corresponding local graphs. Hierarchical Relation Graph Network (HRGN) then performs relational reasoning and predicts the likelihood of linkages among graph nodes of different levels. It efficiently splits the close text instances and grouping node classification results into the arbitrary-shaped text area. We introduce a novel dataset with stroke-level annotations, namelySynthStroke, for offline pre-training of widespread text detectors. Experiments on benchmarks verify the State-of-the-Art performance of our method. Lei Li 0051, Kai Fan 0002, Chun Yuan 0003 |
IEEE Trans. Multim. | 1 |
| 2022 | Cross-modal Representation Learning and Relation Reasoning for Bidirectional Adaptive ManipulationabstractSince single-modal controllable manipulation typically requires supervision of information from other modalities or cooperation with complex software and experts, this paper addresses the problem of cross-modal adaptive manipulation (CAM). The novel task performs cross-modal semantic alignment from mutual supervision and implements bidirectional exchange of attributes, relations, or objects in parallel, benefiting both modalities while significantly reducing manual effort. We introduce a robust solution for CAM, which includes two essential modules, namely Heterogeneous Representation Learning (HRL) and Cross-modal Relation Reasoning (CRR). The former is designed to perform representation learning for cross-modal semantic alignment on heterogeneous graph nodes. The latter is adopted to identify and exchange the focused attributes, relations, or objects in both modalities. Our method produces pleasing cross-modal outputs on CUB and Visual Genome. Lei Li 0051, Kai Fan 0002, Chun Yuan 0003 |
IJCAI | 1 |
| 2022 | PTS: A Prompt-based Teacher-Student Network for Weakly Supervised Aspect DetectionabstractMost existing weakly supervised aspect detection algorithms utilize pre-trained language models as their backbone networks by constructing discriminative tasks with seed words. Once the number of seed words decreases, the performance of current models declines significantly. Recently, prompt tuning has been proposed to bridge the gap of objective forms in pre-training and fine-tuning, which is hopeful of alleviating the above challenge. However, directly applying the existing prompt-based methods to this task not only fails to effectively use large amounts of unlabeled data, but also may cause serious over-fitting problems. In this paper, we propose a lightweight teacher-student network (PTS) based on prompts to solve the above two problems. Concretely, the student network is a hybrid prompt-based classification model to detect aspects, which innovatively compounds hand-crafted prompts and auto-generated prompts. The teacher network comprehensively considers the representation of the sentence and the masked aspect token in the template to guide classification. To utilize unlabeled data and seed words intelligently, we train the teacher and student network alternately. Furthermore, in order to solve the problem that the uneven quality of training data obviously affects the iterative efficiency of PTS, we design a general dynamic data selection strategy to feed the most pertinent data into the current model. Experimental results show that even given the minimum seed words, PTS significantly outperforms previous state-of-the-art methods on three widely used benchmarks. Lingyu Yang, Lei Li 0051, Chengyin Xu, Shutao Xia, Chun Yuan 0003 |
IJCNN | 3 |
| 2022 | ACEs: Unsupervised Multi-label Aspect Detection with Aspect-category ExpertsabstractUnsupervised aspect detection (UAD) aims to identify the aspect categories mentioned in product reviews automatically. Existing unsupervised methods mainly focus on single-label prediction which cannot work well for the realistic application scene since a review usually contain multiple aspect categories. Recent attempts alleviate this issue by setting a threshold. However, the imbalance of the number of category-related words in review segments makes these methods difficult to find optimal thresholds to recall all categories, which is a common but neglected phenomenon. In this paper, we propose a novel unsupervised method termed Aspect-Category Experts (ACEs) to address this problem. Our goal is to train a set of aspect-category experts to encode the sentence in parallel, where experts and aspect categories correspond one-to-one. Experts in different aspects weight the embedding of representative words with aspect-specific attention to avoid the negative impact of accumulation. Besides, to enhance the complementarity between different experts to reduce inter-class feature entanglement, we construct a novel mutual exclusion loss (ME loss) to improve the aspect detection performance. Extensive experimental results on four datasets demonstrate that our proposed ACEs model outperforms the previous state-of-the-art methods. Lingyu Yang, Lei Li 0051, Chun Yuan 0003, Shutao Xia |
IJCNN | 3 |
| 2021 | Explore Hierarchical Relations Reasoning and Global Information Aggregation
Lei Li 0051, Chun Yuan 0003, Kai Fan 0002 |
ICDAR (1) | 1 |
| 2021 | VQMG: Hierarchical Vector Quantised and Multi-hops Graph Reasoning for Explicit Representation LearningabstractVector Quantized Variational AutoEncoder (VQ-VAE) models realize fast image generation by encoding and quantifying the raw input in the single-level or hierarchical compressed latent space. However, the learned representations are not expert in capturing complex relations existed, while one usually adopts domain-specific autoregressive models to fit a prior distribution for two stages of learning. In this work, we propose VQMG, a novel and unified framework for multi-hops relational reasoning and explicit representation learning. By introducing Multi-hops Graph Convolution Networks (MGCN), complicated relations from hierarchical latent space are effectively captured by Inner Graph, while the fitting of autoregressive prior are performed coherently by Outer Graph to promote the performance. Experiments on multimedia tasks including Point cloud segementation, Stroke-level text detection and Image generation verify the efficiency and applicability of our approach. Lei Li 0051, Chun Yuan 0003 |
ACM Multimedia | 1 |
| 2020 | Feature Augmented Memory with Global Attention Network for VideoQAabstractRecently, Recurrent Neural Network (RNN) based methods and Self-Attention (SA) based methods have achieved promising performance in Video Question Answering (VideoQA). Despite the success of these works, RNN-based methods tend to forget the global semantic contents due to the inherent drawbacks of the recurrent units themselves, while SA-based methods cannot precisely capture the dependencies of the local neighborhood, leading to insufficient modeling for temporal order. To tackle these problems, we propose a novel VideoQA framework which progressively refines the representations of videos and questions from fine to coarse grain in a sequence-sensitive manner. Specifically, our model improves the feature representations via the following two steps: (1) introducing two fine-grained feature-augmented memories to strengthen the information augmentation of video and text which can improve memory capacity by memorizing more relevant and targeted information. (2) appending the self-attention and co-attention module to the memory output thus the module is able to capture global interaction between high-level semantic informations. Experimental results show that our approach achieves state-of-the-art performance on VideoQA benchmark datasets. Jiayin Cai, Chun Yuan 0003, Lei Li 0051, Yangyang Cheng, Ying Shan |
IJCAI | 4 |