EDBT 2026 Demo / reviewers in the wild / expert
Lei Zhu 0016
dblp:99/549-16
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2024
0000-0002-0191-7086ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 48% Efficient and distributed learning · 29% Trustworthy machine learning · 7% | |
| Computer graphics and multimedia
5 papers |
Rendering · 51% Image and video processing · 33% Visual content generation and editing · 16% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 100% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › transformer
vision transformer |
1.4 | 2 | 2024 | Revisiting the Integration of Convolution and Attention for Vision Backbone · NeurIPS 2024 BiFormer: Vision Transformer with Bi-Level Routing Attention · CVPR 2023 |
Machine learning › Efficient and distributed learning
attention computation |
0.8 | 1 | 2024 | RelayAttention for Efficient Large Language Model Serving with Long System Prompts · ACL (1) 2024 |
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention |
0.8 | 1 | 2024 | Revisiting the Integration of Convolution and Attention for Vision Backbone · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.8 | 1 | 2024 | RelayAttention for Efficient Large Language Model Serving with Long System Prompts · ACL (1) 2024 |
Machine learning › Deep learning architectures and training
vision backbone |
0.8 | 1 | 2024 | Revisiting the Integration of Convolution and Attention for Vision Backbone · NeurIPS 2024 |
Rendering › gaussian splatting
3d gaussian splatting |
0.8 | 1 | 2024 | Analytic-Splatting: Anti-Aliased 3D Gaussian Splatting via Analytic Integration · ECCV (17) 2024 |
Rendering
antialiasing |
0.8 | 1 | 2024 | Analytic-Splatting: Anti-Aliased 3D Gaussian Splatting via Analytic Integration · ECCV (17) 2024 |
Rendering
inverse rendering |
0.8 | 1 | 2024 | Inverse Rendering of Glossy Objects via the Neural Plenoptic Function and Radiance Fields · CVPR 2024 |
Rendering › inverse rendering
reflectance and illumination estimation |
0.8 | 1 | 2024 | Inverse Rendering of Glossy Objects via the Neural Plenoptic Function and Radiance Fields · CVPR 2024 |
Rendering
volume rendering |
0.8 | 1 | 2024 | Analytic-Splatting: Anti-Aliased 3D Gaussian Splatting via Analytic Integration · ECCV (17) 2024 |
Memory systems › cache
key-value cache |
0.8 | 1 | 2024 | RelayAttention for Efficient Large Language Model Serving with Long System Prompts · ACL (1) 2024 |
Memory systems
memory access optimization |
0.8 | 1 | 2024 | RelayAttention for Efficient Large Language Model Serving with Long System Prompts · ACL (1) 2024 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.7 | 1 | 2023 | BiFormer: Vision Transformer with Bi-Level Routing Attention · CVPR 2023 |
Image and video processing
image enhancement |
0.7 | 1 | 2023 | Neural Preset for Color Style Transfer · CVPR 2023 |
Image and video processing › image enhancement
low-light image enhancement |
0.7 | 1 | 2023 | Neural Preset for Color Style Transfer · CVPR 2023 |
Visual content generation and editing
image and video editing |
0.6 | 1 | 2022 | Harmonizer: Learning to Perform White-Box Image and Video Harmonization · ECCV (15) 2022 |
Visual content generation and editing › image editing › image compositing
image harmonization |
0.6 | 1 | 2022 | Harmonizer: Learning to Perform White-Box Image and Video Harmonization · ECCV (15) 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Mitigating Intensity Bias in Shadow Detection via Feature Decomposition and Reweighting · ICCV 2021 |
Image and video processing › image enhancement › shadow detection and removal
shadow detection |
0.5 | 1 | 2021 | Mitigating Intensity Bias in Shadow Detection via Feature Decomposition and Reweighting · ICCV 2021 |
Computer vision › 3D vision
3d reconstruction |
0.2 | 1 | 2024 | Inverse Rendering of Glossy Objects via the Neural Plenoptic Function and Radiance Fields · CVPR 2024 |
Computer vision › 3D vision › 3d reconstruction
geometric reconstruction |
0.2 | 1 | 2024 | Inverse Rendering of Glossy Objects via the Neural Plenoptic Function and Radiance Fields · CVPR 2024 |
Computer vision › Segmentation and scene understanding
dense prediction |
0.2 | 1 | 2023 | BiFormer: Vision Transformer with Bi-Level Routing Attention · CVPR 2023 |
Machine learning › Generative modeling › generative adversarial network
image-to-image translation |
0.2 | 1 | 2023 | Neural Preset for Color Style Transfer · CVPR 2023 |
Machine learning › Representation and self-supervised learning › feature transformation
feature decomposition |
0.1 | 1 | 2021 | Mitigating Intensity Bias in Shadow Detection via Feature Decomposition and Reweighting · ICCV 2021 |
Methods — techniques the papers use, named apart from their topics
ray tracing · 1.5pre-filtered radiance fields · 1.5neural plenoptic function · 1.5memory access reduction · 1.5material-aware cone sampling · 1.5attention reformulation · 1.5self-supervised learning · 1.3image-adaptive color mapping matrix · 1.3soft clustering · 0.8multi-head self-attention · 0.8gaussian splatting · 0.8analytic integration · 0.8white-box learning · 0.6feature decomposition · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RelayAttention for Efficient Large Language Model Serving with Long System PromptsabstractA practical large language model (LLM) service may involve a long system prompt, which specifies the instructions, examples, and knowledge documents of the task and is reused across requests. However, the long system prompt causes throughput/latency bottlenecks as the cost of generating the next token grows w.r.t the sequence length. This paper aims to improve the efficiency of LLM services that involve long system prompts. Our key observation is that handling these system prompts requires heavily redundant memory accesses in existing causal attention computation algorithms. Specifically, for batched requests, the cached hidden states (i.e., key-value pairs) of system prompts are transferred from off-chip DRAM to on-chip SRAM multiple times, each corresponding to an individual request. To eliminate such a redundancy, we propose RelayAttention, an attention algorithm that allows reading these hidden states from DRAM exactly once for a batch of input tokens. RelayAttention is a free lunch: it maintains the generation quality while requiring no model retraining, as it is based on a mathematical reformulation of causal attention. We have observed significant performance improvements to a production-level system, vLLM, through integration with RelayAttention. The improvements are even more profound with longer system prompts. © 2024 Association for Computational Linguistics Lei Zhu 0016, Xinjiang Wang, Wayne Zhang 0001, Rynson W. H. Lau |
ACL (1) | 1 |
| 2024 | Inverse Rendering of Glossy Objects via the Neural Plenoptic Function and Radiance FieldsabstractInverse rendering aims at recovering both geometry and materials of objects. It provides a more compatible re-construction for conventional rendering engines, compared with the neural radiance fields (NeRFs). On the other hand, existing NeRF-based inverse rendering methods can-not handle glossy objects with local light interactions well, as they typically oversimplify the illumination as a 2D environmental map, which assumes infinite lights only. Observing the superiority of NeRFs in recovering radiance fields, we propose a novel 5D Neural Plenoptic Function (NeP) based on NeRFs and ray tracing, such that more accurate lighting-object interactions can be formulated via the ren-dering equation. We also design a material-aware cone sampling strategy to efficiently integrate lights inside the BRDF lobes with the help of pre-filtered radiance fields. Our method has two stages: the geometry of the target ob-ject and the pre-filtered environmental radiance fields are reconstructed in the first stage, and materials of the target object are estimated in the second stage with the proposed NeP and material-aware cone sampling strategy. Exten-sive experiments on the proposed real-world and synthetic datasets demonstrate that our method can reconstruct high-fidelity geometry/materials of challenging glossy objects with complex lighting interactions from nearby objects. Project webpage: https://whyy.si.te/paper/nep Lei Zhu 0016, Rynson W. H. Lau |
CVPR | 3 |
| 2024 | Analytic-Splatting: Anti-Aliased 3D Gaussian Splatting via Analytic Integration
Zhihao Liang 0002, Qi Zhang 0029, Wenbo Hu 0002, Lei Zhu 0016, Kui Jia |
ECCV (17) | 4 |
| 2024 | Revisiting the Integration of Convolution and Attention for Vision BackboneabstractConvolutions (Convs) and multi-head self-attentions (MHSAs) are typically considered alternatives to each other for building vision backbones. Although some works try to integrate both, they apply the two operators simultaneously at the finest pixel granularity. With Convs responsible for per-pixel feature extraction already, the question is whether we still need to include the heavy MHSAs at such a fine-grained level. In fact, this is the root cause of the scalability issue w.r.t. the input resolution for vision transformers. To address this important problem, we propose in this work to use MSHAs and Convs in parallel \textbf{at different granularity levels} instead. Specifically, in each layer, we use two different ways to represent an image: a fine-grained regular grid and a coarse-grained set of semantic slots. We apply different operations to these two representations: Convs to the grid for local features, and MHSAs to the slots for global features. A pair of fully differentiable soft clustering and dispatching modules is introduced to bridge the grid and set representations, thus
enabling local-global fusion. Through extensive experiments on various vision tasks, we empirically verify the potential of the proposed integration scheme, named \textit{GLMix}: by offloading the burden of fine-grained features to light-weight Convs, it is sufficient to use MHSAs in a few (e.g., 64) semantic slots to match the performance of recent state-of-the-art backbones, while being more efficient. Our visualization results also demonstrate that the soft clustering module produces a meaningful semantic grouping effect with only IN1k classification supervision, which may induce better interpretability and inspire new weakly-supervised semantic segmentation approaches. Code will be available at \url{https://github.com/rayleizhu/GLMix}. Lei Zhu 0016, Xinjiang Wang, Wayne Zhang 0001, Rynson W. H. Lau |
NeurIPS | 1 |
| 2024 | scCaT: An explainable capsulating architecture for sepsis diagnosis transferring from single-cell RNA sequencingabstractSepsis is a life-threatening condition characterized by an exaggerated immune response to pathogens, leading to organ damage and high mortality rates in the intensive care unit. Although deep learning has achieved impressive performance on prediction and classification tasks in medicine, it requires large amounts of data and lacks explainability, which hinder its application to sepsis diagnosis. We introduce a deep learning framework, called scCaT, which blends the capsulating architecture with Transformer to develop a sepsis diagnostic model using single-cell RNA sequencing data and transfers it to bulk RNA data. The capsulating architecture effectively groups genes into capsules based on biological functions, which provides explainability in encoding gene expressions. The Transformer serves as a decoder to classify sepsis patients and controls. Our model achieves high accuracy with an AUROC of 0.93 on the single-cell test set and an average AUROC of 0.98 on seven bulk RNA cohorts. Additionally, the capsules can recognize different cell types and distinguish sepsis from control samples based on their biological pathways. This study presents a novel approach for learning gene modules and transferring the model to other data types, offering potential benefits in diagnosing rare diseases with limited subjects. Xubin Zheng, Dian Meng, Wan-Ki Wong, Ka-Ho To, Lei Zhu 0016, Jiafei Wu, Yining Liang, Kwong-Sak Leung, Man Hon Wong 0001, Lixin Cheng |
PLoS Comput. Biol. | 6 |
| 2023 | Neural Preset for Color Style TransferabstractIn this paper, we present a Neural Preset technique to address the limitations of existing color style transfer methods, including visual artifacts, vast memory requirement, and slow style switching speed. Our method is based on two core designs. First, we propose Deterministic Neural Color Mapping (DNCM) to consistently operate on each pixel via an image-adaptive color mapping matrix, avoiding artifacts and supporting high-resolution inputs with a small memory footprint. Second, we develop a two-stage pipeline by dividing the task into color normalization and stylization, which allows efficient style switching by extracting color styles as presets and reusing them on normalized input images. Due to the unavailability of pairwise datasets, we describe how to train Neural Preset via a self-supervised strategy. Various advantages of Neural Preset over existing methods are demonstrated through comprehensive evaluations. Besides, we show that our trained model can naturally support multiple applications without fine-tuning, including low-light image enhancement, underwater image correction, image dehazing, and image harmonization. The project page is: https://ZHKKKe.github.io/NeuralPreset. Zhanghan Ke, Yuhao Liu 0001, Lei Zhu 0016, Nanxuan Zhao, Rynson W. H. Lau |
CVPR | 3 |
| 2023 | BiFormer: Vision Transformer with Bi-Level Routing AttentionabstractAs the core building block of vision transformers, attention is a powerful tool to capture long-range dependency. However, such power comes at a cost: it incurs a huge computation burden and heavy memory footprint as pairwise token interaction across all spatial locations is computed. A series of works attempt to alleviate this problem by introducing handcrafted and content-agnostic sparsity into attention, such as restricting the attention operation to be inside local windows, axial stripes, or dilated windows. In contrast to these approaches, we propose a novel dynamic sparse attention via bi-level routing to enable a more flexible allocation of computations with content awareness. Specifically, for a query, irrelevant key-value pairs are first filtered out at a coarse region level, and then fine-grained token-to-token attention is applied in the union of remaining candidate regions (i.e., routed regions). We provide a simple yet effective implementation of the proposed bilevel routing attention, which utilizes the sparsity to save both computation and memory while involving only GPU-friendly dense matrix multiplications. Built with the proposed bi-level routing attention, a new general vision transformer, named BiFormer, is then presented. As BiFormer attends to a small subset of relevant tokens in a query adaptive manner without distraction from other irrelevant ones, it enjoys both good performance and high computational efficiency, especially in dense prediction tasks. Empirical results across several computer vision tasks such as image classification, object detection, and semantic segmentation verify the effectiveness of our design. Code is available at https://github.com/rayleizhu/BiFormer. Lei Zhu 0016, Xinjiang Wang, Zhanghan Ke, Wayne Zhang 0001, Rynson W. H. Lau |
CVPR | 1 |
| 2022 | Harmonizer: Learning to Perform White-Box Image and Video Harmonization
Zhanghan Ke, Chunyi Sun, Lei Zhu 0016, Ke Xu 0010, Rynson W. H. Lau |
ECCV (15) | 3 |
| 2021 | Mitigating Intensity Bias in Shadow Detection via Feature Decomposition and ReweightingabstractAlthough CNNs have achieved remarkable progress on the shadow detection task, they tend to make mistakes in dark non-shadow regions and relatively bright shadow regions. They are also susceptible to brightness change. These two phenomenons reveal that deep shadow detectors heavily depend on the intensity cue, which we refer to as intensity bias. In this paper, we propose a novel feature decomposition and reweighting scheme to mitigate this intensity bias, in which multi-level integrated features are decomposed into intensity-variant and intensity-invariant components through self-supervision. By reweighting these two types of features, our method can reallocate the attention to the corresponding latent semantics and achieves balanced exploitation of them. Extensive experiments on three popular datasets show that the proposed method outperforms state-of-the-art shadow detectors. Lei Zhu 0016, Ke Xu 0010, Zhanghan Ke, Rynson W. H. Lau |
ICCV | 1 |