EDBT 2026 Demo / reviewers in the wild / expert
Zechuan Li
dblp:350/0441
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2027
0009-0003-4715-703XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
3D vision · 49% Deep learning architectures and training · 22% Language models and text generation · 14% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
positional encoding |
1.6 | 2 | 2025 | A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation · EMNLP 2025 Improved MLP Point Cloud Processing with High-Dimensional Positional Encoding · AAAI 2024 |
Computer vision › 3D vision
3d object detection |
1.5 | 2 | 2025 | GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector · CVPR 2025 AShapeFormer : Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformers · CVPR 2023 |
Machine learning › Deep learning architectures and training
transformer |
1.1 | 2 | 2025 | A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation · EMNLP 2025 AShapeFormer : Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformers · CVPR 2023 |
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector · CVPR 2025 |
Natural language and speech › Language models and text generation › compositional generalization › length generalization
length extrapolation |
0.9 | 1 | 2025 | A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation · EMNLP 2025 |
Natural language and speech › Language models and text generation › language modeling
long-context language modeling |
0.9 | 1 | 2025 | A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation · EMNLP 2025 |
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
multi-view 3d object detection |
0.9 | 1 | 2025 | GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector · CVPR 2025 |
Computer vision › 3D vision
neural radiance field |
0.9 | 1 | 2025 | GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector · CVPR 2025 |
Computer vision › 3D vision › point cloud analysis
point cloud classification and segmentation |
0.8 | 1 | 2024 | Improved MLP Point Cloud Processing with High-Dimensional Positional Encoding · AAAI 2024 |
Computer vision › 3D vision
point cloud processing |
0.8 | 1 | 2024 | Improved MLP Point Cloud Processing with High-Dimensional Positional Encoding · AAAI 2024 |
Computer vision › Video understanding and tracking
video classification |
0.8 | 1 | 2024 | OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition · CVPR 2024 |
Computer vision › Vision and language › vision-language model
vision-language model adaptation |
0.8 | 1 | 2024 | OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition · CVPR 2024 |
Computer vision › Video understanding and tracking › video classification
zero-shot video classification |
0.8 | 1 | 2024 | OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition · CVPR 2024 |
Computer vision › 3D vision
3d scene understanding |
0.7 | 1 | 2023 | AShapeFormer : Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformers · CVPR 2023 |
Computer vision › 3D vision › 3d shape representation
shape encoding |
0.7 | 1 | 2023 | AShapeFormer : Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformers · CVPR 2023 |
Machine learning › Deep learning architectures and training › feedforward neural network
MLP-based architecture |
0.2 | 1 | 2024 | Improved MLP Point Cloud Processing with High-Dimensional Positional Encoding · AAAI 2024 |
Natural language and speech › Language models and text generation
prompting |
0.2 | 1 | 2024 | OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition · CVPR 2024 |
Machine learning › Deep learning architectures and training › attention mechanism
multi-head attention |
0.2 | 1 | 2023 | AShapeFormer : Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformers · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
voxel optimization · 0.9training-free interpolation · 0.9opacity optimization · 0.9importance sampling · 0.9greedy attention logit interpolation · 0.9positional encoding · 0.8optimal matching · 0.8large language model · 0.8abstraction and refinement · 0.8object-scene positional encoding · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | SFA-DiffNet: Spatial-frequency aware diffusion network for robust medical image segmentation
Tongtong Xie, Hongshan Yu, Yan Zheng 0003, Yong He 0012, Zhengeng Yang, Zechuan Li, Naveed Akhtar |
Expert Syst. Appl. | 6 |
| 2026 | Soft-Masked Transformer for Point Cloud Processing With Skip Attention-Based UpsamplingabstractPoint cloud processing methods leverage local and global point features to cater to downstream tasks, yet they often overlook the task-level context inherent in point clouds during the encoding stage. We argue that integrating task-level information into the encoding stage significantly enhances performance. To that end, we propose SMTransformer which incorporates task-level information into a vector-based transformer by utilizing a soft mask generated from task-level queries and keys to learn the attention weights. Additionally, to facilitate effective communication between features from the encoding and decoding layers in high-level tasks such as segmentation, we introduce a skip-attention-based up-sampling block. This block dynamically fuses features from various resolution points across the encoding and decoding layers. To mitigate the increase in network parameters and training time resulting from the complexity of the aforementioned blocks, we propose a novel shared point position encoding strategy. This strategy allows various transformer blocks to share the same position information over the same resolution points, thereby reducing network parameters and training time without compromising accuracy. Experimental comparisons with existing methods on multiple datasets demonstrate the efficacy of SMTransformer and skip-attention-based up-sampling for semantic segmentation task. In particular, we achieve state-of-the-art semantic segmentation results of 73.9% mIoU on S3DIS Area 5 and 62.4% mIoU on SWAN dataset. Note to Practitioners—Point cloud processing underpins automation tasks such as robotic perception, navigation, and inspection, where accurate 3D understanding is essential. Existing methods often prioritize vision benchmarks while overlooking automation needs like efficiency on limited hardware and robustness in real-world environments. The proposed SMTransformer embeds task-level guidance into feature learning and employs skip-attention up-sampling to improve segmentation accuracy with practical efficiency. It is well-suited for robotic manipulation, autonomous driving, and inspection applications. Current limitations include reliance on GPUs and sensitivity to extreme density variations. Future work will target edge-device deployment and multi-task extensions. Yong He 0012, Hongshan Yu, Chaoxu Mu, Mingtao Feng, Tongjia Chen, Zechuan Li, Anwaar Ulhaq, Ajmal Mian |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Multi-Modal Shape Encoding for 3-D Object Detection
Zechuan Li, Hongshan Yu, Niu Zhang, Jinhao Qiao, Wei Sun 0028, Naveed Akhtar |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | LPL3D: LVLM-Driven Pseudo-Labeling for 3D Object DetectionabstractEffective 3D object detection requires large-scale annotated datasets, which are expensive and time-consuming to produce - especially in indoor environments containing dense object arrangements. To address this, we propose an Large Vision-Language Model (LVLM)-driven automatic high-quality pseudo-label generation technique for 3D object detection in single- and multi-view scenarios. We propose an Auto3DLabeler that introduces the first-ever text-to-3D Bounding Box transformation. Its pipeline employs a text-based detector, a segmenter and an LVLM to generate annotation estimates, which are further refined by our IoU-guided iterative Box Aggregator and layout-aware prompt Class Refiner modules. We also introduce a semantic-enhanced multi-modal fusion module that integrates image-level semantic information into point cloud representations for precise detections. Collectively, our contributions provide a remarkable boost to the 3D object detection state-of-the-art. Extensive experiments on SUN RGB-D and ScanNet datasets show our unsupervised detector variant outperforming existing semi-supervised detectors, and our semi-supervised variant achieving up to 28.2% absolute gain in challenging scenarios - all this while maintaining considerable compute advantage over existing label-efficient methods. Our code and models will be made public for the community. Our code and model will be made public after acceptance. Zechuan Li, Hongshan Yu, Yihao Ding, Shuai Yuan 0013, Naveed Akhtar |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object DetectorabstractWe propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is challenging. Addressing that, we introduce a unique 3D positional information embedded voxel optimization mechanism to fuse multi-view features. To prioritize neural field reconstruction in object regions, we also devise a double importance sampling scheme for the NeRF branch of our detector. We additionally propose an opacity optimization module for precise voxel opacity prediction by enforcing multi-view consistency constraints. Moreover, to further improve voxel density consistency across multiple perspectives, we incorporate ray distance as a weighting factor to minimize cumulative ray errors. Our unique modules synergetically form an end-to-end neural model that establishes new state-of-the-art in NeRF-based multi-view 3D detection, verified with extensive experiments on ScanNet and ARKITScenes. Code will be available at https://github.com/ZechuanLi/GO-N3RDet. Zechuan Li, Hongshan Yu, Yihao Ding, Jinhao Qiao, Basim Azam, Naveed Akhtar |
CVPR | 1 |
| 2025 | A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit InterpolationabstractTransformer-based Large Language Models (LLMs) struggle with inputs exceeding their training context window due to positional outof-distribution (O.O.D.) issues that disrupt attention.Existing solutions, including finetuning and training-free methods, face challenges like inefficiency, redundant interpolation, logit outliers, or loss of local positional information.We propose Greedy Attention Logit Interpolation (GALI), a training-free method that improves length extrapolation by greedily reusing pretrained positional intervals and interpolating attention logit to eliminate outliers.GALI achieves stable and superior performance across a wide range of long-context tasks without requiring input-length-specific tuning.Our analysis further reveals that LLMs interpret positional intervals unevenly and that restricting interpolation to narrower ranges improves performance, even on short-context tasks.GALI represents a step toward more robust and generalizable long-text processing in LLMs.Our implementation of GALI, along with the experiments from our paper, is open-sourced at https://github.com/adlnlp/Gali. Yan Li 0186, Zechuan Li, Soyeon Caren Han |
EMNLP | 3 |
| 2024 | Improved MLP Point Cloud Processing with High-Dimensional Positional EncodingabstractMulti-Layer Perceptron (MLP) models are the bedrock of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength. We first develop an “abstraction and refinement” (ABS-REF) view for the neural modeling of point clouds. This view elucidates that whereas the early models focused on the ABS stage, the more recent techniques devise sophisticated REF stages to attain performance advantage in point cloud processing. We then borrow the concept of “positional encoding” from transformer literature, and propose a High-dimensional Positional Encoding (HPE) module, which can be readily deployed to MLP based architectures. We leverage our module to develop a suite of HPENet, which are MLP networks that follow ABS-REF paradigm, albeit with a sophisticated HPE based REF stage. The developed technique is extensively evaluated for 3D object classification, object part segmentation, semantic segmentation and object detection. We establish new state-of-the-art results of 87.6 mAcc on ScanObjectNN for object classification, and 85.5 class mIoU on ShapeNetPart for object part segmentation, and 72.7 and 78.7 mIoU on Area-5 and 6-fold experiments with S3DIS for semantic segmentation. The source code for this work is available at https://github.com/zouyanmei/HPENet. Yanmei Zou, Hongshan Yu, Zhengeng Yang, Zechuan Li, Naveed Akhtar |
AAAI | 4 |
| 2024 | OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video RecognitionabstractDue to the resource-intensive nature of training vision- language models on expansive video data, a majority of studies have centered on adapting pre-trained image- language models to the video domain. Dominant pipelines propose to tackle the visual discrepancies with additional temporal learners while overlooking the substantial discrepancy for web-scaled descriptive narratives and concise action category names, leading to less distinct semantic space and potential performance limitations. In this work, we prioritize the refinement of text knowledge to facilitate generalizable video recognition. To address the limitations of the less distinct semantic space of category names, we prompt a large language model (LLM) to augment action class names into Spatio-Temporal Descriptors thus bridging the textual discrepancy and serving as a knowledge base for general recognition. Moreover, to assign the best descriptors with different video instances, we propose Optimal Descriptor Solver, forming the video recognition problem as solving the optimal matching flow across frame-level representations and descriptors. Comprehensive evaluations in zero-shot, few-shot, and fully supervised video recognition highlight the effectiveness of our approach. Our best model achieves a state-of-the-art zero-shot accuracy of 75.1% on Kinetics-600. Tom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Zechuan Li, Wei Sun 0028, Chen Chen 0001 |
CVPR | 4 |
| 2024 | MVDetector: Malicious Vehicles Detection Under Sybil Attacks in VANETs
Weiye Qi, Zechuan Li, Yufan Han, Yingxu Lai |
ISC (2) | 3 |
| 2024 | EFRNet-VL: An end-to-end feature refinement network for monocular visual localization in dynamic environments
Jingwen Wang 0009, Hongshan Yu, Xuefei Lin, Zechuan Li, Wei Sun 0028, Naveed Akhtar |
Expert Syst. Appl. | 4 |
| 2023 | AShapeFormer : Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformersabstract3D object detection techniques commonly follow a pipeline that aggregates predicted object central point features to compute candidate points. However, these candidate points contain only positional information, largely ignoring the object-level shape information. This eventually leads to sub-optimal 3D object detection. In this work, we propose AShapeFormer, a semantics-guided object-level shape encoding module for 3D object detection. This is a plug-n-play module that leverages multi-head attention to encode object shape information. We also propose shape tokens and object-scene positional encoding to ensure that the shape information is fully exploited. Moreover, we introduce a semantic guidance sub-module to sample more foreground points and suppress the influence of background points for a better object shape perception. We demonstrate a straightforward enhancement of multiple existing methods with our AShapeFormer. Through extensive experiments on the popular SUN RGB-D and ScanNetV2 dataset, we show that our enhanced models are able to outperform the baselines by a considerable absolute margin of up to 8.1%. Code will be available at https://github.com/ZechuanLi/AShapeFormer Zechuan Li, Hongshan Yu, Zhengeng Yang, Tom Tongjia Chen, Naveed Akhtar |
CVPR | 1 |