VLDB 2026 Research / reviewers in the wild / expert
Li Jun
dblp:98/772
· DBLP profile ↗
6ranked-venue papers
2as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
1.0 | 1 | 2026 | LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model · ACL (1) 2026 |
Natural language and speech › Language models and text generation
tokenization |
1.0 | 1 | 2026 | LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
parallel tokenization · 1.0character-position-based matching · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language ModelabstractLong context inference scenarios have become increasingly important for large language models, yet they introduce significant computational latency. While prior research has optimized long-sequence inference through operators, model architectures, and system frameworks, tokenization remains an overlooked bottleneck. Existing parallel tokenization methods accelerate processing through text segmentation and multi-process tokenization, but they suffer from inconsistent results due to boundary artifacts that occur after merging. To address this, we propose LoPT, a novel Lossless Parallel Tokenization framework that ensures output identical to standard sequential tokenization. Our approach employs character-position-based matching and dynamic chunk length adjustment to align and merge tokenized segments accurately. Extensive experiments across diverse long-text datasets demonstrate that LoPT achieves significant speedup while guaranteeing lossless tokenization. We also provide theoretical proof of consistency and comprehensive analytical studies to validate the robustness of our method. Lingchao Zheng, Peizhen Zheng, Li Jun |
ACL (1) | 5 |
| 2025 | Deterministic Communications in Edge AI: 5G-TSN Integration Solutions and ImplementationabstractThe rapid growth of artificial intelligence (AI) and edge computing is transforming industries by enabling innovative applications and improving operational efficiency. End-to-end deterministic communication is crucial for ensuring reliable and predictable data transmission, particularly in the edge AI applications requiring immediate responses, such as real-time digital twin, robotics and automation systems. This paper addresses the challenges of integrating 5G with Time-Sensitive Networking (TSN) in Edge AI environments by proposing two "lightweight" integration solutions that utilize loose and tight coupling methods. The tight coupling method utilizes the RAN Intelligent Controller (RIC) to achieve tight scheduling between 5G and TSN, significantly reducing latency and improving responsiveness to dynamic wireless environments. We developed a testbed using commercial 5G network equipment to evaluate both integration solutions, focusing on key performance metrics such as transmission latency and jitter. Additionally, based on the proposed 5G and TSN solutions, we constructed a real-time digital twin system to track the movements of a robotic arm, and the results demonstrate a significant improvement in positioning accuracy. Zhuyan Zhao, Shen Jun, Li Jun |
INDIN | 3 |
| 2013 | Design of a monocular simultaneous localisation and mapping system with ORB featureabstractVision-based Simultaneous Localisation and Mapping (Visual SLAM) is a new hot topic in intelligent robotic applications. A new method for the implementation of a visual SLAM system with monocular vision is proposed in this paper. The general framework of our system is first displayed, and then all the main sub-processes are described step by step. In our design we use the ORB feature to represent each natural landmark with an improved map management and modified covariance extended Kalman filter (MVEKF) to estimate the 6D pose of a free-moving camera. In order to validate and demonstrate the performance of the system, some related experiments are carried out. The experimental results show that our method is feasible, robust and efficient. Li Jun, Tien-Szu Pan, Kuo-Kun Tseng, Jeng-Shyang Pan 0001 |
ICME | 1 |
| 2012 | A novel high rate transmission scheme for space time coding with low decoding complexityabstractIn this paper, we propose a new method to transmit one more information bit by using orthogonal Space-Time Block Codes (STBC) for 4 antennas. Using two different STBCs matrices transmit one additional bit to achieve high rate-9/8. To maintain full rank and full diversity for the coding gain matrix, a new STBC code with full rate and full diversity is proposed in this letter. In order to implement a fast Maximum-likelihood (ML) decoding, a property of Frobenius norms of received signals is considered to reduce its complexity by checking the non-null values of the Frobenius at the transmitter side. Simulation results show that this method achieves better bit error rate (BER) performance and throughputs in the high SNR region without losing diversity gain. Yier Yan, Xueqin Jiang 0001, Li Jun, Duan Wei, TaeChol Shin, Moon Ho Lee |
ISCAS | 3 |
| 2010 | The Diagnosability of Augmented Cubes under PMC ModelabstractDiagnosability have played an important role in measuring the reliability of multiprocessor systems. As an enhancement on the hypercube, the augmented cube, not only retains all the favorable properties of the hypercube but also possesses some properties which are not shared by the hypercube and its variations. In this paper, under the PMC model, we prove that the n-dimensional augmented cube AQn is (2n-1)-diagnosable under the precise diagnosis strategy and (4n-8)/(4n-8)-diagnosable under the pessimistic strategy (n>;6). Li Jun, Tan Xuegong |
PDCAT | 1 |
| 2008 | Approximate Validity of XML Streaming DataabstractWe present a SAX implementation of the statistical embedding associated with XML data, introduced in [1], [2], which allows to efficiently decide eps-validity to any DTD or Schema, for the Edit Distance with Moves. It associates a generalized k-gram to unranked labelled trees (with k = 1/epsiv) from which any regular property can be approximately decided. We show how to exactly compute the k-gram with a SAX implementation using a memory of size d, the depth of the tree, and an approximate k-gram with queues of size M = 2kand a global memory of size 2kin the worst-case. Experiments on large XML files from the XML benchmark project confirm the error analysis for various values of M. Huang Cheng, Li Jun, Michel de Rougemont |
WAIM | 2 |