EDBT 2026 Demo / reviewers in the wild / expert
Yuntao Wei
dblp:227/5257
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting MLLM Token Technology through the Lens of Classical Visual CodingabstractClassical visual coding and Multimodal Large Language Model (MLLM) token technology share the core objective - maximizing information fidelity while minimizing computational cost. Therefore, this paper reexamines MLLM token technology, including tokenization, token compression, and token reasoning, through the established principles of long-developed visual coding area. From this perspective, we (1) establish a unified formulation bridging token technology and visual coding, enabling a systematic, module-by-module comparative analysis; (2) synthesize bidirectional insights, exploring how visual coding principles can enhance MLLM token techniques' efficiency and robustness, and conversely, how token technology paradigms can inform the design of next-generation semantic visual codecs; (3) prospect for promising future research directions and critical unsolved challenges. In summary, this study presents the first comprehensive and structured technology comparison of MLLM token and visual coding, paving the way for more efficient multimodal models and more powerful visual codecs simultaneously. Jinming Liu 0001, Junyan Lin, Yuntao Wei, Kele Shao, Keda Tao, Jianguo Huang, Zhibo Chen 0001, Huan Wang 0014, Xin Jin 0014 |
ISCAS | 3 |
| 2025 | Quadtree Partitioning-based Visual Token Pruning for MLLMs Considering Information DensityabstractMultimodal Large Language Models (MLLMs) excel at comprehensive understanding by integrating visual and textual information. However, their inference speed is often bottlenecked by redundant visual token inputs. Existing methods tend to alleviate this issue with a heuristic pruning strategy based on token importance, tailored to certain commonly adopted vision encoders like CLIP. In this paper, we propose a novel training-free token pruning method based on a well-designed metric of information density, where we decide which tokens are retained according to their entropy, following the classic information theory. Based on that, we further propose a quadtree partitioning strategy, in which we retain these tokens with higher entropy so as to preserve the visual spatial structure while allocating more tokens to more informative regions. Experiments on LLaVA-v1.5-7B and 13B across six benchmarks show our method achieves state-of-the-art performance—retaining over 90% of full-token accuracy even at a 6.25% token budget—while cutting TFLOPs by up to 20% compared to FastV and by 81% compared to the original LLaVA-v1.5. Yuntao Wei, Jinming Liu 0001, Shengyang Zhao, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014 |
VCIP | 1 |
| 2024 | PPGNN: Fast and Accurate Privacy-Preserving Graph Neural Network Inference via Parallel and Pipelined Arithmetic-and-Logic FHE AcceleratorabstractGraph Neural Networks (GNNs) are increasingly used in fields like social media and bioinformatics, promoting the prosperity of cloud-based GNN inference services. Nevertheless, data privacy becomes a critical issue when handling sensitive information. Fully Homomorphic Encryption (FHE) enables computations on encrypted data, while privacy-preserving GNN inference generally necessitates ensuring graph structure data confidentiality and maintaining computation precision, both of which are computationally expensive in FHE. Existing schemes of GNNs inference with FHE are deterred by either computational overhead, accuracy degradation, or incomplete data protection. This paper presents PPGNN to address these challenges all at once. We first propose a novel privacy-preserving GNN inference algorithm utilizing a high-accuracy arithmetic-and-logic FHE approach, meanwhile only need much smaller parameters, substantially reducing computational complexity and facilitating parallel processing. Correspondingly, a dedicated hardware architecture has been designed to implement these innovations, with featured specialized units for arithmetic and logic FHE operations in a pipelined manner. Collectively, PPGNN achieves 2.7× and 1.5× speedup over state-of-the-art Arithmetic FHE and Logic FHE accelerators while ensuring high accuracy, simultaneously with about 18× energy reduction on average. Yuntao Wei, Song Bian 0001, Weisheng Zhao 0001, Yier Jin |
DAC | 1 |
| 2024 | CLCP: Realtime Text-Image Retrieval for Retailing via Pre-trained Clustering and Priority QueueabstractReal-time matching between customer demands and product information via text-image retrieval remains a fundamental problem in intelligent retailing. However, this process involves challenges covering data quality, multi-modal retrieval strategies and performing efficiency. To alleviate the case, we propose a cross-modality retrieval pipeline leveraging contrastive loss and a novel sampling strategy. We also address text-image retrieval as a two-stage process, involving unsupervised clustering and contrastive feature representation. Additionally, we create an image-caption matching dataset by expanding the Grocery Store Dataset using a fundamental visual-language model. Our experiments demonstrate the effectiveness of our method on both an expanded new dataset and the well-known cross-modality retrieval benchmark, Flicker30k. Liangwu Wei, Yuntao Wei, Yanzhi Song |
ICMR | 4 |
| 2024 | Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMsabstractWe present a new image compression paradigm to achieve "intelligently coding for machine" by cleverly leveraging the common sense of Large Multimodal Models (LMMs). We are motivated by the evidence that large language/multimodal models are powerful general-purpose semantics predictors for understanding the real world. Different from traditional image compression typically optimized for human eyes, the image coding for machines (ICM) framework we focus on requires the compressed bitstream to more comply with different downstream intelligent analysis tasks. To this end, we employ LMM to${\text{tell codec what to compress}}$: 1) first utilize the powerful semantic understanding capability of LMMs w.r.t object grounding, identification, and importance ranking via prompts, to disentangle image content before compression, 2) and then based on these semantic priors we accordingly encode and transmit objects of the image in order with a structured bitstream. In this way, diverse vision benchmarks including image classification, object detection, instance segmentation, etc., can be well supported with such a semantically structured bitstream. We dub our method "SDComp" for "Semantically Disentangled Compression", and compare it with state-of-the-art codecs on a wide variety of different vision tasks. SDComp codec leads to more flexible reconstruction results, promised decoded visual quality, and a more generic/satisfactory intelligent task-supporting ability. Jinming Liu 0001, Yuntao Wei, Junyan Lin, Shengyang Zhao, Heming Sun, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014 |
VCIP | 2 |
| 2023 | A Partitioned Detection Architecture for Oriented Objects
Yuntao Wei |
ICANN (3) | 2 |
| 2023 | THE-V: Verifiable Privacy-Preserving Neural Network via Trusted Homomorphic ExecutionabstractPrivacy-preserving machine learning (PPML) schemes aim at protecting client-side data privacy in two-party secure computing tasks such as private deep neural network (DNN) inference. While fully homomorphic encryption (FHE) can provide provable security for client data privacy, efficiently verifying that such homomorphic DNN inference protocol is honestly executed on the server presents to be challenging. In this work, we propose THE-V, a novel DNN inference framework that combines FHE and Trusted Execution Environment (TEE) to achieve data privacy, verifiable execution and efficient computation all at once. We first point out that, while the trivial solution of executing FHE entirely within TEE can ensure both private and verifiable computing, the limited resource within TEE becomes a severe computational bottleneck. To solve such dilemma, we devise a new strategy of securely outsourcing computation-heavy tasks in TEE to untrusted environments. By rigorous experiments, we show that we can achieve verifiable and private DNN inference with up to$15\times$speedup compared with the state-of-the-art solution. Yuntao Wei, Song Bian 0001, Weisheng Zhao 0001, Yier Jin |
ICCAD | 1 |
| 2023 | CANAMRF: An Attention-Based Model for Multimodal Depression Detection
Yuntao Wei, Yuzhe Zhang 0002, Hone Zhang |
PRICAI (2) | 1 |
| 2023 | IMGA: Efficient In-Memory Graph Convolution Network Aggregation With Data Flow OptimizationsabstractAggregating features from neighbor vertices is a fundamental operation in graph convolution network (GCN). However, the sparsity in graph data creates poor spatial and temporal locality, causing dynamic and irregular memory access patterns and limiting the performance of aggregation on the Von Neumann architecture. The emerging processing-in-memory (PIM) architecture is based on emerging nonvolatile memory (NVM), like spin-orbit torque magnetic RAM (SOT-MRAM), and demonstrates promising prospects in alleviating the Von Neumann bottleneck. However, the limited memory capacity of PIM medium still incurs non-negligible data movements between PIM architecture and external memory. To solve this challenge, we propose an SOT-MRAM-based in-memory computing architecture, called IMGA, for efficient in-situ graph aggregation. Specifically, we design adaptive data flow management strategies that reuse vertex data in MRAM when processing graphs of different scales and adopt edge data as the control signal source to utilize the graph’s structural information. A reordering optimization strategy leveraging hardware–software co-design principle is proposed to further reduce the costly data movement. Experimental results demonstrate that IMGA achieves an average$2523\times $and$21\times $speedup, and 1.03E+6 and 1.04E+3 energy efficiency compared with CPU and GPU, respectively. Yuntao Wei, Shangtong Zhang, Jianlei Yang 0001, Xiaotao Jia, Zhaohao Wang, Gang Qu 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Finite-Time Synchronization Under Aperiodically Intermittent Control and its Application on Spatially Coupled Reaction-Diffusion Neural NetworksabstractWe investigate a general way to achieve finite-time synchronization (FnTSYN) for networks via aperiodically intermittent control (AInC). At first, we propose a general sufficient criteria under AInC for finite-time stability (FnTSta). Then, a general model for complex networks under this sufficient condition is given. Furthermore, we apply the proposed condition to the spatially coupled reaction-diffusion neural networks (RDNNs) to realize its FnTSYN. Finally, simulations are given to verify the results. Xiwei Liu, Yuntao Wei |
IJCNN | 2 |
| 2021 | An Improved Image Segmentation Algorithm CT Superpixel Grid Using Active ContourabstractThe traditional CT image segmentation algorithm is easy to ignore image contour initialization, which leads to the problem of long time consuming and low accuracy. A superpixel mesh CT image improved segmentation algorithm using active contour was proposed. CT image superpixel gridding was carried out first; secondly, on the basis of gridding, the region growth criterion was improved by superpixel processing, the region growth graph was established, the image edge salient graph was calculated based on the growth graph, and the target edge was obtained as the initial contour; finally, the Mumford‐Shah model in the active contour model was improved; the energy functional was constructed based on the improved model and transformed into the symbol distance function. The results show that the proposed algorithm takes less time to mesh superpixels, the accuracy of image edge calculation is high, the correct classification coefficient is as high as 0.9, and the accuracy of CT image segmentation is always higher than 90%, which has superiority. Yuntao Wei |
Wirel. Commun. Mob. Comput. | 1 |