EDBT 2026 Demo / reviewers in the wild / expert
Tianshi Xu
dblp:163/8470
· DBLP profile ↗
19ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Cross-Modality Interaction for Unified Video-Text Retrieval Modeling
Tianshi Xu, Zhengzheng Sun, Yizheng Hu, Junyuan Shang, Si Wu 0002 |
MMM (2) | 1 |
| 2025 | Octopus: Fast Homomorphic Convolution for Secure Neural Network InferenceabstractSecure two-party neural network (2PC-NN) inference is a privacy-preserving inference method that protects the client's input and the server's model parameters. While addressing privacy concerns, it also incurs considerable over-heads. In this work, we propose Octopus, a faster and more communication-efficient 2PC-NN system than prior works. Octopus designs an optimized encoding method for fast homomorphic convolution, and further constructs homomorphic encryption-based convolutional computation protocol. Compared with the original coefficient encoding proposed by Cheetah, our method significantly reduces the resulting ciphertexts through packing output channels, thereby saving the communication cost and end-to-end execution time. Moreover, Octopus proposes an encoding-motivated fine tuning technique for convolutional neural networks, which fully utilizes the feature of coefficient encoding to adaptively adjust the neural network structure to maximize performance with negligible accuracy loss. We apply Octopus to the widely used model ResNet on CIFAR-10 and ImageNet dataset. Experiments illustrate that Octopus has obvious improvement compared with the state-of-the-art approaches, achieving a speedup of up to 2.75×, and reduces communication overhead by up to 7.19× for convolutions. As for secure inference, compared with Cheetah (resp., CrypTFlow2), Octopus demonstrates 1.41× (resp., 13.20×) lower communication cost and 1.25× (resp., 7.03×) faster execution time under a WAN setting. Yu Fu 0007, Tianshi Xu, Cheng Hong 0001, Meng Li 0004, Wei Wang 0314, Dengguo Feng, Jingqiang Lin 0001 |
ACSAC | 3 |
| 2025 | IP-KGQA: Intent-Aware Prompt Learning for Knowledge Graph Question AnsweringabstractKnowledge Graph Question Answering (KGQA) addresses natural language questions by leveraging structured information stored in knowledge graphs. However, existing KGQA methods are overly concerned with improving the quality of responses by retrieving information, neglecting to identify which type of knowledge is truly useful to optimize the performance of the KGQA system, resulting in redundant retrieval. At the same time, these methods have limitations in aligning user intent and insufficient semantic richness in responses. In this work, we propose IP-KGQA, which introduce a intent-aware prompt learning scheme for KGQA framework. A Selection-Driven Efficient Retrieval (SER) module is incorporated in the framework, which classifies user questions to ensure that only long-tail questions are directed to the knowledge graph retrieval to enhance system efficiency. To filter and select the most relevant triplets, aligning retrieved information more closely with user intent, we introduce the User Intent-aware Filtering (UIF) module, where Monte Carlo sampling is applied to obtain the optimal triplets. The Domain-specific Context Prompt Extension (DCPE) module is utilized in collaboration with a fine-tuned large language model (LLM) to integrate domain-specific knowledge into the responses, ensuring that the answers are enriched in terms of semantic quality. Extensive experiments have been conducted on the CommonSenseQA and TriviaQA datasets, which demonstrate that IP-KGQA outperforms the existing methods in terms of retrieval efficiency, answer accuracy and user intent alignment. Zheng Dai, Chun Ding, Si Wu 0002, Yong Xu 0007, Runzhe Liang, Tianshi Xu, Yedong Li, Dapeng Oliver Wu |
ICME | 7 |
| 2025 | Ironman: Accelerating Oblivious Transfer Extension for Privacy-Preserving AI with Near-Memory Processing
Chenqi Lin, Kang Yang 0002, Tianshi Xu, Ling Liang 0003, Runsheng Wang, Mingyu Gao 0001, Meng Li 0004 |
MICRO | 3 |
| 2025 | EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and RetrievalabstractObject-goal navigation (ObjNav) tasks an agent with navigating to the location of a specific object in an unseen environment.
Embodied agents equipped with large language models (LLMs) and online constructed navigation maps can perform ObjNav in a zero-shot manner. However, existing agents heavily rely on giant LLMs on the cloud, e.g., GPT-4, while directly switching to small LLMs, e.g., LLaMA3.2-11b, suffer from significant success rate drops due to limited model capacity for understanding complex navigation maps, which prevents deploying ObjNav on local devices.
At the same time, the long prompt introduced by the navigation map description will cause high planning latency on local devices.
In this paper, we propose EfficientNav to enable on-device efficient LLM-based zero-shot ObjNav. To help the smaller LLMs better understand the environment, we propose semantics-aware memory retrieval to prune redundant information in navigation maps.
To reduce planning latency, we propose discrete memory caching and attention-based memory clustering to efficiently save and re-use the KV cache.
Extensive experimental results demonstrate that EfficientNav
achieves 11.1\% improvement in success rate on HM3D benchmark over GPT-4-based baselines,
and demonstrates 6.7$\times$ real-time latency reduction and 4.7$\times$ end-to-end latency reduction over GPT-4 planner. Our code is available on https://github.com/PKU-SEC-Lab/EfficientNav. Sunjian Zheng, Tong Xie, Tianshi Xu, Bo Yu 0014, Fan Wang 0021, Jie Tang 0003, Shaoshan Liu |
NeurIPS | 4 |
| 2025 | CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert RoutingabstractPrivate large language model (LLM) inference based on cryptographic primitives offers a promising path towards privacy-preserving deep learning. However, existing frameworks only support dense LLMs like LLaMA-1 and struggle to scale to mixture-of-experts (MoE) architectures. The key challenge comes from securely evaluating the dynamic routing mechanism in MoE layers, which may reveal sensitive input information if not fully protected. In this paper, we propose CryptoMoE, the first framework that enables private, efficient, and accurate inference for MoE-based models. CryptoMoE balances expert loads to protect expert routing information and proposes novel protocols for secure expert dispatch and combine. CryptoMoE also develops a confidence-aware token selection strategy and a batch matrix multiplication protocol to improve accuracy and efficiency further. Extensive experiments on DeepSeekMoE-16.4B, OLMoE-6.9B, and QWenMoE-14.3B show that CryptoMoE achieves $2.8\sim3.5\times$ end-to-end latency reduction and $3\sim6\times$ communication reduction over a dense baseline with minimum accuracy loss. We also adapt CipherPrune (ICLR'25) for MoE inference and demonstrate CryptoMoE can reduce the communication by up to $4.3 \times$. Tianshi Xu, Jue Hong |
NeurIPS | 2 |
| 2025 | Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC
Tianshi Xu, Jiangrui Yu, Chenqi Lin, Runsheng Wang, Meng Li 0004 |
USENIX Security Symposium | 1 |
| 2025 | Swift: Fast Secure Neural Network Inference With Fully Homomorphic EncryptionabstractWith the widespread use of machine learning (ML), privacy concerns during neural network inference are attracting growing attention. Secure two-party neural network (2PC-NN) inference is the privacy-preserving inference method, which allows client to obtain the inference result without disclosing client’s input to the server. The server’s model parameters are also confidential to the client. However, current 2PC-NN inference schemes still have large overhead, especially for non-linear functions. In this paper, we present Swift, a fast secure 2PC-NN inference scheme based on fully homomorphic encryption (FHE) and secret sharing (SS). FHE protects the input and model parameters in linear functions, while SS is integrated to protect the non-linear functions. Concretely, Swift integrates FHE and SS to design secure and efficient non-linear protocols used for ReLU and max pooling. To further optimize performance, Swift employs FHE with computation-friendly coefficient encoding for fast execution of linear functions, and SIMD encoding for non-linear functions. Swift constructs efficient encoding conversion protocol between the coefficient-encoded ciphertext and the SIMD-encoded ciphertext. Finally, Swift achieves secure neural network inference framework for MNIST dataset. Compared with Cheetah (USENIX 2022), the execution time of ReLU, max pooling, secure inference under a WAN setting improves$7.4\times $,$13.3\times $,$1.9\times $, respectively. Yu Fu 0007, Yijing Ning, Tianshi Xu, Meng Li 0004, Jingqiang Lin 0001, Dengguo Feng |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | FastQuery: Communication-efficient Embedding Table Query for Private LLMs inferenceabstractWith the fast evolution of large language models (LLMs), privacy concerns with user queries arise as they may contain sensitive information. Private inference based on homomorphic encryption (HE) has been proposed to protect user query privacy. However, private embedding table query has to be formulated as a HE-based matrix-vector multiplication problem and suffers from enormous computation and communication overhead. We observe the overhead mainly comes from the neglect of 1) the one-hot nature of user queries and 2) the robustness of the embedding table to low bit-width quantization noise. Hence, in this paper, we propose a private embedding table query optimization framework, dubbed FastQuery. FastQuery features a communication-aware embedding table quantization algorithm and a one-hot-aware dense packing algorithm to simultaneously reduce both the computation and communication costs. Compared to prior-art HE-based frameworks, e.g., Cheetah, Iron, and Bumblebee, FastQuery achieves more than 4.3×, 2.7×, 1.3× latency reduction, respectively and more than 75.7×, 60.2×, 20.2× communication reduction, respectively, on both LLAMA-7B and LLAMA-30B. Chenqi Lin, Tianshi Xu, Runsheng Wang, Ru Huang 0001, Meng Li 0004 |
DAC | 2 |
| 2024 | PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-OptimizationabstractPrivate deep neural network (DNN) inference based on secure two-party computation (2PC) enables secure privacy protection for both the server and the client. However, existing secure 2PC frameworks suffer from a high inference latency due to enormous communication. As the communication of both linear and non-linear DNN layers reduces with the bit widths of weight and activation, in this paper, we propose PrivQuant, a framework that jointly optimizes the 2PC-based quantized inference protocols and the network quantization algorithm, enabling communication-efficient private inference. PrivQuant proposes DNN architecture-aware optimizations for the 2PC protocols for communication-intensive quantized operators and conducts graph-level operator fusion for communication reduction. Moreover, PrivQuant also develops a communication-aware mixed precision quantization algorithm to improve the inference efficiency while maintaining high accuracy. The network/protocol co-optimization enables PrivQuant to outperform prior-art 2PC frameworks. With extensive experiments, we demonstrate PrivQuant reduces communication by 11×, 2.5 × and 2.8×, which results in 8.7×, 1.8 × and 2.4× latency reduction compared with SiRNN, COINN, and CoPriv, respectively. Tianshi Xu, Shuzhang Zhong, Wenxuan Zeng, Runsheng Wang, Meng Li 0004 |
ICCAD | 1 |
| 2024 | FlexHE: A flexible Kernel Generation Framework for Homomorphic Encryption-Based Private InferenceabstractSecure two-party computation (2PC) based on homomorphic encryption (HE) achieves formal data privacy protection and gets increasing adoption for private deep neural network (DNN) inference. As modern HE schemes usually operate on polynomials, existing works rely on manually-designed HE kernels for representative DNN operations. However, this is not only unscalable considering the diverse operator types, shapes, polynomial orders, etc, but also misses important optimization opportunities. In this paper, we introduce FlexHE, a flexible kernel generation framework to enable automatic generation and optimization of HE kernels for 2PC-based private inference. Given a high-level description of DNN operations, FlexHE can systematically define the HE kernel design space considering various optimization dimensions, including loop tiling, reordering, etc. We also analyze the communication and computation impact of different optimization dimensions for design space reduction. To search for the best kernel design, a two-level optimization problem is formulated and iteratively solved with an integer linear programming (ILP) formulation. With extensive experimental results, we not only demonstrate a better coverage of DNN operations including depth-wise Conv3D and dilated Conv3D, but also achieve more than 100×, 7.9×, and 4.2× latency reduction compared to prior-art HElayers, Cheetah, and Falcon, respectively. Jiangrui Yu, Wenxuan Zeng, Tianshi Xu, Renze Chen, Yun Liang 0001, Runsheng Wang, Ru Huang 0001, Meng Li 0004 |
ICCAD | 3 |
| 2024 | Balanced Active Sampling for Person Re-identificationabstractActive learning is attracting more and more attention in person re-identification (Re-ID), as it is promising in the scalability of Re-ID models to satisfy performance with reduced labeling cost. Active sampling of pair-wise images in Re-ID is a highly imbalanced problem, where negative pairs are the vast majority. To avoid sampled pairs being dominated by the negative relationship, previous works tend to sample pairs with confident positive relationships in various ways. However, it is a waste of the labeling budget as most sampled pairs will be positive and already have a very close distance. Thus, there is no significant improvement in the model performance. In this paper, we first argue that balanced sampling is the key to active learning for Re-ID. Along this line, we propose a naïve balanced sampling method based on the global estimation of the most confusing distance. It is further improved by the label-wise estimation and diversity measurement. We also formulate the training of Re-ID models as a constrained clustering problem, where labeled positive and negative pairs are as must-link and cannot-link. Then the model training is based on the pseudo labels. Extensive experiments on benchmarks evaluate the effectiveness and superiority of the proposed methods. Specifically, it achieves comparable performance with supervised counterparts with less than 0.1% pair-wise annotation, which significantly surpasses the state-of-the-art. Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005 |
ICME | 5 |
| 2024 | Camera Bias Regularization for Person Re-identificationabstractPerson re-identification (Re-ID) is to match persons captured by non-overlapping cameras. Due to the discrepancies between cameras caused by illumination, background, or viewpoint, the underlying difficulty for Re-ID is the camera bias problem, which leads to the large gap of within-identity features from different cameras. With limited cross-camera annotation, Re-ID models tend to learn camera-related features, instead of identity-related features. Consequently, Re-ID models suffer from poor transfer ability from seen to unseen domains. In this paper, we investigate the camera bias problem in both supervised and unsupervised learning. In particular, we propose a novel Camera Bias Regularization (CBR) term to reduce the feature distribution gap between cameras. The CBR works by simultaneously enlarging the distance of intra-camera distributions between positive and negative pairs, and reducing the distance of positive pairs’ distributions between intra-camera and cross-camera. In addition, a Cross-Camera (CC) clustering method is also designed for unsupervised learning, which puts more emphasis on cross-camera pairs than intra-camera ones during the clustering process. Extensive experiments are conducted to validate the effectiveness of the proposed CBR and CC. Specifically, with only a plain ResNet-50, it achieves 56.7% mAP and 40.7% mAP on the challenging MSMT17 dataset in supervised and unsupervised settings respectively, which surpasses most state-of-the-arts. Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005 |
ICME | 5 |
| 2024 | X-ReID: Cross-Instance Transformer for Identity-Level Person Re-IdentificationabstractCurrently, most existing person re-identification methods use instance-level features, which are extracted only from a single image. However, these instance-level features can easily ignore the discriminative information because the appearance of each identity varies greatly in different images. Thus, it is necessary to exploit identity-level features, which can be shared across different images of each identity. In this paper, we propose a novel training framework, named X-ReID, to promote instance-level features to identity-level features by employing cross-attention to incorporate information from one image to another of the same identity, thus more unified and discriminative pedestrian information can be obtained. Extensive experiments on benchmark datasets show the superiority of our method over existing works. Particularly, on the challenging MSMT17, our proposed method gains 1.1% mAP improvements when compared to the second place. Leqi Shen, Sicheng Zhao, Zhelun Shen, Tianshi Xu, Guiguang Ding |
ICME | 6 |
| 2024 | PrivCirNet: Efficient Private Inference via Block Circulant TransformationabstractHomomorphic encryption (HE)-based deep neural network (DNN) inference protects data and model privacy but suffers from significant computation overhead. We observe transforming the DNN weights into circulant matrices converts general matrix-vector multiplications into HE-friendly 1-dimensional convolutions, drastically reducing the HE computation cost. Hence, in this paper, we propose PrivCirNet, a protocol/network co-optimization framework based on block circulant transformation. At the protocol level, PrivCirNet customizes the HE encoding algorithm that is fully compatible with the block circulant transformation and reduces the computation latency in proportion to the block size. At the network level, we propose a latency-aware formulation to search for the layer-wise block size assignment based on second-order information. PrivCirNet also leverages layer fusion to further reduce the inference cost. We compare PrivCirNet with the state-of-the-art HE-based framework Bolt (IEEE S\&P 2024) and HE-friendly pruning method SpENCNN (ICML 2023). For ResNet-18 and Vision Transformer (ViT) on Tiny ImageNet, PrivCirNet reduces latency by $5.0\times$ and $1.3\times$ with iso-accuracy over Bolt, respectively, and improves accuracy by $4.1$\% and $12$\% over SpENCNN, respectively. For MobileNetV2 on ImageNet, PrivCirNet achieves $1.7\times$ lower latency and $4.2$\% better accuracy over Bolt and SpENCNN, respectively. Our code and checkpoints are available on Git Hub. Tianshi Xu, Lemeng Wu, Runsheng Wang, Meng Li 0004 |
NeurIPS | 1 |
| 2023 | Falcon: Accelerating Homomorphically Encrypted Convolutions for Efficient Private Mobile Network InferenceabstractEfficient networks, e.g., MobileNetV2, EfficientNet, etc, achieves state-of-the-art (SOTA) accuracy with lightweight computation. However, existing homomorphic encryption (HE)-based two-party computation (2PC) frameworks are not op-timized for these networks and suffer from a high inference overhead. We observe the inefficiency mainly comes from the packing algorithm, which ignores the computation character-istics and the communication bottleneck of homomorphically encrypted depthwise convolutions. Therefore, in this paper, we propose Falcon, an effective dense packing algorithm for HE-based 2PC frameworks. Falcon features a zero-aware greedy packing algorithm and a communication-aware operator tiling strategy to improve the packing density for depth wise convo-lutions. Compared to SOTA HE-based 2PC frameworks, e.g., CrypTFlow2, Iron and Cheetah, Falcon achieves more than 15.6 x, 5.1 x and 1.8 x latency reduction, respectively, at operator level. Meanwhile, at network level, Falcon allows for 1.4 % and 4.2% accuracy improvement over Cheetah on CIFAR-100 and Tiny Imagenet datasets with iso-communication, respecitvely. Tianshi Xu, Meng Li 0004, Runsheng Wang, Ru Huang 0001 |
ICCAD | 1 |
| 2022 | parGeMSLR: A parallel multilevel Schur complement low-rank preconditioning and solution package for general sparse matrices
Tianshi Xu, Vassilis Kalantzis, Yuanzhe Xi, Geoffrey Dillon, Yousef Saad |
Parallel Comput. | 1 |
| 2021 | Adaptively Fusing Complete Multi-resolution Features for Human Pose Estimation
Yuezhen Huang, Xiaofeng Jin, Yuheng Huang 0005, Tianshi Xu |
ICIG (2) | 6 |
| 2021 | Dual Gated Learning for Visible-Infrared Person Re-identification
Yuheng Huang 0005, Jincai Xian, Xiaofeng Jin, Tianshi Xu |
ICIG (2) | 5 |