EDBT 2026 Demo / reviewers in the wild / expert
Jinchao Zhang 0002
dblp:127/3143-2
· DBLP profile ↗
15ranked-venue papers
2as first author
13since 2021 · last 2026
0009-0005-1835-1700ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Security and privacy · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection
Sa Zhu, Wanqian Zhang, Lin Wang 0108, Jinchao Zhang 0002, Bo Li 0063 |
SIGIR | 4 |
| 2026 | Reducing Write Amplification in LSM-Trees Through Orchestrated Fine-Grained Compaction and Persistent MemoryabstractThis paper investigates how to leverage emerging non-volatile memory (NVM) to enhance the performance of Log-Structure Merge (LSM) tree based key-value (KV) stores. We propose KVFG-DB, which efficiently integrates fine-block granularity and single-NVM-level compaction, to deliver high write and read performance with minimal write amplification and reduced write stalls. KVFG-DB leverages a streamlined SSTable layout called CSSTable to manage data blocks (each containing ordered KV pairs) and their indexes (i.e., the minimal and maximal keys of every block), optimizing both storage costs and compaction performance. It organizes the LSM-tree data as CSSTables into a single level on NVM, with the left area serving as a buffer to receive flushed data with substantially greater capacity, and the right area storing compacted data in a global order among CSSTables. A fine-grained compaction is then performed under various conditions to select the most relevant data blocks with intersecting keys from CSSTables in both areas, facilitating byte-addressable, fast parallel execution across multiple threads. As a result, KVFG-DB adaptively compacts data and quickly moves them from the left area to the right area to enhance performance efficiency, significantly reducing write amplification and further mitigating write stalls. Our extensive experimental studies demonstrate that KVFG-DB achieves 1.2× and 2× lower write amplification, compared with state-of-the-art KV stores MioDB and SLM-DB. Accordingly, KVFG-DB shows a 1.1× and 7.1× improvement in random write performance compared to them, with tail latency reduced by 1.1× and 7×, respectively. Licheng Shan, Jinchao Zhang 0002, Youyou Lu, Bo Li 0063, Xiaoyan Gu 0001, Weiping Wang 0005 |
IEEE Trans. Computers | 3 |
| 2026 | A Practical Dynamic Searchable Symmetric Encryption Scheme With Trusted Hardware-Assisted
Yang Yang 0192, Tianming Hou, Haihui Fan, Bo Li 0063, Xiaoyan Gu 0001, Jinchao Zhang 0002, Hui Ma 0002, Weiping Wang 0005 |
IEEE Trans. Computers | 6 |
| 2026 | MithrilRB: Resource-Efficient Redactable Blockchain With Single-Use AuthorizationabstractRedactable blockchains preserve the integrity of hash links while enabling authorized redactions to comply with regulatory requirements. However, existing permissioned solutions suffer from three severe issues. First, fine-grained privilege control incurs significant storage overhead, especially in the case of single-use authorization. Second, the reliance on bilinear pairings in chameleon hash leads to significant performance degradation when handling large-scale redaction requests. Finally, multiple incorporated components often unconsciously introduce centralized entities, undermining the decentralized nature. In this paper, we propose MithrilRB, a resource-efficient and decentralized redactable blockchain with single-use authorization. Specifically, we introduce a privilege control mechanism with our proposed multi-authority attribute-based signature (MA-ABS) and the threshold BLS signature, achieving fine-grained single-use authorization and direct user revocation without extra ciphertext storage. We also design a pairing-free non-interactive threshold chameleon hash (PNITCH), which enhances efficiency and is better suited for large-scale redaction requests. Moreover, MithrilRB eliminates centralized trust points that hold secret information, ensuring fully decentralization-compatible functional integration in redactable blockchains. Finally, we implement MithrilRB, and the experimental results demonstrate that MithrilRB significantly outperforms existing solutions in both computational efficiency and storage requirements. Tianming Hou, Hui Ma 0002, Jinchao Zhang 0002, Yang Li 0192, Bo Li 0063, Weiping Wang 0005 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Dangerous Language Habits! Exploiting Code-Mixing for Backdoor Attacks on NLP ModelsabstractBackdoor attacks threaten the reliability of NLP models by embedding hidden behaviors during training, which are activated by specific inputs at inference time. Traditional backdoor triggers often rely on explicit content alterations-such as token insertion or stylistic modification-which may compromise semantic coherence and be easily detected.In this work, we propose a novel backdoor attack strategy that leverages the linguistic properties of code-mixing(a language form that combines elements from two or more languages) as implicit triggers. Drawing inspiration from natural code-mixing communication, we design three types of linguistically grounded triggers: inter-word mixing, intra-sentential mixing, and inter-sentential mixing. These forms reflect realistic language usage patterns in bilingual communities, enhancing the stealthiness of the attack. The experiment results show that existing NLP models perform poorly when faced with backdoor attacks based on code-mixing triggers. We are the first to focus on code-mixing as a trigger for text backdoor attacks. We hope this research raises awareness of the vulnerability of models during training when faced with code-mixing. Haotian Jin, Haihui Fan, Jinchao Zhang 0002, Yang Li 0192, Bo Li 0063, Junhao Zhou |
CIKM | 3 |
| 2025 | Retrieval-Augmented Image Captioning via Synthesized Entity-Aware Knowledge RepresentationsabstractRetrieval-Augmented Image Captioning enhances the model's understanding of real-world images by retrieving external knowledge. Existing methods mainly use original captions or isolated entities related to the query image to help generate captions. However, these methods make the model either imitate the caption style or fail to capture the relationship between entities, resulting in a lack of diversity or inaccuracy in the generated captions. To address these issues, we propose SEAR, a novel framework that utilizes external Synthesized Entity-Aware knowledge Representations to improve captioning performance. Specifically, SEAR clusters images based on scene-level and entity-level features, and synthesizes each clustered images into representative images as retrieval indexes, and simultaneously utilizes a large model to extract and supplement structured knowledge graphs from the corresponding cluster captions. Furthermore, we design a knowledge-graph pruner to prune the knowledge graph by retaining the most relevant subgraphs to the query image. By undertaking these steps in an integrated manner, SEAR enables the model to acquire non-redundant and structured information for generating captions and avoid data-related privacy issues. Extensive experiments on MSCOCO, Flickr30k, and NoCaps demonstrate the effectiveness of our method both in-domain and out-of-domain, outperforming existing lightweight RAIC methods and remaining competitive with heavyweight models. Chenxu Cui, Jinchao Zhang 0002, Haihui Fan, Haotian Jin, Bo Li 0063 |
CIKM | 3 |
| 2025 | An Efficient and Privacy-Preserving Cross-Modal Retrieval Scheme for Encrypted Data in the CloudabstractThe rapid development and widespread adoption of cloud computing have made privacy-preserving data retrieval in the cloud a hot topic in privacy computing research. While numerous schemes have been proposed for privacy-preserving text-to-text and image-to-image retrieval, there is a lack of research on cross-modal retrieval with privacy protection in the cloud, leaving ample room for enhancing the security and efficiency of existing schemes. To address these challenges, this paper proposes a scheme for privacy-preserving cross-modal retrieval. The proposed scheme leverages a cross-modal pretrained model to extract features from both text and images, which are then mapped into hash buckets using locality-sensitive hashing (LSH), thereby enhancing retrieval efficiency. Additionally, a data structure is employed to generate secure indexes for these hash buckets via symmetric searchable encryption (SSE). To further improve retrieval accuracy, a secure k-nearest neighbor (kNN) algorithm is applied for second-round ranking of the retrieval results. Analysis and extensive experiments on widely-used real-world datasets have demonstrated the scheme's effectiveness in preserving privacy while maintaining efficiency and feasibility. Yang Li 0192, Bo Li 0063, Jinchao Zhang 0002, Chuanrong Li |
CSCWD | 5 |
| 2025 | Uneven Event Modeling for Partially Relevant Video RetrievalabstractGiven a text query, partially relevant video retrieval (PRVR) aims to retrieve untrimmed videos containing relevant moments, wherein event modeling is crucial for partitioning the video into smaller temporal events that partially correspond to the text. Previous methods typically segment videos into a fixed number of equal-length clips, resulting in ambiguous event boundaries. Additionally, they rely on mean pooling to compute event representations, inevitably introducing undesired misalignment. To address these, we propose an Uneven Event Modeling (UEM) framework for PRVR. We first introduce the Progressive-Grouped Video Segmentation (PGVS) module, to iteratively formulate events in light of both temporal dependencies and semantic similarity between consecutive frames, enabling clear event boundaries. Furthermore, we also propose the Context-Aware Event Refinement (CAER) module to refine the event representation conditioned the text’s cross-attention. This enables event representations to focus on the most relevant frames for a given text, facilitating more precise text-video alignment. Extensive experiments demonstrate that our method achieves state-of-the-art performance on two PRVR benchmarks. Code is available at https://github.com/Sasa77777779/UEM.git. Sa Zhu, Huashan Chen, Wanqian Zhang, Jinchao Zhang 0002, Zexian Yang, Xiaoshuai Hao, Bo Li 0063 |
ICME | 4 |
| 2025 | ShieldIR: Privacy-Preserving Unsupervised Cross-Domain Image Retrieval via Dual Protection TransformationabstractUnsupervised cross-domain image retrieval (UCDIR) aims to retrieve images across different domains without the guidance of labels. However, existing UCDIR methods assume that data can be shared across domains in plaintext, which is often impractical due to strict data privacy protection policies. In this paper, we propose ShieldIR, a novel privacy-preserving unsupervised cross-domain image retrieval framework that enhances the retrieval performance across two domains while safeguarding data privacy. ShieldIR unifies intra-domain and cross-domain representation learning through a Dual Protection Transformation (DPT) module, which introduces a structured feature space via orthogonal projection and ensures data privacy by adding calibrated differential privacy noise. This transformation allows ShieldIR to preserve semantic structure while formally protecting private data. For intra-domain representation learning, ShieldIR enhances discriminability by using DPT to map prototypes into an independent feature space and subsequently aligning the resulting dual-protected prototypes with instance-level features. For cross-domain alignment, ShieldIR maps intra-domain features and cross-domain prototypes into a shared structured space using DPT, achieving semantic alignment under privacy constraints. Extensive experiments on real-world datasets demonstrate that our ShieldIR outperforms state-of-the-art methods while effectively protecting data privacy. Haihui Fan, Jinchao Zhang 0002, Hui Ma 0002, Xiaoyan Gu 0001, Bo Li 0063, Weiping Wang 0005 |
ACM Multimedia | 3 |
| 2025 | CIRAG: Retrieval-Augmented Language Model with Collective IntelligenceabstractRetrieval-augmented generation (RAG) paradigms can integrate external knowledge to enhance and validate the output of Large Language Models (LLMs) thereby mitigating generative hallucinations and broadening the model's knowledge scope. Despite advancements, existing RAG methods still suffer from uncertainty of prediction during the multi-round retrieval-generation process, and a lack of the ability to balance the adequacy and redundancy of retrieved information. To address these challenges, we propose CIRAG, an approach that combines the RAG process with collective intelligence. Inspired by the crowd of wisdom, CIRAG simulates individual independent decision-making and information aggregation within a crowd. Specifically, CIRAG first enhances retrieval diversity by expanding queries based on extracted entities, then combines frequency-based and semantic-based reranking to form a multi granularity fusion reranking thereby assessing better relevance, and integrate multiple information sources for accurate content generation. By undertaking these steps in an integrated manner, CIRAG enables the model to acquire comprehensive and non-redundant information for generating responses. We conduct extensive experiments with HotPotQA and 2WikiMultihopQA datasets, popular benchmark for retrieval-based, multi-step question-answering. Experimental results show that our approach surpasses existing advanced RAG framework while providing high portability in query expansion as well as strong comprehensiveness exhibited in the collective intelligence. Chenxu Cui, Haihui Fan, Jinchao Zhang 0002, Bo Li 0063, Weiping Wang 0005 |
SIGIR | 3 |
| 2025 | SMSSE: Size-Pattern Mitigation Searchable Symmetric EncryptionabstractSearchable Symmetric Encryption (SSE) enables clients to make confidential queries over encrypted data while revealing some formally-defined leakage profiles. Despite the promising performance and application prospects of SSE, the recent leakage-abuse attacks show that a passive adversary can recover queries by exploiting patterns about data disclosed from leakage profiles. Among those attacks, the size pattern is a frequently exploited leakage. Although several countermeasures have been proposed, they can provide neither sufficient protection to mitigate size pattern leakage, nor sufficient scalability for large-scale databases. To address those challenges, we present an SGX-based size-pattern mitigation SSE schemeSMSSEwith two tailored response padding approaches and an I/O efficient disk-based index construction. In addition, we evaluate the size pattern leakage after padding through conditional entropy and differential privacy. Furthermore, we demonstrate the scalability robustness ofSMSSEon different databases by theoretically deducing the approximate boundary of index reading efficiency under a reasonable query distribution. Experiment results on representative real-world datasets show thatSMSSEcan provide high utility and strong protection against newly size pattern-based leakage-abuse attacks. Yang Yang 0192, Haihui Fan, Jinchao Zhang 0002, Bo Li 0063, Hui Ma 0002, Xiaoyan Gu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | T-ABE: A practical ABE scheme to provide trustworthy key hosting on untrustworthy cloudabstractAttribute-based encryption (ABE) is a highly promising mechanism for secure fine-grained access control over encrypted data, particularly in cloud storage applications. However, the integration of built-in key escrow and susceptibility to key exfiltration poses a prevalent challenge during the implementation of ABE. In this paper, we present a novel ABE scheme called T-ABE to address the above challenges. Specifically, T-ABE utilizes a Trusted Execution Environment (i.e. Intel SGX) to provide a precise security guarantee for the entire process, from establishing trust between users and ABE services to key generation and management, ensuring that the entire scheme can provide users with secure and reliable ABE key hosting services even when deployed in untrustworthy cloud service providers. Through comprehensive security analysis and experimental demonstration, we have demonstrated the advantages of T-ABE in practical applications. We are confident that this new scheme has the potential to effectively tackle the outsourced key escrow challenges encountered by ABE. Shuaishuai Chang, Yuzhe Li 0001, Jinchao Zhang 0002, Bo Li 0063 |
TrustCom | 3 |
| 2024 | Toward Automated Field Semantics Inference for Binary Protocol Reverse EngineeringabstractNetwork protocol reverse engineering is the basis for many security applications. A common class of protocol reverse engineering methods is based on the analysis of network message traces. After performing message field identification by segmenting messages into multiple fields, a key task is to infer the semantics of the fields. One of the limitations of existing field semantics inference methods is that they usually infer semantics for only a few fields and often require a lot of manual effort. In this paper, we propose an automated field semantics inference method for binary protocol reverse engineering (FSIBP). FSIBP aims to automatically learn semantics inference knowledge from known protocols and use it to infer the semantics of any field of an unknown protocol. To achieve this goal, we design a feature extraction method that can extract features of the field itself and of the field context. We also propose a semantic category aggregation method that abstracts the fine-grained semantics of all fields of known protocols into aggregated semantic categories. Moreover, we make FSIBP infer semantics based on the similarity of fields to semantic categories. The above design enables FSIBP to utilize the semantic knowledge of all fields of known protocols and infer the semantics of any fields of unknown protocols. The whole process of FSIBP does not require any expert knowledge or manual parameter setting. We conduct extensive experiments to demonstrate the effectiveness of FSIBP. Moreover, we find a utility for FSIBP besides field semantics inference, its output can help to detect the mis-segmented fields generated during the message field identification. Mengqi Zhan, Yang Li 0192, Bo Li 0063, Jinchao Zhang 0002, Chuanrong Li, Weiping Wang 0005 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | Efficient Algorithms of Parallel Skyline Join over Data Streams
Jinchao Zhang 0002, Jingzi Gu, Shuai Cheng 0002, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002 |
ICA3PP (1) | 1 |
| 2015 | Skyline Query on Anti-correlated Distributions: From the Perspective of Spatial Index
Jinchao Zhang 0002, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002 |
ICA3PP (4) | 1 |