Jiheng Zhang

dblp:13/7602 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
18since 2021 · last 2026
0000-0003-3025-1495ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 OR-R1: Automating Modeling and Solving of Operations Research Optimization Problem via Test-Time Reinforcement Learning
abstract
Optimization modeling and solving are fundamental to the application of Operations Research (OR) in real-world decision making, yet the process of translating natural language problem descriptions into formal models and solver code remains highly expertise intensive. While recent advances in large language models (LLMs) have opened new opportunities for automation, the generalization ability and data efficiency of existing LLM-based methods are still limited, asmost require vast amounts of annotated or synthetic data, resulting in high costs and scalability barriers. In this work, we present OR-R1, a data-efficient training framework for automated optimization modeling and solving. OR-R1 first employs supervised fine-tuning (SFT) to help the model acquire the essential reasoning patterns for problem formulation and code generation from limited labeled data. In addition, it improves the capability and consistency through Test-Time Group Relative Policy Optimization (TGRPO). This two-stage design enables OR-R1 to leverage both scarce labeled and abundant unlabeled data for effective learning. Experiments show that OR-R1 achieves state-of-the-art performance with an average solving accuracy of 67.7%, using only 1/10 the synthetic data required by prior methods such as ORLM, exceeding ORLM’s solving accuracy by up to 4.2%. Remarkably, OR-R1 outperforms ORLM by over 2.4% with just 100 synthetic samples. Furthermore, TGRPO contributes an additional 3.1%–6.4% improvement in accuracy, significantly narrowing the gap between single-attempt (Pass@1) and multi-attempt (Pass@8) performance from 13% to 7%. Extensive evaluations across diverse real-world benchmarks demonstrate that OR-R1 provides a robust, scalable, and cost-effective solution for automated OR optimization problem modeling and solving, lowering the expertise and data barriers for industrial OR applications.
Zezhen Ding, Zhen Tan 0001, Jiheng Zhang, Tianlong Chen 0001
AAAI3
2026 Privacy-preserving and verifiable aggregation for multi-task and multi-dimensional data
abstract
Abstract With the widespread adoption of smart devices, vast amounts of multidimensional data are continuously generated across various sectors. Integrating and evaluating this data can provide strong support for decision-making. However, existing aggregation schemes face several challenges when handling multidimensional data tasks from multiple requesters, including single points of failure, privacy leakage, lack of result validation, and inefficiency. This paper mitigates the challenges by integrating packed secret sharing, Kate, Zaverucha and Goldberg (KZG) commitments and the Chinese Remainder Theorem to design a data aggregation scheme. For the first time, we achieve both multi-task and multi-dimensional data aggregation within a single data transmission, ensuring data integrity and privacy, while keeping the overhead minimal. Users can provide data for only certain dimensions, thereby safeguarding personal data privacy while emphasizing the user-centric approach to data control. By incorporating multiple fog nodes (FN) and blockchain, the data aggregation scheme is implemented in a decentralized manner. Both data users and FNs are accountable, ensuring the data authenticity and the correctness of the aggregation process. Moreover, it is fault-tolerant, making the system robust. Rigorous security analysis and extensive experimental evaluations demonstrate that the proposed approach excels in computational efficiency.
Liang Zhang 0043, Jiheng Zhang
Comput. J.4
2026 Attribute-based publicly verifiable secret sharing
abstract
Abstract Can a dealer share a secret without knowing the shareholders? We provide a positive answer to this question by introducing the concept of an attribute-based secret sharing (AB-SS) scheme.With AB-SS, a dealer can distribute a secret based on attributes rather than specific individuals or shareholders. Only authorized users whose attributes satisfy a given access structure can recover the secret. Furthermore, we introduce the concept of attribute-based publicly verifiable secret sharing (AB-PVSS). An AB-PVSS scheme allows external users to verify the correctness of all broadcast messages from the dealer and shareholders, similar to a traditional PVSS scheme. Additionally, AB-SS (or AB-PVSS) distinguishes itself from traditional SS (or PVSS) by enabling a dealer to generate shares according to an arbitrary monotone access structure.To build an AB-PVSS scheme, we first implement a decentralized ciphertext-policy attribute-based encryption (CP-ABE) scheme, though not a fully-fledged one.We then incorporate non-interactive zero-knowledge (NIZK) proofs to enable public verification of the CP-ABE ciphertext. Based on the CP-ABE and NIZK proofs, we construct an AB-PVSS primitive.Finally, we conduct security analysis and comprehensive experiments on the proposed CP-ABE and AB-PVSS schemes. The results demonstrate that both schemes exhibit plausible performance compared to related works.
Liang Zhang 0043, Qiuling Yue, Haibin Kan, Jiheng Zhang
Cybersecur.5
2026 A Blockchain-Envisioned Mailing System
abstract
Traditional email systems rely on centralized servers for message storage and routing, making them vulnerable to single points of failure and privacy breaches. While blockchain technology has emerged as a decentralized alternative for internet infrastructure, its application to email systems remains underexplored. This paper fills the gaps by proposing a blockchain-based mailing system that eliminates trusted intermediaries while enhancing privacy. Our approach integrates four key components: (1) the ECIES scheme to ensure end-to-end confidentiality of email content, (2) BIP-32-derived addresses to achieve sender anonymity, (3) stealth addresses to provide receiver k-anonymity, and (4) a broadcast encryption scheme to enable efficient broadcast messaging. Moreover, our design allows the sender to prove membership in the broadcast mailing while preserving anonymity. Further, the proposed system introduces an incentive mechanism that allows recipients to charge fees for incoming emails, effectively discouraging spam. We present a proof-of-concept implementation on the Ethereum blockchain and evaluate its on-chain performance in terms of gas consumption. Experimental results demonstrate the practicality and efficiency of the proposed system.
Liang Zhang 0043, Haibin Kan, Jiheng Zhang
IEEE Trans. Cloud Comput.3
2025 Efficient Graph Bandit Learning with Side-Observations and Switching Constraints
abstract
This paper presents a novel framework for multi-armed bandit problems with side-observations and switching constraints, which arises in a range of real-world applications such as robotic. To address the challenges of effectively utilizing graph-structured observations while adhering to graph constraints, we design graph-agnostic and graph-aware algorithms tailored to this new setting. Specifically, our graph-agnostic algorithm selects nodes with the highest upper confidence bound without prior knowledge of feedback probabilities, while minimizing switching costs using offline shortest path planning and the doubling trick. If the graph structure and associated probability matrix are known, our graph-aware algorithm plans the exploration step using a linear programming approach and eliminates suboptimal nodes iteratively. We rigorously analyze the performance of our proposed algorithms, providing near-optimal minimax and instance-dependent regret upper bounds. Our analysis shows that our algorithms outperform generic reinforcement learning methods in terms of both regret and computational efficiency. Extensive numerical experiments on various types of graphs, including two real-world datasets, demonstrate the efficacy of our proposed methods and their advantages over benchmark methods in graph bandit settings.
Xueping Gong, Jiheng Zhang
AAAI2
2025 Annotation-Free Mask-Guided Transformer for Real-Time Multi-Aspect Blastocyst Grading
abstract
Accurate blastocyst grading is essential for maximizing single-embryo transfer success in in vitro fertilization (IVF). We present BlastocystMask-DINO (BM-DINO), a two-stage framework. In Stage I, a deterministic mask generator produces four-class masks from raw day-5 microscopy images at zero labeling cost. In Stage II, these masks drive a parameter-free Region of Interest (ROI)-guided pooling and are also processed by a maskencoder CNN to extract structural priors. Therefore, prompt tokens, ROI-guided features, and mask-derived structural priors are fused additively to integrate global context, local focus, and morphological cues. Particularly, we leverage pretrained DINOv2, an unsupervised vision transformer, as the backbone and fine-tune the last four blocks to adapt the model efficiently to blastocyst grading. When evaluated on the public Human Blastocyst Dataset (HBD), BM-DINO achieves macro-F1 scores of 0.78 for expansion (EXP), 0.80 for inner cell mass (ICM), and 0.73 for trophectoderm (TE). Our model exceeds expert consensus by 6% on ICM and 2% on TE, achieving state-of-the-art performance. Also, it outperforms prior works in delivering the most balanced performance across all grading tasks. The entire approach adds only 0.07 M trainable parameters and runs end-to-end in$\approx 0.24\mathrm{s}$on commodity hardware (233ms CPU +8ms GPU), making the system suitable for real-time deployment in IVF laboratory. The implementation source code and underlying dataset are available at https://github.com/WFJKh/BM-DINO.
Liang Zhang 0043, Honglan Huang, Jiheng Zhang
BIBM4
2025 RESA: RLWE-Based Efficient Secure Aggregation For Federated Learning
abstract
With the widespread application of federated learning in sensitive domains such as healthcare and finance, preserving the privacy of participants’ local data while maintaining model performance has become a critical challenge. Existing secure aggregation schemes based on homomorphic encryption or pairwise masking still suffer from substantial computational and communication overhead. To address this problem, this paper presents RESA, an efficient secure aggregation protocol based on Ring Learning with Errors (RLWE) encryption. By replacing pairwise masking with the RLWE-based encryption scheme, gradients can be efficiently encrypted and decrypted. Furthermore, by combining Paillier homomorphic encryption with the homomorphic pseudorandom generator (HPRG) to aggregate all clients’ RLWE keys, our scheme effectively protects the privacy of honest clients. Also, the proposed protocol is resilient to client dropouts due to the application of secret sharing. The experimental results show that the proposed protocol achieves 5-10× speedup in the server-side computation, compared to Bell et al. (USENIX Security’23).
Liang Zhang 0043, Haibin Kan, Jiheng Zhang
TrustCom5
2025 Unveiling Discrete Clues: Superior Healthcare Predictions for Rare Diseases
abstract
Accurate healthcare prediction is essential for improving patient outcomes. Existing work primarily leverages advanced frameworks like attention or graph networks to capture the intricate collaborative (CO) signals in electronic health records. However, prediction for rare diseases remains challenging due to limited co-occurrence and inadequately tailored approaches. To address this issue, this paper proposes UDC, a novel method that unveils discrete clues to bridge consistent textual knowledge and CO signals within a unified semantic space, thereby enriching the representation semantics of rare diseases. Specifically, we focus on addressing two key sub-problems: (1) acquiring distinguishable discrete encodings for precise disease representation and (2) achieving semantic alignment between textual knowledge and the CO signals at the code level. For the first sub-problem, we refine the standard vector quantized process to include condition awareness. Additionally, we develop an advanced contrastive approach in the decoding stage, leveraging synthetic and mixed-domain targets as hard negatives to enrich the perceptibility of the reconstructed representation for downstream tasks. For the second sub-problem, we introduce a novel codebook update strategy using co-teacher distillation. This approach facilitates bidirectional supervision between textual knowledge and CO signals, thereby aligning semantically equivalent information in a shared discrete latent space. Extensive experiments on three datasets demonstrate our superiority.
Chuang Zhao 0002, Jiheng Zhang, Xiaomeng Li 0001
WWW3
2025 A closer look at the explainability of Contrastive language-image pre-training
Yi Li 0050, Hualiang Wang, Yiqun Duan, Jiheng Zhang, Xiaomeng Li 0001
Pattern Recognit.4
2025 Data Sharing in the Metaverse With Key Abuse Resistance Based on Decentralized CP-ABE
abstract
Data sharing is ubiquitous in the metaverse, which adopts blockchain as its foundation. Blockchain is employed because it enables data transparency, achieves tamper resistance, and supports smart contracts. However, securely sharing data based on blockchain necessitates further consideration. Ciphertext-policy attribute-based encryption (CP-ABE) is a promising primitive to provide confidentiality and fine-grained access control. Nonetheless, authority accountability and key abuse are critical issues that practical applications must address. Few studies have considered CP-ABE key confidentiality and authority accountability simultaneously. To our knowledge, we are the first to fill this gap by integrating non-interactive zero-knowledge (NIZK) proofs into CP-ABE keys and outsourcing the verification process to a smart contract. To meet the decentralization requirement, we incorporate a decentralized CP-ABE scheme into the proposed data sharing system. Additionally, we provide an implementation based on smart contract to determine whether an access control policy is satisfied by a set of CP-ABE keys. We also introduce an open incentive mechanism to encourage honest participation in data sharing. Hence, the key abuse issue is resolved through the NIZK proof and the incentive mechanism. We provide a theoretical analysis and conduct comprehensive experiments to demonstrate the feasibility and efficiency of the data sharing system. Based on the proposed accountable approach, we further illustrate an application in GameFi, where players can play to earn or contribute to an accountable DAO, fostering a thriving metaverse ecosystem.
Liang Zhang 0043, Zhanrong Ou, Changhui Hu 0002, Haibin Kan, Jiheng Zhang
IEEE Trans. Computers5
2024 RL in Markov Games with Independent Function Approximation: Improved Sample Complexity Bound under the Local Access Model
abstract
Efficiently learning equilibria with large state and action spaces in general-sum Markov games while overcoming the curse of multi-agency is a challenging problem. Recent works have attempted to solve this problem by employing independent linear function classes to approximate the marginal $Q$-value for each agent. However, existing sample complexity bounds under such a framework have a suboptimal dependency on the desired accuracy $\varepsilon$ or the action space. In this work, we introduce a new algorithm, Lin-Confident-FTRL, for learning coarse correlated equilibria (CCE) with local access to the simulator, i.e., one can interact with the underlying environment on the visited states. Up to a logarithmic dependence on the size of the state space, Lin-Confident-FTRL learns $\epsilon$-CCE with a provable optimal accuracy bound $O(\epsilon^{-2})$ and gets rids of the linear dependency on the action space, while scaling polynomially with relevant problem parameters (such as the number of agents and time horizon). Moreover, our analysis of Linear-Confident-FTRL generalizes the virtual policy iteration technique in the single-agent local planning literature, which yields a new computationally efficient algorithm with a tighter sample complexity bound when assuming random access to the simulator.
Junyi Fan, Jialin Zeng, Jian-Feng Cai 0001, Yang Wang 0020, Jiheng Zhang
AISTATS7
2024 Single-Trajectory Distributionally Robust Reinforcement Learning
abstract
To mitigate the limitation that the classical reinforcement learning (RL) framework heavily relies on identical training and test environments, Distributionally Robust RL (DRRL) has been proposed to enhance performance across a range of environments, possibly including unknown test environments. As a price for robustness gain, DRRL involves optimizing over a set of distributions, which is inherently more challenging than optimizing over a fixed distribution in the non-robust case. Existing DRRL algorithms are either model-based or fail to learn from a single sample trajectory. In this paper, we design a first fully model-free DRRL algorithm, called distributionally robust Q-learning with single trajectory (DRQ). We delicately design a multi-timescale framework to fully utilize each incrementally arriving sample and directly learn the optimal distributionally robust policy without modeling the environment, thus the algorithm can be trained along a single trajectory in a model-free fashion. Despite the algorithm’s complexity, we provide asymptotic convergence guarantees by generalizing classical stochastic approximation tools.Comprehensive experimental results demonstrate the superior robustness and sample complexity of our proposed algorithm, compared to non-robust methods and other robust RL algorithms.
Xiaoteng Ma, Jose H. Blanchet, Jun Yang 0028, Jiheng Zhang, Zhengyuan Zhou
ICML5
2023 Optimal Contextual Bandits with Knapsacks under Realizability via Regression Oracles
abstract
We study the stochastic contextual bandit with knapsacks (CBwK) problem, where each action, taken upon a context, not only leads to a random reward but also costs a random resource consumption in a vector form. The challenge is to maximize the total reward without violating the budget for each resource. We study this problem under a general realizability setting where the expected reward and expected cost are functions of contexts and actions in some given general function classes $\mathcal{F}$ and $\mathcal{G}$, respectively. Existing works on CBwK are restricted to the linear function class since they use UCB-type algorithms, which heavily rely on the linear form and thus are difficult to extend to general function classes. Motivated by online regression oracles that have been successfully applied to contextual bandits, we propose the first universal and optimal algorithmic framework for CBwK by reducing it to online regression. We also establish the lower regret bound to show the optimality of our algorithm for a variety of function classes.
Jialin Zeng, Yang Wang 0020, Jiheng Zhang
AISTATS5
2023 Debiasing Recommendation by Learning Identifiable Latent Confounders
abstract
Recommendation systems aim to predict users' feedback on items not exposed to them yet. Confounding bias arises due to the presence of unmeasured variables (e.g., the socio-economic status of a user) that can affect both a user's exposure and feedback. Existing methods either (1) make untenable assumptions about these unmeasured variables or (2) directly infer latent confounders from users' exposure. However, they cannot guarantee the identification of counterfactual feedback, which can lead to biased predictions. In this work, we propose a novel method, i.e., identifiable deconfounder (iDCF), which leverages a set of proxy variables (e.g., observed user features) to resolve the aforementioned non-identification issue. The proposed iDCF is a general deconfounded recommendation framework that applies proximal causal inference to infer the unmeasured confounders and identify the counterfactual feedback with theoretical guarantees. Extensive experiments on various real-world and synthetic datasets verify the proposed method's effectiveness and robustness.
Yang Liu 0018, Hongning Wang, Min Gao 0001, Jiheng Zhang, Ruocheng Guo
KDD6
2022 Private Streaming SCO in ℓp geometry with Applications in High Dimensional Online Decision Making
Zhicong Liang, Yang Wang 0020, Yuan Yao 0011, Jiheng Zhang
ICML6
2022 A Reduction from Linear Contextual Bandit Lower Bounds to Estimation Lower Bounds
Jiheng Zhang, Rachel Q. Zhang
ICML2
2021 Generalized Linear Bandits with Local Differential Privacy
abstract
Contextual bandit algorithms are useful in personalized online decision-making. However, many applications such as personalized medicine and online advertising require the utilization of individual-specific information for effective learning, while user's data should remain private from the server due to privacy concerns. This motivates the introduction of local differential privacy (LDP), a stringent notion in privacy, to contextual bandits. In this paper, we design LDP algorithms for stochastic generalized linear bandits to achieve the same regret bound as in non-privacy settings. Our main idea is to develop a stochastic gradient-based estimator and update mechanism to ensure LDP. We then exploit the flexibility of stochastic gradient descent (SGD), whose theoretical guarantee for bandit problems is rarely explored, in dealing with generalized linear bandits. We also develop an estimator and update mechanism based on Ordinary Least Square (OLS) for linear bandits. Finally, we conduct experiments with both simulation and real-world datasets to demonstrate the consistently superb performance of our algorithms under LDP constraints with reasonably small parameters $(\varepsilon, \delta)$ to ensure strong privacy protection.
Yang Wang 0020, Jiheng Zhang
NeurIPS4
2021 Consensus mechanism design based on structured directed acyclic graphs
abstract
Capacity limit is a bottleneck for broader applications of blockchain systems. Scaling up capacity while preserving security and decentralization are major challenges in blockchain infrastructure design. In this paper, we design a proof of work-based mechanism by endowing directed acyclic graphs (DAG) with a novel structure so that peers can reach consensus at a large scale. At a high level, we break large blocks into smaller ones to improve utilization of broadcast network and embed a Nakamoto chain inside the DAG in a decent way to ensure security. We further exploit the DAG structure and design a mempool transaction assignment method. The method reduces the probability that a transaction is processed by multiple miners and hence improves processing efficiency. Without sacrificing security and decentralization, our design significant scales up capacity and also addresses important issues such as high latency and mining power concentration in existing blockchain systems.
Guangju Wang, Guangyuan Zhang, Jiheng Zhang
Blockchain Res. Appl.4