Jiayi Lei

dblp:300/7763 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CRB minimization for PASS Assisted ISAC
Haochen Li 0007, Ruikang Zhong, Jiayi Lei, Zhiwen Pan, Yuanwei Liu
ICC3
2026 Joint Sensing and Covert Communications in RIS-NOMA Systems
abstract
A reconfigurable intelligent surface (RIS)-assisted non-orthogonal multiple access (NOMA) system is investigated, where the transmitter (Alice) is a dual-functional radar-communication (DFRC) base station (BS) that aims to sense the location of a potential warden (Willie), while simultaneously transmitting public and covert signals to the legitimate users, Carol and Bob, respectively. Both cases of known and unknown Willie locations are considered. For the known-location case, assuming perfect channel state information (CSI) at Willie, a covert rate maximization is formulated with the joint optimization of active and passive beamforming, which is solved using successive convex approximation (SCA), penalty method, and semidefinite relaxation (SDR). For the unknown-location case, we propose to estimate Willie’s location via radar sensing and develop a sensing-based imperfect CSI model. In particular, the CSI error uncertainty is bounded by the sensing accuracy, which is characterized by the Cramér-Rao bound (CRB). Subsequently, a robust communication rate maximization problem is formulated under the constraints on quality-of-service (QoS) of Carol, sensing accuracy, and covertness level. The Schur complement and S-procedure are employed to handle the non-convex constraints. Numerical results compare the system performance under the two cases, and demonstrate the significant covert performance superiority of the sensing-based imperfect CSI model and NOMA over the general norm-bounded imperfect CSI model and the orthogonal multiple access scheme. Furthermore, the dual yet contradictory effects of sensing on covert communications are revealed. It is also found that Alice primarily utilizes Carol’s signal for sensing, while allocating almost all of Bob’s signal for communication.
Jiayi Lei, Xidong Mu, Tiankui Zhang, Wenjun Xu 0001, Ping Zhang 0003
IEEE J. Sel. Areas Commun.1
2026 Pinching Antenna Systems for Integrated Sensing and Communications
abstract
In this work, a multiple waveguide pinching antenna system (PASS) assisted integrated sensing and communication (ISAC) system is proposed, where the base station (BS) is equipped with transmitting pinching antennas (PAs) and receiving uniform linear array (ULA) antennas. The PASS-transmitting- ULA-receiving (PTUR) BS transmits the communication and sensing signals through the PAs on waveguides and collects the echo sensing signals with the mounted ULA. Based on this configuration, a target sensing Cramèr–Rao Bound (CRB) minimization problem is formulated under communication quality-of-service (QoS) constraints, power budget constraint, and PA deployment constraints. To tackle the resulting non-convex problem, an alternating optimization (AO) framework is developed, which decomposes the problem into a digital beamforming sub-problem and a pinching beamforming sub-problem. The digital beamforming design is optimized via semidefinite relaxation (SDR), while the PA deployment is updated using penalty-based method. Simulation results demonstrate that: 1) the proposed PASS assisted ISAC framework achieves superior performance over benchmark schemes; and 2) the PASS assisted ISAC is less affected by stringent communication constraints compared to conventional MIMO-ISAC, and benefits from increasing the number of waveguides and PAs per waveguide.
Haochen Li 0007, Ruikang Zhong, Zhiwen Pan, Chao Dong 0001, Jiayi Lei, Yuanwei Liu
IEEE Trans. Wirel. Commun.5
2025 MMSearch: Unveiling the Potential of Large Models as Multi-modal Search Engines
abstract
The advent of Large Language Models (LLMs) has paved the way for AI search engines, e.g., SearchGPT, showcasing a new paradigm in human-internet interaction. However, most current AI search engines are limited to text-only settings, neglecting the multimodal user queries and the text-image interleaved nature of website information. Recently, Large Multimodal Models (LMMs) have made impressive strides. Yet, whether they can function as AI search engines remains under-explored, leaving the potential of LMMs in multimodal search an open question. To this end, we first design a delicate pipeline, MMSearch-Engine, to empower any LMMs with multimodal search capabilities. On top of this, we introduce MMSearch, a comprehensive evaluation benchmark to assess the multimodal search performance of LMMs. The curated dataset contains 300 manually collected instances spanning 14 subfields, which involves no overlap with the current LMMs' training data, ensuring the correct answer can only be obtained within searching. By using MMSearch-Engine, the LMMs are evaluated by performing three individual tasks (requery, rerank, and summarization), and one challenging end-to-end task with a complete searching process. We conduct extensive experiments on closed-source and open-source LMMs. Among all tested models, GPT-4o with MMSearch-Engine achieves the best results, which surpasses the commercial product, Perplexity Pro, in the end-to-end task, demonstrating the effectiveness of our proposed pipeline. We further present error analysis to unveil current LMMs still struggle to fully grasp the multimodal search tasks, and conduct ablation study to indicate the potential of scaling test-time computation for AI search engine. We hope MMSearch may provide unique insights to guide the future development of multimodal AI search engine.
Dongzhi Jiang, Renrui Zhang, Yanmin Wu, Jiayi Lei, Pengshuo Qiu, Pan Lu, Guanglu Song, Peng Gao 0007, Yu Liu 0015, Chunyuan Li, Hongsheng Li 0001
ICLR5
2025 Fixed-Time Distributed Optimization of High-Order Nonlinear Multiagent Systems Under Switching Topologies: A Two-Layer Control Framework
abstract
This paper investigates the fixed-time distributed optimization of nonlinear multi-agent systems (MASs) under switching topologies. In contrast to existing optimization strategies, this paper considers uncertain high-order dynamics and realizes fixed-time stability. To overcome the challenges brought by the coexistence of the switching topologies and high-order uncertain dynamics, a novel distributed optimization control method via two-layer framework is constructed, which consists of a Network Layer-optimization estimator design and a Physical Layer-Agents’ reference-tracking control law design. In Network Layer, the fixed-time optimal signal generator is constructed by using information interaction between agent and real-time feedback value of local gradient information over switching topologies. Then, in Physical Layer, the fixed-time fuzzy adaptive tracking control strategy is designed via backstepping technology to track the virtual signal generated from the Network-Layer. A fast fixed-time filter (FFTF) is introduced to avoid taking the derivation of discontinuous gradient functions in the process of backstepping. Furthermore, fuzzy logic systems (FLSs) are used to handle unknown nonlinear functions. By using the convex optimization theory, Lyapunov stability theory, and the fixed-time stability criterion, we analyze the convergence of the system and fixed-time stability. Finally, a simulation example is given to validate the control strategy.
Jiayi Lei, Yuan-Xin Li 0001, Choon Ki Ahn, Heng Wang 0002
IEEE Internet Things J.1
2025 Equivariance-Based Markov Decision Process for Unsupervised Point Cloud Registration
abstract
Unsupervised point cloud registration is crucial in 3D computer vision. However, most unsupervised methods struggle to construct effective optimization objectives and reliable unsupervised signals to enhance the performance of the model. To address these issues, with the observation of the significant alignment between the registration process and the Markov Decision Process (MDP), we model point cloud registration as MDP, which can provide more reliable unsupervised signals through the reward. We propose a colored noise based cross-entropy method, which introduces colored noise into sampling process, regulating the power spectral density of the action sequence and expanding the search space, improving the registration effect. Particularly, to strengthen constraints on MDP and training in the transformation space, we utilize equivariance theory to construct transformation equivariant constraint as a new optimization objective and derive equivariant constraint solutions for optimization, providing more reliable unsupervised signals. Extensive experiments demonstrate the superior performance of our method on benchmark datasets.
Yue Wu 0004, Jiayi Lei, Yongzhe Yuan, Xiaolong Fan, Maoguo Gong, Wenping Ma 0001, Qiguang Miao, Mingyang Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2024 MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
abstract
Large Vision-Language Models (LVLMs) show significant strides in general-propose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, and reasoning. MMT-Bench comprises $31,325$ meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering $32$ core meta-tasks and $162$ subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving $20$ publicly available LVLMs such as the proprietary GeminiProVision model, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence.
Kaining Ying, Fanqing Meng, Zhiqian Li, Hao Zhang 0117, Wenbo Zhang 0009, Yuqi Lin, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu 0035, Renrui Zhang, Haozhe Zhang 0002, Peng Gao 0007, Yali Wang 0001, Yu Qiao 0001, Ping Luo 0002, Kaipeng Zhang, Wenqi Shao
ICML11
2024 NOMA for STAR-RIS Assisted UAV Networks
abstract
This paper proposes a novel simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) assisted unmanned aerial vehicle (UAV) non-orthogonal multiple access (NOMA) emergency communication network. Multiple STAR-RISs are deployed to provide additional and intelligent transmission links between trapped users and UAV-mounted base station (BS). Each user selects the nearest STAR-RIS for uploading data, and NOMA is employed for users located at the same side of the same STAR-RIS. Considering practical requirements of post-disaster emergency communications, we formulate a throughput maximization problem subject to constraints on minimum average rate and maximum energy consumption, where the UAV trajectory, STAR-RIS passive beamforming, and time and power allocation are jointly optimized. Furthermore, we propose a Lagrange based reward constrained proximal policy optimization (LRCPPO) algorithm, which provides an adaptive method for solving the long-term optimization problem with cumulative constraints. Specifically, using Lagrange relaxation, the original problem is transformed into an unconstrained problem with a two-layer structure. The inner layer is solved by penalized reward based proximal policy optimization (PPO) algorithm. In the outer layer, Lagrange multipliers are updated by gradient descent. Numerical results show the proposed algorithm can effectively improve network performance while satisfying the constraints well. It also demonstrates the superiority of the proposed STAR-RIS assisted UAV NOMA network architecture over the benchmark schemes employing reflecting-only RISs and orthogonal multiple access.
Jiayi Lei, Tiankui Zhang, Xidong Mu, Yuanwei Liu
IEEE Trans. Commun.1
2024 Fuzzy Adaptive Distributed Optimization of Uncertain Multiagent Systems With Time-Varying Delays
abstract
The distributed optimal problem is investigated in this article for uncertain nonlinear time-varying delay multiagent systems (MASs) under unknown virtual control coefficients. The primary goal is to ensure the output states converge to the ideal point of the global objective function. In comparison to the preceding results on distributed optimization, the considered MASs are subject to time delays and unknown nonlinear functions, which brought great challenges to the controller design process. To overcome this challenge, a distributed optimization algorithm is proposed via a two-stage strategy, which consists of optimization estimator design and tracking controller design. First, in the first-stage, a novel optimal signal generator is constructed by using information interaction between agents, real-time feedback value of local gradient information, so that the optimization problem is transformed into a problem of tracking control. Second, in second-stage, the fuzzy adaptive tracking control strategy is designed via backstepping technology to track the virtual signal accurately. In addition, the Lyapunov–Krasovskii functional (LKF) is applied to counteract influences brought by time-varying delays, and fuzzy logic systems are employed for handling the uncertain function in the derivative of LKF. Furthermore, an assumption is made for the time-varying delays to settle the singularity problem, which often occurs in time-delay problems. In the spirit of the convex analysis and the Barbalats lemma, the stability of the proposed algorithm is illustrated, and all variables in the closed-loop systems are bounded. Lastly, simulations are presented to validate the efficacy of the scheme.
Jiayi Lei, Yuan-Xin Li 0001, Shaocheng Tong
IEEE Trans. Fuzzy Syst.1
2023 SACAT: Student-Adaptive Computerized Adaptive Testing
abstract
Computerized Adaptive Testing refers to an online student ability testing system. It consists of Selection Algorithm and Cognitive Diagnosis Module (CDM). The Selection Algorithm aims to carefully select a small set of questions for the newly arrived student to answer, and the CDM aims to infer the ability of the newly arrived student based on his answers to the selected questions. The CDM faces the challenge of inferring student ability based on the short question-answer sequence. The state-of-the-art approach employs a meta-learning framework to solve this challenge. Specifically, in the outer-layer of the meta-learning framework, it learns an average student ability from all students as the initial ability for the newly arrived student, while in the inner-layer of the meta-learning framework, it infers the newly arrived student’s ability from the short sequence. However, this approach suffers from two major limitations: Firstly, collaborative information between the newly arrived student and all students was not utilized. Secondly, the meta-learning framework has a slow convergence. This paper introduces Student-Adaptive Computerized Adaptive Testing (SACAT), which addresses these issues by: Firstly, use the short sequence of the newly arrived student to retrieve similar students from all students. Secondly, design a fast adaptive function that can adaptively generate an ability representation suitable for the newly arrived student from abilities of similar students. Through extensive experiments on three real-world student response datasets, we validate that SACAT can outperform existing CAT methods. Furthermore, through the fast adaptation, we show that SACAT achieves fast model convergence.
Jingwei Yu, Zhenyu Mu, Jiayi Lei, Wei Xia 0001, Yong Yu 0001, Ting Long
DAI3
2023 C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
abstract
New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present C-Eval, the first comprehensive Chinese evaluation suite designed to assess advanced knowledge and reasoning abilities of foundation models in a Chinese context. C-Eval comprises multiple-choice questions across four difficulty levels: middle school, high school, college, and professional. The questions span 52 diverse disciplines, ranging from humanities to science and engineering. C-Eval is accompanied by C-Eval Hard, a subset of very challenging subjects in C-Eval that requires advanced reasoning abilities to solve. We conduct a comprehensive evaluation of the most advanced LLMs on C-Eval, including both English- and Chinese-oriented models. Results indicate that only GPT-4 could achieve an average accuracy of over 60%, suggesting that there is still significant room for improvement for current LLMs. We anticipate C-Eval will help analyze important strengths and shortcomings of foundation models, and foster their development and growth for Chinese users.
Yuzhen Huang 0002, Yuzhuo Bai, Junlei Zhang, Jinghan Zhang 0006, Tangjun Su, Junteng Liu, Chuancheng Lv, Jiayi Lei, Maosong Sun 0001, Junxian He
NeurIPS10
2023 A Multi-Modal Behavior Quantitative Analysis Model for Autism Early Screening
abstract
Human-Computer Interaction (HCI) and Machine Learning (ML) technologies have potential for the behavioral screening of autistic children but how to design a tool and analyse behavior reliably is challenging. Based on psychophysiological computation, this paper proposes an interactive behavior perception analytical model for autism screening. We presented the multi-scenario reactive behavior paradigms that designed based on the atypical characteristics of autistic children. We recorded the eye movement data and facial data of 91 participants, and performed multi-modal feature extraction, used machine learning to train classification model. We conducted comparative experiments, and the experimental results verified the advantages of multi-scenario paradigms and multi-modal feature groups, which indicates that our analysis methods and screening models are effective and reliable and have real research significance.
Jiayi Lei, Erping Zhang, Yingying She, Yuhan Liao, Bin Hu 0001, Minqiang Yang, Jiajia Tian
SMC1
2022 Using a social robot for children with autism: A therapist-robot interactive model
abstract
Abstract Social robots have great potential for the therapy of children with autism spectrum disorder, but the practical use of them is challenging. In this article, we presented a social robot YANG, which can interact with children with autism in their daily training. We propose a therapist‐robot interactive (TRI) model, which integrates with the practice of discrete trial training (DTT), a basic method utilized in autism training. To evaluate the TRI model, we implemented the model in YANG and conducted a single‐subject experiment in a rehabilitation training center. Data were collected on three children (ages 3–4) and their three therapists as they interacted with YANG in training sessions. Results showed that the children's learning ability significantly improved. YANG formed natural and friendly relationships with the children and delivered substantial support to therapists. Our research brings insight into using social robots for children with autism.
Yingying She, E (Alice) Zhang, Jiayi Lei, Jufeng Li
Comput. Animat. Virtual Worlds7