Zefan Wang

dblp:246/2190 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Optimizing Resource Allocation and Secure Wireless Communication in Large Model Based Mobile Edge Computing Systems
abstract
With the rapid advancement of large models and mobile edge computing, transfer learning through fine-tuning has become essential for adapting models to downstream tasks. Traditionally, users must share their data with model owners, which is costly and raises privacy risks. In addition, fine-tuning large-scale models is computationally intensive and often impractical for many users. To address these challenges, we propose a model that combines offsite-tuning with physical-layer security. Local data owners are given a lightweight adapter and a compressed emulator extracted from the original model. They fine-tune the adapter locally and securely send it back to the model owner through a confidential channel for integration, ensuring privacy and resource conservation. Our work focuses on optimizing computational resource allocation between data owners and the large model owner at the edge, while also optimizing the adapter compression ratio to improve efficiency. We integrate a secrecy uplink channel to maximize the defined utility while minimizing system costs such as energy consumption and delay. The optimization process employs the Dinkelbach algorithm, fractional programming, successive convex approximation, branch-and-bound algorithm, and alternating optimization. Experimental results validate the superiority of our algorithm over several baseline methods.
Zefan Wang, Jun Zhao 0007
IEEE Trans. Mob. Comput.1
2025 GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
abstract
Ultra-high-resolution (UHR) remote sensing (RS) imagery offers valuable data for Earth observation but pose challenges for existing multimodal foundation models due to two key bottlenecks: (1) limited availability of UHR training data, and (2) token explosion caused by the large image size. To address data scarcity, we introduce **SuperRS-VQA** (avg. 8,376$\times$8,376) and **HighRS-VQA** (avg. 2,000$\times$1,912), the highest-resolution vision-language datasets in RS to date, covering 22 real-world dialogue tasks. To mitigate token explosion, our pilot studies reveal significant redundancy in RS images: crucial information is concentrated in a small subset of object-centric tokens, while pruning background tokens (e.g., ocean or forest) can even improve performance. Motivated by these findings, we propose two strategies: *Background Token Pruning* and *Anchored Token Selection*, to reduce the memory footprint while preserving key semantics. Integrating these techniques, we introduce **GeoLLaVA-8K**, the first RS-focused multimodal large language model capable of handling inputs up to 8K$\times$8K resolution, built on the LLaVA framework. Trained on SuperRS-VQA and HighRS-VQA, GeoLLaVA-8K sets a new state-of-the-art on the XLRS-Bench. Datasets and code were released at https://github.com/MiliLab/GeoLLaVA-8K.
Fengxiang Wang 0004, Mingshuo Chen, Di Wang 0023, Haotian Wang 0001, Zonghao Guo, Zefan Wang, Boqi Shan, Long Lan, Yulin Wang 0002, Hongzhen Wang, Wenjing Yang 0002, Bo Du 0001, Jing Zhang 0037
NeurIPS7
2025 A novel framework for assessing determinant risk factors on cyber (dis)trust behaviors of netizens in deepfakes
abstract
Nowadays, Generative Artificial Intelligence (GenAI) tools or trainable agents can craft synthetic media (hereafter referred to as deepfakes) in the form of realistic texts, images, videos, and audios, incorporating events or things that never occurred in real life. These GenAI tools empower marketers and malicious actors to create deepfakes, both authorized and weaponized multimedia, which allows them to include celebrities without appearing in front of cameras or creating seductive phishing scams. Although GenAI tools can reduce the cost of content construction, they enable new risky opportunities (e.g., deepfake phishing and cyberbullying) that negatively impact netizens’ learning and (dis)trust behaviors in cyberspace. To address such risks, this study proposes a Multi-Criteria-Multi-Decision-Makers (MCMDM)-based Deepfake Risk Assessment Framework (DeepFakeR-MF) to evaluate determinant factors that impact the cyber (dis)trust behaviors of netizens in deepfakes. Moreover, DeepFakeR-MF deploys a combination of a novel optimized spherical fuzzy analytic hierarchy process method and a game theory-based MCMDM approach to prioritize and recommend alternative strategies that can be taken by five management sectors (e.g., industrial enterprises, governmental organizations, media outlets, social non-profit, and educational institutes) to mitigate GenAI-associated risks. Then, we collect 100 experts’ judgments by analyzing their responses to our questionnaire and prioritize the importance of determinant factors considering their preferences. To validate the prioritized factors on the performance of DeepFakeR-MF, we conduct a sensitivity analysis applying Monte Carlo statistical modeling. Finally, our results confirm that DeepFakeR-MF provides effective strategic alternatives for policymakers, educators, media professionals, engineers, and netizens, hopefully reducing the socio-economic risks of deepfakes.
Milad Taleby Ahvanooey, Wojciech Mazurczyk, Zefan Wang, Jun Zhao 0007
Eng. Appl. Artif. Intell.3
2024 RCAgent: Cloud Root Cause Analysis by Autonomous Agents with Tool-Augmented Large Language Models
abstract
Large language model (LLM) applications in cloud root cause analysis (RCA) have been actively explored recently. However, current methods are still reliant on manual workflow settings and do not unleash LLMs' decision-making and environment interaction capabilities. We present RCAgent, a tool-augmented LLM autonomous agent framework for practical and privacy-aware industrial RCA usage. Running on an internally deployed model rather than GPT families, RCAgent is capable of free-form data collection and comprehensive analysis with tools. Our framework combines a variety of enhancements, including a unique Self-Consistency for action trajectories, and a suite of methods for context management, stabilization, and importing domain knowledge. Our experiments show RCAgent's evident and consistent superiority over ReAct across all aspects of RCA--predicting root causes, solutions, evidence, and responsibilities--and tasks covered or uncovered by current rules, as validated by both automated metrics and human evaluations. Furthermore, RCAgent has already been integrated into the diagnosis and issue discovery workflow of the Real-time Compute Platform for Apache Flink of Alibaba Cloud.
Zefan Wang, Zichuan Liu, Aoxiao Zhong, Jihong Wang 0003, Fengbin Yin, Lunting Fan, Lingfei Wu 0001, Qingsong Wen
CIKM1
2024 Explaining Time Series via Contrastive and Locally Sparse Perturbations
abstract
Explaining multivariate time series is a compound challenge, as it requires identifying important locations in the time series and matching complex temporal patterns. Although previous saliency-based methods addressed the challenges, their perturbation may not alleviate the distribution shift issue, which is inevitable especially in heterogeneous samples. We present ContraLSP, a locally sparse model that introduces counterfactual samples to build uninformative perturbations but keeps distribution using contrastive learning. Furthermore, we incorporate sample-specific sparse gates to generate more binary-skewed and smooth masks, which easily integrate temporal trends and select the salient features parsimoniously. Empirical studies on both synthetic and real-world datasets show that ContraLSP outperforms state-of-the-art models, demonstrating a substantial improvement in explanation quality for time series data. The source code is available at \url{https://github.com/zichuan-liu/ContraLSP}.
Zichuan Liu, Tianchun Wang, Zefan Wang, Mengnan Du, Min Wu 0008, Yi Wang 0022, Lunting Fan, Qingsong Wen
ICLR4
2024 Resource Allocation and Secure Wireless Communication in the Large Model based Mobile Edge Computing System
abstract
With the rapid advancement of large models and mobile edge computing, transfer learning, particularly through fine-tuning, has become crucial for adapting models to downstream tasks. Traditionally, this requires users to share their data with model owners for fine-tuning, which is not only costly but also raises significant privacy concerns. Furthermore, fine-tuning large-scale models is computationally intensive and often impractical for many users. To tackle these challenges, we introduce a system that combines offsite-tuning with physical-layer security, which provides local data owners with a lightweight adapter and a compressed emulator. Data owners then fine-tune the adapter locally and securely send it back to the model owners through a confidential channel for integration, ensuring privacy and resource conservation. Our paper focuses on optimizing computational resource allocation among data owners and the large model owner deployed on edge, and on the compression ratio of adapters. We incorporate a secrecy uplink channel to maximize the utility that we defined while minimizing system costs like energy consumption and delay. The optimization uses the Dinkelbach algorithm, fractional programming, successive convex approximation and alternating optimization. Experiments demonstrate our algorithm's superiority over baseline methods.
Zefan Wang, Jun Zhao 0007
MobiHoc1
2024 Protecting Your LLMs with Information Bottleneck
abstract
The advent of large language models (LLMs) has revolutionized the field of natural language processing, yet they might be attacked to produce harmful content. Despite efforts to ethically align LLMs, these are often fragile and can be circumvented by jailbreaking attacks through optimized or manual adversarial prompts. To address this, we introduce the Information Bottleneck Protector (IBProtector), a defense mechanism grounded in the information bottleneck principle, and we modify the objective to avoid trivial solutions. The IBProtector selectively compresses and perturbs prompts, facilitated by a lightweight and trainable extractor, preserving only essential information for the target LLMs to respond with the expected answer. Moreover, we further consider a situation where the gradient is not visible to be compatible with any LLM. Our empirical evaluations show that IBProtector outperforms current defense methods in mitigating jailbreak attempts, without overly affecting response quality or inference speed. Its effectiveness and adaptability across various attack methods and target LLMs underscore the potential of IBProtector as a novel, transferable defense that bolsters the security of LLMs without requiring modifications to the underlying models.
Zichuan Liu, Zefan Wang, Linjie Xu, Lei Song 0001, Tianchun Wang, Wei Cheng 0002, Jiang Bian 0002
NeurIPS2
2024 The Convergence of Artificial Intelligence Foundation Models and 6G Wireless Communication Networks
abstract
This review paper explores the powerful convergence of AI foundation models and 6G wireless communication networks, emphasizing their symbiotic relationship and transformative potential. It investigates the advancements within 6G networks, highlighting key areas such as federated learning, blockchain integration, and mobile edge computing. The paper discusses how AI foundation models can enhance 6G communications and vice versa, outlining applications such as the Internet of Vehicles (IoVs) and the Metaverse, aiming to open challenges and future research directions, and underscoring the profound impact of integrating these technologies.
Mohamed R. Shoaib, Zefan Wang, Jun Zhao 0007
VTC Spring2
2024 EViT: Privacy-Preserving Image Retrieval via Encrypted Vision Transformer in Cloud Computing
abstract
Image retrieval systems help users to browse and search among extensive images in real time. With the rise of cloud computing, retrieval tasks are usually outsourced to cloud servers. However, the cloud scenario brings a daunting challenge of privacy protection as cloud servers cannot be fully trusted. To this end, image-encryption-based privacy-preserving image retrieval (PPIR) schemes have been developed, which first extract features from cipher-images, and then build retrieval models based on these features. Yet, most existing PPIR approaches extract shallow features and design trivial unsupervised retrieval models, resulting in insufficient expressiveness for the cipher-images. In this paper, we propose a novel paradigm named Encrypted Vision Transformer (EViT), which advances the discriminative representations capability of cipher-images. First, to capture comprehensive ruled information, we extract multi-level local length sequence and global Huffman-Code frequency features from the cipher-images which are encrypted by permutation encryption, sign encryption, and stream cipher during the JPEG compression process. Second, we design the modified self-supervised Vision Transformer with Huffman-embedding and propose two robust data augmentations on cipher-images to improve representation power of the retrieval model. Moreover, our proposal can be easily adapted to unsupervised or supervised settings. Extensive experiments reveal that EViT achieves both excellent encryption and retrieval performance, outperforming current schemes in terms of retrieval accuracy by large margins while protecting image privacy effectively. Code is publicly available at https://github.com/onlinehuazai/EViT.
Qihua Feng, Peiya Li, Zhixun Lu, Chaozhuo Li, Zefan Wang, Zhiquan Liu 0001, Chunhui Duan, Feiran Huang, Jian Weng 0001, Philip S. Yu
IEEE Trans. Circuits Syst. Video Technol.5
2023 Utility-Oriented Communications for 6G Mobile Networks and the Metaverse: Semantic, Task-Oriented, Goal-Oriented, and More
abstract
Utility-oriented communications consider communication quality metrics traditionally used in bit-oriented communication, such as latency, bit error rate (BER), quality of experience (QoE), and quality of service (QoS), and achieves the ideal utility by encompassing evaluation methods for communication quality in existing communication paradigms such as semantic similarity in semantic communication, task completion in Task-Oriented communication (TOC) and quality of collaboration in Goal-Oriented communications (GOC). We present a utility-oriented communication system for leverages 6G-enabled technologies to support real-time synchronization of IoT-collected data to the metaverse, ensuring rapid and accurate updates of digital twins and optimizing network performance.
Zefan Wang, Jun Zhao 0007
ICDCS1
2023 Multi-factor Sequential Re-ranking with Perception-Aware Diversification
abstract
Feed recommendation systems, which recommend a sequence of items for users to browse and interact with, have gained significant popularity in practical applications. In feed products, users tend to browse a large number of items in succession, so the previously viewed items have a significant impact on users' behavior towards the following items. Therefore, traditional methods that mainly focus on improving the accuracy of recommended items are suboptimal for feed recommendations because they may recommend highly similar items. For feed recommendation, it is crucial to consider both the accuracy and diversity of the recommended item sequences in order to satisfy users' evolving interest when consecutively viewing items. To this end, this work proposes a general re-ranking framework named Multi-factor Sequential Re-ranking with Perception-Aware Diversification~(MPAD) to jointly optimize accuracy and diversity for feed recommendation in a sequential manner. Specifically, MPAD first extracts users' different scales of interests from their behavior sequences through graph clustering-based aggregations. Then, MPAD proposes two sub-models to respectively evaluate the accuracy and diversity of a given item by capturing users' evolving interest due to the ever-changing context and users' personal perception of diversity from an item sequence perspective. This is consistent with the browsing nature of the feed scenario. Finally, MPAD generates the return list by sequentially selecting optimal items from the candidate set to maximize the joint benefits of accuracy and diversity of the entire list. MPAD has been implemented in Taobao's homepage feed to serve the main traffic and provide services to recommend billions of items to hundreds of millions of users every day.
Hao Chen 0062, Zefan Wang, Jianwen Yin, Qijie Shen, Dimin Wang, Feiran Huang, Lixiang Lai, Junfeng Ge, Xia Ben Hu
KDD3
2023 Aligning Distillation For Cold-start Item Recommendation
abstract
Recommending cold items in recommendation systems is a longstanding challenge due to the inherent differences between warm items, which are recommended based on user behavior, and cold items, which are recommended based on content features. To tackle this, generative models generate synthetic embeddings from content features, while dropout models enhance the robustness of the recommendation system by randomly dropping behavioral embeddings during training. However, these models primarily focus on handling the recommendation of cold items, but do not effectively address the differences between warm and cold recommendations. As a result, generative models may over-recommend either warm or cold items, neglecting the other type, and dropout models may negatively impact warm item recommendations. To address this, we propose the Aligning Distillation (ALDI) framework, which leverages warm items as "teachers" to transfer their behavioral information to cold items, referred to as "students". ALDI aligns the students with the teachers by comparing the differences in their recommendation characters, using tailored rating distribution aligning, ranking aligning, and identification aligning losses to narrow these differences. Furthermore, ALDI incorporates a teacher-qualifying weighting structure to prevent students from learning inaccurate information from unreliable teachers. Experiments on three datasets show that our approach outperforms state-of-the-art baselines in terms of overall, warm, and cold recommendation performance with three different recommendation backbones.
Feiran Huang, Zefan Wang, Xiao Huang 0001, Yufeng Qian, Zhetao Li, Hao Chen 0062
SIGIR2
2023 Utility-Oriented Wireless Communications for 6G Networks: Semantic Information Transfer for IRS aided Vehicular Metaverse
abstract
This paper introduces the novel utility-oriented communications (UOC) concept and identifies its importance for 6G wireless technology. UOC encompasses existing communication paradigms and includes emerging human-centric and task-oriented communications concepts. The authors investigate semantic communications and semantic information transfer for vehicular metaverse as a case study of UOC. Consider the Internet of Vehicles (IoV) users access real-time virtual world updates from the base station (BS) wirelessly using semantic communication, and an intelligent reflecting surface (IRS) is deployed to impair co-channel interference. This paper formulates an optimization problem where a novel utility expression for semantic communications is incorporated. The proposed system model jointly considers latency and power in wireless communication and the utility of semantic communication. The proposed alternative optimization algorithm balances system efficiency and economics and outperforms existing optimization algorithms under the same channel conditions.
Zefan Wang, Jun Zhao 0007
VTC2023-Spring1
2022 Generative Adversarial Framework for Cold-Start Item Recommendation
abstract
The cold-start problem has been a long-standing issue in recommendation. Embedding-based recommendation models provide recommendations by learning embeddings for each user and item from historical interactions. Therefore, such embedding-based models perform badly for cold items which haven't emerged in the training set. The most common solutions are to generate the cold embedding for the cold item from its content features. However, the cold embeddings generated from contents have different distribution as the warm embeddings are learned from historical interactions. In this case, current cold-start methods are facing an interesting seesaw phenomenon, which improves the recommendation of either the cold items or the warm items but hurts the opposite ones. To this end, we propose a general framework named Generative Adversarial Recommendation (GAR). By training the generator and the recommender adversarially, the generated cold item embeddings can have similar distribution as the warm embeddings that can even fool the recommender. Simultaneously, the recommender is fine-tuned to correctly rank the "fake'' warm embeddings and the real warm embeddings. Consequently, the recommendation of the warms and the colds will not influence each other, thus avoiding the seesaw phenomenon. Additionally, GAR could be applied to any off-the-shelf recommendation model. Experiments on two datasets present that GAR has strong overall recommendation performance in cold-starting both the CF-based model (improved by over 30.18%) and the GNN-based model (improved by over 17.78%).
Hao Chen 0062, Zefan Wang, Feiran Huang, Xiao Huang 0001, Yishi Lin, Zhoujun Li 0001
SIGIR2
2021 Robust Logo Detection in E-Commerce Images by Data Augmentation
abstract
Logo detection is an important task in the intellectual property protection in e-commerce. In the paper, we introduce our solution for the ACM MM2021 Robust Logo Detection Grand Challenge. The competition requires the detection of logos (515 categories) in e-commerce images. This competition is challenged by long-tail distribution, small objects, and different types of noises. To overcome these challenges, we built a highly optimized and robust detector. We first tested many effective techniques for general object detection and then focused on data augmentation. We found that data augmentation was effective in improving the performance and robustness of logo detection. Based on the combination of these techniques, we achieved APs of 64.6% and 61.3% on the clean and noisy datasets respectively, which were improved by 8.1% and 19.5% relative to the official baseline. We ranked 5th among 36489 teams in the competition.
Hang Chen 0004, Xiao Li 0028, Zefan Wang, Xiaolin Hu 0001
ACM Multimedia3
2019 SignSpeaker: A Real-time, High-Precision SmartWatch-based Sign Language Translator
abstract
Sign language is a natural and fully-formed communication method for deaf or hearing-impaired people. Unfortunately, most of the state-of-the-art sign recognition technologies are limited by either high energy consumption or expensive device costs and have a difficult time providing a real-time service in a daily-life environment. Inspired by previous works on motion detection with wearable devices, we propose Sign Speaker - a real-time, robust, and user-friendly American sign language recognition (ASLR) system with affordable and portable commodity mobile devices. SignSpeaker is deployed on a smartwatch along with a smartphone; the smartwatch collects the sign signals and the smartphone outputs translation through an inbuilt loudspeaker. We implement a prototype system and run a series of experiments that demonstrate the promising performance of our system. For example, the average translation time is approximately $1.1$ seconds for a sentence with eleven words. The average detection ratio and reliability of sign recognition are 99.2% and 99.5%, respectively. The average word error rate of continuous sentence recognition is 1.04% on average.
Jiahui Hou, Xiang-Yang Li 0001, Peide Zhu, Zefan Wang, Yu Wang 0003, Jianwei Qian, Panlong Yang
MobiCom4