Peng Yuan Zhou

dblp:192/6936 · also Pengyuan Zhou · DBLP profile ↗
← Back
39ranked-venue papers
3as first author
36since 2021 · last 2026
0000-0002-7909-4059ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 15 since 2021Artificial intelligence and machine learning · 13 · 13 since 2021Systems, architecture and hardware · 5 · 4 since 2021Computer networks · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
YearPublicationVenuePosition
2026 QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis
abstract
Yitong Zhu, Yuxuan Jiang, Guanxuan Jiang, Bojing Hou, Peng Yuan Zhou, Ge Lin, Yuyang Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yitong Zhu, Guanxuan Jiang, Bojing Hou, Peng Yuan Zhou, Ge Lin 0001
ACL (1)5
2026 DiSCo: Disrupting Semantic Consistency for Transferable Cross-Modal Adversarial Attacks
Peirou Liang, Meng Yang 0029, Zhiqian Wu, Peng Yuan Zhou, Yong Liao 0003
MMM (2)4
2025 GraphCheck: Breaking Long-Term Text Barriers with Extracted Knowledge Graph-Powered Fact-Checking
abstract
, a fact-checking framework that uses extracted knowledge graphs to enhance text representation. Graph Neural Networks further process these graphs as a soft prompt, enabling LLMs to incorporate structured knowledge more effectively. Enhanced with graph-based reasoning, GraphCheck captures multihop reasoning chains that are often overlooked by existing methods, enabling precise and efficient fact-checking in a single inference call. Experimental results on seven benchmarks spanning both general and medical domains demonstrate up to a 7.1% overall improvement over baseline models. Notably, GraphCheck outperforms existing specialized fact-checkers and achieves comparable performance with state-of-the-art LLMs, such as DeepSeek-V3 and OpenAI-o1, with significantly fewer parameters.
Yingjian Chen, Yinhong Liu, Jinxiang Xie, Rui Yang 0016, Yanran Fu, Peng Yuan Zhou, Qingyu Chen 0001, James Caverlee, Irene Li
ACL (1)8
2025 A Unified Framework to BRIDGE Complete and Incomplete Deep Multi-View Clustering Under Non-IID Missing Patterns
Xiaorui Jiang, Buyun He, Peng Yuan Zhou, Jingcai Guo, Yong Liao 0003
ICCV3
2025 Knowledge Rumination for Client Utility Evaluation in Heterogeneous Federated Learning
abstract
Federated Learning (FL) allows several clients to cooperatively train machine learning models without disclosing the raw data. In practical applications, asynchronous FL (AFL) can address the straggler effect compared to synchronous FL. However, Non-Iid data and stale models pose significant challenges to AFL, as they can diminish the practicality of the global model and even lead to training failures. In this work, we propose a novel AFL framework called Federated Historical Learning (FedHist), which effectively addresses the challenges posed by both Non-Iid data and gradient staleness based on the concept of knowledge rumination. FedHist enhances the stability of local gradients by performing weighted fusion with historical global gradients cached on the server. Relying on hindsight, it assigns aggregation weights to each participant in a multi-dimensional manner during each communication round. To further enhance the efficiency and stability of the training process, we introduce an intelligent ℓ2-norm amplification scheme, which dynamically regulates the learning progress based on the ℓ2-norms of the submitted gradients. Extensive experiments indicate FedHist outperforms state-of-the-art methods in terms of convergence performance and test accuracy.
Xiaorui Jiang, Hengwei Xu, Qi Zhang 0001, Yong Liao 0003, Peng Yuan Zhou
ICME6
2025 Detecting AI-Generated Video via Frame Consistency
abstract
The increasing realism of AI-generated videos has raised potential security concerns, making it difficult for humans to distinguish them from the naked eye. Despite these concerns, limited research has been dedicated to detecting such videos effectively. To this end, we propose an open-source AI-generated video detection dataset. Our dataset spans diverse objects, scenes, behaviors, and actions by organizing input prompts into independent dimensions. It also includes various generation models with different generative models, featuring popular commercial models such as OpenAI’s Sora, Google’s Veo, and Kwai’s Kling. Furthermore, we propose a simple yet effective Detection model based on Concistency of Frame (DeCoF), which learns robust temporal artifacts across different generation methods. Extensive experiments demonstrate the generality and efficacy of the proposed DeCoF in detecting AI-generated videos, including those from nowadays’ mainstream commercial generators.
Qinglang Guo, Yong Liao 0003, Haiyang Yu 0003, Peng Yuan Zhou
ICME6
2025 $\lambda$-SecAgg: Partial Vector Freezing for Lightweight Secure Aggregation in Federated Learning
abstract
Secure aggregation of user update vectors (e.g. gradients) has become a critical issue in the field of federated learning. Many Secure Aggregation Protocols (SAPs) face exorbitant computation costs, severely constraining their applicability. Given the observation that a considerable portion of SAP's computation burden stems from processing each entry in the private vectors, we propose Partial Vector Freezing (PVF), a portable module for compressing computation costs without introducing additional communication overhead.$\lambda$-SecAgg, which integrates SAP with PVF, “freezes” a substantial portion of the private vector through specific transformations, requiring only$\frac{1}{\lambda}$of the original vector to participate in SAP. Eventually, users can “thaw” the public sum of the “frozen entries” by the result of SAP. To avoid potential privacy leakage, we devise Disrupting Variables Extension for PVF. We demonstrate that PVF can seamlessly integrate with various SAPs and it poses no threat to user privacy in the semi-honest and active adversary settings. We include 7 baselines, encompassing 5 distinct types of masking schemes, and explore the acceleration effects of PVF on these SAPs. Empirical investigations indicate that when$\lambda=100$, PVF yields up to$99.5 \times$speedup and up to$32.3 \times$communication reduction.
Siqing Zhang 0002, Yong Liao 0003, Peng Yuan Zhou
ICPADS3
2025 Privacy-Preserving Orthogonal Aggregation for Guaranteeing Gender Fairness in Federated Recommendation
abstract
Under stringent privacy constraints, whether federated recommendation systems can achieve group fairness remains an inadequately explored question. Taking gender fairness as a representative issue, we identify three phenomena in federated recommendation systems: performance difference, data imbalance, and preference disparity. We discover that the state-of-the-art methods only focus on the first phenomenon. Consequently, their imposition of inappropriate fairness constraints detrimentally affects the model training. Moreover, due to insufficient sensitive attribute protection of existing works, we can infer the gender of all users with 99.90% accuracy even with the addition of maximal noise. In this work, we propose Privacy-Preserving Orthogonal Aggregation (PPOA), which employs the secure aggregation scheme and quantization technique, to prevent the suppression of minority groups by the majority and preserve the distinct preferences for better group fairness. PPOA can assist different groups in obtaining their respective model aggregation results through a designed orthogonal mapping while keeping their attributes private. Experimental results on three real-world datasets demonstrate that PPOA enhances recommendation effectiveness for both females and males by up to 8.25% and 6.36%, respectively, with a maximum overall improvement of 7.30%, and achieves optimal fairness in most cases. Extensive ablation experiments and visualizations indicate that PPOA successfully maintains preferences for different gender groups.
Siqing Zhang 0002, Yuchen Ding, Wei Tang 0015, Yong Liao 0003, Peng Yuan Zhou
WSDM6
2025 360SFUDA++: Towards Source-Free UDA for Panoramic Segmentation by Learning Reliable Category Prototypes
abstract
In this paper, we address the challenging source-free unsupervised domain adaptation (SFUDA) for pinhole-to-panoramic semantic segmentation, given only a pinhole image pre-trained model (i.e., source) and unlabeled panoramic images (i.e., target). Tackling this problem is non-trivial due to three critical challenges: 1) semantic mismatches from the distinct Field-of-View (FoV) between domains, 2) style discrepancies inherent in the UDA problem, and 3) inevitable distortion of the panoramic images. To tackle these problems, we propose 360SFUDA++ that effectively extracts knowledge from the source pinhole model with only unlabeled panoramic images and transfers the reliable knowledge to the target panoramic domain. Specifically, we first utilize Tangent Projection (TP) as it has less distortion and meanwhile slits the equirectangular projection (ERP) to patches with fixed FoV projection (FFP) to mimic the pinhole images. Both projections are shown effective in extracting knowledge from the source model. However, as the distinct projections make it less possible to directly transfer knowledge between domains, we then propose Reliable Panoramic Prototype Adaptation Module (RPAM) to transfer knowledge at both prediction and prototype levels. RPAM selects the confident knowledge and integrates panoramic prototypes for reliable knowledge adaptation. Moreover, we introduce Cross-projection Dual Attention Module (CDAM), which better aligns the spatial and channel characteristics across projections at the feature level between domains. Both knowledge extraction and transfer processes are synchronously updated to reach the best performance. Extensive experiments on the synthetic and real-world benchmarks, including outdoor and indoor scenarios, demonstrate that our 360SFUDA++ achieves significantly better performance than prior SFUDA methods.
Xu Zheng 0002, Peng Yuan Zhou, Athanasios V. Vasilakos, Lin Wang 0025
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Distilling efficient Vision Transformers from CNNs for semantic segmentation
Xu Zheng 0002, Yunhao Luo 0001, Peng Yuan Zhou, Lin Wang 0025
Pattern Recognit.3
2025 Viewport Prediction for Volumetric Video Streaming by Exploring Video Saliency and User Trajectory Information
abstract
Volumetric video, also referred to as hologram video, is an emerging medium that represents 3D content in extended reality. As a next-generation video technology, it is poised to become a key application in 5G and future wireless communication networks. Because each user generally views only a specific portion of the volumetric video, known as the viewport, accurate prediction of the viewport is crucial for ensuring an optimal streaming performance. Despite its significance, research in this area is still in the early stages. To this end, this paper introduces a novel approach called Saliency and Trajectory-based Viewport Prediction (STVP), which enhances the accuracy of viewport prediction in volumetric video streaming by effectively leveraging both video saliency and viewport trajectory information. In particular, we first introduce a novel sampling method, Uniform Random Sampling (URS), which efficiently preserves video features while minimizing computational complexity. Next, we propose a saliency detection technique that integrates both spatial and temporal information to identify visually static and dynamic geometric and luminance-salient regions. Finally, we fuse saliency and trajectory information to achieve more accurate viewport prediction. Extensive experimental results validate the superiority of our method over existing state-of-the-art schemes. To the best of our knowledge, this is the first comprehensive study of viewport prediction in volumetric video streaming. We also make the source code of this work publicly available.
Jie Li 0015, Zhi Liu 0002, Peng Yuan Zhou, Richang Hong, Qiyue Li 0001, Han Hu 0003
IEEE Trans. Circuits Syst. Video Technol.4
2025 VPFormer: Leveraging Transformer with Voxel Integration for Viewport Prediction in Volumetric Video
abstract
With the continuous advancement of computer vision, image processing technologies, volumetric video, represented by point cloud videos, holds the potential for extensive applications in areas such as Virtual Reality (VR) and Augmented Reality (AR). Viewport prediction, also referred to as Field of View (FoV) prediction, is a crucial component in emerging VR and AR applications, playing a vital role in the transmission of point cloud videos. Currently, models for viewpoint prediction that integrate feature extraction and FoV information heavily rely on the spatial-temporal features extracted by convolutional neural networks. However, the drawback of 3D convolution lies in its inability to effectively capture long-term spatial-temporal dependencies within videos. Moreover, the temporal contrast layer used for time feature extraction only compares features within each block, leading to matching errors and inaccurate temporal feature extraction, consequently diminishing predictive performance. To address these limitations, we propose a Transformer-based Volumetric Point Cloud Video Viewport Prediction Network (VPFormer) that can efficiently extract spatial-temporal features from point cloud videos. VPFormer constitutes a viewport prediction framework that combines the spatial-temporal features of point cloud videos with user trajectory information. Specifically, we introduce a novel sampling method that effectively preserves spatial-temporal information while reducing computational complexity. Additionally, we incorporate context-aware dynamic positional encoding to capture inter-frame spatial-temporal context information. Subsequently, we introduce a voxel-based temporal contrast layer and partition the point cloud into smaller voxel blocks during feature matching, significantly reducing matching errors and enhancing the analysis and extraction of temporal features. Finally, by combining the spatial-temporal features of point cloud videos with user head trajectory information, we successfully predict future user viewpoints. Experimental results demonstrate that this approach outperforms other solutions in terms of performance.
Jie Li 0015, Zhixia Zhao, Qiyue Li 0001, Peng Yuan Zhou, Zhi Liu 0002, Hao Zhou 0001, Zhu Li 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Semantics, Distortion, and Style Matter: Towards Source-Free UDA for Panoramic Segmentation
abstract
This paper addresses an interesting yet challenging problem-source-free unsupervised domain adaptation (SFUDA) for pinhole-to-panoramic semantic segmentation-given only a pinhole image-trained model (i.e., source) and unlabeled panoramic images (i.e., target). Tackling this problem is nontrivial due to the semantic mismatches, style discrepancies, and inevitable distortion of panoramic images. To this end, we propose a novel method that utilizes Tangent Projection (TP) as it has less distortion and meanwhile slits the equirectangular projection (ERP) with a fixed FoV to mimic the pinhole images. Both projections are shown effective in extracting knowledge from the source model. However, the distinct projection discrepancies between source and target domains impede the direct knowledge transfer; thus, we propose a panoramic prototype adaptation module (PPAM) to integrate panoramic prototypes from the extracted knowledge for adaptation. We then impose the loss constraints on both predictions and prototypes and propose a cross-dual attention module (CDAM) at the feature level to better align the spatial and channel characteristics across the domains and projections. Both knowledge extraction and transfer processes are synchronously updated to reach the best performance. Extensive experiments on the synthetic and real-world benchmarks, including outdoor and indoor scenarios, demonstrate that our method achieves significantly better performance than prior SFUDA methods for pinhole-to-panoramic adaptation.
Xu Zheng 0002, Peng Yuan Zhou, Athanasios V. Vasilakos, Lin Wang 0025
CVPR2
2024 3D-GOI: 3D GAN Omni-Inversion for Multifaceted and Multi-object Editing
Haoran Li 0020, Haolin Shi, Yanbin Hao, Yong Liao 0003, Lechao Cheng, Peng Yuan Zhou
ECCV (62)7
2024 DreamScene: 3D Gaussian-Based Text-to-3D Scene Generation via Formation Pattern Sampling
Haoran Li 0020, Haolin Shi, Yong Liao 0003, Lin Wang 0025, Lik-Hang Lee, Peng Yuan Zhou
ECCV (74)8
2024 Noise-NeRF: Hide Information in Neural Radiance Field Using Trainable Noise
abstract
Neural Radiance Field (NeRF) has been proposed as an innovative advancement in 3D reconstruction techniques. However, little research has been conducted on the issues of information confidentiality and security to NeRF, such as steganography. Existing NeRF steganography solutions have shortcomings in low steganography quality, model weight damage, and limited amount of steganographic information. This paper proposes Noise-NeRF, a novel NeRF steganography method employing Adaptive Pixel Selection strategy and Pixel Perturbation strategy to improve the quality and efficiency of steganography via trainable noise. Extensive experiments validate the state-of-the-art performances of Noise-NeRF on both steganography quality and rendering quality, as well as effectiveness in super-resolution image steganography.
Qinglong Huang, Haoran Li 0020, Yong Liao 0003, Yanbin Hao, Peng Yuan Zhou
ICANN (2)5
2024 Dynamicity-aware Social Bot Detection with Dynamic Graph Transformers
Buyun He, Yingguang Yang, Qi Wu 0021, Hao Liu 0007, Renyu Yang, Hao Peng 0001, Xiang Wang 0010, Yong Liao 0003, Peng Yuan Zhou
IJCAI9
2024 FacGNN: Multi-faceted Fairness Enhancement for GNN through Adversarial and Contrastive Learning
abstract
Albeit presenting great capabilities for graph node representations, Graph neural networks (GNNs) face biases and discrimination issues. Existing solutions mostly focus on one aspect of fairness, failing to thoroughly consider the multiple dimensions of fairness. To break this limitation, we propose a novel fairness-aware framework, FacGNN, which establishes connections between group fairness, counterfactual fairness, and stability fairness for the first time. Specifically, FacGNN enhances both group fairness and counterfactual fairness while extending stability metrics beyond prior works. FacGNN employs a phased nested iterative training process with self-supervised contrastive learning, to model fairness relationships and improve discriminatory detection in adversarial frameworks. Experimental results show that FacGNN outperforms previous methods across group fairness, counterfactual fairness, and model stability while maintaining model performance.
Hao Liu 0007, Yingguang Yang, Qi Wu 0021, Buyun He, Yong Liao 0003, Peng Yuan Zhou
IJCNN6
2024 Heterogeneity-Aware Federated Deep Multi-View Clustering towards Diverse Feature Representations
abstract
Multi-view clustering has proven to be highly effective in exploring consistency information across multiple views/modalities when dealing with large-scale unlabeled data. However, in the real world, multi-view data is often distributed across multiple entities, and due to privacy concerns, federated multi-view clustering solutions have emerged. Existing federated multi-view clustering algorithms often result in misalignment in feature representations among clients, difficulty in integrating information across multiple views, and poor performance in heterogeneous scenarios. To address these challenges, we propose HFMVC, a heterogeneity-aware federated deep multi-view clustering method. Specifically, HFMVC adaptively perceives the degree of heterogeneity in the environment and employs contrastive learning to explore consistency and complementarity information across clients' multi-view data. Besides, we seek consensus among clients where local data originates from the same view, incorporating a contrastive loss between local models and the global model during local training to adjust consistency among local models. Furthermore, we elucidate the sample representation logic for local clustering in different heterogeneous environments, identifying the degree of heterogeneity by computing the within-cluster sum of squares (WCSS) and the average inter-cluster distance (AICD). Extensive experiments verify the superior performance of HFMVC across both IID and Non-IID settings.
Xiaorui Jiang, Zhongyi Ma, Yulin Fu, Yong Liao 0003, Peng Yuan Zhou
ACM Multimedia5
2024 FedLoCA: Low-Rank Coordinated Adaptation with Knowledge Decoupling for Federated Recommendations
abstract
Privacy protection in recommendation systems is gaining increasing attention, for which federated learning has emerged as a promising solution. Current federated recommendation systems grapple with high communication overhead due to sharing dense global embeddings, and also poorly reflect user preferences due to data heterogeneity. To overcome these challenges, we propose a two-stage Federated Low-rank Coordinated Adaptation (FedLoCA) framework to decouple global and client-specific knowledge into low-rank embeddings, which significantly reduces communication overhead while enhancing the system’s ability to capture individual user preferences amidst data heterogeneity. Further, to tackle gradient estimation inaccuracies stemming from data sparsity in federated recommendation systems, we introduce an adversarial gradient projected descent approach in low-rank spaces, which significantly boosts model performance while maintaining robustness. Remarkably, FedLoCA also alleviates performance loss even under the stringent constraints of differential privacy. Extensive experiments on various real-world datasets demonstrate that FedLoCA significantly outperforms existing methods in both recommendation accuracy and communication efficiency.
Yuchen Ding, Siqing Zhang 0002, Boyu Fan, Yong Liao 0003, Peng Yuan Zhou
RecSys6
2024 The Jade Gateway to Exergaming: How Socio-Cultural Factors Shape Exergaming Among East Asian Older Adults
abstract
Exergaming, blending exercise and gaming, improves the physical and mental health of older adults. We currently do not fully know the factors that drive older adults to either engage in or abstain from exergaming. Large-scale studies investigating this are still scarce, particularly those studying East Asian older adults. To address this, we interviewed 64 older adults from China, Japan, and South Korea about their attitudes toward exergames. Most participants viewed exergames with a positive inquisitiveness. However, socio-cultural factors can obstruct this curiosity. Our study shows that perceptions of aging, lifestyle, the presence of support networks, and the cultural relevance of game mechanics are the crucial factors influencing their exergame engagement. Thus, we stress the value of socio-cultural sensitivity in game design and urge the HCI community to adopt more diverse design practices. We provide several design suggestions for creating more culturally approachable exergames.
Reza Hadi Mogavi, Juhyung Son, Simin Yang, Derrick M. Wang, Lydia Choong, Ahmad Yousef Alhilal, Peng Yuan Zhou, Pan Hui 0001, Lennart E. Nacke
Proc. ACM Hum. Comput. Interact.7
2024 Unleashing GPT on the Metaverse: Savior or Destroyer? [Point of View]
abstract
Incorporating artificial intelligence (AI) technology, particularly large language models (LLMs), is becoming increasingly vital for developing immersive and interactive metaverse experiences. GPT, a representative LLM developed by OpenAI, is leading LLM development and gaining attention for its potential in building the metaverse. This article delves into the pros and cons of utilizing GPT for metaverse-based education, entertainment, personalization, and support. Dynamic and personalized experiences are possible with this technology, but there are also legitimate privacy, bias, and ethical issues to consider. This article aims to help readers understand the possible influence of GPT, according to its unique technological advantages, on the metaverse and how it may be used to effectively create a more immersive and engaging virtual environment by evaluating these opportunities and obstacles.
Peng Yuan Zhou
Proc. IEEE1
2024 Efficient Unsupervised Video Hashing With Contextual Modeling and Structural Controlling
abstract
The most important effect of the video hashing technique is to support fast retrieval, which is benefiting from the high efficiency of binary calculation. Current video hash approaches are thus mainly targeted at learning compact binary codes to represent video content accurately. However, they may overlook the generation efficiency for hash codes, i.e., designing lightweight neural networks. This paper proposes anEfficientUnsupervisedVideoHashing (EUVH)method, which is not only for computing compact hash codes but also for designing a lightweight deep model. Specifically, we present an MLP-based model, where the video tensor is split into several groups and multiple axial contexts are explored to separately refine them in parallel. The axial contexts are referred to as the dynamics aggregated from different axial scales, including long/middle/short-range dependencies. The group operation significantly reduces the computational cost of the MLP backbone. Moreover, to achieve compact video hash codes, three structural losses are utilized. As demonstrated by the experiment, the three structures are highly complementary for approximating the real data structure. We conduct extensive experiments on three benchmark datasets for the unsupervised video hashing task and show the superior trade-off between performance and computational cost of our EUVH to the state of the arts.
Jingru Duan, Yanbin Hao, Bin Zhu 0006, Lechao Cheng, Peng Yuan Zhou, Xiang Wang 0010
IEEE Trans. Multim.5
2024 Dream360: Diverse and Immersive Outdoor Virtual Scene Creation via Transformer-Based 360° Image Outpainting
abstract
360° images, with a field-of-view (FoV) of $180^{\circ}\times 360^{\circ}$, provide immersive and realistic environments for emerging virtual reality (VR) applications, such as virtual tourism, where users desire to create diverse panoramic scenes from a narrow FoV photo they take from a viewpoint via portable devices. It thus brings us to a technical challenge: 'How to allow the users to freely create diverse and immersive virtual scenes from a narrow FoV image with a specified viewport?' To this end, we propose a transformer-based 360° image outpainting framework called Dream360, which can generate diverse, high-fidelity, and high-resolution panoramas from user-selected viewports, considering the spherical properties of 360° images. Compared with existing methods, e.g., [3], which primarily focus on inputs with rectangular masks and central locations while overlooking the spherical property of 360° images, our Dream360 offers higher outpainting flexibility and fidelity based on the spherical representation. Dream360 comprises two key learning stages: (I) codebook-based panorama outpainting via Spherical-VQGAN (S-VQGAN), and (II) frequency-aware refinement with a novel frequency-aware consistency loss. Specifically, S-VQGAN learns a sphere-specific codebook from spherical harmonic (SH) values, providing a better representation of spherical data distribution for scene modeling. The frequency-aware refinement matches the resolution and further improves the semantic consistency and visual fidelity of the generated results. Our Dream360 achieves significantly lower Frechet Inception Distance (FID) scores and better visual fidelity than existing methods. We also conducted a user study involving 15 participants to interactively evaluate the quality of the generated results in VR, demonstrating the flexibility and superiority of our Dream360 framework.
Hao Ai, Zidong Cao, Haonan Lu, Chen Chen 0015, Jian Ma 0010, Peng Yuan Zhou, Tae-Kyun Kim 0001, Pan Hui 0001, Lin Wang 0025
IEEE Trans. Vis. Comput. Graph.6
2023 Training ChatGPT-like Models with In-network Computation
abstract
ChatGPT shows the enormous potential of large language models (LLMs). These models can easily reach the size of billions of parameters and create training difficulties for the majority. We propose a paradigm to train LLMs using distributed in-network computation on routers. Our preliminary result shows that our design allows LLMs to be trained at a reasonable learning rate without demanding extensive GPU resources.
Shuhao Fu, Yong Liao 0003, Peng Yuan Zhou
APNet3
2023 Hierarchical Privacy-Preserved Knowledge Graph
abstract
The knowledge graphs have found widespread use in numerous areas. However, as its applications expand, privacy concerns have been raised due to its ability to reveal links between entities in the real world. This brief work introduces an innovative approach to address the concern by considering the hierarchical structure of the knowledge graph to safeguard the privacy of information presented in knowledge graphs.
Xiaolu Chen, Peng Yuan Zhou, Yong Liao 0003
ICDCS3
2023 Cross-Domain Data Extraction and Knowledge Graph Construction for Dispute Analysis
abstract
This study aims to establish a comprehensive knowledge graph that spans domains and networks, with a specific focus on legal cases and their applications. The proposed methodology enables efficient collection and storage of large volumes of structured, semi-structured, and unstructured data related to cases from various sources including organizations, the government, and the internet. To analyze the relationships between roles in cases, a multimodal model is proposed to process and collect data for domain-specific knowledge graphs. Furthermore, to support social governance and public safety, a knowledge-driven intelligent recommendation algorithm is proposed in the form of question-answering, providing multiple strategies such as causal analysis, similar case matching and pre-disaster response. This work contributes to the field of artificial intelligence and natural language processing, with potential applications in legal and governmental domains, as well as in disaster response and prevention.
Qinglang Guo, Xiaolu Chen, Peng Yuan Zhou, Yong Liao 0003
ICDCS3
2023 Mercury: Fast and Optimal Device Placement for Large Deep Learning Models
abstract
The rapidly expanding neural network models are becoming increasingly challenging to run on a single device. Hence, model parallelism over multiple devices is critical to guaranteeing the efficiency of training large models. Recent proposals either have long processing time or poor performance. Therefore, we propose Mercury, a fast framework for optimizing device placement for large models. Mercury employs a simple but efficient model parallelization strategy in the baseline measurement, and generates placement policies through a series of scheduling algorithms. We conduct experiments to deploy and evaluate Mercury on numerous large models. The results show that Mercury not only reduces the placement policy generation time by 26.4% but also improves the model throughput by 218.5% compared to the most advanced methods.
Hengwei Xu, Peng Yuan Zhou, Haiyong Xie 0001, Yong Liao 0003
ICPP2
2023 Demo: Near Real-time ChatGPT-AR
abstract
Augmented reality (AR) applications based on conventional approaches lack the adaptability to cater to different scene requirements and address users' personalized demands effectively. This demon presents ChatGPT-AR, a ChatGPT-powered near real-time voice-to-AR mobile application system, that can create different 3D-augmented spaces using voice commands. Further, ChatGPT-AR enables near real-time contextual editions in AR fashion.
Yuchen Ding, Peng Yuan Zhou
MobiSys2
2023 FedACK: Federated Adversarial Contrastive Knowledge Distillation for Cross-Lingual and Cross-Model Social Bot Detection
abstract
Social bot detection is of paramount importance to the resilience and security of online social platforms. The state-of-the-art detection models are siloed and have largely overlooked a variety of data characteristics from multiple cross-lingual platforms. Meanwhile, the heterogeneity of data distribution and model architecture make it intricate to devise an efficient cross-platform and cross-model detection framework. In this paper, we propose FedACK, a new federated adversarial contrastive knowledge distillation framework for social bot detection. We devise a GAN-based federated knowledge distillation mechanism for efficiently transferring knowledge of data distribution among clients. In particular, a global generator is used to extract the knowledge of global data distribution and distill it into each client’s local model. We leverage local discriminator to enable customized model design and use local generator for data enhancement with hard-to-decide samples. Local training is conducted as multi-stage adversarial and contrastive learning to enable consistent feature spaces among clients and to constrain the optimization direction of local models, reducing the divergences between local and global models. Experiments demonstrate that FedACK outperforms the state-of-the-art approaches in terms of accuracy, communication efficiency, and feature space consistency.
Yingguang Yang, Renyu Yang, Hao Peng 0001, Tong Li 0013, Yong Liao 0003, Peng Yuan Zhou
WWW7
2023 Toward Optimal Real-Time Volumetric Video Streaming: A Rolling Optimization and Deep Reinforcement Learning Based Approach
abstract
Volumetric video provides users with a good viewing experience of six degrees of freedom (DoF) and has wide applications in many fields such as teleconferencing and online games. However, the huge data volume and strict latency requirements of point cloud video, the most popular representative of volumetric video, pose a challenge to its transmission. Existing point cloud video transmission algorithms usually segment a long video by every one or several group of frames, predict network bandwidth and field of view (FoV) information, then perform adaptive transmission by solving the quality of experience (QoE) optimization problem. However, such segmentation neglects the impact of current optimization decisions on the subsequent video streaming process, as well as the accumulated prediction error across a long interval, severely degrading user’s QoE. Moreover, the complex constrained optimization problem makes the solution time too long to meet the real-time video streaming requirements. To this end, in this paper, we propose a rolling prediction-optimization-transmission (POT) framework, which makes predictions of network bandwidth and FoV in each short rolling window to reduce prediction error. And our framework takes into account the upper bounded QoE contribution of the subsequent point cloud video to improve the system performance. In addition, we design a deep reinforcement learning based real-time solver to make decisions for the fixed structure optimization problem in each roll, allowing our system to run in real-time. We have performed simulations and experiments, and the results show that our solution outperforms existing methods.
Jie Li 0015, Zhi Liu 0002, Peng Yuan Zhou, Xianfu Chen, Qiyue Li 0001, Richang Hong
IEEE Trans. Circuits Syst. Video Technol.4
2022 Spatial-Temporal Attention Network for Crime Prediction with Adaptive Graph Learning
Mingjie Sun, Peng Yuan Zhou, Yong Liao 0003, Haiyong Xie 0001
ICANN (2)2
2022 Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure Preservation
abstract
Unsupervised video hashing typically aims to learn a compact binary vector to represent complex video content without using manual annotations. Existing unsupervised hashing methods generally suffer from incomplete exploration of various perspective dependencies (e.g., long-range and short-range) and data structures that exist in visual contents, resulting in less discriminative hash codes. In this paper, we propose aMulti-granularity Contextualized and Multi-Structure preserved Hashing (MCMSH) method, exploring multiple axial contexts for discriminative video representation generation and various structural information for unsupervised learning simultaneously. Specifically, we delicately design three self-gating modules to separately model three granularities of dependencies (i.e., long/middle/short-range dependencies) and densely integrate them into MLP-Mixer for feature contextualization, leading to a novel model MC-MLP. To facilitate unsupervised learning, we investigate three kinds of data structures, including clusters, local neighborhood similarity structure, and inter/intra-class variations, and design a multi-objective task to train MC-MLP. These data structures show high complementarities in hash code learning. We conduct extensive experiments using three video retrieval benchmark datasets, demonstrating that our MCMSH not only boosts the performance of the backbone MLP-Mixer significantly but also outperforms the competing methods notably. Code is available at: https://github.com/haoyanbin918/MCMSH.
Yanbin Hao, Jingru Duan, Hao Zhang 0047, Bin Zhu 0006, Peng Yuan Zhou, Xiangnan He 0001
ACM Multimedia5
2022 EdgeXAR: A 6-DoF Camera Multi-target Interaction Framework for MAR with User-friendly Latency Compensation
abstract
The computational capabilities of recent mobile devices enable the processing of natural features for Augmented Reality (AR), but the scalability is still limited by the devices' computation power and available resources. In this paper, we propose EdgeXAR, a mobile AR framework that utilizes the advantages of edge computing through task offloading to support flexible camera-based AR interaction. We propose a hybrid tracking system for mobile devices that provides lightweight tracking with 6 Degrees of Freedom and hides the offloading latency from users' perception. A practical, reliable and unreliable communication mechanism is used to achieve fast response and consistency of crucial information. We also propose a multi-object image retrieval pipeline that executes fast and accurate image recognition tasks on the cloud and edge servers. Extensive experiments are carried out to evaluate the performance of EdgeXAR by building mobile AR apps upon it. Regarding the Quality of Experience (QoE), the mobile AR apps powered by EdgeXAR framework run on average at the speed of 30 frames per second with precise tracking of only 1-2 pixel errors and accurate image recognition of at least 97% accuracy. As compared to Vuforia, one of the leading commercial AR frameworks, EdgeXAR transmits 87% less data while providing a stable 30FPS performance and reducing the offloading latency by 50 to 70% depending on the transmission medium. Our work facilitates the large-scale deployment of AR as the next generation of ubiquitous interfaces.
Sikun Lin, Farshid Hassani Bijarbooneh, Hao Fei Cheng, Tristan Braud, Peng Yuan Zhou, Lik-Hang Lee, Pan Hui 0001
Proc. ACM Hum. Comput. Interact.6
2022 AICP: Augmented Informative Cooperative Perception
abstract
Connected vehicles, whether equipped with advanced driver-assistance systems or fully autonomous, require human driver supervision and are currently constrained to visual information in their line-of-sight. A cooperative perception system among vehicles increases their situational awareness by extending their perception range. Existing solutions focus on improving perspective transformation and fast information collection. However, such solutions fail to filter out large amounts of less relevant data and thus impose significant network and computation load. Moreover, presenting all this less relevant data can overwhelm the driver and thus actually hinder them. To address such issues, we present Augmented Informative Cooperative Perception (AICP), the first fast-filtering system which optimizes the informativeness of shared data at vehicles to improve the fused presentation. To this end, an informativeness maximization problem is presented for vehicles to select a subset of data to display to their drivers. Specifically, we propose (i) a dedicated system design with custom data structure and lightweight routing protocol for convenient data encapsulation, fast interpretation and transmission, and (ii) a comprehensive problem formulation and efficient fitness-based sorting algorithm to select the most valuable data to display at the application layer. We implement a proof-of-concept prototype of AICP with a bandwidth-hungry, latency-constrained real-life augmented reality application. The prototype adds only 12.6 milliseconds of latency to a current informativeness-unaware system. Next, we test the networking performance of AICP at scale and show that AICP effectively filters out less relevant packets and decreases the channel busy time.
Peng Yuan Zhou, Pranvera Kortoçi, Yui-Pan Yau, Benjamin Finley, Xiujun Wang, Tristan Braud, Lik-Hang Lee, Sasu Tarkoma, Jussi Kangasharju, Pan Hui 0001
IEEE Trans. Intell. Transp. Syst.1
2021 DRLE: Decentralized Reinforcement Learning at the Edge for Traffic Light Control in the IoV
abstract
The Internet of Vehicles (IoV) enables real-time data exchange among vehicles and roadside units and thus provides a promising solution to alleviate traffic jams in the urban area. Meanwhile, better traffic management via efficient traffic light control can benefit the IoV as well by enabling a better communication environment and decreasing the network load. As such, IoV and efficient traffic light control can formulate a virtuous cycle. Edge computing, an emerging technology to provide low-latency computation capabilities at the edge of the network, can further improve the performance of this cycle. However, while the collected information is valuable, an efficient solution for better utilization and faster feedback has yet to be developed for edge-empowered IoV. To this end, we propose a Decentralized Reinforcement Learning at the Edge for traffic light control in the IoV (DRLE). DRLE exploits the ubiquity of the IoV to accelerate traffic data collection and interpretation towards better traffic light control and congestion alleviation. Operating within the coverage of the edge servers, DRLE aggregates data from neighboring edge servers for city-scale traffic light control. DRLE decomposes the highly complex problem of large area control into a decentralized multi-agent problem. We prove its global optima with concrete mathematical reasoning and demonstrate its superiority over several state-of-the-art algorithms via extensive evaluations.
Peng Yuan Zhou, Xianfu Chen, Zhi Liu 0002, Tristan Braud, Pan Hui 0001, Jussi Kangasharju
IEEE Trans. Intell. Transp. Syst.1
2020 Bricklayer: Resource Composition on the Spot Market
abstract
AWS offers discounted transient virtual instances as a way to sell unused resources in their data-centers, and users can enjoy up to 90% discount as compared to the regular on-demand pricing. Despite the economic incentives to purchase these transient instances, they do not come with regular availability SLAs, meaning that they can be evicted at any moment. Hence, the user is responsible for managing the instance availability to meet the application requirements. In this paper, we present Bricklayer, a software tool that assists users to better use transient resources in the cloud, reducing costs for the same amount of resources, and increasing the overall instance availability. Bricklayer searches for possible combinations of smaller and cheaper instances to compose the requested amount of resources while deploying them into different spot markets to reduce the risk of eviction. We implemented and evaluated Bricklayer using 3 months of historical data from AWS and found out that it can reduce up 54% of the regular spot price and up to 95% compared to the standard on-demand pricing.
Walter Wong, Lorenzo Corneo, Aleksandr Zavodovski, Peng Yuan Zhou, Nitinder Mohan, Jussi Kangasharju
ICC4
2020 Multipath Computation Offloading for Mobile Augmented Reality
abstract
Mobile Augmented Reality (MAR) applications employ computationally demanding vision algorithms on resource-limited devices. In parallel, communication networks are becoming more ubiquitous. Offloading to distant servers can thus overcome the device limitations at the cost of network delays. Multipath networking has been proposed to overcome network limitations but it is not easily adaptable to edge computing due to the server proximity and networking differences. In this article, we extend the current mobile edge offloading models and present a model for multi-server device-to-device, edge, and cloud offloading. We then introduce a new task allocation algorithm exploiting this model for MAR offloading. Finally, we evaluate the allocation algorithm against naive multipath scheduling and single path models through both a real-life experiment and extensive simulations. In case of sub-optimal network conditions, our model allows reducing the latency compared to single-path offloading, and significantly decreases packet loss compared to random task allocation. We also display the impact of the variation of WiFi parameters on task completion. We finally demonstrate the robustness of our system in case of network instability. With only 70% WiFi availability, our system keeps the excess latency below 9 ms. We finally evaluate the capabilities of the upcoming 5G and 802.11ax.
Tristan Braud, Peng Yuan Zhou, Jussi Kangasharju, Pan Hui 0001
PerCom2
2019 DeCloud: Truthful Decentralized Double Auction for Edge Clouds
abstract
The sharing economy has made great inroads with services like Uber or Airbnb enabling people to share their unused resources with those needing them. The computing world, however, despite its abundance of excess computational resources has remained largely unaffected by this trend, save for few examples like SETI@home. We present DeCloud, a decentralized market framework bringing the sharing economy to on-demand computing where the offering of pay-as-you-go services will not be limited to large companies, but ad hoc clouds can be spontaneously formed on the edge of the network. We design incentive compatible double auction mechanism targeted specifically for distributed ledger trust model instead of relying on third-party auctioneer. DeCloud incorporates innovative matching heuristic capable of coping with the level of heterogeneity inherent for large-scale open systems. Evaluating DeCloud on Google cluster-usage data, we demonstrate that the system has a near-optimal performance from an economic point of view, additionally enhanced by the flexibility of matching.
Aleksandr Zavodovski, Suzan Bayhan, Nitinder Mohan, Peng Yuan Zhou, Walter Wong, Jussi Kangasharju
ICDCS4