Haihan Duan

dblp:253/1007 · DBLP profile ↗
← Back
32ranked-venue papers
8as first author
29since 2021 · last 2026
0000-0001-6438-3790ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 13 since 2021Computer networks · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators
abstract
Quadrupedal robots with manipulators offer strong mobility and adaptability for grasping in unstructured, dynamic environments through coordinated whole-body control. However, existing research has predominantly focused on static-object grasping, neglecting the challenges posed by dynamic targets and thus limiting applicability in dynamic scenarios such as logistics sorting and human–robot collaboration. To address this, we introduce DQ-Bench, a new benchmark that systematically evaluates dynamic grasping across varying object motions, velocities, heights, object types, and terrain complexities, along with comprehensive evaluation metrics. Building upon this benchmark, we propose DQ-Net, a compact teacher–student framework designed to infer grasp configurations from limited perceptual cues. During training, the teacher network leverages privileged information to holistically model both the static geometric properties and dynamic motion characteristics of the target, and integrates a grasp fusion module to deliver robust guidance for motion planning. Concurrently, we design a lightweight student network that performs dual-viewpoint temporal modeling using only the target mask, depth map, and proprioceptive state, enabling closed-loop action outputs without reliance on privileged data. Extensive experiments on DQ-Bench demonstrate that DQ-Net achieves robust dynamic objects grasping across multiple task settings, substantially outperforming baseline methods in both success rate and responsiveness. We will release our codebase and benchmark publicly.
Qiwei Liang, Boyang Cai, Rongyi He, Tao Teng, Haihan Duan, Changxin Huang, Runhao Zeng
AAAI6
2026 PolyGnosis: An AI Framework for Forecasting Public Opinion from Polymarket Dynamics
Daren Wang, Haihan Duan
INFOCOM3
2026 DRLLMS: Network-Adaptive Reasoning Control for Interactive LLM Streaming
Tao Lyu 0005, Cong Zhang 0002, Haihan Duan, Xiaoyi Fan 0001, Xiping Hu, Laizhong Cui
NOSSDAV3
2026 PivotSketch: Control-Ready Semantic Ranking for Adaptive Video Streaming
Sirui Zhang, Cong Zhang 0002, Xiaoyi Fan 0001, Xiping Hu, Haihan Duan
NOSSDAV6
2026 FedSM: Semantic-Guided Feature Mixup for Bias Reduction in Federated Learning With Long-Tail Data
abstract
Federated Learning (FL) has emerged as a promising paradigm for decentralized machine learning, where a central server coordinates distributed clients to collaboratively train a global model without direct access to raw data. Despite its advantages, heterogeneous and long-tail data distributions across clients remain a major bottleneck, particularly in IoT scenarios with diverse devices and sensing modalities. To address these challenges, we propose FedSM, a novel framework that integrates multimodal semantic knowledge with balanced pseudo features to enhance global model optimization. Unlike conventional approaches that rely on single-modal information, FedSM leverages CLIP’s cross-modal representations and open-vocabulary priors to guide semantic-aware data augmentation. A probabilistic selection mechanism further refines local features by mixing them with global prototypes, ensuring pseudo features are semantically reliable and reducing bias caused by skewed client distributions. Almost all computations are performed locally at the client side, thereby alleviating server overhead and improving scalability in resource-constrained IoT environments. Extensive experiments on long-tail benchmarks including CIFAR-10-LT, CIFAR-100-LT, and ImageNet-LT demonstrate the superiority of FedSM over state-of-the-art baselines, highlighting its potential for robust communication-efficient FL in IoT networks.
Jingrui Zhang, Shujie Li 0001, Feng Liang 0004, Haihan Duan, Yanjie Dong 0003, Victor C. M. Leung, Xiping Hu
IEEE Internet Things J.5
2026 User-Generated Content and Editors in Games: A Comprehensive Survey
abstract
User-Generated Content (UGC) refers to any form of content, such as posts and images, created by users rather than by professionals. In recent years, UGC has become an essential part of the evolving video game industry, influencing both game culture and community dynamics. The ability for users to actively contribute to the games they engage with has shifted the landscape of gaming from a one-directional entertainment experience into a collaborative, user-driven ecosystem. Therefore, this growing trend highlights the urgent need for summarizing the current UGC development in game industry. Our conference paper has systematically classified the existing UGC in games and the UGC editors separately into four types. However, the previous survey lacks the depth and precision necessary to capture the wide-ranging and increasingly complex nature of UGC. To this end, as an extension of previous work, this article presents a refined and expanded classification of UGC and UGC editors within video games, offering a more robust and comprehensive framework with representative cases that better reflects the diversity and nuances of contemporary user-generated contributions. Moreover, we provide our insights on the future of UGC, involving application of artificial intelligence, potential ethical considerations, relationship between games, users and communities, game culture, and game genre, and user creative tendencies.
Haihan Duan, Yuyue Liu, Wei Cai 0002
ACM Trans. Multim. Comput. Commun. Appl.1
2026 UniqueNFT: Uniqueness Protection of Digital Assets in Decentralized Web
abstract
With the rapid evolution of the Decentralized Web (DWeb), decentralized technologies have paved new avenues for Web3 applications and the authentication of digital assets. Among them, Non-Fungible Tokens (NFTs) have gained significant popularity due to their immutability and uniqueness, reshaping the landscape of artistic creation, marketing, and intellectual property protection. However, current blockchain-based NFT implementations still face core challenges within decentralized architecture: how to maintain decentralization while ensuring the visual uniqueness of digital assets and reducing storage costs. The rampant issue of duplication undermines the scarcity of digital art and erodes market confidence in copyright authenticity. Moreover, high gas fees and energy consumption further hinder the widespread adoption of NFTs, while reliance on external storage solutions like InterPlanetary File System (IPFS) introduces risks of data instability and loss. To address these challenges, this article presents the UniqueNFT framework, a novel architecture that deeply integrates blockchain oracles with decentralized storage verification mechanisms. The framework achieves three key technological breakthroughs: Using image inversion and generation techniques based on Encoder for Editing (E4E) and StyleGAN3, it extracts compact and expressive semantic features from NFT images, enabling efficient data compression and significantly reducing on-chain storage volume; The Crypto-Mask algorithm, by utilizing the hash value of blockchain user information (user-controlled SHA-256 digest of Ethereum address, user nickname, and registration time), ensures the visual uniqueness of NFTs; A smart contract extension compatible with the ERC721 standard, demonstrating UniqueNFT’s seamless integration within the blockchain ecosystem. By leveraging the technologies of the Decentralized Web, our framework represents an important step forward in enhancing the security and uniqueness of digital assets. It not only innovatively resolves the issues of NFT duplication and homogenization but also injects new vitality and long-term momentum into the creation of a trusted, sustainable blockchain-based digital asset ecosystem.
Kun Yang 0010, Haihan Duan, Runhao Zeng, Xiping Hu
ACM Trans. Web2
2025 Blockchain-Enabled Market Clearing Mechanism for Peer-to-Peer Energy Storage Sharing
Haihan Duan, Hengming Dai, Xiaoyi Fan 0001, Cong Zhang 0002, Xiping Hu
IEEE Big Data2
2025 Sparse Manifold Retrieval Network for ICESat-2 Photon Point Cloud Denoising
abstract
The photon point clouds acquired by ICESat-2/ATLAS offer unprecedented potential for Earth observation but are heavily contaminated by noise photons, posing a significant challenge for downstream applications. Traditional denoising methods, which often rely on local density statistics, struggle with complex terrains and varying signal-to-noise ratios. While deep learning presents a promising alternative, existing approaches often inefficiently process the inherently sparse data via 2D projections or non-optimized 3D networks. To address these limitations, this paper introduces a novel deep learning framework for ICESat-2 photon denoising, termed Sparse Manifold Retrieval Network (SMRNet). We propose a Manifold-Aware Convolution (MAC) module to capture the continuous manifold structures of signal photons through multi-scale dilated sparse convolutions, and a Cross-Scale Pyramid Enhancement (CSPE) module to effectively refine multi-level features extracted from the encoder. Evaluated on a manually annotated dataset covering southeastern coastal regions of China, SMRNet demonstrates superior performance over traditional denoising method and data-driven baselines across multiple metrics. The results underscore the effectiveness of SMRNet in enhancing denoising accuracy, particularly in challenging environments with sparse signals and rugged topography.
Hengming Dai, Haihan Duan, Cong Zhang 0002, Xiaoyi Fan 0001, Zhifang Zhao
CloudCom2
2025 Blockchain-Enabled Pricing Mechanism in Energy Markets: Survey and Vision
abstract
The growth of distributed energy resources and local energy markets heightens the need for price formation that is transparent, privacy preserving, and compatible with network constraints. Blockchain provides a trust-minimized substrate for auditable clearing and settlement through consensus, tamperevident ledgers, and smart contracts. This survey organizes blockchain-enabled pricing into three families, namely auction-based, game-theoretic, and optimization-based, and links them to enabling techniques such as metering oracles, secure multiparty computation, zero-knowledge proofs, and verifiable optimality certificates. Applications span wholesale electricity, carbon and green certificates, distributed energy trading, ancillary services, and electric vehicles. Evidence indicates gains in auditability, privacy, network awareness, and automated settlement, alongside challenges in scalability, data protection, grid integration, and regulation. The survey distills design patterns and research directions toward verifiable, interoperable, and governable pricing modules that complement system-operator markets.
Xiaoyi Fan 0001, Cong Zhang 0002, Hengming Dai, Haihan Duan
CloudCom5
2025 Airdrop Hunter Detection via PageRank-Augmented Multimodal Graph Neural Networks
abstract
Airdrops are a widely used mechanism in Web3 ecosystems to incentivize early users by distributing governance tokens. However, these mechanisms are increasingly targeted by airdrop hunters—malicious actors who exploit token distribution systems through address farming, automated scripts, and behavioral camouflage. While prior work such as ARTEMIS leverages multimodal features and local transaction patterns to detect such behavior, it lacks a global understanding of wallet influence in the transaction graph. In this paper, we propose an enhanced detection framework that augments the ARTEMIS by incorporating PageRank-based global centrality as an additional structural feature. This allows the model to better distinguish superficially active wallets from those with broader influence in the network. We evaluate our method on real-world Non-Fungible Token (NFT) data from the Blur marketplace and achieve state-of-the-art performance. Furthermore, a feature substitution experiment reveals that simple degree-based features alone can achieve near-perfect performance, even outperforming PageRank, suggesting that the labels are strongly coupled with topological properties. These findings highlight both the effectiveness of structural augmentation and the potential risks of shortcut learning in graph-based detection systems.
Jiajie Shi, Yuyang Qin, Hengming Dai, Xiaoyi Fan 0001, Haihan Duan
CloudCom5
2025 A Lyapunov Optimization Framework for Green Satellite Communications
abstract
Satellite communication plays a crucial role in future networks, but traditional systems face significant challenges in interference management and energy efficiency. To address these issues, this paper proposes a green satellite communication framework that combines Lyapunov optimization with a one-dimensional golden-section search method. The framework builds a two-timescale frame-slot model that jointly captures fast-varying channel fading and slow-varying renewable energy dynamics. By introducing a drift-plus-penalty optimization method, the system ensures queue stability and minimizes long-term grid energy expenditure, while employing the golden-section search to optimize beamforming parameters with reduced computational complexity. Simulation results show that the proposed framework effectively balances energy efficiency and communication performance, demonstrating good scalability for large-scale satellite networks.
Qilu Wu, Xiaoyi Fan 0001, Haihan Duan
CloudCom6
2025 ASimp: Automatic High-Poly 3D Mesh Simplification for Preprocessing Based on QoE
abstract
Mesh simplification of 3D models can accelerate rendering, reduce storage space, and improve performance. However, for high-poly 3D models, there are ongoing concerns about potentially compromising the Quality of Experience (QoE), the need to set simplification ratios or parameters, and the time-consuming nature of the simplification process. To address these issues, we proposed a new mesh simplification for the preprocessing step. Based on the Quadratic Error Metric (QEM) simplification algorithm, we conducted human-centered 3D model comparison experiments to determine the optimal simplification ratio for high-poly 3D models in full body shots. From experimental data, we proposed and implemented ASimp, an automatic 3D mesh simplification scheme. In evaluation experiments, ASimp demonstrated rapid preprocessing speeds while ensuring QoE and the effectiveness of its simplification products. We hope that ASimp will contribute to the optimization of 3D models and find applications in fields such as cultural heritage, archaeology, visual effects, video games, medicine, metaverse, and beyond.
Lehao Lin, Hong Kang, Yuqi Shi, Haihan Duan, Abdulmotaleb El Saddik, Wei Cai 0002
ICME4
2025 Navigating the Deployment Dilemma and Innovation Paradox: Open-Source versus Closed-source Models
abstract
Recent advances in Artificial Intelligence (AI) have introduced a popular paradigm in Machine Learning (ML) model development: pre-training and domain adaptation. As both closed-source developers and open-source community lead in pre-training foundation models, domain deployers face the dilemma about whether to use closed-source models via API access or to host open-source models on proprietary hardware. Using closed-source models incurs recurring costs, while hosting open-source models requires substantial hardware investments and may lead to potentially lagging advancements. This paper presents a game-theoretical model to examine the economic incentives behind the deployment choice and the impact of open-source engagement strategies on technological innovation. We find that deployers consistently opt for closed-source APIs when the open-source community engages reactively by maintaining a fixed performance ratio relative to closed-source advancements. However, open-source models can become preferable when a proactive open-source community produces high-performance models independently. Furthermore, we identify conditions under which the engagement and competitiveness of the open-source community can either foster or inhibit technological progress. These insights offer valuable implications for market regulation and the future of technology innovation.
Yanxuan Wu, Haihan Duan, Xitong Li, Xiping Hu
WWW2
2025 LR-ASD: Lightweight and Robust Network for Active Speaker Detection
Junhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao, Yanbing Yang 0001, Liangyin Chen, Yanru Chen 0001
Int. J. Comput. Vis.2
2025 DeRelayL: Sustainable Decentralized Relay Learning
abstract
In the era of Big Data, large-scale machine learning models have revolutionized various fields, driving significant advancements. However, large-scale model training demands high financial and computational resources, which are only affordable by a few technological giants and well-funded institutions. In this case, common users like mobile users, the real creators of valuable data, are often excluded from fully benefiting due to the barriers, while the current methods for accessing largescale models either limit user ownership or lack sustainability. This growing gap highlights the urgent need for a collaborative model training approach, allowing common users to train and share models. However, existing collaborative model training paradigms, especially federated learning (FL), primarily focus on data privacy and group-based model aggregation. To this end, this paper intends to address this issue by proposing a novel training paradigm named decentralized relay learning (DeRelayL), a sustainable learning system where permissionless participants can contribute to model training in a relay-like manner and share the model. In detail, this paper presents the architecture and workflow of DeRelayL, designs incentive mechanisms to ensure sustainability, and conducts theoretical analysis and numerical simulations to demonstrate its effectiveness
Haihan Duan, Yuyang Qin, Runhao Zeng, Wei Cai 0002, Victor C. M. Leung, Xiping Hu
IEEE Trans. Mob. Comput.1
2024 Incentive Mechanism Design Toward a Win-Win Situation for Generative Art Trainers and Artists
abstract
The recent development of generative art, a typical category of artificial intelligence-generated content (AIGC), is essentially beneficial for social good, which can help amateurs to create artwork and improve experts’ efficiency. However, some artists are opposed to generative art technologies due to the copyright infringement and influence of the artists’ way of earning a living, which makes the artists protest against generative art technologies, causing a lose–lose situation. Adversarial attacks against generative model training are potential solutions to address this issue, while the lose–lose situation cannot be improved. To build a win–win situation, a feasible method is to incentivize the artists to actively contribute their artworks to generative model training without influencing their living or infringing copyright, such as data crowdsourcing, but traditional data crowdsourcing methods cannot well fit the generative art area. Therefore, this article builds a blockchain-based trading system for generative model training data collection and generated artwork circulation. Specifically, this article formulates a social welfare maximization problem based on the reverse auction and designs a corresponding incentive mechanism. The conducted theoretical analysis and numerical evaluation demonstrate the effectiveness of the proposed incentive mechanism toward a win–win situation for generative art model trainers and artists.
Haihan Duan, Abdulmotaleb El Saddik, Wei Cai 0002
IEEE Trans. Comput. Soc. Syst.1
2024 A Video Shot Occlusion Detection Algorithm Based on the Abnormal Fluctuation of Depth Information
abstract
To make the video more attractive, original video materials usually need postprocessing by video editors, especially to eliminate low-quality abnormal clips, which seriously affect the visual effect. One of the main reasons for the low-quality abnormal clips is that there are occluders that accidentally break into the shot to occlude the protagonist, resulting in the loss of the video protagonist’s information. However, it is time-consuming and laborious to manually find shot occlusion clips, so computer vision technology can be used to assist editors in completing this work. The previous solutions directly utilize neural networks to detect shot occlusion, so their performance is affected by the size and quality of the dataset. In contrast, inspired by the change of depth information in the frame caused by the occluder breaking into the shot, we propose an algorithm for video shot occlusion detection based on the fluctuation of depth information. This algorithm does not need occlusion data training and can detect shot occlusion well only by capturing the abnormal fluctuations of the frame depth information. Additionally, to overcome the defect in that the first video shot occlusion detection (VSOD) dataset released in our conference publication can only verify the sensitivity of detection methods, we expand the VSOD dataset to evaluate the comprehensive performance of detection algorithms. The plentiful experimental results show that, compared with state-of-the-art occlusion detection methods and self-designed baseline methods, our algorithm significantly improves the comprehensive performance of video shot occlusion detection. Furthermore, through verification on datasets with different data types and distributions, our shot occlusion detection algorithm can maintain an occlusion event recall of over 95%, while the false positive rate does not exceed 3%, demonstrating good generalization ability. To promote reproducible research, the code and dataset are available athttps://github.com/Junhua-Liao/VSOD.
Junhua Liao, Haihan Duan, Wanbing Zhao, Kanghui Feng, Yanbing Yang 0001, Liangyin Chen
IEEE Trans. Circuits Syst. Video Technol.2
2024 Web3 Metaverse: State-of-the-Art and Vision
abstract
The metaverse, as a rapidly evolving socio-technical phenomenon, exhibits significant potential across diverse domains by leveraging Web3 (a.k.a. Web 3.0) technologies such as blockchain, smart contracts, and non-fungible tokens (NFTs). This survey aims to provide a comprehensive overview of the Web3 metaverse from a human-centered perspective. We (i) systematically review the development of the metaverse over the past 30 years, highlighting the balanced contributions from its core components: Web3, immersive convergence, and crowd intelligence communities, (ii) define the metaverse that integrates the Web3 community as the Web3 metaverse and propose an analysis framework from the community, society, and human layers to describe the features, missions, and relationships for each community and their overlapping sections, (iii) survey the state-of-the-art of the Web3 metaverse from a human-centered perspective, namely, the identity, field, and behavior aspects, and (iv) provide supplementary technical reviews. To the best of our knowledge, this work represents the first systematic, interdisciplinary survey on the Web3 metaverse. Specifically, we commence by discussing the potential for establishing decentralized identities (DID) utilizing mechanisms such as profile picture (PFP) NFTs, domain name NFTs, and soulbound tokens (SBTs). Subsequently, we examine land, utility, and equipment NFTs within the Web3 metaverse, highlighting interoperable and full on-chain solutions for existing centralization challenges. Lastly, we spotlight current research and practices about individual, intra-group, and inter-group behaviors within the Web3 metaverse, such as Creative Commons Zero license (CC0) NFTs, decentralized education, decentralized science (DeSci), and decentralized autonomous organizations (DAO). Furthermore, we share our insights into several promising directions, encompassing three key socio-technical facets of Web3 metaverse development.
Hongzhou Chen, Haihan Duan, Maha Abdallah, Yufeng Zhu, Yonggang Wen 0001, Abdulmotaleb El Saddik, Wei Cai 0002
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Meetor: A Human-Centered Automatic Video Editing System for Meeting Recordings
abstract
Widely adopted digital cameras and smartphones have generated a large number of videos, which have brought a tremendous workload to video editors. Recently, a variety of automatic/semi-automatic video editing methods have been proposed to tackle these issues in some specific areas. However, for the production of meeting recordings, the existing studies highly depend on extra equipment in the conference venues, such as the infrared camera or special microphone, which are not practical. In this article, we design and implement Meetor, a human-centered automatic video editing system for meeting recordings. The Meetor mainly contains three parts: an audio-based video synchronization algorithm, human-centered video content flaw detection algorithms, and an automatic video editing algorithm. Two main experiments are conducted from both objective and subjective aspects to evaluate the performance of the Meetor. The experimental results on a testbed illustrate that the proposed algorithms could achieve state-of-the-art (SOTA) performance in video content flaw detection. However, the conducted user study demonstrates that Meetor could generate meeting recordings with a satisfactory quality compared with professional video editors. Moreover, we also present a practical application of the Meetor in a university campus prototype, in which the Meetor is applied in the automatic editing of lecture recordings. All in all, the proposed Meetor can be utilized in practical applications to release the workload of professional video editors.
Haihan Duan, Junhua Liao, Lehao Lin, Abdulmotaleb El Saddik, Wei Cai 0002
ACM Trans. Multim. Comput. Commun. Appl.1
2023 A Light Weight Model for Active Speaker Detection
abstract
Active speaker detection is a challenging task in audiovisual scenarios, with the aim to detect who is speaking in one or more speaker scenarios. This task has received considerable attention because it is crucial in many applications. Existing studies have attempted to improve the performance by inputting multiple candidate information and designing complex models. Although these methods have achieved excellent performance, their high memory and computational power consumption render their application to resource-limited scenarios difficult. Therefore, in this study, a lightweight active speaker detection architecture is constructed by reducing the number of input candidates, splitting 2D and 3D convolutions for audio-visual feature extraction, and applying gated recurrent units with low computational complexity for cross-modal modeling. Experimental results on the AVA-ActiveSpeaker dataset reveal that the proposed framework achieves competitive mAP performance (94.1% vs. 94.2%), while the resource costs are significantly lower than the state-of-the-art method, particularly in model parameters (1.0M vs. 22.5M, approximately 23×) and FLOPs (0.6G vs. 2.6G, approximately 4×). Additionally, the proposed framework also performs well on the Columbia dataset, thus demonstrating good robustness. The code and model weights are available at https://github.com/Junhua-Liao/Light-ASD.
Junhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao, Yanbing Yang 0001, Liangyin Chen
CVPR2
2023 MetaCast: A Self-Driven Metaverse Announcer Architecture Based on Quality of Experience Evaluation Model
abstract
Metaverse provides users with a novel experience through immersive multimedia technologies. Along with the rapid user growth, numerous events bursting in the metaverse necessitate an announcer to help catch and monitor ongoing events. However, systems on the market primarily serve for esports competitions and rely on human directors, making it challenging to provide 24-hour delivery in the metaverse persistent world. To fill the blank, we proposed a three-stage architecture for metaverse announcers, which is designed to identify events, position cameras, and blend between shots. Based on the architecture, we introduced a Metaverse Announcer User Experience (MAUE) model to identify the factors affecting the users' Quality of Experience (QoE) from a human-centered perspective. In addition, we implemented MetaCast, a practical self-driven metaverse announcer in a university campus metaverse prototype, to conduct user studies for MAUE model. The experimental results have effectively achieved satisfactory announcer settings that align with the preferences of most users, encompassing parameters such as video transition rate, repetition rate, importance threshold value, and image composition.
Zhonghao Lin, Haihan Duan, Xinyao Sun, Wei Cai 0002
ACM Multimedia2
2023 Web3DP: A Crowdsourcing Platform for 3D Models Based on Web3 Infrastructure
abstract
Recently, the concept of metaverse has been rapidly emerging, which highly expands the human living space. Specifically, 3D models are at the heart of building a vast metaverse space, so a massive number of 3D models are needed. Existing 3D model libraries and platforms have achieved great results. However, most of them are unscalable, insufficiently open, inefficient to collect, and at risk of service disruption and data corruption. Therefore, we propose and implement Web3DP, a crowdsourcing platform for 3D models based on Web3 (a.k.a. Web 3.0) infrastructure. By using the decentralized blockchain technology, Web3DP has the advantages of transparency, auditability, traceability, data tamper-proof, high file transfer efficiency, and service stability. Experiments are conducted to validate the performance of the proposed platform. It illustrates that Web3DP shows better file transmission capabilities with an acceptable transaction fee to facilitate 3D model collecting and managing for metaverse, games, cultural heritage, etc.
Lehao Lin, Haihan Duan, Wei Cai 0002
MMSys2
2022 User-Generated Content and Editors in Video Games: Survey and Vision
abstract
User-generated content (UGC) is any form of content that has been created by users rather than the developers of online platforms. The UGC has been playing a very important role in video games. For instance, Counter-Strike (CS) and Defense of the Ancients (DOTA) originated from modifications of Half-Life and Warcraft III: Reign of Chaos respectively. As a promising trend, the UGC will be increasingly developed and extended to wider scopes in virtual worlds, such as the metaverse, which is highly promising for further study in both academia and industry. However, there are few existing surveys that systematically discuss the UGC in video games. In this paper, we systematically review the representative UGC in video games and their corresponding UGC editors based on a decision tree-style classification method. Then we enumerate the propagation methods of UGC in video games. Moreover, we propose our vision of the future development of UGC in the metaverse.
Haihan Duan, Wei Cai 0002
CoG1
2022 A Light Weight Model for Video Shot Occlusion Detection
abstract
The popularity of video social platforms (TikTok, etc.) shows that video is a popular information carrier at present. However, shot occlusion frequently occurs when people are shooting videos to record information. Since the shot occlusion seriously affects the viewers’ experience, the video editors need to find and delete such segments from the video material during post-processing. However, finding the shot occlusion from the video is a time-consuming and laborious task. To reduce the workload of editors, previous researchers proposed a shot occlusion detection algorithm using deep learning technology, which has promotion space in both recognition accuracy and computational efficiency. In this paper, we propose a neural network module, named SAT module, which can effectively extract spatio-temporal information with fewer parameters. We apply SAT module to construct a novel occlusion detection model, and improve the existing occlusion detection loss function for model training. The experimental results on the public dataset show that our method achieves the state-of-the-art performance of 88.25% accuracy and FPS of 130 with the least parameters. Code and models will be available at https://github.com/Junhua-Liao/ICASSP22-OcclusionDetection.
Junhua Liao, Haihan Duan, Wanbin Zhao, Yanbing Yang 0001, Liangyin Chen
ICASSP2
2022 Crypto-Dropout: To Create Unique User-Generated Content Using Crypto Information in Metaverse
abstract
In a blockchain-driven metaverse, user-generated content (UGC) is the core power for building the metaverse, so an easy-to-use UGC editor is imperative. Specifically, using artificial intelligence (AI) to simplify the UGC creation procedure is promising, e.g., generating images from sketches using generative adversarial networks (GANs). However, the simplicity of these UGC creation methods would lead to weak distinctions between the generated UGC, since the users' created drafts may be very similar. In this paper, we propose Crypto-dropout, a specially designed dropout used in the generative neural networks, which could cause pseudo-random disturbance based on the hash value of user information to generate unique results. With a pilot study, the experimental results demonstrate that the participants have different preferences for the generated images when setting Crypto-dropout in the different layers. Accordingly, we implement a practical profile pictures (PFPs) creation prototype. The proposed Crypto-dropout can provide a novel and general insight for creating unique UGC using generative neural networks.
Haihan Duan, Wei Cai 0002
MMSP1
2022 FLAD: a human-centered video content flaw detection system for meeting recordings
abstract
Widely adopted digital cameras and smartphones have generated a large number of videos, which have brought a tremendous workload to video editors. Recently, a variety of automatic/semi-automatic video editing methods have been proposed to tackle this issue in some specific areas. However, for the production of meeting recordings, the existing studies highly depend on additional conditions of conference venues, like infrared camera or special microphone, which are not practical. Moreover, current video quality assessment works mainly focus on the quality loss after compression or encoding rather than the human-centered video content flaws. In this paper, we design and implement FLAD, a human-centered video content flaw detection system for meeting recordings, which could build a bridge between subjective sense and objective measures from a human-centered perspective. The experimental results illustrate the proposed algorithms could achieve the state-of-the-art video content flaw detection performance for meeting recordings.
Haihan Duan, Junhua Liao, Lehao Lin, Wei Cai 0002
NOSSDAV1
2022 An Energy-efficient and Privacy-aware Decomposition Framework for Edge-assisted Federated Learning
abstract
Deep Learning (DL) is an essential technology for modern intelligent sensor network and interactive multimedia applications, having problems with user data privacy when training on a central cloud. While Federated Learning (FL) motivates to preserve user privacy, it also causes new problems of lower user terminal usability and training efficiency, which caused substantial energy consumption. This article proposes a novel energy-efficient and privacy-aware decomposition framework to improve user-side FL efficiency under pre-defined privacy requirements with the assistance of Mobile Edge Computing (MEC) and Software Decomposition. It takes the propagation of each neural layer as the migrating unit and considers the tradeoff relationship between privacy and efficiency. We also propose an online scheduling algorithm to optimize the framework’s training performance. Furthermore, we summarize eight privacy-sensitive information classes on which existing privacy attacks base and design configurable privacy preservation mechanisms for each class. Simulations and experiments prove the effectiveness of our framework and algorithm in FL efficiency improvement and the effects of different privacy constraints on the overall training efficiency.
Yimin Shi 0001, Haihan Duan, Lei Yang 0024, Wei Cai 0002
ACM Trans. Sens. Networks2
2021 Metaverse for Social Good: A University Campus Prototype
abstract
In recent years, the metaverse has attracted enormous attention from around the world with the development of related technologies. The expected metaverse should be a realistic society with more direct and physical interactions, while the concepts of race, gender, and even physical disability would be weakened, which would be highly beneficial for society. However, the development of metaverse is still in its infancy, with great potential for improvement. Regarding metaverse's huge potential, industry has already come forward with advance preparation, accompanied by feverish investment, but there are few discussions about metaverse in academia to scientifically guide its development. In this paper, we highlight the representative applications for social good. Then we propose a three-layer metaverse architecture from a macro perspective, containing infrastructure, interaction, and ecosystem. Moreover, we journey toward both a historical and novel metaverse with a detailed timeline and table of specific attributes. Lastly, we illustrate our implemented blockchain-driven metaverse prototype of a university campus and discuss the prototype design and insights.
Haihan Duan, Sizheng Fan, Zhonghao Lin, Wei Cai 0002
ACM Multimedia1
2020 A Dynamic Partitioning Framework for Edge-Assisted Cloud Computing
Zhengjia Cao, Haihan Duan, Lei Yang 0024, Wei Cai 0002
ICA3PP (2)3
2020 Edge-Assisted Federated Learning: An Empirical Study from Software Decomposition Perspective
Yimin Shi 0001, Haihan Duan, Yuanfang Chi, Keke Gai, Wei Cai 0002
ICA3PP (2)2
2020 Occlusion Detection for Automatic Video Editing
abstract
Videos have become the new preference comparing with images in recent years. However, during the recording of videos, the cameras are inevitably occluded by some objects or persons that pass through the cameras, which would highly increase the workload of video editors for searching out such occlusions. In this paper, for releasing the burden of video editors, a frame-level video occlusion detection method is proposed, which is a fundamental component of automatic video editing. The proposed method enhances the extraction of spatial-temporal information based on C3D yet only using around half amount of parameters, with an occlusion correction algorithm for correcting the prediction results. In addition, a novel loss function is proposed to better extract the characterization of occlusion and improve the detection performance. For performance evaluation, this paper builds a new large scale dataset, containing 1,000 video segments from seven different real-world scenarios, which could be available at: https://junhua-liao.github.io/Occlusion-Detection/. All occlusions in video segments are annotated frame by frame with bounding-boxes so that the dataset could be utilized in both frame-level occlusion detection and precise occlusion location. The experimental results illustrate that the proposed method could achieve good performance on video occlusion detection compared with the state-of-the-art approaches. To the best of our knowledge, this is the first study which focuses on occlusion detection for automatic video editing.
Junhua Liao, Haihan Duan, Yanbing Yang 0001, Wei Cai 0002, Yanru Chen 0001, Liangyin Chen
ACM Multimedia2