EDBT 2026 Demo / reviewers in the wild / expert
Yixue Hao
dblp:163/7333
· DBLP profile ↗
56ranked-venue papers
8as first author
41since 2021 · last 2026
0000-0001-7296-2522ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 16 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revealing Procedural Reasoning Structures in Chain-of-Thought Training via Span-Level Gradient OrganizationabstractJia Liu, Jiaxin Luo, Weiwen Xu, Jonathan M. Garibaldi, Xiao-Kun Wu, Yixue Hao, Min Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jia Liu 0009, Jiaxin Luo, Weiwen Xu, Jonathan M. Garibaldi, Xiaokun Wu 0004, Yixue Hao, Min Chen 0003 |
ACL (1) | 6 |
| 2026 | ACM: Defending Against Label-Only Membership Inference Attacks via Dynamic Adversarial Confidence Modification
Zishan Li, Yongfeng Qian, Yixue Hao |
DASFAA (5) | 3 |
| 2026 | Multi-source sensing adaptation for human behavior modeling in fabric space
Haodong Yi, Xiaokun Wu 0004, Yixue Hao, Min Chen 0003 |
Inf. Sci. | 3 |
| 2026 | Multi-modal model partition strategy for end-edge collaborative inference
Dongkun Huo, Yingting Zhou, Yixue Hao, Long Hu, Yijun Mo, Min Chen 0003, Iztok Humar |
J. Parallel Distributed Comput. | 3 |
| 2026 | Joint Fine-Grained Representation Learning and Masked Relational Modeling for EEG-Based Automatic Sleep Staging in Fabric SpaceabstractSleep staging is a crucial method for the evaluation of sleep quality and the diagnosis of sleep disorders. In recent years, rapid progress has been made in sleep research through the application of fabric computing and neural networks. Flexible fabric sensors introduced by fabric computing minimize the discomfort of data collection devices on individuals, while neural network-based algorithms can automatically perform sleep staging based on the collected signals. However, there are two key challenges hinder the integration of automatic sleep staging networks with fabric computing: (1) signals in fabric-based environments exhibit strong heterogeneity due to the wide range of individuals, and (2) interactions between individuals and the fabric space introduce behavioral dynamics to the system. In this paper, we propose a masked autoencoder-based sleep staging neural networks (MAESleepNet), designed to integrate automatic sleep staging algorithm with fabric space. Specifically, MAESleepNet addresses the challenge of signal heterogeneity by learning fine-grained representations from local signals. Furthermore, MAESleepNet tackle the challenge of behavioral dynamics through stochastic masking and reconstruction pre-training. Experiments were conducted on three public datasets: (1) Sleep-EDF-20, (2) Sleep-EDF-78 and (3) SHHS. MAESleepNet achieves overall accuracies of 88.9%, 85.5%, and 87.3%, respectively, outperforming other state-of-the-art models. Furthermore, feature visualization and reconstruction visualization experiments were also conducted. The results demonstrates that MAESleepNet is an effective solution to the aforementioned challenges, paving the way for seamless integration into the fabric space. Lejun Ai, Yixue Hao, Xiaoli Li 0002, Min Chen 0003, Xiaokun Wu 0004 |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Rethinking Point Cloud Representation Learning for Freeing Transformer to Perceive LocalabstractTransformers are widely utilized in the point cloud domain. However, existing methods tend to overburden Transformer with the dual task of local geometric perception and global feature extraction, limiting its ability to capture highlevel semantic knowledge. To address this issue, we present Representation Decoder (R-Decoder), a novel representation extraction module compatible with various point cloud Transformer methods, enabling the Transformer to focus on its excellent local perception. The R-Decoder iteratively extracts multiple global features from tokens generated by Transformer, refining them to construct an overall representation of point cloud. To ensure full adaptation of the R-Decoder to the knowledge of pre-trained Transformers, we design a cross-modal representation alignment task that leverages multimodal knowledge to specifically pre-train the R-Decoder. As a post-processing module, the R-Decoder seamlessly integrates with Transformers, while decoupling local perception and global representation. This design allows the Transformer to focus on the semantic encoding role for point tokens. Extensive experiments show that our RDecoder significantly boosts the capabilities of 3D representation learning in various point cloud Transformer methods. Notably, it achieves impressive classification accuracies of 95.1% on the ScanObjectNN dataset and 95.3% on the ModelNet40 dataset. Moreover, our method obtains new SOTA on all benchmarks of few-shot and zero-shot classification, while enhancing the multimodal task capabilities of pre-trained Transformers. Code and weights are available athttps://github.com/TangYuan96/RDecoder. Yunlong Yu 0002, Xianzhi Li 0001, Rui Wang 0077, Jinfeng Xu 0002, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
IEEE Trans. Multim. | 7 |
| 2026 | DTSNet: Dynamic Transformer Slimming for Efficient Vision RecognitionabstractTransformer-based models have recently adopted increasingly complex structure (e.g., deeper or wider stacked network) to promote the representation learning capabilities of vision recognition. However, progressively deeper or wider stacked network cause the expensive computation cost, which hinders their effective deployment in resource-constrained edge clouds or end devices. In this paper, we propose DTSNet, a dynamic transformer slimming model, which scales vision transformers (ViTs) down across layers from both of the model depth and input width. This is the first time to explore the joint reduction of input tokens and model parameters for ViTs under maintaining performance. Specifically, DTSNet adopts a diversity-enhanced weight sharing module to reduce network parameters, where the weight knowledge of multiple adjacent blocks is effectively integrated into one block. Furthermore, DTSNet designs a unified and massively scalable token pruning mechanism that dynamically discarding less important tokens with a model-driven manner, by introducing a series of discriminant parameters, which is a simple change to the common architecture of vision transformers. Extensive experiments are conducted to verify that DTSNet is able to yield high efficacy in compressing parameter space and accelerating model inference. DTSNet-T/-S/-B on ImageNet achieves 3.0M/11.1M/42.9M parameters and 0.8/2.9/13.7 GFLOPs, where number of parameters are reduced by 48%$\sim$51% and inference speed are improved by 1.3$\times \sim 1.5\times$. Experiments results on semantic segmentation and object detection dataset further demonstrate the potential of DTSNet on complex dense prediction tasks. Code will be available upon publication. Wenjing Xiao, Xianzhi Li 0001, Long Hu, Yixue Hao, Min Chen 0003 |
IEEE Trans. Multim. | 4 |
| 2026 | Generative Aspect-Based Sentiment Quadruple Prediction Based on Multi-Order PromptingabstractRecently, generative aspect-level sentiment quadruple prediction (ASQP) methods based on pre-trained language models have made significant progress. However, some challenges remain in extracting and recognizing complex sentiment elements from semantically rich sentences, limiting the generalization and adaptability of unidirectional generative models in aspect-level sentiment analysis. To overcome this limitation, this article proposes a Generative Aspect-Based Sentiment Quadruple Prediction Model based on Multi-Order Prompting (GenMOP). The model draws on the concept of prompt learning and introduces a multi-order prompting strategy, which breaks the traditional framework of a single generative order and enhances the flexibility and adaptability of the model. Furthermore, we integrate a quadruple quantity-aware module and a multi-view uncertainty-aware module based on a basic generative architecture, not only providing the model with more fine-grained information about the quadruple quantity but also improving the prediction accuracy through uncertainty estimation. The extensive experiments show that the GenMOP method achieves excellent performance in the ASQP task. On the four benchmark datasets including Rest15, Rest16, Rest and Lap, our model achieves F1 score improvements of 1.26%, 0.23%, 0.28%, and 2.07%, respectively, compared to existing state-of-the-arts, demonstrating its effectiveness and superiority in dealing with the joint extraction of multiple sentiment elements of the ASQP model. Rui Wang 0077, Muyao He, Yixue Hao, Long Hu, Min Chen 0003, Baoru Huang |
ACM Trans. Inf. Syst. | 4 |
| 2026 | Context-Aware AIGC Service Migration in Edge Intelligence Networks via Transformer DRLabstractWith the increasing demand for artificial intelligence generated content (AIGC) services across diverse applications, AIGC service migration is essential to ensuring continuous service for mobile users in edge intelligence networks. However, AIGC service migration can lead to decreased inference accuracy due to the discarding of contextual memory. Furthermore, migrating large-scale AIGC models incurs high migration costs and latency. In this paper, we propose a context-aware AIGC service migration scheme to address the trade-off among inference accuracy, latency, and migration cost. Specifically, we focus on migrating historical AIGC context rather than large-scale AIGC models to achieve cost-efficient service provisioning. To improve service migration performance, we propose a Value of Context (VoC) metric to quantify the relevance and freshness of historical AIGC context. Based on the VoC, we formulate an optimization problem to jointly optimize inference accuracy, latency, and migration cost. To solve this problem, we develop a TransFormer-based Soft actor-critic algorithm for Context-aware AIGC service Migration (TFSCM) that leverages long-term dependencies in historical decisions for optimizing the migration process. Extensive experiments on real-world datasets demonstrate that the proposed TFSCM algorithm significantly enhances system performance compared to baseline solutions. Yixue Hao, Rui Wang 0077, Long Hu, Kaibin Huang, Dusit Niyato, Min Chen 0003 |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | More Text, Less Point: Towards 3D Data-Efficient Point-Language UnderstandingabstractEnabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge. Due to the lack of large-scale 3D-text pair datasets, the success of LLMs has yet to be replicated in 3D understanding. In this paper, we rethink this issue and propose a new task: 3D Data-Efficient Point-Language Understanding. The goal is to enable LLMs to achieve robust 3D object understanding with minimal 3D point cloud and text data pairs. To address this task, we introduce GreenPLM, which leverages more text data to compensate for the lack of 3D data. First, inspired by using CLIP to align images and text, we utilize a pre-trained point cloud-text encoder to map the 3D point cloud space to the text space. This mapping leaves us to seamlessly connect the text space with LLMs. Once the point-text-LLM connection is established, we further enhance text-LLM alignment by expanding the intermediate text space, thereby reducing the reliance on 3D point cloud data. Specifically, we generate 6M free-text descriptions of 3D objects, and design a three-stage training strategy to help LLMs better explore the intrinsic connections between different modalities. To achieve efficient modality alignment, we design a zero-parameter cross-attention module for token pooling. Extensive experimental results show that GreenPLM requires only 12% of the 3D training data used by existing state-of-the-art models to achieve superior 3D understanding. Remarkably, GreenPLM also achieves competitive performance using text-only data. Xu Han 0016, Xianzhi Li 0001, Qiao Yu 0002, Jinfeng Xu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
AAAI | 6 |
| 2025 | SASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point CloudsabstractRecent advancements in deep learning have greatly enhanced 3D object recognition, but most models are limited to closed-set scenarios, unable to handle unknown samples in real-world applications. Open-set recognition (OSR) addresses this limitation by enabling models to both classify known classes and identify novel classes. However, current OSR methods rely on global features to differentiate known and unknown classes, treating the entire object uniformly and overlooking the varying semantic importance of its different parts. To address this gap, we propose Salience-Aware Structured Separation (SASep), which includes (i) a tunable semantic decomposition (TSD) module to semantically decompose objects into important and unimportant parts, (ii) a geometric synthesis strategy (GSS) to generate pseudo-unknown objects by combining these unimportant parts, and (iii) a synth-aided margin separation (SMS) module to enhance feature-level separation by expanding the feature distributions between classes. Together, these components improve both geometric and feature representations, enhancing the model’s ability to effectively distinguish known and unknown classes. Experimental results show that SASep achieves superior performance in 3D OSR, outperforming existing state-of-the-art methods. The codes are available at https://github.com/JinfengX/SASep. Jinfeng Xu 0002, Xianzhi Li 0001, Xu Han 0016, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
CVPR | 6 |
| 2025 | Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play DeformationabstractGenerating 3D meshes from a single image is an important but ill-posed task. Existing methods mainly adopt 2D multiview diffusion models to generate intermediate multiview images, and use the Large Reconstruction Model (LRM) to create the final meshes. However, the multiview images exhibit local inconsistencies, and the meshes often lack fidelity to the input image or look blurry. We propose Fancy123, featuring two enhancement modules and an unprojection operation to address the above three issues, respectively. The appearance enhancement module deforms the 2D multiview images to realign misaligned pixels for better multiview consistency. The fidelity enhancement module deforms the 3D mesh to match the input image. The unprojection of the input image and deformed multiview images onto LRM’s generated mesh ensures high clarity, discarding LRM’s predicted blurry-looking mesh colors. Extensive qualitative and quantitative experiments verify Fancy123’s SoTA performance with significant improvement. Also, the two enhancement modules are plug-and-play and work at inference time, allowing seamless integration into various existing single-image-to-3D methods. Project page: https://github.com/YuQiao0303/Fancy123. Qiao Yu 0002, Xianzhi Li 0001, Xu Han 0016, Long Hu, Yixue Hao, Min Chen 0003 |
CVPR | 6 |
| 2025 | Fusion-PSRO: Nash Policy Fusion for Policy Space Response OraclesabstractFor solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown that the Policy Space Response Oracles (PSRO) algorithm is an effective framework for solving such games. However, current methods initialize a new policy from scratch or inherit a single historical policy for Best Response (BR), missing the opportunity to leverage past policies to generate a better BR. In this paper, we propose Fusion-PSRO, which employs Nash Policy Fusion to initialize a new policy for BR training. Nash Policy Fusion serves as an implicit guiding policy that starts exploration on the current Meta-NE, thus providing a closer approximation to BR. Moreover, it insightfully captures a weighted moving average of past policies, dynamically adjusting these weights based on the Meta-NE in each iteration. This cumulative process further enhances the policy population. Empirical results on classic benchmarks show that Fusion-PSRO achieves lower exploitability, thereby mitigating the shortcomings of previous research on policy initialization in BR. Jiesong Lian, Yucong Huang, Chengdong Ma, Ying Wen 0001, Long Hu, Yixue Hao |
ECAI | 7 |
| 2025 | Membership Inference Attacks against Fine-Tuned Stable Diffusion Models Based on Semantic Loss TrajectoryabstractWith the increasing popularity of text-to-image generation models, numerous researchers have recently released their fine-tuned models, accompanied by automatically saved and publicly available checkpoints, which serve as essential resources for model replication and practical applications across various tasks. However, these checkpoints introduce a significant risk of leaking sensitive information, which serves as a primary target for membership inference attacks (MIAs). Despite this, most existing MIAs methods remain limited in their ability to effectively exploit these checkpoints, particularly in cross-modal contexts. In this paper, we propose a membership inference method based on semantic loss trajectories, which analyzes the checkpoints of shadow models to assess the degree of semantic alignment between generated images and text during the model training process. These trajectory changes are then used to represent membership information. To more accurately capture the semantic matching between generated images and text, we introduce a semantic similarity calculation method that, unlike traditional pixel-level comparisons, focuses on high-level semantic information. Experimental results demonstrate that semantic loss trajectories can accurately infer membership with high confidence, achieving an attack success rate approaching 0.82, approximately 30% higher than existing methods, while maintaining robust performance under low false positive rates. Jing Lai, Yongfeng Qian, Yixue Hao |
IJCNN | 3 |
| 2025 | JIMR: Joint Semantic and Geometry Learning for Point Scene Instance Mesh ReconstructionabstractPoint scene instance mesh reconstruction is a challenging task since it requires both scene-level instance segmentation and instance-level mesh reconstruction from partial observations simultaneously. Previous works either adopt a detection backbone or a segmentation one, and then directly employ a mesh reconstruction network to produce complete meshes from incomplete instance point clouds. To further boost the mesh reconstruction quality with both local details and global smoothness, in this work, we propose JIMR, a joint framework with two cascaded stages for semantic and geometry understanding. In the first stage, we propose to perform both instance segmentation and object detection simultaneously. By making both tasks promote each other, this design facilitates subsequent mesh reconstruction by providing more precisely-segmented instance points and better alignment benefiting from predicted complete bounding boxes. In the second stage, we propose a complete-then-reconstruct procedure, where the completion module explicitly disentangles completion from reconstruction, and enables the usage of pre-trained weights of existing powerful completion and reconstruction networks. Moreover, we propose a comprehensive confidence score to filter proposals considering the quality of instance segmentation, bounding box detection, semantic classification, and mesh reconstruction at the same time. Experiments show that our proposed JIMR outperforms state-of-the-art methods regarding instance reconstruction qualitatively and quantitatively. Qiao Yu 0002, Xianzhi Li 0001, Jinfeng Xu 0002, Long Hu, Yixue Hao, Min Chen 0003 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | SE-DCFN: Semantic-Enhanced Dual Cross-modal Fusion Network for Depression RecognitionabstractAutomatic multi-modal depression recognition using artificial intelligence technology is crucial to advance early diagnosis and treatment. Existing methods suffer from a weak performance in detecting depression due to incomplete unimodal semantic information and insufficient fusion effects. To address these challenges, we propose a novel Semantic-Enhanced Dual Cross-modal Fusion Network (SE-DCFN) for multi-modal depression recognition, specifically designed for text-audio data. Firstly, we utilize a prompt learning-based text encoder and a language-audio pertaining-based audio encoder to capture specific information to enhance the semantic representation. Then, we introduce a dual cross-modal fusion module based on self-attention and cross-attention mechanisms to effectively explore linguistic and acoustic representation, facilitating inter-modal and intra-modal interaction and fusion. Additionally, a triplet contrastive loss is formulated to optimize the training process of the SE-DCFN. Experimental results on the EATD-Corpus dataset and AVEC-2017 dataset demonstrate the effectiveness and superiority of our proposed SE-DCFN on multi-modal depression recognition, outperforming existing methods. Long Hu, Qingyi Yang, Rui Wang 0077, Yixue Hao, Min Chen 0003, Yijun Mo |
BIBM | 4 |
| 2024 | PDF: A Probability-Driven Framework for Open World 3D Point Cloud Semantic SegmentationabstractExisting point cloud semantic segmentation networks cannot identify unknown classes and update their knowledge, due to a closed-set and static perspective of the real world, which would induce the intelligent agent to make bad decisions. To address this problem, we propose a Probability-Driven Framework (PDF)11Code available at: https://github.com/JinfengX/PointCloudPDF. for open world semantic segmentation that includes (i) a lightweight U-decoder branch to identify unknown classes by estimating the uncertainties, (ii) a flexible pseudo-labeling scheme to supply geometry features along with probability distribution features of unknown classes by generating pseudo labels, and (iii) an incremental knowledge distillation strategy to incorporate novel classes into the existing knowledge base gradually. Our framework enables the model to behave like human beings, which could recognize unknown objects and incrementally learn them with the corresponding knowledge. Experimental results on the S3DIS and ScanNetv2 datasets demonstrate that the proposed PDF outperforms other methods by a large margin in both important tasks of open world semantic segmentation. Jinfeng Xu 0002, Xianzhi Li 0001, Yixue Hao, Long Hu, Min Chen 0003 |
CVPR | 5 |
| 2024 | Multimodal Physiological Signals Representation Learning via Multiscale Contrasting for Depression Recognition
Kai Shao, Rui Wang 0077, Yixue Hao, Long Hu, Min Chen 0003, Hans-Arno Jacobsen |
ACM Multimedia | 3 |
| 2024 | MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D PriorsabstractLarge 2D vision-language models (2D-LLMs) have gained significant attention by bridging Large Language Models (LLMs) with images using a simple projector. Inspired by their success, large 3D point cloud-language models (3D-LLMs) also integrate point clouds into LLMs. However, directly aligning point clouds with LLM requires expensive training costs, typically in hundreds of GPU-hours on A100, which hinders the development of 3D-LLMs. In this paper, we introduce MiniGPT-3D, an efficient and powerful 3D-LLM that achieves multiple SOTA results while training for only 27 hours on one RTX 3090. Specifically, we propose to align 3D point clouds with LLMs using 2D priors from 2D-LLMs, which can leverage the similarity between 2D and 3D visual information. We introduce a novel four-stage training strategy for modality alignment in a cascaded way, and a mixture of query experts module to adaptively aggregate features with high efficiency. Moreover, we utilize parameter-efficient fine-tuning methods LoRA and Norm fine-tuning, resulting in only 47.8M learnable parameters, which is up to 260x fewer than existing methods. Extensive experiments show that MiniGPT-3D achieves SOTA on 3D object classification and captioning tasks, with significantly cheaper training costs. Notably, MiniGPT-3D gains an 8.12 increase on GPT-4 evaluation score for the challenging object captioning task compared to ShapeLLM-13B, while the latter costs 160 total GPU-hours on 8 A800. We are the first to explore the efficient 3D-LLM, offering new insights to the community. Code and weights are available at https://github.com/TangYuan96/MiniGPT-3D. Xu Han 0016, Xianzhi Li 0001, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
ACM Multimedia | 5 |
| 2024 | Dynamic differential privacy-based dataset condensation
Zhaoxuan Wu, Yongfeng Qian, Yixue Hao, Min Chen 0003 |
Neurocomputing | 4 |
| 2024 | Big Fiber Slicing for Dynamic Multimodal Multipreference Applications of Smart FabricsabstractIn recent years, significant breakthroughs have been achieved in smart fabric technology within the healthcare sector, providing an impetus for the smart integration of wearable devices and equipment in medical applications. However, the tight coupling between fabric hardware devices and software solutions, tailored for various scenarios, has led to inefficient utilization of hardware resources and led to challenges for device upgrades and iterations. This paper focuses on the virtualization technology of smart fabric hardware resources and introduces a novel approach, termed “Big Fiber Slicing”. First, we outline the design of novel fiber devices customized for two major application scenarios: health monitoring and protection. Subsequently, we delve into the process of partitioning hardware resources into multiple “fiber slices” to better meet the unique requirements of various application scenarios and services. Next, we built a smart fabric platform, combined with 5 real multi-modal applications with different preferences, to verify the performance of the system when resources are limited and demand changes dynamically. Lastly, we explore the potential challenges that smart fabric technology may encounter in future application scenarios and provide insights into the future direction of this field. Jia Liu 0009, Huanke Zheng, Dongkun Huo, Yixue Hao, Dusit Niyato, Salman AlQahtani, Min Chen 0003 |
IEEE Internet Things J. | 4 |
| 2024 | Efficient Crowd Counting via Dual Knowledge DistillationabstractMost researchers focus on designing accurate crowd counting models with heavy parameters and computations but ignore the resource burden during the model deployment. A real-world scenario demands an efficient counting model with low-latency and high-performance. Knowledge distillation provides an elegant way to transfer knowledge from a complicated teacher model to a compact student model while maintaining accuracy. However, the student model receives the wrong guidance with the supervision of the teacher model due to the inaccurate information understood by the teacher in some cases. In this paper, we propose a dual-knowledge distillation (DKD) framework, which aims to reduce the side effects of the teacher model and transfer hierarchical knowledge to obtain a more efficient counting model. First, the student model is initialized with global information transferred by the teacher model via adaptive perspectives. Then, the self-knowledge distillation forces the student model to learn the knowledge by itself, based on intermediate feature maps and target map. Specifically, the optimal transport distance is utilized to measure the difference of feature maps between the teacher and the student to perform the distribution alignment of the counting area. Extensive experiments are conducted on four challenging datasets, demonstrating the superiority of DKD. When there are only approximately 6% of the parameters and computations from the original models, the student model achieves a faster and more accurate counting performance as the teacher model even surpasses it. Rui Wang 0077, Yixue Hao, Long Hu, Xianzhi Li 0001, Min Chen 0003, Yiming Miao, Iztok Humar |
IEEE Trans. Image Process. | 2 |
| 2024 | Spotlighter: Backup Age-Guaranteed Immersive Virtual Vehicle Service Provisioning in Edge-Enabled Vehicular MetaverseabstractEdge-enabled Vehicular Metaverse (EVM) is a new paradise supported by various compute-intensive Virtual Vehicle Services (VVSs), where users can immerse and enjoy their spiritual world. User immersion is critical during VVS provisioning in the EVM, yet it can be weakened or curtailed by a sense of disengagement caused by unknown failures. Providing redundant backups VVSs (BVVSs) and keeping the Age of Backup Information (AoBI) could effectively resist and avoid this disengagement when failures occur. However, the trajectories of mobile vehicles are unknown and dynamic, which makes it challenging to optimally migrate VVSs and BVVSs or adjust the update frequency of backup information in real-time, so as to ensure service reliability and AoBI while minimizing the cost of accepting VVS-based metaverse services. In this paper, the above long-term issue is first decomposed into discrete single-slot sub-problems that are modeled as integer linear programming problems. Then, a comprehensive resource explorer named spotlighter is designed, where the first and second parts are a metaverse service home prediction algorithm based on deep learning and a VVS migration algorithm based on randomized rounding, respectively. By tracking the dynamical locations of service homes based on current and historical information, the former can help the latter to adaptively minimize migration costs on VVS re-instantiation and traffic transmission among services and moving vehicles. Finally, a cost-adaptive AoBI guarantee algorithm is merged in spotlighter to ensure the freshness of backup status, by trading-off synchronization cost on BVVS migration, backup update, and backup synchronization. Theoretical analyses and experiments based on real databases show that our algorithms are promising compared with baseline algorithms. Min Chen 0003, Hebin Huang, Weifa Liang, Junbin Liang, Yixue Hao, Dusit Niyato |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Reliable or Green? Continual Individualized Inference Provisioning in Fabric Metaverse via Multi-Exit AccelerationabstractFabric metaverse employs intelligence fibers embedded with flexible sensors to unknowingly gather and transmit massive hypermodal data around humans to a deep neural network-based metaverse inference service (DMS) for continual and real-time analysis. Each DMS has one primary branch and multiple side branches that allow early termination of service with differential accuracy and energy consumption. However, the continual provisioning of compute-intensive DMS with varying requirements for service model, accuracy, delay, and reliability poses a challenge for edge servers characterized by restricted computing resources and intermittent green energy. In this paper, we focus on a continual individualized DMS provisioning problem in the fabric metaverse consisting of a side branch insertion subproblem and a server activation and service deployment subproblem, and formulate them as Integer linear Programming and Markov Decision Process, respectively. Then, we propose a green continual inference (GCI) system, where a pruner with provable approximation ratios trims superfluous branches of every model to the given number$K$to minimize total overflow accuracy between accuracy demands and reserved branches assigned to users. Based on this exit result, each DMS is further divided into several blocks with dependencies to exploit constrained resources of computing and energy in a fine-grained manner. Finally, a learning-based scheduler is merged into GCI to maximize request throughput while minimizing the activation number of edge servers on different demand scenarios, by adaptively activating suitable servers and deploying required blocks and their corresponding backups on selected servers. Theoretical analyses, simulations, and experiments demonstrate that the GCI is promising compared with baseline algorithms. Min Chen 0003, Weifa Liang, Dusit Niyato, Yue Wang 0092, Victor C. M. Leung, Yixue Hao, Long Hu, Yin Zhang 0002 |
IEEE Trans. Mob. Comput. | 8 |
| 2024 | Point-LGMask: Local and Global Contexts Embedding for Point Cloud Pre-Training With Multi-Ratio MaskingabstractSelf-supervised learning has achieved great success in both natural language processing and 2D vision, where masked modeling is a quite popular pre-training scheme. However, extending masking to 3D point cloud understanding that combines local and global features poses a new challenge. In our work, we present Point-LGMask, a novel method to embed both local and global contexts with multi-ratio masking, which is quite effective for self-supervised feature learning of point clouds but is unfortunately ignored by existing pre-training works. Specifically, to avoid fitting to a fixed masking ratio, we first propose multi-ratio masking, which prompts the encoder to fully explore representative features thanks to tasks of different difficulties. Next, to encourage the embedding of both local and global features, we formulate a compound loss, which consists of (i) a global representation contrastive loss to encourage the cluster assignments of the masked point clouds to be consistent to that of the completed input, and (ii) a local point cloud prediction loss to encourage accurate prediction of masked points. Equipped with our Point-LGMask, we show that our learned representations transfer well to various downstream tasks, including few-shot classification, shape classification, object part segmentation, as well as real-world scene-based 3D object detection and 3D semantic segmentation. Particularly, our model largely advances existing pre-training methods on the difficult few-shot classification task using the real-captured ScanObjectNN dataset by surpassing over 4% to the second-best method. Also, our Point-LGMask achieves 0.4%$AP_{25}$and 0.8%$AP_{50}$gains on 3D object detection task over the second-best method. 0.4% mAcc and 0.5% mIoU. Codes have been released athttps://github.com/TangYuan96/Point-LGMask. Xianzhi Li 0001, Jinfeng Xu 0002, Qiao Yu 0002, Long Hu, Yixue Hao, Min Chen 0003 |
IEEE Trans. Multim. | 6 |
| 2024 | Immersive Multimedia Service Caching in Edge Cloud with Renewable EnergyabstractImmersive service caching, based on the intelligent edge cloud, can meet delay-sensitive service requirements. Although numerous service caching solutions for edge clouds have been designed, they have not been well explored. Moreover, to the best of our knowledge, there is no work to consider the immersive service caching scheme under the supply of renewable energy. In this article, we investigate the service caching problem under the renewable energy supply to minimize service latency while making full use of renewable energy. Specifically, we formulate the service caching and renewable energy harvesting problem, which considers the dynamic renewable energy, unknown service requests, and limited capacity of the edge cloud. To solve this problem, we propose an effective algorithm, called OSCRE. Our algorithm first uses Lyapunov optimization to convert the time-average problem into time-independence optimization and thus realizes optimal renewable energy harvesting. Then, it realizes the service caching scheme using data-driven combinatorial multi-armed bandit learning. The simulation results show that the OSCRE scheme can save service latency while making sufficient use of renewable energy. M. Shamim Hossain, Yixue Hao, Long Hu, Jia Liu 0009, Min Chen 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | RT3C: Real-Time Crowd Counting in Multi-Scene Video Streams via Cloud-Edge-Device CollaborationabstractRecently, the advancements in edge computing have boosted the deployment of video analysis systems based on deep learning, which breaks the limitation of the constrained communication and computing resources of local devices. However, processing multi-scene high-resolution video streams in crowd surveillance remains a significant challenge since it is difficult to formulate dynamic video content and communication environments to support offloading decisions. To bridge the gap between applications and modeling, this paper presents aReal-TimeCloud-edge-deviceCollaboration framework, which enables fast and accurateCrowd counting (RT3C) on the real dataset. RT3C comprises key frame detection, adaptive patch partition, patch encoder and decoder and computation offloading decision, designed to divide key frames into a minimum number of patches and determine the offloading location of patches. A Real-Time Multi-Agent Actor-Critic (RTMAAC) algorithm based on multi-agent reinforcement learning is proposed to decide whether to compute patches with a lightweight model on edge or a large model on cloud. Unlike traditional approaches ignoring the contents, RTMAAC is a dynamic online decision algorithm based on context of the network and video. Extensive experiments demonstrate that RT3C effectively discriminate the valid frames and optimizes offloading decisions in complex environments, outperforming other baseline algorithms on the two crowd counting datasets. In summary, RT3C provides a promising framework for multi-scene video streams, which can be extended to other applications to realize video computation based on deep models. Rui Wang 0077, Yixue Hao, Yiming Miao, Long Hu, Min Chen 0003 |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | CasFusionNet: A Cascaded Network for Point Cloud Semantic Scene Completion by Dense Feature FusionabstractSemantic scene completion (SSC) aims to complete a partial 3D scene and predict its semantics simultaneously. Most existing works adopt the voxel representations, thus suffering from the growth of memory and computation cost as the voxel resolution increases. Though a few works attempt to solve SSC from the perspective of 3D point clouds, they have not fully exploited the correlation and complementarity between the two tasks of scene completion and semantic segmentation. In our work, we present CasFusionNet, a novel cascaded network for point cloud semantic scene completion by dense feature fusion. Specifically, we design (i) a global completion module (GCM) to produce an upsampled and completed but coarse point set, (ii) a semantic segmentation module (SSM) to predict the per-point semantic labels of the completed points generated by GCM, and (iii) a local refinement module (LRM) to further refine the coarse completed points and the associated labels from a local perspective. We organize the above three modules via dense feature fusion in each level, and cascade a total of four levels, where we also employ feature fusion between each level for sufficient information usage. Both quantitative and qualitative results on our compiled two point-based datasets validate the effectiveness and superiority of our CasFusionNet compared to state-of-the-art methods in terms of both scene completion and semantic segmentation. The codes and datasets are available at: https://github.com/JinfengX/CasFusionNet. Jinfeng Xu 0002, Xianzhi Li 0001, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
AAAI | 5 |
| 2023 | Data Augmentation and Pseudo-sequence of fNIRS for Depression RecognitionabstractDepression is a mental disorder caused by factors such as genetics, life events and social influences, and has become a major public health problem worldwide. Previous studies have demonstrated the potential of functional near-infrared spectroscopy (fNIRS) in the diagnosis of depression. However, in the real medical scene, fNIRS data are difficult to obtain, limited in number and suffer from class imbalance. To overcome these problems, in this paper, we propose a novel model for depression identification based on data augmentation and pseudo-sequence of fNIRS. Specifically, the data augmentation using the time masking and warping method generates richer data. Then, a stimulation task-driven data pseudo-sequence method is designed to map the sequence data into pseudo-sequence activation images. Finally, a depression recognition model is established based on the class imbalance loss function. Experiments show that the precision of our depression recognition model reaches 0.94. This scheme transforms fNIRS data into image sequences, which provides a new solution idea for subsequent research. Kai Shao, Yixue Hao, Long Hu, Xiaofen Zong, Min Chen 0003 |
BIBM | 2 |
| 2023 | Crowd Intelligent Grouping Collaboration Evacuation via Multi-agent Reinforcement LearningabstractThe crowd evacuation strategy seeks to arrange crowd evacuation in an orderly manner to protect people’s lives and reduce property damage in case of sudden emergencies in crowded and complex places. In recent years, there have been several works to apply deep learning to crowd evacuation to make evacuation strategies more intelligent. However, existing researches rarely consider the integration of scene perception and crowd evacuation, which leads to evacuation methods that are detached from the scene and also ignore the crowd collaboration in the evacuation process. To this end, we propose Intelligent Crowd Evacuation Architecture based on Visual features using Multi-Agent Reinforcement Learning (ICEA-VMARL). Subsequently, we present modeling analysis on the crowd grouping and group evacuation modules of the architecture. First, we propose the Population Grouping algorithm based on Continuous Spatiotemporal individual Similarity (PGCSS), which combines crowd features to group crowds. Then, we propose a Group Collaborative Evacuation algorithm based on Multi-Agent Reinforcement Learning (GCE-MARL), which considers group collaboration while evacuating to achieve global optimal evacuation. Finally, we build an experimental crowd simulation system, and the results demonstrate that the crowd grouping algorithm and group evacuation algorithm proposed have better performance compared with other methods. Rui Wang 0077, Jinfeng Xu 0002, Long Hu, Yixue Hao |
CSCWD | 6 |
| 2023 | Joint Sensing Adaptation and Model Placement in 6G Fabric ComputingabstractSensing and computing based on intelligent fabrics can meet the ultra-reliable and low-latency communication (URLLC) needs of sixth-generation wireless (6G) by integrating sensing units into fabric fibers to perceive user data. Although some researchers have designed sensing or computing solutions, such solutions have not been well explored. In this paper, we consider the joint sensing adaptation and model placement in a 6G fabric space. We first propose an intelligent-fiber-driven 6G fabric computing network to minimize acquisition latency while ensuring accuracy. Then, we formulate an optimization model that takes the fabric sampling rate, sampling density, and model placement as variables. To solve the model, we propose an effective learning algorithm based on deep reinforcement learning. That is, by transforming the optimization problem into a state space, action space, and reward function, we design an optimal sensing and placement scheme. The simulation results show that our proposed scheme can achieve optimal sensing and computing compared with several baseline algorithms. Yixue Hao, Long Hu, Min Chen 0003 |
IEEE J. Sel. Areas Commun. | 1 |
| 2023 | Digital Twin-Assisted URLLC-Enabled Task Offloading in Mobile Edge Network via Robust Combinatorial OptimizationabstractDigital twin (DT)-assisted mobile edge network can achieve energy-efficient task offloading by optimizing the decision-making in real time. Although many DT-assisted task offloading solutions in mobile edge networks have been designed, stochastic asynchronizations between the DTs and physical entities are still ignored. In this paper, we investigate a task offloading problem in a DT-assisted URLLC-enabled mobile edge network which considered the uncertain deviation between DT estimated values and physical actual values. Specifically, we formulate a latency and energy consumption minimization problem by optimizing task offloading, resource allocation, and power management. To solve this problem, we propose a DT-assisted robust task offloading scheme (DTRTO) based on learning composed of decision and deviation networks. The deviation network predicts the worst-case deviations based on the pre-decision, and the decision network optimize the decision considered the worst-case deviation. The simulation results show that, compared to the baseline algorithms, the DTRTO scheme can realize low latency and energy consumption in task offloading while maintaining high robustness. Yixue Hao, Dongkun Huo, Nadra Guizani, Long Hu, Min Chen 0003 |
IEEE J. Sel. Areas Commun. | 1 |
| 2023 | Drone Swarm Path Planning for Mobile Edge Computing in Industrial Internet of ThingsabstractDrone-swarm-assisted mobile edge computing (MEC) provides extra computation and storage capacity for smart city applications and the Industrial Internet of Things. To solve the problems of traditional fixed base stations in a complex terrain, including cost of deployment, transmission loss of telecommunication, and limited coverage, this article brings forward the unmanned aerial vehicles (UAVs) as MEC nodes in the air. For the purpose of matching the dynamic mobile devices and UAV trajectory, this article raises a multi-UAVs-assisted MEC offloading algorithm based on global and local path planning controlled by ground station and onboard computer. Firstly, this article considers a drone swarm scheduling and allocation strategy based on the priority of monitoring areas, UAVs residual energy and distance to target points, so as to minimize the global flight length and energy consumption. Secondly, based on user mobility, this article calculates the optimal communication coverage of a UAV, and jointly optimizes the local path planning and computing offloading, so as to maximize the number of offloading services and minimize the total latency in completing the computation task. Finally, based on the total latency and energy consumption of path planning and computation offloading, a UAV cluster computation offloading strategy with optimized energy efficiency is realized. Experimental results prove that the proposed algorithm can provide more offloading services while obtaining shorter path length and greater energy efficiency. Yiming Miao, Kai Hwang 0001, Di Wu 0001, Yixue Hao, Min Chen 0003 |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Self-Supervised Learning With Data-Efficient Supervised Fine-Tuning for Crowd CountingabstractDue to the expensive and laborious annotations of labeled data required by fully-supervised learning in the crowd counting task, it is desirable to explore a method to reduce the labeling burden. There exists a large number of unlabeled images in the wild that can be easily obtained compared to labeled datasets. Based on the characteristics of consistent spatial transformation with the annotations of heads and image, this paper proposes a self-supervised learning framework with unlabeled and limited labeled data for pre-training and fine-tuning crowd counting model (SSL-FT). It includes an online network and a target network that receive the same images but are randomly processed by two defined augmentation transformations. We leverage unlabeled data to pre-train the online network based on a self-supervised loss and small-scale labeled data to transfer the model to a specific domain based on a fully-supervised loss. We demonstrate the effectiveness of the SSL-FT on four public datasets including ShanghaiTech PartA, PartB, UCF-QNRF and WorldExpo'10 utilizing a classical counting model. Experimental results show that our approach performs better than state-of-art semi-supervised methods. Rui Wang 0077, Yixue Hao, Long Hu, Jincai Chen, Min Chen 0003, Di Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | An End-to-End Human Abnormal Behavior Recognition Framework for Crowds With Mentally Disordered IndividualsabstractAbnormal or violent behavior by people with mental disorders is common. When individuals with mental disorders exhibit abnormal behavior in public places, they may cause physical and mental harm to others as well as to themselves. Thus, it is necessary to monitor their behavior using visual surveillance systems. However, it is challenging to automatically detect human abnormal behavior (especially for individuals with mental disorders) based on motion recognition technologies. To address these issues, in the current work, we propose an end-to-end abnormal behaviour detection framework from a new perspective in conjunction with the Graph Convolutional Network (GCN) and a 3D Convolutional Neural Network (3DCNN). Specifically, we first train a one-class classifier to extract features and estimate abnormality scores. To improve the performance of abnormal behavior detection, GCN is used to model the similarity between video clips for the correction of noisy labels. Then, based on this framework, GCN recognizes the normal behavior clips in the abnormal video and removes them, while the clips identified as abnormal behavior are retained. Finally, a 3D CNN is used to extract spatiotemporal features to classify different abnormal behaviors. In order to better detect the violent behavior of individuals with mental disorders, the paper focuses on the UCF-Crime dataset with various types of violent behaviors. By experimenting with this dataset, the classification accuracy reaches 37.9%, which is significantly better than that of the current state-of-the-art approaches. Yixue Hao, Zaiyang Tang, Bander A. Alzahrani, Reem Alotaibi, Reem Alharthi, Miaomiao Zhao, Arif Mahmood |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | TIF: Trajectory and Information Flow Coupling Mechanism for Behavior Analysis in Autonomous DrivingabstractThe significant achievements have been made in crowd detection and tracking due to the advancement of artificial intelligence in the autonomous driving. However, the image-based methods have strict requirements for the collection conditions of video, and the development of the new generation of flexible fabrics has become potential sensors to perceive context. In this paper, an intelligent fabric space enabled by multi-sensing sensors is established to track the motion objects. We propose a behavior analysis pipeline including the modules of data preparation, trajectory coupling, motion scenario segmentation, and motion pattern measurement to capture the crowd information from micro-level and macro-level over the intelligent fabric space. After making preprocess for the multi-sensing data, a coupling mechanism is formulated to fuse the video-based trajectory and fabric-based trajectory. And an automatic motion scenario segmentation model divides the surrounding scenario into main-crowd, sub-crowd, and background according to the motion behavior. Further, we define measurement metrics to analyze the motion pattern for the different crowds. Extensive experiments prove that our proposed methods effectively fuse multiple trajectories and realize the crowd segmentation and the motion description. This will greatly help autonomous vehicles and control system perceive the surrounding pedestrians and the environment to make precise driving decisions. Rui Wang 0077, Jinfeng Xu 0002, Jia Liu 0009, Di Wu 0001, Yixue Hao, Xianzhi Li 0001, Min Chen 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Negative Information Measurement at AI Edge: A New Perspective for Mental Health MonitoringabstractThe outbreak of the corona virus disease 2019 (COVID-19) has caused serious harm to people’s physical and mental health. Due to the serious situation of the epidemic, a lot of negative energy information increases people’s psychological burden. However, effective interventions against mental health problems are not in abundance. To address such challenges, in this article, we propose the concept of negative information to describe information that has a negative impact on people’s mental health. To achieve the measurement of negative information, the level of mental health inversely measures the degree of negative information. Specifically, we design a system to measure the negative information used to monitor the mental health state of the user under the impact of negative information. The cognition of mental health is realized based on the intelligent algorithm deployed on the edge cloud, and the needs of users can be responded to in real time in practical applications. Finally, we use real collected dataset to verify the influence of negative information. The experiments show that the system can achieve negative information measurement and provide an effective countermeasure for solving mental health problems during a pandemic situation. Min Chen 0003, Ke Shen 0004, Rui Wang 0077, Yiming Miao, Kai Hwang 0001, Yixue Hao, Guangming Tao, Long Hu, Zhongchun Liu |
ACM Trans. Internet Techn. | 7 |
| 2022 | A Multi-feature and Time-aware-based Stress Evaluation Mechanism for Mental Status AdjustmentabstractWith the rapid economic development, the prominent social competition has led to increasing psychological pressure of people felt from each aspect of life. Driven by the Internet of Things and artificial intelligence, intelligent psychological pressure detection systems based on deep learning and wearable devices have acquired some good results in practical application. However, existing studies argue that the psychological stress state is influenced by the current environment. They put much attention on the momentary features but ignore the dynamic change process of mental status in the time dimension. Besides, the lack of research in the general laws of psychological stress makes it difficult to quantitatively evaluate the stress status, resulting in the inability to perceive the stress state of users effectively. Thus, this article proposes an evaluation mechanism of psychological stress for adjusting the mental status of users. Specifically, we design a multi-dimensional feature space and a time-aware feature encoder, which integrate various stress features and capture time characteristics of stress state change. Moreover, a novel mental state model is proposed, which uses the pressure features with time characteristics to evaluate the pressure stress level. This model also quantifies the internal relationship between pressure features. Last, we establish a practicable testbed to demonstrate how to evaluate and adjust mental state of users by the proposed evaluation mechanism of psychological stress. Min Chen 0003, Wenjing Xiao, Yixue Hao, Long Hu, Guangming Tao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Online Vehicle Selection for Task Replication via Bandit LearningabstractTask replication in a vehicular environment, which exploits vehicular spare computation and storage resources to process tasks of others, has attracted significant attention recently. Although there are many approaches to design optimal schemes, the vehicle selection for task replication has not been well explored. Moreover, to the best of our knowledge, none of the existing vehicle selection mechanisms for task replication can provide quality-guaranteed selection under unknown vehicles with dynamic and uncertain costs. In this paper, we propose a novel online vehicle selection strategy for task replication in vehicular networks with unknown information, with the objective to minimize task processing latency. We model such an unknown vehicle selection process as a budgeted multi-arm bandit problem, and propose a quality-aware vehicle selection (QVS) scheme. Furthermore, we extend this problem to the case where vehicular costs for task processing are dynamic and uncertain, and design a quality- and cost-aware vehicle selection (QCVS) scheme to dynamically select suitable vehicles. Moreover, we analyze the regrets of two algorithms. Finally, we conduct extensive experiments to demonstrate the effectiveness of our proposed mechanisms. Yongfeng Qian, Zhoutong Zuo, Yixue Hao |
COMPSAC | 3 |
| 2021 | Deep Reinforcement Learning for Edge Service Placement in Softwarized Industrial Cyber-Physical SystemabstractFuture industrial cyber-physical system (CPS) devices are expected to request a large amount of delay-sensitive services that need to be processed at the edge of a network. Due to limited resources, service placement at the edge of the cloud has attracted significant attention. Although there are many methods of design schemes, the service placement problem in industrial CPS has not been well studied. Furthermore, none of existing schemes can optimize service placement, workload scheduling, and resource allocation under uncertain service demands. To address these issues, we first formulate a joint optimization problem of service placement, workload scheduling, and resource allocation in order to minimize service response delay. We then propose an improved deep Q-network (DQN)-based service placement algorithm. The proposed algorithm can achieve an optimal resource allocation by means of convex optimization where the service placement and workload scheduling decisions are assisted by means of DQN technology. The experimental results verify that the proposed algorithm, compared with existing algorithms, can reduce the average service response time by 8-10%. Yixue Hao, Min Chen 0003, Hamid Gharavi, Yin Zhang 0002, Kai Hwang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Depression Analysis and Recognition Based on Functional Near-Infrared SpectroscopyabstractDepression is the result of a complex interaction of social, psychological and physiological elements. Research into the brain disorders of patients suffering from depression can help doctors to understand the pathogenesis of depression and facilitate its diagnosis and treatment. Functional near-infrared spectroscopy (fNIRS) is a non-invasive approach to the detection of brain functions and activities. In this paper, a comprehensive fNIRS-based depression-processing architecture, including the layers of source, feature and model, is first established to guide the deep modeling for fNIRS. In view of the complexity of depression, we propose a methodology in the time and frequency domains for feature extraction and deep neural networks for depression recognition combined with current research. It is found that compared to non-depression people, patients with depression have a weaker encephalic area connectivity and lower level of activation in the prefrontal lobe during brain activity. Finally, based on raw data, manual features and channel correlations, the AlexNet model shows the best performance, especially in terms of the correlation features and presents an accuracy rate of 0.90 and a precision rate of 0.91, which is higher than ResNet18 and machine-learning algorithms on other data. Therefore, the correlation of brain regions can effectively recognize depression (from cases of non-depression), making it significant for the recognition of brain functions in the clinical diagnosis and treatment of depression. Rui Wang 0077, Yixue Hao, Qiao Yu 0002, Min Chen 0003, Iztok Humar, Giancarlo Fortino |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Human-Like Hybrid Caching in Software-Defined Edge CloudabstractWith the development of Internet of Things (IoT) and communication technology, the number of next-generation IoT devices has increased explosively, and the delay requirement for content requests is becoming progressively higher. Fortunately, the edge-caching scheme can satisfy users' demands for low latency of content. However, the existing caching schemes are not smart enough. To address these challenges, we propose a human-like hybrid caching architecture based on the software-defined edge cloud, which simultaneously considers the content popularity and the fine-grained user characteristics. Then, an optimization problem with a caching hit ratio as an optimization objective is formulated. To solve this problem, using reinforcement learning, we design a human-like hybrid caching algorithm. The extensive experiments show that compared with popular caching schemes, human-like hybrid caching schemes can improve the cache hit ratio by 20%. Yixue Hao, Di Wu 0001, Min Chen 0003, Mohammad Mehedi Hassan, Giancarlo Fortino |
IEEE Internet Things J. | 1 |
| 2020 | Learning for Smart Edge: Cognitive Learning-Based Computation Offloading
Yixue Hao, Yinging Jiang, M. Shamim Hossain, Mohammed F. Alhamid, Syed Umar Amin |
Mob. Networks Appl. | 1 |
| 2020 | Label-less Learning for Emotion CognitionabstractIn this paper, we propose a label-less learning for emotion cognition (LLEC) to achieve the utilization of a large amount of unlabeled data. We first inspect the unlabeled data from two perspectives, i.e., the feature layer and the decision layer. By utilizing the similarity model and the entropy model, this paper presents a hybrid label-less learning that can automatically label data without human intervention. Then, we design an enhanced hybrid label-less learning to purify the automatic labeled data. To further improve the accuracy of emotion detection model and increase the utilization of unlabeled data, we apply enhanced hybrid label-less learning for multimodal unlabeled emotion data. Finally, we build a real-world test bed to evaluate the LLEC algorithm. The experimental results show that the LLEC algorithm can improve the accuracy of emotion detection significantly. Min Chen 0003, Yixue Hao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Recurrent convolutional neural network based multimodal disease risk prediction
Yixue Hao, Mohd Usama, Jun Yang 0014, M. Shamim Hossain, Ahmed Ghoneim |
Future Gener. Comput. Syst. | 1 |
| 2019 | Cognitive information measurements: A new perspective
Min Chen 0003, Yixue Hao, Hamid Gharavi, Victor C. M. Leung |
Inf. Sci. | 2 |
| 2019 | Profit Maximization for Video Caching and Processing in Edge CloudabstractWith the development of communication technology and the explosive growth of video traffic brought by the rapid growth of mobile devices (such as smartphones and wearable devices), great business opportunities have been brought to video service providers. In this paper, we make full use of the cache and computing capacity of edge cloud. Considering the multi bitrate of video, we design the video caching and processing model that offers maximized profit to video service provider. Specifically, we model this problem as the 0-1 optimization problem and design the learning-based online upper confidence bound algorithm based on multi-arm bandit theory. This algorithm can design the corresponding cache and process strategy in real time according to the users' request to video. Furthermore, this strategy can maximize the profit of video provider and satisfy the service quality for users. Finally, experimental results show that our proposed video caching and processing scheme is superior to other schemes. Yixue Hao, Long Hu, Yongfeng Qian, Min Chen 0003 |
IEEE J. Sel. Areas Commun. | 1 |
| 2019 | A Dynamic Service Migration Mechanism in Edge Cognitive ComputingabstractDriven by the vision of edge computing and the success of rich cognitive services based on artificial intelligence, a new computing paradigm, edge cognitive computing (ECC), is a promising approach that applies cognitive computing at the edge of the network. ECC has the potential to provide the cognition of users and network environmental information, and further to provide elastic cognitive computing services to achieve a higher energy efficiency and a higher Quality of Experience (QoE) compared to edge computing. This article first introduces our architecture of the ECC and then describes its design issues in detail. Moreover, we propose an ECC-based dynamic service migration mechanism to provide insight into how cognitive computing is combined with edge computing. In order to evaluate the proposed mechanism, a practical platform for dynamic service migration is built up, where the services are migrated based on the behavioral cognition of a mobile user. The experimental results show that the proposed ECC architecture has ultra-low latency and a high user experience, while providing better service to the user, saving computing resources, and achieving a high energy efficiency. Min Chen 0003, Wei Li 0061, Giancarlo Fortino, Yixue Hao, Long Hu, Iztok Humar |
ACM Trans. Internet Techn. | 4 |
| 2018 | Edge cognitive computing based smart healthcare system
Min Chen 0003, Wei Li 0061, Yixue Hao, Yongfeng Qian, Iztok Humar |
Future Gener. Comput. Syst. | 3 |
| 2018 | Task Offloading for Mobile Edge Computing in Software Defined Ultra-Dense NetworkabstractWith the development of recent innovative applications (e.g., augment reality, self-driving, and various cognitive applications), more and more computation-intensive and data-intensive tasks are delay-sensitive. Mobile edge computing in ultra-dense network is expected as an effective solution for meeting the low latency demand. However, the distributed computing resource in edge cloud and energy dynamics in the battery of mobile device makes it challenging to offload tasks for users. In this paper, leveraging the idea of software defined network, we investigate the task offloading problem in ultra-dense network aiming to minimize the delay while saving the battery life of user's equipment. Specifically, we formulate the task offloading problem as a mixed integer non-linear program which is NP-hard. In order to solve it, we transform this optimization problem into two sub-problems, i.e., task placement sub-problem and resource allocation sub-problem. Based on the solution of the two sub-problems, we propose an efficient offloading scheme. Simulation results prove that the proposed scheme can reduce 20% of the task duration with 30% energy saving, compared with random and uniform task offloading schemes. Min Chen 0003, Yixue Hao |
IEEE J. Sel. Areas Commun. | 2 |
| 2018 | Opportunistic Task Scheduling over Co-Located Clouds in Mobile EnvironmentabstractWith the growing popularity of mobile devices, a new type of peer-to-peer communication mode for mobile cloud computing has been introduced. By applying a variety of short-range wireless communication technologies to establish connections with nearby mobile devices, we can construct a mobile cloudlet in which each mobile device can either works as a computing service provider or a service requester. Although the paradigm of mobile cloudlet is cost-efficient in handling computation-intensive tasks, the understanding of its corresponding service mode from a theoretic perspective is still in its infancy. In this paper, we first propose a new mobile cloudlet-assisted service mode named Opportunistic task Scheduling over Co-located Clouds (OSCC), which achieves flexible cost-delay tradeoffs between conventional remote cloud service mode and mobile cloudlets service mode. Then, we perform detailed analytic studies for OSCC mode, and solve the energy minimization problem by compromising among remote cloud mode, mobile cloudlets mode and OSCC mode. We also conduct extensive simulations to verify the effectiveness of the proposed OSCC mode, and analyze its applicability. Moreover, experimental results show that when the ratio of data size after task execution over original data size associated with the task is smaller than 1 (i.e.,r<; 1) and the average meeting rate of two mobile devices λ is larger than 0:00014, our proposed OSCC mode outperforms existing service modes. Min Chen 0003, Yixue Hao, Chin-Feng Lai, Di Wu 0001, Yong Li 0008, Kai Hwang 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2017 | Cloud-assisted hugtive robot for affective interaction
Yixue Hao, Jun Yang 0014, Wei Li 0061, Yiming Miao, Jeungeun Song 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Green and Mobility-Aware Caching in 5G NetworksabstractWith the drastic increase of mobile devices, there are more and more mobile traffic and repeated requests for content. In 5G networks, small cell base stations (SBSs) caching and caching in wireless device-to-device network can effectively decrease the mobile traffic during peak hours. Currently, most of the related work is focused on how to cache content on SBSs and on mobile devices, and it is assumed that the user can download the entire requested content through the connected SBSs and mobile devices. However, few works have taken user mobility and the randomness of contact duration into consideration. How to improve the caching strategy by exploiting user mobility is still a challenging problem. Thus, in this paper, we first investigate the problem of how to conduct caching placement on SBS and on mobile devices leveraging user mobility, aiming to maximize the cache hit ratio. Specifically, the caching placement on SBSs and on mobile devices is formulated as an integer programming problem, and submodular optimization is adopted to solve the formulated problem. Then, we give the optimal transmission power of SBSs and mobile devices to deliver the caching content in order to reduce the energy cost. Simulation results prove that our caching strategy is more efficient than other existing caching strategies in terms of both cache hit ratio and energy efficiency. Min Chen 0003, Yixue Hao, Long Hu, Kaibin Huang, Vincent K. N. Lau |
IEEE Trans. Wirel. Commun. | 2 |
| 2016 | User Intent-Oriented Video QoE with Emotion Detection NetworkingabstractWith the ever-growing number of users enjoying online video service in mobile environments, video streaming services have been dominating the mobile traffic. It can be predicted that a small improvement in the user's watching experience will cause a substantial leap in profitability in terms of content providers and distributors, network operators and service providers for mobile videos. Though recent years have witnessed effective efforts to improve a user's video quality of experience (QoE) by the use of big data for analyzing users' viewing behaviors based on large-scale, video- viewing history datasets, it is very challenging to precisely analyze users' hidden intents and feelings when they are watching online videos. In addition to obtain a better video QoE, we propose to introduce user's emotional reactions into QoE assessment. In this scheme, first, the user's mood is detected in a real time fashion via emotion detection networking. Then, a mood matching process is performed to gain the similarity of the user's intent and the video content property in terms of emotion design. Finally, a novel, decision tree-based adjustment model is proposed to characterize the relationship between QoE and various factors, including buffer ratio, average bitrate, and the user's emotions. Our study opens a road for improving video QoE based on emotion detection networking. Min Chen 0003, Yixue Hao, Shiwen Mao, Di Wu 0001 |
GLOBECOM | 2 |
| 2016 | Cloud-Assisted Mood Fatigue Detection System
Xiaobo Shi, Yixue Hao, Delu Zeng, M. Shamim Hossain, Sk. Md. Mizanur Rahman, Abdulhameed Alelaiwi |
Mob. Networks Appl. | 2 |
| 2015 | Demo: LIVES: Learning through Interactive Video and Emotion-aware SystemabstractIn order to improve the accuracy and efficiency of emotion recognition, we design a novel system called Learning through Interactive Video and Emotion-aware System (LIVES). LIVES includes data collection, emotion recognition, and result validation, as well as emotion feedback. We adopt transfer learning to label and validate moods in LIVES, while the emotion can be classified into six types of mood in a reasonable accuracy. Through transfer learning, the time-consuming and labor-intensive processing cost on data collection and labeling can also be greatly reduced. In our prototype system, LIVES is used to enhance an emotion-aware robot's intelligence provided by cloud. LIVES-based emotion recognition is executed in the remote cloud while corresponding result is sent to the robot for emotion feedback. The experimental results demonstrate LIVES significantly improves the accuracy and effective of emotion classification. Min Chen 0003, Yixue Hao, Yong Li 0008, Di Wu 0001, Dijiang Huang |
MobiHoc | 2 |