EDBT 2026 Demo / reviewers in the wild / expert
Long Hu
dblp:130/4705
· DBLP profile ↗
63ranked-venue papers
9as first author
39since 2021 · last 2026
0000-0001-6496-6793ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 20 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 18 since 2021Systems, architecture and hardware · 11 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-modal model partition strategy for end-edge collaborative inference
Dongkun Huo, Yingting Zhou, Yixue Hao, Long Hu, Yijun Mo, Min Chen 0003, Iztok Humar |
J. Parallel Distributed Comput. | 4 |
| 2026 | Rethinking Point Cloud Representation Learning for Freeing Transformer to Perceive LocalabstractTransformers are widely utilized in the point cloud domain. However, existing methods tend to overburden Transformer with the dual task of local geometric perception and global feature extraction, limiting its ability to capture highlevel semantic knowledge. To address this issue, we present Representation Decoder (R-Decoder), a novel representation extraction module compatible with various point cloud Transformer methods, enabling the Transformer to focus on its excellent local perception. The R-Decoder iteratively extracts multiple global features from tokens generated by Transformer, refining them to construct an overall representation of point cloud. To ensure full adaptation of the R-Decoder to the knowledge of pre-trained Transformers, we design a cross-modal representation alignment task that leverages multimodal knowledge to specifically pre-train the R-Decoder. As a post-processing module, the R-Decoder seamlessly integrates with Transformers, while decoupling local perception and global representation. This design allows the Transformer to focus on the semantic encoding role for point tokens. Extensive experiments show that our RDecoder significantly boosts the capabilities of 3D representation learning in various point cloud Transformer methods. Notably, it achieves impressive classification accuracies of 95.1% on the ScanObjectNN dataset and 95.3% on the ModelNet40 dataset. Moreover, our method obtains new SOTA on all benchmarks of few-shot and zero-shot classification, while enhancing the multimodal task capabilities of pre-trained Transformers. Code and weights are available athttps://github.com/TangYuan96/RDecoder. Yunlong Yu 0002, Xianzhi Li 0001, Rui Wang 0077, Jinfeng Xu 0002, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
IEEE Trans. Multim. | 8 |
| 2026 | DTSNet: Dynamic Transformer Slimming for Efficient Vision RecognitionabstractTransformer-based models have recently adopted increasingly complex structure (e.g., deeper or wider stacked network) to promote the representation learning capabilities of vision recognition. However, progressively deeper or wider stacked network cause the expensive computation cost, which hinders their effective deployment in resource-constrained edge clouds or end devices. In this paper, we propose DTSNet, a dynamic transformer slimming model, which scales vision transformers (ViTs) down across layers from both of the model depth and input width. This is the first time to explore the joint reduction of input tokens and model parameters for ViTs under maintaining performance. Specifically, DTSNet adopts a diversity-enhanced weight sharing module to reduce network parameters, where the weight knowledge of multiple adjacent blocks is effectively integrated into one block. Furthermore, DTSNet designs a unified and massively scalable token pruning mechanism that dynamically discarding less important tokens with a model-driven manner, by introducing a series of discriminant parameters, which is a simple change to the common architecture of vision transformers. Extensive experiments are conducted to verify that DTSNet is able to yield high efficacy in compressing parameter space and accelerating model inference. DTSNet-T/-S/-B on ImageNet achieves 3.0M/11.1M/42.9M parameters and 0.8/2.9/13.7 GFLOPs, where number of parameters are reduced by 48%$\sim$51% and inference speed are improved by 1.3$\times \sim 1.5\times$. Experiments results on semantic segmentation and object detection dataset further demonstrate the potential of DTSNet on complex dense prediction tasks. Code will be available upon publication. Wenjing Xiao, Xianzhi Li 0001, Long Hu, Yixue Hao, Min Chen 0003 |
IEEE Trans. Multim. | 3 |
| 2026 | Generative Aspect-Based Sentiment Quadruple Prediction Based on Multi-Order PromptingabstractRecently, generative aspect-level sentiment quadruple prediction (ASQP) methods based on pre-trained language models have made significant progress. However, some challenges remain in extracting and recognizing complex sentiment elements from semantically rich sentences, limiting the generalization and adaptability of unidirectional generative models in aspect-level sentiment analysis. To overcome this limitation, this article proposes a Generative Aspect-Based Sentiment Quadruple Prediction Model based on Multi-Order Prompting (GenMOP). The model draws on the concept of prompt learning and introduces a multi-order prompting strategy, which breaks the traditional framework of a single generative order and enhances the flexibility and adaptability of the model. Furthermore, we integrate a quadruple quantity-aware module and a multi-view uncertainty-aware module based on a basic generative architecture, not only providing the model with more fine-grained information about the quadruple quantity but also improving the prediction accuracy through uncertainty estimation. The extensive experiments show that the GenMOP method achieves excellent performance in the ASQP task. On the four benchmark datasets including Rest15, Rest16, Rest and Lap, our model achieves F1 score improvements of 1.26%, 0.23%, 0.28%, and 2.07%, respectively, compared to existing state-of-the-arts, demonstrating its effectiveness and superiority in dealing with the joint extraction of multiple sentiment elements of the ASQP model. Rui Wang 0077, Muyao He, Yixue Hao, Long Hu, Min Chen 0003, Baoru Huang |
ACM Trans. Inf. Syst. | 5 |
| 2026 | Context-Aware AIGC Service Migration in Edge Intelligence Networks via Transformer DRLabstractWith the increasing demand for artificial intelligence generated content (AIGC) services across diverse applications, AIGC service migration is essential to ensuring continuous service for mobile users in edge intelligence networks. However, AIGC service migration can lead to decreased inference accuracy due to the discarding of contextual memory. Furthermore, migrating large-scale AIGC models incurs high migration costs and latency. In this paper, we propose a context-aware AIGC service migration scheme to address the trade-off among inference accuracy, latency, and migration cost. Specifically, we focus on migrating historical AIGC context rather than large-scale AIGC models to achieve cost-efficient service provisioning. To improve service migration performance, we propose a Value of Context (VoC) metric to quantify the relevance and freshness of historical AIGC context. Based on the VoC, we formulate an optimization problem to jointly optimize inference accuracy, latency, and migration cost. To solve this problem, we develop a TransFormer-based Soft actor-critic algorithm for Context-aware AIGC service Migration (TFSCM) that leverages long-term dependencies in historical decisions for optimizing the migration process. Extensive experiments on real-world datasets demonstrate that the proposed TFSCM algorithm significantly enhances system performance compared to baseline solutions. Yixue Hao, Rui Wang 0077, Long Hu, Kaibin Huang, Dusit Niyato, Min Chen 0003 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | More Text, Less Point: Towards 3D Data-Efficient Point-Language UnderstandingabstractEnabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge. Due to the lack of large-scale 3D-text pair datasets, the success of LLMs has yet to be replicated in 3D understanding. In this paper, we rethink this issue and propose a new task: 3D Data-Efficient Point-Language Understanding. The goal is to enable LLMs to achieve robust 3D object understanding with minimal 3D point cloud and text data pairs. To address this task, we introduce GreenPLM, which leverages more text data to compensate for the lack of 3D data. First, inspired by using CLIP to align images and text, we utilize a pre-trained point cloud-text encoder to map the 3D point cloud space to the text space. This mapping leaves us to seamlessly connect the text space with LLMs. Once the point-text-LLM connection is established, we further enhance text-LLM alignment by expanding the intermediate text space, thereby reducing the reliance on 3D point cloud data. Specifically, we generate 6M free-text descriptions of 3D objects, and design a three-stage training strategy to help LLMs better explore the intrinsic connections between different modalities. To achieve efficient modality alignment, we design a zero-parameter cross-attention module for token pooling. Extensive experimental results show that GreenPLM requires only 12% of the 3D training data used by existing state-of-the-art models to achieve superior 3D understanding. Remarkably, GreenPLM also achieves competitive performance using text-only data. Xu Han 0016, Xianzhi Li 0001, Qiao Yu 0002, Jinfeng Xu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
AAAI | 7 |
| 2025 | SASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point CloudsabstractRecent advancements in deep learning have greatly enhanced 3D object recognition, but most models are limited to closed-set scenarios, unable to handle unknown samples in real-world applications. Open-set recognition (OSR) addresses this limitation by enabling models to both classify known classes and identify novel classes. However, current OSR methods rely on global features to differentiate known and unknown classes, treating the entire object uniformly and overlooking the varying semantic importance of its different parts. To address this gap, we propose Salience-Aware Structured Separation (SASep), which includes (i) a tunable semantic decomposition (TSD) module to semantically decompose objects into important and unimportant parts, (ii) a geometric synthesis strategy (GSS) to generate pseudo-unknown objects by combining these unimportant parts, and (iii) a synth-aided margin separation (SMS) module to enhance feature-level separation by expanding the feature distributions between classes. Together, these components improve both geometric and feature representations, enhancing the model’s ability to effectively distinguish known and unknown classes. Experimental results show that SASep achieves superior performance in 3D OSR, outperforming existing state-of-the-art methods. The codes are available at https://github.com/JinfengX/SASep. Jinfeng Xu 0002, Xianzhi Li 0001, Xu Han 0016, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
CVPR | 7 |
| 2025 | Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play DeformationabstractGenerating 3D meshes from a single image is an important but ill-posed task. Existing methods mainly adopt 2D multiview diffusion models to generate intermediate multiview images, and use the Large Reconstruction Model (LRM) to create the final meshes. However, the multiview images exhibit local inconsistencies, and the meshes often lack fidelity to the input image or look blurry. We propose Fancy123, featuring two enhancement modules and an unprojection operation to address the above three issues, respectively. The appearance enhancement module deforms the 2D multiview images to realign misaligned pixels for better multiview consistency. The fidelity enhancement module deforms the 3D mesh to match the input image. The unprojection of the input image and deformed multiview images onto LRM’s generated mesh ensures high clarity, discarding LRM’s predicted blurry-looking mesh colors. Extensive qualitative and quantitative experiments verify Fancy123’s SoTA performance with significant improvement. Also, the two enhancement modules are plug-and-play and work at inference time, allowing seamless integration into various existing single-image-to-3D methods. Project page: https://github.com/YuQiao0303/Fancy123. Qiao Yu 0002, Xianzhi Li 0001, Xu Han 0016, Long Hu, Yixue Hao, Min Chen 0003 |
CVPR | 5 |
| 2025 | Fusion-PSRO: Nash Policy Fusion for Policy Space Response OraclesabstractFor solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown that the Policy Space Response Oracles (PSRO) algorithm is an effective framework for solving such games. However, current methods initialize a new policy from scratch or inherit a single historical policy for Best Response (BR), missing the opportunity to leverage past policies to generate a better BR. In this paper, we propose Fusion-PSRO, which employs Nash Policy Fusion to initialize a new policy for BR training. Nash Policy Fusion serves as an implicit guiding policy that starts exploration on the current Meta-NE, thus providing a closer approximation to BR. Moreover, it insightfully captures a weighted moving average of past policies, dynamically adjusting these weights based on the Meta-NE in each iteration. This cumulative process further enhances the policy population. Empirical results on classic benchmarks show that Fusion-PSRO achieves lower exploitability, thereby mitigating the shortcomings of previous research on policy initialization in BR. Jiesong Lian, Yucong Huang, Chengdong Ma, Ying Wen 0001, Long Hu, Yixue Hao |
ECAI | 6 |
| 2025 | PointDreamer: Zero-Shot 3D Textured Mesh Reconstruction From Colored Point CloudabstractFaithfully reconstructing textured meshes is crucial for many applications. Compared to text or image modalities, leveraging 3D colored point clouds as input (colored-PC-to-mesh) offers inherent advantages in comprehensively and precisely replicating the target object's 360$^{\circ }$∘ characteristics. While most existing colored-PC-to-mesh methods suffer from blurry textures or require hard-to-acquire 3D training data, we propose PointDreamer, a novel framework that harnesses 2D diffusion prior for superior texture quality. Crucially, unlike prior 2D-diffusion-for-3D works driven by text or image inputs, PointDreamer successfully adapts 2D diffusion models to 3D point cloud data by a novel project-inpaint-unproject pipeline. Specifically, it first projects the point cloud into sparse 2D images and then performs diffusion-based inpainting. After that, diverging from most existing 3D reconstruction or generation approaches that predict texture in 3D/UV space thus often yielding blurry texture, PointDreamer achieves high-quality texture by directly unprojecting the inpainted 2D images to the 3D mesh. Furthermore, we identify for the first time a typical kind of unprojection artifact appearing in occlusion borders, which is common in other multiview-image-to-3D pipelines but less-explored. To address this, we propose a novel solution named the Non-Border-First (NBF) unprojection strategy. Extensive qualitative and quantitative experiments on various synthetic and real-scanned datasets demonstrate that PointDreamer, though zero-shot, exhibits SoTA performance (30% improvement on LPIPS score from 0.118 to 0.068), and is robust to noisy, sparse, or even incomplete input data. Qiao Yu 0002, Xianzhi Li 0001, Xu Han 0016, Jinfeng Xu 0002, Long Hu, Min Chen 0003 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | JIMR: Joint Semantic and Geometry Learning for Point Scene Instance Mesh ReconstructionabstractPoint scene instance mesh reconstruction is a challenging task since it requires both scene-level instance segmentation and instance-level mesh reconstruction from partial observations simultaneously. Previous works either adopt a detection backbone or a segmentation one, and then directly employ a mesh reconstruction network to produce complete meshes from incomplete instance point clouds. To further boost the mesh reconstruction quality with both local details and global smoothness, in this work, we propose JIMR, a joint framework with two cascaded stages for semantic and geometry understanding. In the first stage, we propose to perform both instance segmentation and object detection simultaneously. By making both tasks promote each other, this design facilitates subsequent mesh reconstruction by providing more precisely-segmented instance points and better alignment benefiting from predicted complete bounding boxes. In the second stage, we propose a complete-then-reconstruct procedure, where the completion module explicitly disentangles completion from reconstruction, and enables the usage of pre-trained weights of existing powerful completion and reconstruction networks. Moreover, we propose a comprehensive confidence score to filter proposals considering the quality of instance segmentation, bounding box detection, semantic classification, and mesh reconstruction at the same time. Experiments show that our proposed JIMR outperforms state-of-the-art methods regarding instance reconstruction qualitatively and quantitatively. Qiao Yu 0002, Xianzhi Li 0001, Jinfeng Xu 0002, Long Hu, Yixue Hao, Min Chen 0003 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | SE-DCFN: Semantic-Enhanced Dual Cross-modal Fusion Network for Depression RecognitionabstractAutomatic multi-modal depression recognition using artificial intelligence technology is crucial to advance early diagnosis and treatment. Existing methods suffer from a weak performance in detecting depression due to incomplete unimodal semantic information and insufficient fusion effects. To address these challenges, we propose a novel Semantic-Enhanced Dual Cross-modal Fusion Network (SE-DCFN) for multi-modal depression recognition, specifically designed for text-audio data. Firstly, we utilize a prompt learning-based text encoder and a language-audio pertaining-based audio encoder to capture specific information to enhance the semantic representation. Then, we introduce a dual cross-modal fusion module based on self-attention and cross-attention mechanisms to effectively explore linguistic and acoustic representation, facilitating inter-modal and intra-modal interaction and fusion. Additionally, a triplet contrastive loss is formulated to optimize the training process of the SE-DCFN. Experimental results on the EATD-Corpus dataset and AVEC-2017 dataset demonstrate the effectiveness and superiority of our proposed SE-DCFN on multi-modal depression recognition, outperforming existing methods. Long Hu, Qingyi Yang, Rui Wang 0077, Yixue Hao, Min Chen 0003, Yijun Mo |
BIBM | 1 |
| 2024 | PDF: A Probability-Driven Framework for Open World 3D Point Cloud Semantic SegmentationabstractExisting point cloud semantic segmentation networks cannot identify unknown classes and update their knowledge, due to a closed-set and static perspective of the real world, which would induce the intelligent agent to make bad decisions. To address this problem, we propose a Probability-Driven Framework (PDF)11Code available at: https://github.com/JinfengX/PointCloudPDF. for open world semantic segmentation that includes (i) a lightweight U-decoder branch to identify unknown classes by estimating the uncertainties, (ii) a flexible pseudo-labeling scheme to supply geometry features along with probability distribution features of unknown classes by generating pseudo labels, and (iii) an incremental knowledge distillation strategy to incorporate novel classes into the existing knowledge base gradually. Our framework enables the model to behave like human beings, which could recognize unknown objects and incrementally learn them with the corresponding knowledge. Experimental results on the S3DIS and ScanNetv2 datasets demonstrate that the proposed PDF outperforms other methods by a large margin in both important tasks of open world semantic segmentation. Jinfeng Xu 0002, Xianzhi Li 0001, Yixue Hao, Long Hu, Min Chen 0003 |
CVPR | 6 |
| 2024 | Multimodal Physiological Signals Representation Learning via Multiscale Contrasting for Depression Recognition
Kai Shao, Rui Wang 0077, Yixue Hao, Long Hu, Min Chen 0003, Hans-Arno Jacobsen |
ACM Multimedia | 4 |
| 2024 | MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D PriorsabstractLarge 2D vision-language models (2D-LLMs) have gained significant attention by bridging Large Language Models (LLMs) with images using a simple projector. Inspired by their success, large 3D point cloud-language models (3D-LLMs) also integrate point clouds into LLMs. However, directly aligning point clouds with LLM requires expensive training costs, typically in hundreds of GPU-hours on A100, which hinders the development of 3D-LLMs. In this paper, we introduce MiniGPT-3D, an efficient and powerful 3D-LLM that achieves multiple SOTA results while training for only 27 hours on one RTX 3090. Specifically, we propose to align 3D point clouds with LLMs using 2D priors from 2D-LLMs, which can leverage the similarity between 2D and 3D visual information. We introduce a novel four-stage training strategy for modality alignment in a cascaded way, and a mixture of query experts module to adaptively aggregate features with high efficiency. Moreover, we utilize parameter-efficient fine-tuning methods LoRA and Norm fine-tuning, resulting in only 47.8M learnable parameters, which is up to 260x fewer than existing methods. Extensive experiments show that MiniGPT-3D achieves SOTA on 3D object classification and captioning tasks, with significantly cheaper training costs. Notably, MiniGPT-3D gains an 8.12 increase on GPT-4 evaluation score for the challenging object captioning task compared to ShapeLLM-13B, while the latter costs 160 total GPU-hours on 8 A800. We are the first to explore the efficient 3D-LLM, offering new insights to the community. Code and weights are available at https://github.com/TangYuan96/MiniGPT-3D. Xu Han 0016, Xianzhi Li 0001, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
ACM Multimedia | 6 |
| 2024 | DCTracker: Rethinking MOT in soccer events under dual views via cascade association
Long Hu, Junjie Zhang 0002, Weiyi Lv, Yongshun Gong, Jingya Wang 0001, Jian Zhang 0002, Dan Zeng 0001 |
Knowl. Based Syst. | 1 |
| 2024 | GaitASMS: gait recognition by adaptive structured spatial representation and multi-scale temporal aggregation
Long Hu, Xueling Feng, Mark S. Nixon |
Neural Comput. Appl. | 2 |
| 2024 | Efficient Crowd Counting via Dual Knowledge DistillationabstractMost researchers focus on designing accurate crowd counting models with heavy parameters and computations but ignore the resource burden during the model deployment. A real-world scenario demands an efficient counting model with low-latency and high-performance. Knowledge distillation provides an elegant way to transfer knowledge from a complicated teacher model to a compact student model while maintaining accuracy. However, the student model receives the wrong guidance with the supervision of the teacher model due to the inaccurate information understood by the teacher in some cases. In this paper, we propose a dual-knowledge distillation (DKD) framework, which aims to reduce the side effects of the teacher model and transfer hierarchical knowledge to obtain a more efficient counting model. First, the student model is initialized with global information transferred by the teacher model via adaptive perspectives. Then, the self-knowledge distillation forces the student model to learn the knowledge by itself, based on intermediate feature maps and target map. Specifically, the optimal transport distance is utilized to measure the difference of feature maps between the teacher and the student to perform the distribution alignment of the counting area. Extensive experiments are conducted on four challenging datasets, demonstrating the superiority of DKD. When there are only approximately 6% of the parameters and computations from the original models, the student model achieves a faster and more accurate counting performance as the teacher model even surpasses it. Rui Wang 0077, Yixue Hao, Long Hu, Xianzhi Li 0001, Min Chen 0003, Yiming Miao, Iztok Humar |
IEEE Trans. Image Process. | 3 |
| 2024 | Reliable or Green? Continual Individualized Inference Provisioning in Fabric Metaverse via Multi-Exit AccelerationabstractFabric metaverse employs intelligence fibers embedded with flexible sensors to unknowingly gather and transmit massive hypermodal data around humans to a deep neural network-based metaverse inference service (DMS) for continual and real-time analysis. Each DMS has one primary branch and multiple side branches that allow early termination of service with differential accuracy and energy consumption. However, the continual provisioning of compute-intensive DMS with varying requirements for service model, accuracy, delay, and reliability poses a challenge for edge servers characterized by restricted computing resources and intermittent green energy. In this paper, we focus on a continual individualized DMS provisioning problem in the fabric metaverse consisting of a side branch insertion subproblem and a server activation and service deployment subproblem, and formulate them as Integer linear Programming and Markov Decision Process, respectively. Then, we propose a green continual inference (GCI) system, where a pruner with provable approximation ratios trims superfluous branches of every model to the given number$K$to minimize total overflow accuracy between accuracy demands and reserved branches assigned to users. Based on this exit result, each DMS is further divided into several blocks with dependencies to exploit constrained resources of computing and energy in a fine-grained manner. Finally, a learning-based scheduler is merged into GCI to maximize request throughput while minimizing the activation number of edge servers on different demand scenarios, by adaptively activating suitable servers and deploying required blocks and their corresponding backups on selected servers. Theoretical analyses, simulations, and experiments demonstrate that the GCI is promising compared with baseline algorithms. Min Chen 0003, Weifa Liang, Dusit Niyato, Yue Wang 0092, Victor C. M. Leung, Yixue Hao, Long Hu, Yin Zhang 0002 |
IEEE Trans. Mob. Comput. | 9 |
| 2024 | Point-LGMask: Local and Global Contexts Embedding for Point Cloud Pre-Training With Multi-Ratio MaskingabstractSelf-supervised learning has achieved great success in both natural language processing and 2D vision, where masked modeling is a quite popular pre-training scheme. However, extending masking to 3D point cloud understanding that combines local and global features poses a new challenge. In our work, we present Point-LGMask, a novel method to embed both local and global contexts with multi-ratio masking, which is quite effective for self-supervised feature learning of point clouds but is unfortunately ignored by existing pre-training works. Specifically, to avoid fitting to a fixed masking ratio, we first propose multi-ratio masking, which prompts the encoder to fully explore representative features thanks to tasks of different difficulties. Next, to encourage the embedding of both local and global features, we formulate a compound loss, which consists of (i) a global representation contrastive loss to encourage the cluster assignments of the masked point clouds to be consistent to that of the completed input, and (ii) a local point cloud prediction loss to encourage accurate prediction of masked points. Equipped with our Point-LGMask, we show that our learned representations transfer well to various downstream tasks, including few-shot classification, shape classification, object part segmentation, as well as real-world scene-based 3D object detection and 3D semantic segmentation. Particularly, our model largely advances existing pre-training methods on the difficult few-shot classification task using the real-captured ScanObjectNN dataset by surpassing over 4% to the second-best method. Also, our Point-LGMask achieves 0.4%$AP_{25}$and 0.8%$AP_{50}$gains on 3D object detection task over the second-best method. 0.4% mAcc and 0.5% mIoU. Codes have been released athttps://github.com/TangYuan96/Point-LGMask. Xianzhi Li 0001, Jinfeng Xu 0002, Qiao Yu 0002, Long Hu, Yixue Hao, Min Chen 0003 |
IEEE Trans. Multim. | 5 |
| 2024 | Immersive Multimedia Service Caching in Edge Cloud with Renewable EnergyabstractImmersive service caching, based on the intelligent edge cloud, can meet delay-sensitive service requirements. Although numerous service caching solutions for edge clouds have been designed, they have not been well explored. Moreover, to the best of our knowledge, there is no work to consider the immersive service caching scheme under the supply of renewable energy. In this article, we investigate the service caching problem under the renewable energy supply to minimize service latency while making full use of renewable energy. Specifically, we formulate the service caching and renewable energy harvesting problem, which considers the dynamic renewable energy, unknown service requests, and limited capacity of the edge cloud. To solve this problem, we propose an effective algorithm, called OSCRE. Our algorithm first uses Lyapunov optimization to convert the time-average problem into time-independence optimization and thus realizes optimal renewable energy harvesting. Then, it realizes the service caching scheme using data-driven combinatorial multi-armed bandit learning. The simulation results show that the OSCRE scheme can save service latency while making sufficient use of renewable energy. M. Shamim Hossain, Yixue Hao, Long Hu, Jia Liu 0009, Min Chen 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | RT3C: Real-Time Crowd Counting in Multi-Scene Video Streams via Cloud-Edge-Device CollaborationabstractRecently, the advancements in edge computing have boosted the deployment of video analysis systems based on deep learning, which breaks the limitation of the constrained communication and computing resources of local devices. However, processing multi-scene high-resolution video streams in crowd surveillance remains a significant challenge since it is difficult to formulate dynamic video content and communication environments to support offloading decisions. To bridge the gap between applications and modeling, this paper presents aReal-TimeCloud-edge-deviceCollaboration framework, which enables fast and accurateCrowd counting (RT3C) on the real dataset. RT3C comprises key frame detection, adaptive patch partition, patch encoder and decoder and computation offloading decision, designed to divide key frames into a minimum number of patches and determine the offloading location of patches. A Real-Time Multi-Agent Actor-Critic (RTMAAC) algorithm based on multi-agent reinforcement learning is proposed to decide whether to compute patches with a lightweight model on edge or a large model on cloud. Unlike traditional approaches ignoring the contents, RTMAAC is a dynamic online decision algorithm based on context of the network and video. Extensive experiments demonstrate that RT3C effectively discriminate the valid frames and optimizes offloading decisions in complex environments, outperforming other baseline algorithms on the two crowd counting datasets. In summary, RT3C provides a promising framework for multi-scene video streams, which can be extended to other applications to realize video computation based on deep models. Rui Wang 0077, Yixue Hao, Yiming Miao, Long Hu, Min Chen 0003 |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | CasFusionNet: A Cascaded Network for Point Cloud Semantic Scene Completion by Dense Feature FusionabstractSemantic scene completion (SSC) aims to complete a partial 3D scene and predict its semantics simultaneously. Most existing works adopt the voxel representations, thus suffering from the growth of memory and computation cost as the voxel resolution increases. Though a few works attempt to solve SSC from the perspective of 3D point clouds, they have not fully exploited the correlation and complementarity between the two tasks of scene completion and semantic segmentation. In our work, we present CasFusionNet, a novel cascaded network for point cloud semantic scene completion by dense feature fusion. Specifically, we design (i) a global completion module (GCM) to produce an upsampled and completed but coarse point set, (ii) a semantic segmentation module (SSM) to predict the per-point semantic labels of the completed points generated by GCM, and (iii) a local refinement module (LRM) to further refine the coarse completed points and the associated labels from a local perspective. We organize the above three modules via dense feature fusion in each level, and cascade a total of four levels, where we also employ feature fusion between each level for sufficient information usage. Both quantitative and qualitative results on our compiled two point-based datasets validate the effectiveness and superiority of our CasFusionNet compared to state-of-the-art methods in terms of both scene completion and semantic segmentation. The codes and datasets are available at: https://github.com/JinfengX/CasFusionNet. Jinfeng Xu 0002, Xianzhi Li 0001, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003 |
AAAI | 6 |
| 2023 | Data Augmentation and Pseudo-sequence of fNIRS for Depression RecognitionabstractDepression is a mental disorder caused by factors such as genetics, life events and social influences, and has become a major public health problem worldwide. Previous studies have demonstrated the potential of functional near-infrared spectroscopy (fNIRS) in the diagnosis of depression. However, in the real medical scene, fNIRS data are difficult to obtain, limited in number and suffer from class imbalance. To overcome these problems, in this paper, we propose a novel model for depression identification based on data augmentation and pseudo-sequence of fNIRS. Specifically, the data augmentation using the time masking and warping method generates richer data. Then, a stimulation task-driven data pseudo-sequence method is designed to map the sequence data into pseudo-sequence activation images. Finally, a depression recognition model is established based on the class imbalance loss function. Experiments show that the precision of our depression recognition model reaches 0.94. This scheme transforms fNIRS data into image sequences, which provides a new solution idea for subsequent research. Kai Shao, Yixue Hao, Long Hu, Xiaofen Zong, Min Chen 0003 |
BIBM | 3 |
| 2023 | Crowd Intelligent Grouping Collaboration Evacuation via Multi-agent Reinforcement LearningabstractThe crowd evacuation strategy seeks to arrange crowd evacuation in an orderly manner to protect people’s lives and reduce property damage in case of sudden emergencies in crowded and complex places. In recent years, there have been several works to apply deep learning to crowd evacuation to make evacuation strategies more intelligent. However, existing researches rarely consider the integration of scene perception and crowd evacuation, which leads to evacuation methods that are detached from the scene and also ignore the crowd collaboration in the evacuation process. To this end, we propose Intelligent Crowd Evacuation Architecture based on Visual features using Multi-Agent Reinforcement Learning (ICEA-VMARL). Subsequently, we present modeling analysis on the crowd grouping and group evacuation modules of the architecture. First, we propose the Population Grouping algorithm based on Continuous Spatiotemporal individual Similarity (PGCSS), which combines crowd features to group crowds. Then, we propose a Group Collaborative Evacuation algorithm based on Multi-Agent Reinforcement Learning (GCE-MARL), which considers group collaboration while evacuating to achieve global optimal evacuation. Finally, we build an experimental crowd simulation system, and the results demonstrate that the crowd grouping algorithm and group evacuation algorithm proposed have better performance compared with other methods. Rui Wang 0077, Jinfeng Xu 0002, Long Hu, Yixue Hao |
CSCWD | 5 |
| 2023 | TriGait: Aligning and Fusing Skeleton and Silhouette Gait Data via a Tri-Branch NetworkabstractGait recognition is a promising biometric technology for identification due to its non-invasiveness and long-distance. However, external variations such as clothing changes and viewpoint differences pose significant challenges to gait recognition. Silhouette-based methods preserve body shape but neglect internal structure information, while skeleton-based methods preserve structure information but omit appearance. To fully exploit the complementary nature of the two modalities, a novel triple branch gait recognition framework, TriGait, is proposed in this paper. It effectively integrates features from the skeleton and silhouette data in a hybrid fusion manner, including a two-stream network to extract static and motion features from appearance, a simple yet effective module named JSA-TC to capture dependencies between all joints, and a third branch for cross-modal learning by aligning and fusing low-level features of two modalities. Experimental results demonstrate the superiority and effectiveness of TriGait for gait recognition. The proposed method achieves a mean rank-1 accuracy of 96.0% over all conditions on CASIA-B dataset and 94.3% accuracy for CL, significantly outperforming all the state-of-the-art methods. The source code will be available at https://github.com/feng-xueling/TriGait/. Xueling Feng, Liyan Ma, Long Hu, Mark S. Nixon |
IJCB | 4 |
| 2023 | Joint Sensing Adaptation and Model Placement in 6G Fabric ComputingabstractSensing and computing based on intelligent fabrics can meet the ultra-reliable and low-latency communication (URLLC) needs of sixth-generation wireless (6G) by integrating sensing units into fabric fibers to perceive user data. Although some researchers have designed sensing or computing solutions, such solutions have not been well explored. In this paper, we consider the joint sensing adaptation and model placement in a 6G fabric space. We first propose an intelligent-fiber-driven 6G fabric computing network to minimize acquisition latency while ensuring accuracy. Then, we formulate an optimization model that takes the fabric sampling rate, sampling density, and model placement as variables. To solve the model, we propose an effective learning algorithm based on deep reinforcement learning. That is, by transforming the optimization problem into a state space, action space, and reward function, we design an optimal sensing and placement scheme. The simulation results show that our proposed scheme can achieve optimal sensing and computing compared with several baseline algorithms. Yixue Hao, Long Hu, Min Chen 0003 |
IEEE J. Sel. Areas Commun. | 2 |
| 2023 | Digital Twin-Assisted URLLC-Enabled Task Offloading in Mobile Edge Network via Robust Combinatorial OptimizationabstractDigital twin (DT)-assisted mobile edge network can achieve energy-efficient task offloading by optimizing the decision-making in real time. Although many DT-assisted task offloading solutions in mobile edge networks have been designed, stochastic asynchronizations between the DTs and physical entities are still ignored. In this paper, we investigate a task offloading problem in a DT-assisted URLLC-enabled mobile edge network which considered the uncertain deviation between DT estimated values and physical actual values. Specifically, we formulate a latency and energy consumption minimization problem by optimizing task offloading, resource allocation, and power management. To solve this problem, we propose a DT-assisted robust task offloading scheme (DTRTO) based on learning composed of decision and deviation networks. The deviation network predicts the worst-case deviations based on the pre-decision, and the decision network optimize the decision considered the worst-case deviation. The simulation results show that, compared to the baseline algorithms, the DTRTO scheme can realize low latency and energy consumption in task offloading while maintaining high robustness. Yixue Hao, Dongkun Huo, Nadra Guizani, Long Hu, Min Chen 0003 |
IEEE J. Sel. Areas Commun. | 5 |
| 2023 | DecLog: Decentralized Logging in Non-Volatile Memory for Time Series Database SystemsabstractGrowing demands for the efficient processing of extreme-scale time series workloads call for more capable time series database management systems (TSDBMS). Specifically, to maintain consistency and durability of transaction processing, systems employ write-ahead logging (WAL) whereby transactions are committed only after the related log entries are flushed to disk. However, when faced with massive I/O, this becomes a throughput bottleneck. Recent advances in byte-addressable Non-Volatile Memory (NVM) provide opportunities to improve logging performance by persisting logs to NVM instead. Existing studies typically track complex transaction dependencies and use barrier instructions of NVM to ensure log ordering. In contrast, few studies consider the heavy-tailed characteristics of time series workloads, where most transactions are independent of each other. We propose DecLog, a decentralized NVM-based logging system that enables concurrent logging of TSDBMS transactions. Specifically, we propose data-driven log sequence numbering and relaxed ordering strategies to track transaction dependencies and resolve serialization issues. We also propose a parallel logging method to persist logs to NVM after being compressed and aligned. An experimental study on the YCSB-TS benchmark offers insight into the performance properties of DecLog, showing that it improves throughput by up to 4.6× while offering lower recovery time in comparison to the open source TSDBMS Beringei. Bolong Zheng, Yongyong Gao, Jingyi Wan, Lingsen Yan, Long Hu, Yunjun Gao, Xiaofang Zhou 0001, Christian S. Jensen |
Proc. VLDB Endow. | 5 |
| 2023 | Self-Supervised Learning With Data-Efficient Supervised Fine-Tuning for Crowd CountingabstractDue to the expensive and laborious annotations of labeled data required by fully-supervised learning in the crowd counting task, it is desirable to explore a method to reduce the labeling burden. There exists a large number of unlabeled images in the wild that can be easily obtained compared to labeled datasets. Based on the characteristics of consistent spatial transformation with the annotations of heads and image, this paper proposes a self-supervised learning framework with unlabeled and limited labeled data for pre-training and fine-tuning crowd counting model (SSL-FT). It includes an online network and a target network that receive the same images but are randomly processed by two defined augmentation transformations. We leverage unlabeled data to pre-train the online network based on a self-supervised loss and small-scale labeled data to transfer the model to a specific domain based on a fully-supervised loss. We demonstrate the effectiveness of the SSL-FT on four public datasets including ShanghaiTech PartA, PartB, UCF-QNRF and WorldExpo'10 utilizing a classical counting model. Experimental results show that our approach performs better than state-of-art semi-supervised methods. Rui Wang 0077, Yixue Hao, Long Hu, Jincai Chen, Min Chen 0003, Di Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | GNN-Based Depression Recognition Using Spatio-Temporal Information: A fNIRS StudyabstractIn recent years, depression has become an increasingly serious problem globally. Previous studies of automatic depression recognition based on functional near-Infrared spectroscopy (fNIRS) or other brain imaging techniques have shown potential to serve as auxiliary diagnosis methods that provide assistance to medical professionals. Recently, some studies have found that, besides directly using the data themselves (temporal data), the use of functional connectivity among channels (spatial data) also can be effective. In this paper, we propose a method based on Graph Neural Network (GNN) that combines both temporal and spatial features of fNIRS data for automatic depression recognition. Specifically, fNIRS data of 96 subjects were collected and pre-processed. Basic statistical metrics of each channel were extracted as temporal features, and channel connectivity (coherence and correlation) were calculated as spatial features. Point-biserial analysis was conducted on these features and depression labels as a data-driven motivation. For classification, we considered data of each subject as a graph, with temporal features as node features and spatial features as edge weights. The graphs were fed into GNNs for training and testing. Experimental results showed that our GNN-based methods realized the best depression recognition performance compared with classical machine-learning methods regarding accuracy, F1 score, and precision, especially in F1 score for over 10%. Qiao Yu 0002, Rui Wang 0077, Jia Liu 0009, Long Hu, Min Chen 0003, Zhongchun Liu |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Negative Information Measurement at AI Edge: A New Perspective for Mental Health MonitoringabstractThe outbreak of the corona virus disease 2019 (COVID-19) has caused serious harm to people’s physical and mental health. Due to the serious situation of the epidemic, a lot of negative energy information increases people’s psychological burden. However, effective interventions against mental health problems are not in abundance. To address such challenges, in this article, we propose the concept of negative information to describe information that has a negative impact on people’s mental health. To achieve the measurement of negative information, the level of mental health inversely measures the degree of negative information. Specifically, we design a system to measure the negative information used to monitor the mental health state of the user under the impact of negative information. The cognition of mental health is realized based on the intelligent algorithm deployed on the edge cloud, and the needs of users can be responded to in real time in practical applications. Finally, we use real collected dataset to verify the influence of negative information. The experiments show that the system can achieve negative information measurement and provide an effective countermeasure for solving mental health problems during a pandemic situation. Min Chen 0003, Ke Shen 0004, Rui Wang 0077, Yiming Miao, Kai Hwang 0001, Yixue Hao, Guangming Tao, Long Hu, Zhongchun Liu |
ACM Trans. Internet Techn. | 9 |
| 2022 | A Multi-feature and Time-aware-based Stress Evaluation Mechanism for Mental Status AdjustmentabstractWith the rapid economic development, the prominent social competition has led to increasing psychological pressure of people felt from each aspect of life. Driven by the Internet of Things and artificial intelligence, intelligent psychological pressure detection systems based on deep learning and wearable devices have acquired some good results in practical application. However, existing studies argue that the psychological stress state is influenced by the current environment. They put much attention on the momentary features but ignore the dynamic change process of mental status in the time dimension. Besides, the lack of research in the general laws of psychological stress makes it difficult to quantitatively evaluate the stress status, resulting in the inability to perceive the stress state of users effectively. Thus, this article proposes an evaluation mechanism of psychological stress for adjusting the mental status of users. Specifically, we design a multi-dimensional feature space and a time-aware feature encoder, which integrate various stress features and capture time characteristics of stress state change. Moreover, a novel mental state model is proposed, which uses the pressure features with time characteristics to evaluate the pressure stress level. This model also quantifies the internal relationship between pressure features. Last, we establish a practicable testbed to demonstrate how to evaluate and adjust mental state of users by the proposed evaluation mechanism of psychological stress. Min Chen 0003, Wenjing Xiao, Yixue Hao, Long Hu, Guangming Tao |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | A Sustainable Multi-Modal Multi-Layer Emotion-Aware Service at the EdgeabstractLimited by the computational capabilities and battery energy of terminal devices and network bandwidth, emotion recognition tasks fail to achieve good interactive experience for users. The intolerable latency for users also seriously restricts the popularization of emotion recognition applications in the edge environments such as fatigue detection in auto-driving. The development of edge computing provides a more sustainable solution for this problem. Based on edge computing, this article proposes a multi-modal multi-layer emotion-aware service (MULTI-EASE) architecture that considers user’s facial expression and voice as a multi-modal data source of emotion recognition, and employs the intelligent terminal, edge server and cloud as multi-layer execution environment. By analyzing the average delay of each task and the average energy consumption at the mobile device, we formulate a delay-constrained energy minimization problem and perform a task scheduling policy between multiple layers to reduce the end-to-end delay and energy consumption by using an edge-based approach, further to improve the users’ emotion interactive experience and achieve energy saving in edge computing. Finally, a prototype system is also implemented to validate the architecture of MULTI-EASE, the experimental results show that MULTI-EASE is a sustainable and efficient platform for emotion analysis applications, and also provide a valuable reference for dynamic task scheduling under MULTI-EASE architecture. Long Hu, Wei Li 0061, Jun Yang 0014, Giancarlo Fortino, Min Chen 0003 |
IEEE Trans. Sustain. Comput. | 1 |
| 2021 | Ultra Large-Scale Crowd Monitoring System Architecture and Design IssuesabstractThis article proposes a novel ultralarge-scale crowd monitoring system, namely, the ULCM system. The ULCM system enables advanced sensing and networking technologies aimed at collecting and processing multimodal, multiperspective, and real-time crowding data relevant to crowd management. This data will be further analyzed to provide a global realization of evolving events over a large geographical area as they occur in real time. The ULCM is the infrastructure component of an intelligent platform that is being developed by our research group to provide crowd intelligence to decision makers through an interactive digitized visual environment. In order to achieve a full comprehensive scene overview, the ULCM deployment utilizes a multiplicity of unmanned aerial vehicle (UAV) agents in different operational scenarios. The aerial deployment and control are realized by custom multiple UAV networks and airborne LiDAR sensors. The deployment and control on the ground sensory agents are based on multiple subnetworks, including closed-circuit television (CCTV) and infrared gas and ultrasonic sensors networks. Eventually, ULCM employs the software-defined network (SDN) and edge cloud technologies to optimize the networking and data analytics performance from the perspective of infrastructure. Yiming Miao, Bander A. Alzahrani, Ahmed Barnawi, Reem Alotaibi, Long Hu |
IEEE Internet Things J. | 6 |
| 2021 | Smart Micro-GaS: A Cognitive Micro Natural Gas Industrial Ecosystem Based on Mixed Blockchain and Edge ComputingabstractWith the increase in natural gas consumption, distributed natural gas supply and transaction have become new development goals of the industrial Internet of Things (IoT) for natural gas. However, there are obvious disadvantages of the existing natural gas pipeline network in aspects of infrastructure warning, multilevel data transmission, automatic transaction, and security. Emerging technologies, such as blockchain, edge computing, and AI have been introduced to address these shortcomings. This article proposes Smart Micro-GaS, i.e., the concept of a cognitive micro natural gas industrial ecosystem based on mixed blockchain and edge computing. Three aspects, multilevel, multiview, and multidimension, are put forward for its design and deployment. Then, based on the most important smart contract algorithm in blockchain, a mixed transaction model for natural gas is established. Finally, a case analysis is conducted on a smart natural gas testbed for data prediction and the proposed smart contract algorithm. The framework proposed in this article makes the natural gas data have multilevel liquidity and realizes diversified transactions. Yiming Miao, Jeungeun Song 0001, Haoquan Wang, Long Hu, Mohammad Mehedi Hassan, Min Chen 0003 |
IEEE Internet Things J. | 4 |
| 2021 | Medical-Level Suicide Risk Analysis: A Novel Standard and Evaluation ModelabstractThe frequent occurrence of suicides in modern society constitutes a serious public health issue. While the motives, methods, and consequences of suicide are quite complicated, if people at risk of suicide can be identified and intervened in time, the loss of life can be reduced. Through analyses based on combining a large number of suicide texts and professional medical literature, a dictionary of potential suicide risk impact factors has been established in this article. Based on this dictionary, a novel medical-level suicide risk standard is proposed to monitor suicide risk from point-to-surface under the timeline baseline. In order to solve the problem of insufficient Chinese suicide data sets, the manually assisted method based on knowledge perception is adopted to annotate the data set with corresponding to risk level. At the same time, a Bert evaluation model based on knowledge perception was established for the classification of risk level. The experimental results showed that proposed method has a 56% recognition accuracy in the prediction of 10-Label suicide risk level proposed in this article, and the classification performance is better than traditional machine learning algorithms. Therefore, the results showed that the classification standard and evaluation model can be effectively used for the identification and early warning of suicide risk, which can discover high suicide risk groups to reduce the occurrence of suicide. It is of great significance to people’s emotion care monitoring. Rui Wang 0077, Bing Xiang Yang, Yujun Ma, Qiao Yu 0002, Xiaofen Zong, Simeng Ma, Long Hu, Kai Hwang 0001, Zhongchun Liu |
IEEE Internet Things J. | 9 |
| 2021 | Airborne LiDAR Assisted Obstacle Recognition and Intrusion Detection Towards Unmanned Aerial Vehicle: Architecture, Modeling and EvaluationabstractWith the rapid development of wireless communication and flight control technologies, the unmanned aerial vehicles (UAVs) have been widely used in multiple application scenarios. A typical scenario is massive crowd management of the multi-millions annual Hajj Pilgrimage to Mecca where UAVs are widely utilized to conduct crowd monitoring by carrying sensory devices. The safe flight of a UAV is crucial for ensuring the successful execution of missions. With the aim to overcome the disadvantage caused by the ground station intrusion detection, the combination of UAV and airborne LiDAR has been widely studied in the field of UAV obstacle recognition. This article studies the UAV network architecture under a common scenario and proposes an obstacle recognition and intrusion detection algorithm for UAV based on an airborne LiDAR (ALORID). First, the preprocessing of the data obtained by a LiDAR, i.e., the coordinate conversion of LiDAR data in combination with UAV motion parameters, is completed. Then, the LiDAR data graph at the current moment is generated by the image noisy point filtering algorithm. After that, the improved density-based spatial clustering of applications with noise (DBSCAN) algorithm is used for image clustering of intrusions to obtain the LiDAR time-domain cumulative graph in a certain detection time. Finally, the motion recognition and location detection of each cluster are completed. The experiment results verify the effectiveness of the proposed algorithm in identifying the moving state of the intrusions. Yiming Miao, Bander A. Alzahrani, Ahmed Barnawi, Tarik K. Alafif, Long Hu |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Cognitive Wearable Robotics for Autism Perception EnhancementabstractAutism spectrum disorder (ASD) is a serious hazard to the physical and mental health of children, which limits the social activities of patients throughout their lives and places a heavy burden on families and society. The developments of communication techniques and artificial intelligence (AI) have provided new potential methods for the treatment of autism. The existing treatment systems based on AI for children with ASD focus on detecting health status and developing social skills. However, the contradiction between the terminal interaction capability and availability cannot meet the needs for real application scenarios. At the same time, the lack of diverse data cannot provide individualized care for autistic children. To explore this robot-based approach, a novel AI-based first-view-robot architecture is proposed in this article. By providing care from the first-person perspective, the proposed wearable robot overcomes the difficulty of the absence of cognitive ability in the third-view of traditional robotics and improves the social interaction ability of children with ASD. The first-view-robot architecture meets the requirements of dynamic, individualized, and highly immersed interaction services for autistic children. First, the multi-modal and multi-scene data collection processes of standard, static, and dynamic datasets are introduced in detail. Then, to comprehensively evaluate the learning ability of children with ASD through mental states and external performances, a learning assessment model with emotion correction is proposed. Besides, a wearable robot-assisted environment perception and expression enhancement mechanism for children with ASD is realized by reinforcement learning, which can be adapted to interactive environments with optimal action policies. An interactive testbed for children with ASD treatments is demonstrated and experimental cases for test subjects are presented. Last, three open issues are discussed from data processing, robot designing, and service responding perspectives. Min Chen 0003, Wenjing Xiao, Long Hu, Yujun Ma, Yin Zhang 0002, Guangming Tao |
ACM Trans. Internet Techn. | 3 |
| 2020 | AI-based Satellite Ground Communication System with Intelligent Antenna PointingabstractWith the advent of the Internet era, the trend of highly informed society has been becoming more and more obvious, and the requirement of society on communication is also increasing. flexible satellite communication mode has many advantages such as large communication load and no geographic restriction, which cannot be replaced by other communication modes. In the satellite communication system, the most important is the satellite earth station (SES). When receiving signals from the target satellite, the SES terminal must accurately point to the satellite and track it to obtain the maximum receiving signal and reduce the interference with other signals simultaneously. However, the motion of either satellite or terminal can cause a change in signal intensity, so it is necessary to adjust the pointing of the SES antenna in time to maintain optimal signal receiving conditions. In order to satisfy different satellite communication scenarios, in this paper, Artificial intelligent (AI) technology is applied to the satellite communication process, mainly to optimize the optimal antenna angle and time consumption reduction. Firstly, the process of antenna pointing is introduced, and the traditional antenna search algorithm Auto-Acqire algorithm (AA algorithm) is analyzed in detail. Considering that the satellite system needs to adapt to the communication requirements of different terminals, based on AI antenna pointing algorithms are proposed. In order to verify this research, we build an experimental platform and compare the traditional AA algorithm as a benchmark algorithm with FI-GRU and II-DRL algorithms. According to the experimental results, the two algorithms proposed in this paper can improve the efficiency of satellite pointing and tracking tasks. Wenjing Xiao, Rui Wang 0077, Jeungeun Song 0001, Di Wu 0001, Long Hu, Min Chen 0003 |
GLOBECOM | 5 |
| 2020 | Follow me Robot-Mind: Cloud brain based personalized robot service with migration
Long Hu, Yinging Jiang, Fangxin Wang 0001, Kai Hwang 0001, M. Shamim Hossain, Muhammad Ghulam |
Future Gener. Comput. Syst. | 1 |
| 2020 | Power cognition: Enabling intelligent energy harvesting and resource allocation for solar-powered UAVsabstractSolar-powered unmanned aerial vehicles (SUAVs) are a promising solution to increase the flight time of unmanned aerial vehicles (UAVs) in the sky, reducing human interventions for battery charging. Exploiting networked SUAVs for providing long-duration wireless communication cannot only improve the signal transmission reliability but realize energy autonomy. To reap these benefits, in this article, we propose an efficient energy and radio resource management framework based on intelligent power cognition at the SUAVs. Thereby, power-cognitive SUAVs can learn the environment including the spatial distributions of solar energy density, the channel state evolution, and the traffic patterns of wireless communication applications in adaption to the environment changes. These SUAVs intelligently adjust the energy harvesting, information transmission, and flight trajectory to improve the utilization of solar energy for two primary goals: staying aloft over a long time period and achieving high communication performance. We adopt reinforcement learning to compute the optimal decisions for maximization of the total system throughput within the lifetime of the SUAV. Simulation results show that the proposed power cognition scheme can simultaneously improve the communication throughput and the harvested energy for SUAVs. Jing Zhang 0025, Minhao Lou, Lin Xiang 0001, Long Hu |
Future Gener. Comput. Syst. | 4 |
| 2020 | Privacy-preserving based task allocation with mobile edge clouds
Yongfeng Qian, M. Shamim Hossain, Long Hu, Muhammad Ghulam, Syed Umar Amin |
Inf. Sci. | 4 |
| 2020 | Privacy Protection and Intrusion Avoidance for Cloudlet-Based Medical Data SharingabstractWith the popularity of wearable devices, along with the development of clouds and cloudlet technology, there has been increasing need to provide better medical care. The processing chain of medical data mainly includes data collection, data storage and data sharing, etc. Traditional healthcare system often requires the delivery of medical data to the cloud, which involves users' sensitive information and causes communication energy consumption. Practically, medical data sharing is a critical and challenging issue. Thus in this paper, we build up a novel healthcare system by utilizing the flexibility of cloudlet. The functions of cloudlet include privacy protection, data sharing and intrusion detection. In the stage of data collection, we first utilize Number Theory Research Unit (NTRU) method to encrypt user's body data collected by wearable devices. Those data will be transmitted to nearby cloudlet in an energy efficient fashion. Second, we present a new trust model to help users to select trustable partners who want to share stored data in the cloudlet. The trust model also helps similar patients to communicate with each other about their diseases. Third, we divide users' medical data stored in remote cloud of hospital into three parts, and give them proper protection. Finally, in order to protect the healthcare system from malicious attacks, we develop a novel collaborative intrusion detection system (IDS) method based on cloudlet mesh, which can effectively prevent the remote healthcare big data cloud from attacks. Our experiments demonstrate the effectiveness of the proposed scheme. Min Chen 0003, Yongfeng Qian, Jing Chen 0003, Kai Hwang 0001, Shiwen Mao, Long Hu |
IEEE Trans. Cloud Comput. | 6 |
| 2019 | Social Weak-tie Assisted Cross-domain Short Video RecommendationabstractAs a comprehensive information carrier, the short video is gaining increasing attention for a user to spread information, read the news and conduct interpersonal communication. Consequently, short video recommendation problem has been a hot spot in the field of the recommender system. However, current short video recommendation algorithms have to tackle with data sparsity and cold start problem, which is caused by the small amounts of data that the recommender system has accumulated and massive video data while limited users' access.Aiming at the problem of data sparsity and cold start, the paper proposes a short video recommendation algorithm Social Weak-tie Bayesian Personalized Ranking (SWTBPR). Social weak-tie refers to the following relation in online social network, which means they are not real-world friends, somehow it can reflect users' preference. Experimental results demonstrate that SWTBPR outperforms other existing video recommendation algorithms and solves the data sparsity and cold start problem with a real-world dataset collected from Sina Weibo. Xichen Wang, Chen Gao 0001, Jingtao Ding, Long Hu, Yong Li 0008, Depeng Jin |
IWCMC | 4 |
| 2019 | CHPC: A complex semantic-based secured approach to heritage preservation and secure IoT-based museum processes
Anatoly Konev, Rezeda Khaydarova, Maxim Lapaev, Luanye Feng, Long Hu, Min Chen 0003, Igor Bondarenko |
Comput. Commun. | 5 |
| 2019 | Angular beta distribution for 3D vehicle-to-vehicle channel modeling
Derong Du, Xin Jian, Long Hu, Xiaoping Zeng, Xiaoheng Tan |
Future Gener. Comput. Syst. | 3 |
| 2019 | iRobot-Factory: An intelligent robot factory based on cognitive manufacturing and edge computing
Long Hu, Yiming Miao, Gaoxiang Wu, Mohammad Mehedi Hassan, Iztok Humar |
Future Gener. Comput. Syst. | 1 |
| 2019 | Privacy-aware service placement for mobile edge computing via federated learning
Yongfeng Qian, Long Hu, Jing Chen 0003, Xin Guan 0003, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi |
Inf. Sci. | 2 |
| 2019 | Opportunistic computing offloading in edge clouds
Wei Li 0061, Xinghui You, Jun Yang 0014, Long Hu |
J. Parallel Distributed Comput. | 5 |
| 2019 | Profit Maximization for Video Caching and Processing in Edge CloudabstractWith the development of communication technology and the explosive growth of video traffic brought by the rapid growth of mobile devices (such as smartphones and wearable devices), great business opportunities have been brought to video service providers. In this paper, we make full use of the cache and computing capacity of edge cloud. Considering the multi bitrate of video, we design the video caching and processing model that offers maximized profit to video service provider. Specifically, we model this problem as the 0-1 optimization problem and design the learning-based online upper confidence bound algorithm based on multi-arm bandit theory. This algorithm can design the corresponding cache and process strategy in real time according to the users' request to video. Furthermore, this strategy can maximize the profit of video provider and satisfy the service quality for users. Finally, experimental results show that our proposed video caching and processing scheme is superior to other schemes. Yixue Hao, Long Hu, Yongfeng Qian, Min Chen 0003 |
IEEE J. Sel. Areas Commun. | 2 |
| 2019 | A Dynamic Service Migration Mechanism in Edge Cognitive ComputingabstractDriven by the vision of edge computing and the success of rich cognitive services based on artificial intelligence, a new computing paradigm, edge cognitive computing (ECC), is a promising approach that applies cognitive computing at the edge of the network. ECC has the potential to provide the cognition of users and network environmental information, and further to provide elastic cognitive computing services to achieve a higher energy efficiency and a higher Quality of Experience (QoE) compared to edge computing. This article first introduces our architecture of the ECC and then describes its design issues in detail. Moreover, we propose an ECC-based dynamic service migration mechanism to provide insight into how cognitive computing is combined with edge computing. In order to evaluate the proposed mechanism, a practical platform for dynamic service migration is built up, where the services are migrated based on the behavioral cognition of a mobile user. The experimental results show that the proposed ECC architecture has ultra-low latency and a high user experience, while providing better service to the user, saving computing resources, and achieving a high energy efficiency. Min Chen 0003, Wei Li 0061, Giancarlo Fortino, Yixue Hao, Long Hu, Iztok Humar |
ACM Trans. Internet Techn. | 5 |
| 2019 | Photo Crowdsourcing Based Privacy-Protected HealthcareabstractIn this paper, the concept of crowdsourcing is applied to the medical field and a health monitoring mechanism based on photo crowdsourcing is proposed. Specifically, with photo crowdsourcing by many participators, the routine circumstances of users may be represented. However, these photos may include other people than the user, such as the visibility requestor, the invisibility requestor, and the passerby. The visibility and invisibility requestor are the participators in the system, whose identity can be set as visible or invisible, while the passerbys do not participate in the system. Hence, a privacy protection mechanism is proposed for this system, which includes two categories: i) The image fuzzy processing is provided for the invisibility requestor, while the original image is reserved for the visibility requestor. ii) The passerby's image is directly fuzzy processed for privacy protection. Long Hu, Yongfeng Qian, Jing Chen 0003, Xiaobo Shi, Jing Zhang 0025, Shiwen Mao |
IEEE Trans. Sustain. Comput. | 1 |
| 2018 | SCAI-SVSC: Smart clothing for effective interaction with a sustainable vital sign collection
Long Hu, Jun Yang 0014, Min Chen 0003, Yongfeng Qian, Joel J. P. C. Rodrigues |
Future Gener. Comput. Syst. | 1 |
| 2018 | Editorial: Cognitive Industrial Internet of Things
Long Hu, Daxin Tian |
Mob. Networks Appl. | 1 |
| 2017 | ASA: Against statistical attacks for privacy-aware users in Location Based Service
Min Chen 0003, Long Hu, Yongfeng Qian, Mohammad Mehedi Hassan |
Future Gener. Comput. Syst. | 3 |
| 2017 | Localization Based on Social Big Data Analysis in the Vehicular NetworksabstractLocation-based services, especially for vehicular localization, are an indispensable component of most technologies and applications related to the vehicular networks. However, because of the randomness of the vehicle movement and the complexity of a driving environment, attempts to develop an effective localization solution face certain difficulties. In this paper, an overlapping and hierarchical social clustering model (OHSC) is first designed to classify the vehicles into different social clusters by exploring the social relationship between them. By using the results of the OHSC model, we propose a social-based localization algorithm (SBL) that use location prediction to assist in global localization in the vehicular networks. The experiment results validate the performance of the OHSC model and show that the presented SBL algorithm demonstrates superior localization performance compared with the existing methods. Jiming Luo, Long Hu, M. Shamim Hossain, Ahmed Ghoneim |
IEEE Trans. Ind. Informatics | 3 |
| 2017 | Green and Mobility-Aware Caching in 5G NetworksabstractWith the drastic increase of mobile devices, there are more and more mobile traffic and repeated requests for content. In 5G networks, small cell base stations (SBSs) caching and caching in wireless device-to-device network can effectively decrease the mobile traffic during peak hours. Currently, most of the related work is focused on how to cache content on SBSs and on mobile devices, and it is assumed that the user can download the entire requested content through the connected SBSs and mobile devices. However, few works have taken user mobility and the randomness of contact duration into consideration. How to improve the caching strategy by exploiting user mobility is still a challenging problem. Thus, in this paper, we first investigate the problem of how to conduct caching placement on SBS and on mobile devices leveraging user mobility, aiming to maximize the cache hit ratio. Specifically, the caching placement on SBSs and on mobile devices is formulated as an integer programming problem, and submodular optimization is adopted to solve the formulated problem. Then, we give the optimal transmission power of SBSs and mobile devices to deliver the caching content in order to reduce the energy cost. Simulation results prove that our caching strategy is more efficient than other existing caching strategies in terms of both cache hit ratio and energy efficiency. Min Chen 0003, Yixue Hao, Long Hu, Kaibin Huang, Vincent K. N. Lau |
IEEE Trans. Wirel. Commun. | 3 |
| 2015 | Cloud-based Wireless Network: Virtualized, Reconfigurable, Smart Wireless Network to Enable 5G Technologies
Min Chen 0003, Yin Zhang 0002, Long Hu, Tarik Taleb, Zhengguo Sheng |
Mob. Networks Appl. | 3 |
| 2015 | CFSF: On Cloud-Based Recommendation for Large-Scale E-commerce
Long Hu, Mohammad Mehedi Hassan, Atif Alamri, Abdulhameed Alelaiwi |
Mob. Networks Appl. | 1 |
| 2015 | Cross-Layer Software-Defined 5G Network
Mao Yang 0001, Yong Li 0008, Long Hu, Bo Li 0089, Depeng Jin, Sheng Chen 0001, Zhongjiang Yan |
Mob. Networks Appl. | 3 |
| 2014 | COMER: Cloud-based medicine recommendationabstractWith the development of e-commerce, a growing number of people prefer to purchase medicine online for the sake of convenience. However, it is a serious issue to purchase medicine blindly without necessary medication guidance. In this paper, we propose a novel cloud-based medicine recommendation, which can recommend users with top-N related medicines according to symptoms. Firstly, we cluster the drugs into several groups according to the functional description information, and design a basic personalized medicine recommendation based on user collaborative filtering. Then, considering the shortcomings of collaborative filtering algorithm, such as computing expensive, cold start, and data sparsity, we propose a cloud-based approach for enriching end-user Quality of Experience (QoE) of medicine recommendation, by modeling and representing the relationship of the user, symptom and medicine via tensor decomposition. Finally, the proposed approach is evaluated with experimental study based on a real dataset crawled from Internet. Yin Zhang 0002, Long Wang 0012, Long Hu, Xiaofei Wang 0001, Min Chen 0003 |
QSHINE | 3 |
| 2014 | A Cloudlet-Assisted Multiplayer Cloud Gaming System
Wei Cai 0002, Victor C. M. Leung, Long Hu |
Mob. Networks Appl. | 3 |