EDBT 2026 Demo / reviewers in the wild / expert
Bo Wan 0002
dblp:86/4321-2
· DBLP profile ↗
31ranked-venue papers
0as first author
26since 2021 · last 2026
0000-0002-3410-9560ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 9 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Systems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HMEN: A hybrid modular network with dynamic expansion for continual learning
Ziye Fang, Bo Wan 0002, Shangqi Guo, Jian K. Liu |
Knowl. Based Syst. | 2 |
| 2026 | DuetGS: Two-Stage Controllable 3D Human Reconstruction From Dual ImagesabstractCreating realistic and fully detailed 3D human models using a minimal number of views has long been a challenging goal in 3D human reconstruction. Reconstructing a realistic human model from only two images (front and back) is particularly difficult due to the limited 3D information available, leading to two major challenges: (1) it is difficult to establish spatial consistency for reconstruction due to the lack of sufficient images for reliable matching, and (2) incomplete field of view results in missing color information. To address these challenges, we propose DuetGS, a novel pipeline that divides the reconstruction process into two stages: geometry reconstruction and color reconstruction. For geometry reconstruction, we employ a data-driven neural network to recover a full-body mesh from the front and back images, providing the spatial positioning for Gaussians. For color reconstruction, we adapt Gaussian Splatting and integrate our proposed unsupervised color propagation method to establish the color details of the Gaussians. Furthermore, our Gaussians are directly mapped to the mesh, allowing us to control their rotation and translation through mesh manipulation. This mapping ensures compatibility with various animation techniques. Extensive experiments on the THUman, CustomHumans, and PeopleSnapshot datasets demonstrate that our approach outperforms existing methods in terms of reconstruction accuracy and visual quality. Yanxin Chen, Bo Wan 0002, XuanZhu Yao, Nan Luo |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | EvoFormer: Learning Dynamic Graph-Level Representations with Structural and Temporal Bias CorrectionabstractDynamic graph-level embedding aims to capture structural evolution in networks, which is essential for modeling real-world scenarios. However, existing methods face two critical yet under-explored issues: Structural Visit Bias, where random walk sampling disproportionately emphasizes high-degree nodes, leading to redundant and noisy structural representations; and Abrupt Evolution Blindness, the failure to effectively detect sudden structural changes due to rigid or overly simplistic temporal modeling strategies, resulting in inconsistent temporal embeddings. To overcome these challenges, we propose EvoFormer, an evolution-aware Transformer framework tailored for dynamic graph-level representation learning. To mitigate Structural Visit Bias, EvoFormer introduces a Structure-Aware Transformer Module that incorporates positional encoding based on node structural roles, allowing the model to globally differentiate and accurately represent node structures. To overcome Abrupt Evolution Blindness, EvoFormer employs an Evolution-Sensitive Temporal Module, which explicitly models temporal evolution through a sequential three-step strategy: (I) Random Walk Timestamp Classification, generating initial timestamp-aware graph-level embeddings; (II) Graph-Level Temporal Segmentation, partitioning the graph stream into segments reflecting structurally coherent periods; and (III) Segment-Aware Temporal Self-Attention combined with an Edge Evolution Prediction task, enabling the model to precisely capture segment boundaries and perceive structural evolution trends, effectively adapting to rapid temporal shifts. Extensive evaluations on five benchmark datasets confirm that EvoFormer achieves state-of-the-art performance in graph similarity ranking, temporal anomaly detection, and temporal segmentation tasks, validating its effectiveness in correcting structural and temporal biases. Code is available at https://github.com/zlx0823/EvoFormerCode. Haodi Zhong, Liuxin Zou, Di Wang 0011, Bo Wan 0002, Zhenxing Niu, Quan Wang 0006 |
CIKM | 4 |
| 2025 | An Adaptive Dynamic Feature Selection Framework with CNN-Vision Transformer Hybrid Architecture for Continuous Joint Angle EstimationabstractContinuous estimation plays a critical role in Human-Computer Interaction (HCI) by enabling natural and precise motion control. While multimodal sensing (e.g., sEMG and accelerometer signals) theoretically provides richer information than single-modality approaches, existing multimodal methods face inherent information redundancy across modalities and excessive computational overhead. To address these challenges, we propose ADFS-CViT, an Adaptive Dynamic Feature Selection framework with CNN-Vision Transformer hybrid architecture. ADFS-CViT is a two-stage learning architecture that decouples modality-specific spatial-temporal feature extraction (Stage 1) from context-aware model selection (Stage 2), enabling specialized processing while maintaining computational efficiency. In stage 1, a novel spatial-temporal CNN-Vision Transformer design establishes cross-channel correlations in sEMG signals through patch-based attention mechanisms, effectively overcoming the temporal discontinuity limitation of conventional CNN-based approaches. In stage 2, an adaptive dynamic decision network is designed with a cost-aware loss and a joint angle estimation loss that dynamically activates optimal processing pathways (sEMG-only, ACC-only, or multimodal) based on real-time signal characteristics and resource constraints. Experimental results demonstrate that ADFS-CViT achieves the highest estimation performance compared to conventional multimodal fusion baselines, making it highly suitable for practical, real-time HCI systems with limited computational resources. Kejia Su, Hanbing Qiao, Bo Wan 0002, Changhua Jiang |
IJCB | 3 |
| 2025 | Text-Guided Attribute Enhancement Framework for Composed Image RetrievalabstractComposed image retrieval is a challenging multimodal task that refers to the process of retrieving target image by taking advantage of both complementary and synergistic image and text input. Existing efforts often focus on designing interaction models to fuse global query image and text features. However, these approaches struggle to capture fine-grained semantic association information between query image and text, especially when it comes to identifying specific objects or attributes in the query image that need to be modified in the text. In addition, these methods fail to adequately model cross-modal attention when dealing with composed query and target image, resulting in the model's inability to accurately map the semantic information from the composed query to the corresponding regions in the target image. To address these challenges, we propose a Text-Guided Attribute Enhancement Framework for Composed Image Retrieval (TAE-CIR). Our approach consists of three key modules: (a) Multi-granularity vision aggregation module, which extracts multi-granularity visual features and captures fine-grained object-level features related to the query text, refining object and attribute representations for more precise retrieval; (b) Multi-level fusion interaction module, which facilitates deep cross-modal interactions between the composed query and target image features, effectively capturing complex semantic relationships from the composed query to target image; (c) Composed feature alignment, which fuses multi-granularity visual features with the text using a text-guided Q-Former and contrastive learning to ensure accurate alignment between the composed query and the target image. Our extensive experiments on benchmark datasets FashionIQ and CIRR demonstrate the superiority of our proposed method. Yizi Huang, Di Wang 0011, Bo Wan 0002, Lin Zhao 0003, Quan Wang 0006 |
ICMR | 4 |
| 2025 | CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual RationaleabstractVision-language models (VLMs) are prone to hallucinations that critically compromise reliability in medical applications. While preference optimization can mitigate these hallucinations through clinical feedback, its implementation faces challenges such as clinically irrelevant training samples, imbalanced data distributions, and prohibitive expert annotation costs. To address these challenges, we introduce CheXPO, a Chest X-ray Preference Optimization strategy that combines confidence-similarity joint mining with counterfactual rationale. Our approach begins by synthesizing a unified, fine-grained multi-task chest X-ray visual instruction dataset across different question types for supervised fine-tuning (SFT). We then identify hard examples through token-level confidence analysis of SFT failures and use similarity-based retrieval to expand hard examples for balancing preference sample distributions, while synthetic counterfactual rationales provide fine-grained clinical preferences, eliminating the need for additional expert input. Experiments show that CheXPO achieves 8.93% relative performance gain using only 5% of SFT samples, reaching state-of-the-art performance across diverse clinical tasks. 1Code: https://github.com/ResearchGroup-MedVLLM/CheX-Phi35V Di Wang 0011, Lin Zhao 0003, Ronghan Li, Bo Wan 0002, Quan Wang 0006 |
ACM Multimedia | 7 |
| 2025 | A Task Scheduling Method for Minimizing Completion Time in Edge Collaboration EnvironmentabstractIn the edge computing environment, the uneven geographical distribution of tasks may lead to unbalanced load on the edge server. In addition, some larger tasks are difficult to completely offload to edge servers, which cannot fully utilize edge server resources. To solve the above problems, we propose a task scheduling method to minimize the completion time by combining the horizontal edge collaboration and fine-grained task partial offloading technology. First, combining horizontal edge collaboration and fine-grained task partial offloading technology, considering the location relationship between users and edge servers in multiuser multiedge server scenario, a task partial offloading optimization problem is established to minimize task completion time. Second, due to the nonconvex and variables coupling, we decompose the original problem into resource allocation, user-server association, and offloading strategy subproblems. A task scheduling algorithm based on improved teaching-learning-based optimization (ITLBO) is proposed to obtain the best task scheduling decision which includes task offloading location and offloading ratio. Simulation results show that the proposed method can effectively reduce the task completion time in edge collaboration environment. Hui Zhao 0003, Xiaoqin Lu, Jing Wang 0028, Pengfei Yang 0001, Bo Wan 0002, Quan Wang 0006 |
IEEE Internet Things J. | 6 |
| 2025 | Multi-workflow fault-tolerance scheduling strategy considering resources supply delay in WaaS platforms
Hui Zhao 0003, Wentao Zhi, Xiaoqin Lu, Jing Wang 0028, Nan Luo, Bo Wan 0002, Quan Wang 0006 |
Parallel Comput. | 6 |
| 2025 | Transferring Common Model Parameters From Chirp-Modulated to Steady-State Visual Evoked Potentials for Calibration-Efficient BCIs
Bang Xiong, Bo Wan 0002, Jiayang Huang, Pengfei Yang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Object Detector Based on Center Keypoints for Behavior Recognition in Classroom ScenesabstractStudent pose information can reflect learning status, which is significant in teaching management and evaluation. However, the traditional manual behavior recognition and analysis process is complex and slow, so counting massive classroom data is a difficult task. Therefore, using computer vision technology to accurately recognize student behavior is of great significance to teaching management and evaluation. Besides, the performance of existing behavior recognition methods is limited by problems such as dense objects and occlusion in classroom scenes. To address these issues, we propose an anchor-free object detector based on center keypoints for behavior recognition in classroom scenes. Specifically, we design a multiscale convolution neural networks (CNNs) module to alleviate the scale variation of objects through multiscale receptive fields. The head network is based on the anchor-free detection head architecture and integrates the keypoint heatmaps to accurately extract the center keypoint position through the center pooling module (CPM). the CPM can suppress irrelevant background information and effectively alleviate the missed detection of dense objects. Regressing the distances from the center keypoint of the positive region to the four sides can suppress low-quality bounding boxes so that the detector can accurately locate the object. Furthermore, during the inference stage, the heatmap scores of keypoints and center keypoints are linearly combined as novel confidence to mitigate inaccurate measurement of predicted bounding boxes. The extensive experimental performance comparison on the classroom behavior (CB) and SCB-dataset3 benchmark datasets demonstrate that the proposed behavior recognition method can accurately detect objects in classroom scenes. Compared with other current state-of-the-art methods, the proposed behavior recognition method based on anchor-free object detector is able to achieve good performance. Min Dang, Gang Liu 0006, Xike Li, Bo Wan 0002, Rong Pan 0004 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Multimodal transformer with adaptive modality weighting for multimodal sentiment analysis
Yifeng Wang 0004, Di Wang 0011, Quan Wang 0006, Bo Wan 0002, Xuemei Luo |
Neurocomputing | 5 |
| 2024 | Candidate-Heuristic In-Context Learning: A new framework for enhancing medical visual question answering with LLMs
Di Wang 0011, Haodi Zhong, Quan Wang 0006, Ronghan Li, Rui Jia, Bo Wan 0002 |
Inf. Process. Manag. | 7 |
| 2024 | Hierarchical reinforcement learning from imperfect demonstrations through reachable coverage-based subgoal filtering
Yu Tang 0010, Shangqi Guo, Bo Wan 0002, Lingling An, Jian K. Liu |
Knowl. Based Syst. | 4 |
| 2024 | Global semantic enhancement network for video captioning
Xuemei Luo, Xiaotong Luo, Di Wang 0011, Bo Wan 0002, Lin Zhao 0003 |
Pattern Recognit. | 5 |
| 2024 | Anonymity in Attribute-Based Access Control: Framework and MetricabstractAnonymous access is an effective method for preserving privacy in access control. This study assumes that anonymous access control requires both frameworks and policies. Numerous solutions have been proposed for anonymous access at the framework level. In this study, these solutions are analyzed and quantified using a unified attribute-based access control (ABAC) anonymous access reference framework. Anonymous access at the framework level is the first line of defense, and inappropriate policies may undermine subject anonymity. An anonymity metric is proposed at the policy level to prevent authorization authority from re-identification using specific attributes and policies. The anonymity metric evaluates the risk of re-identifying a subject due to inappropriate access requests, as well as subject attribute assignment schemes and policies. This study is the first to focus on anonymity at the policy level in ABAC. Furthermore, a formal definition of anonymity suitable for ABAC is proposed. The feasibility of the proposed anonymity metric is verified through simulations. Runnan Zhang, Gang Liu 0006, Hongzhaoning Kang, Quan Wang 0006, Bo Wan 0002, Nan Luo |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | TIR-Net: Task Integration Based on Rotated Convolution Kernel for Oriented Object Detection in Aerial ImagesabstractThe application of oriented object detection in the field of aerial images has gained substantial attention and made significant progress. However, most one-stage object detectors struggle to extract rotation-invariant features of oriented objects using ordinary convolutions. And the structure of two parallel vision subtasks can result in the inconsistency between the classification and regression. In this article, we propose a Task Integration based on a Rotated convolution kernel Network (TIR-Net) consisting of three modules: selective rotation of the kernel (SRK), regression feature refinement (RFR), and task integration (TI). Specifically, an SRK module enhances classification features by applying rotated convolution kernels selectively, introducing rotational invariance to the features. An RFR module places more emphasis on feature extraction of large aspect ratio objects to improve their perceptibility. A TI module integrates classification and regression features to alleviate the inconsistency between classification and regression. Experimental evaluations on the benchmarks for oriented object detection indicate that our method achieves excellent detection performance. Hao Li 0095, Rong Pan 0004, Gang Liu 0006, Min Dang, Qijie Xu, Xu Wang 0057, Bo Wan 0002 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Gist, Content, Target-Oriented: A 3-Level Human-Like Framework for Video Moment RetrievalabstractVideo moment retrieval (VMR) aims to locate corresponding moments in an untrimmed video via a given natural language query. While most existing approaches treat this task as a cross-modal content matching or boundary prediction problem, recent studies have started to solve the VMR problem from a reading comprehension perspective. However, the cross-modal interaction processes of existing models are either insufficient or overly complex. Therefore, we reanalyze human behaviors in the document fragment location task of reading comprehension, and design a specific module for each behavior to propose a 3-level human-like moment retrieval framework (Tri-MRF). Specifically, we summarize human behaviors such as grasping the general structures of the document and the question separately, cross-scanning to mark the direct correspondences between keywords in the document and in the question, and summarizing to obtain the overall correspondences between document fragments and the question. Correspondingly, the proposed Tri-MRF model contains three modules: 1) a gist-oriented intra-modal comprehension module is used to establish contextual dependencies within each modality; 2) a content-oriented fine-grained comprehension module is used to explore direct correspondences between clips and words; and 3) a target-oriented integrated comprehension module is used to verify the overall correspondence between the candidate moments and the query. In addition, we introduce a biconnected GCN feature enhancement module to optimize query-guided moment representations. Extensive experiments conducted on three benchmarks, TACoS, ActivityNet Captions and Charades-STA demonstrate that the proposed framework outperforms State-of-the-Art methods. Di Wang 0011, Xiantao Lu, Quan Wang 0006, Yumin Tian, Bo Wan 0002, Lihuo He |
IEEE Trans. Multim. | 5 |
| 2023 | A Logic Encryption-Enhanced PUF Architecture to Deceive Machine Learning-Based Modeling AttacksabstractAs a low-cost hardware security primitive, physically unclonable function (PDF) has been widely utilized in secure key generation and identity authentication of physical devices due to the advantages of high reliability and randomness. However, hackers can model a PDF circuit by collecting a small number of challenge- response pairs (CRPs), making it potentially vulnerable to machine learning (ML) attacks. To effectively resist this risk, various ML-resistance methods have been proposed. However, most of them mainly enhance the anti-attack ability of PDFs by structure nonlinear and CRP obfuscation, reducing stability and reliability. Therefore, we present a logic encryption-enhanced PDF (LEE PDF) architecture to resist ML-based attacks without affecting PDF performance. A logic encryption unit is applied to protect the function of original PDF circuits, thus concealing the valid CRPs in a large number of useless ones. Since hackers can only obtain a sparse number of useful CRPs without knowing the correct key, making it impossible to model PDF using ML methods. We have implemented the proposed LEE PDF on FPGA micro-boards. The experimental results demonstrate that the LEE PDF with only 2-bit key can effectively resist various ML- based attacks, average prediction rate is close to 50 %. In addition, the PDF performance remains almost constant. Lirong Zhou, Quan Wang 0006, Bo Wan 0002 |
ATS | 7 |
| 2023 | Language-Guided Visual Aggregation Network for Video Question AnsweringabstractVideo Question Answering (VideoQA) aims to comprehend intricate relationships, actions, and events within video content, as well as the inherent links between objects and scenes, to answer text-based questions accurately. Transferring knowledge from the cross-modal pre-trained model CLIP is a natural approach, but its dual-tower structure hinders fine-grained modality interaction, posing challenges for direct application to VideoQA tasks. To address this issue, we introduce a Language-Guided Visual Aggregation (LGVA) network. It employs CLIP as an effective feature extractor to obtain language-aligned visual features with different granularities and avoids resource-intensive video pre-training. The LGVA network progressively aggregates visual information in a bottom-up manner, focusing on both regional and temporal levels, and ultimately facilitating accurate answer prediction. More specifically, it employs local cross-attention to combine pre-extracted question tokens and region embeddings, pinpointing the object of interest in the question. Then, graph attention is utilized to aggregate regions at the frame level and integrate additional captions for enhanced detail. Following this, global cross-attention is used to merge sentence and frame-level embeddings, identifying the video segment relevant to the question. Ultimately, contrastive learning is applied to optimize the similarities between aggregated visual and answer embeddings, unifying upstream and downstream tasks. Our method conserves resources by avoiding large-scale video pre-training and simultaneously demonstrates commendable performance on the NExT-QA, MSVD-QA, MSRVTT-QA, TGIF-QA, and ActivityNet-QA datasets, even outperforming some end-to-end trained models. Our code is available at https://github.com/ecoxial2007/LGVA_VideoQA. Di Wang 0011, Quan Wang 0006, Bo Wan 0002, Lingling An, Lihuo He |
ACM Multimedia | 4 |
| 2023 | Enhancing CLIP-Based Text-Person Retrieval by Leveraging Negative Samples
Yumin Tian, Di Wang 0011, Bo Wan 0002 |
PRCV (7) | 6 |
| 2023 | VM performance-aware virtual machine migration method based on ant colony optimization in cloud environment
Hui Zhao 0003, Nanzhi Feng, Guobin Zhang, Jing Wang 0028, Quan Wang 0006, Bo Wan 0002 |
J. Parallel Distributed Comput. | 7 |
| 2023 | Bi-Attention enhanced representation learning for image-text matchingabstractImage-text matching has become a research hotspot in recent years. The key point of image-text matching is to accurately measure the similarity between an image and a sentence. However, most existing methods either focus on the inter-modality similarities between regions in image and words in text or the intra-modality similarities within image regions or words, such that they cannot well exploit detailed correlations between images and texts. Furthermore, existing methods typically train their models using a triplet ranking loss, which relies on the similarity of randomly sampled triples. Since the weights of positive and negative samples are not adjusted, it cannot provide enough gradient information for training, resulting in slow convergence and limited performance. To address the above problems, we propose an image-text matching method named Bi-Attention Enhanced Representation Learning (BAERL). It builds a self-attention learning sub-network to exploit intra-modality correlations within image regions or words and a co-attention learning sub-network to exploit inter-modality correlations between image regions and words. Then, representations obtained by two sub-networks capture holistic correlations between images and texts. In addition, BAERL uses the self-similarity polynomial loss instead of triplet ranking loss to train the model. The self-similarity polynomial loss can adaptively assign appropriate weights to different pairs based on their similarity scores so as to further improve the retrieval performance . Experiments on two benchmark datasets demonstrate the superior performance of the proposed BAERL method over several state-of-the-art methods. Yumin Tian, Aqiang Ding, Di Wang 0011, Xuemei Luo, Bo Wan 0002, Yifeng Wang 0004 |
Pattern Recognit. | 5 |
| 2021 | Investigating the Effectiveness of Virtual Reality for Culture LearningabstractPeople who are to live, study and work abroad will face more challenges in the new cultural environment and suffer more acculturative stress. Virtual Reality (VR), by which an immersive learning environment can be built, may help them adapt to a foreign culture at a lower cost of time and money. In order to work out a design method for culture learning in VR, we have designed a VR application so that learners can experience and learn the typical western festival culture – Christmas culture – in an immersive environment. To evaluate the effectiveness of the VR method, 50 EFL Chinese university students were enrolled in our experiments and randomly assigned to the VR group and the non-VR group, the data was drawn from cultural knowledge questionnaire, behavior test and Intercultural Sensitivity Scale (ISS). The ANCOVA revealed no major effect for group factor on knowledge learning. Similarly, the Mixed ANOVA identified no major effect for group factor on behavior learning and attitude learning. There was no interaction effect between time and group in all experiments. Our results show that the VR method is preferred by most of the participants, but it shows no remarkable advantage over the non-VR method. Moreover, regression analysis between the culture learning and the sense of presence in VR shows that presence has the potential to improve the performance of intercultural interaction engagement. Our findings are of practical value for culture learning in VR. Lei Gao 0007, Bo Wan 0002, Gang Liu 0006, Guojun Xie, Jiayang Huang, Guanglan Meng |
Int. J. Hum. Comput. Interact. | 2 |
| 2021 | Zero-watermarking method for resisting rotation attacks in 3D models
Gang Liu 0006, Quan Wang 0006, Lianqin Wu, Rong Pan 0004, Bo Wan 0002, Yumin Tian |
Neurocomputing | 5 |
| 2021 | Retrieving point cloud models of target objects in a scene from photographed images
Nan Luo, Quan Wang 0006, Bo Wan 0002 |
Multim. Tools Appl. | 4 |
| 2021 | Popularity-Based and Version-Aware Caching Scheme at Edge Servers for Multi-Version VoD SystemsabstractRecently, many video-on-demand (VoD) providers have begun storing multiple versions of the same video to offer multiple-quality video services with different bitrates to users, called multi-version VoD. To improve users' quality of experience (QoE), it is a good idea to cache videos at edge servers in multi-version VoD systems. However, determining which versions of which videos should be cached or replaced in an edge server is still a major challenge for a multi-version VoD system because of its limited cache storage. In this paper, we propose a popularity-based and version-aware caching scheme (PVCS) at edge servers for multi-version VoD systems. First, based on video popularity, we formulate cache placement as a knapsack problem under constraints such as the cache storage and transcoding computation of the edge server, which aims to maximize the cache hit ratio. Second, we use the transcoding relations among versions to calculate a version-aware caching profit when caching a certain version or multiple versions of a video. The version-aware caching profit is the basis for the subsequent cache replacement algorithm. Third, we propose two algorithms, the video cache placement (VCP) algorithm and the video cache replacement (VCrP) algorithm, to solve the cache placement and replacement problems respectively. VCP utilizes the Lagrangian relaxation algorithm to decide which video files should be cached initially, and VCrP decides which video files cached at the edge server will be replaced dynamically based on the version-aware profit. In this way, the PVCS can improve the cache hit ratio and decrease the average start-up delay. Our simulation results have shown that the PVCS outperforms the other schemes in terms of the cache hit ratio and the average start-up delay. Hui Zhao 0003, Quan Wang 0006, Jing Wang 0028, Bo Wan 0002, Zili Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | VM Performance Maximization and PM Load Balancing Virtual Machine Placement in CloudabstractVirtual machine placement (VMP) technology is widely used in cloud computing systems. The existing VMP methods mainly aimed at improving the cloud resource utilization, such as load balancing among physical machines (PMs), but they may result in virtual machine (VM) performance degradation because of the great resource contention among VMs running on top of the same PM. In contract to existing VMP algorithms, this paper proposes a virtual machine (VM) Performance maximization and physical machine (PM) Load balancing Virtual Machine Placement method (PLVMP) in cloud, which tries to maximize VM performance and balance PM workload from both users' and cloud providers' perspectives. First, we study the relationship between PM workload and VM performance to train a new and improved VM performance model, which can predict VM performance more accurately and offer help to the following VMP. Second, we take VM performance maximization and PM workload balancing into account to formulate the VMP as an optimization problem, which tries to maximize VM performance for users and make load balancing among PMs for cloud providers. Third, we propose a greedy-based algorithm to solve the VMP problem efficiently. We then evaluate PLVMP with other VMP methods on CloudSim platform and a real OpenStack platform. The results show PLVMP can maximize the VM performance significantly and make a good load balancing among PMs. Hui Zhao 0003, Quan Wang 0006, Jing Wang 0028, Bo Wan 0002, Shangshu Li |
CCGRID | 4 |
| 2020 | Fast and Efficient Facial Expression Recognition Using a Gabor Convolutional NetworkabstractAutomatic facial expression recognition (FER) is a fundamental topic in computer vision. Many studies have indicated that facial emotion changes are strongly related to certain regions of interest (ROIs), such as the mouth, eyes, eyebrows, and nose; therefore, the features of these facial ROIs are very important for identifying expressions. Since Gabor filters are very efficient in extracting visual content, Gabor orientation filters (GoFs) modulated by Gabor kernels and traditional convolutional filters can capture such ROI information better than conventional convolutional filters. Consequently, this letter presents a light Gabor convolutional network (GCN) consisting of only four Gabor convolutional layers and two linear layers for FER tasks. Extensive experiments on the FER2013, FERPlus and Real-world Affective Faces (RAF) databases demonstrate that the proposed method achieves good recognition accuracy and requires very low computational costs. The source code can be found at https://github.com/general515/Facial_Expression_Recognition_Using _GCN. Ping Jiang 0004, Bo Wan 0002, Quan Wang 0006, Jian Wu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2019 | High confidence detection for moving target in aerial videoabstractThe moving target detection and tracking in aerial video is a challenge task because of its moving background, smaller target sizes, lower resolution and limited onboard computing resources. In this study, a high confidence detection method based on background compensation and three‐frame‐difference method is designed, which can detect moving objects in a dynamic background accurately. First, the authors use local feature extraction and matching for image registration and demonstrate that speed‐up robust feature key points are suitable for the stabilisation task. Then, they estimate the global camera motion parameters using affine transformation which are obtained by the random sample consensus algorithm. Finally, they detect moving object by three‐frame‐difference method. As the detection results of the frame‐difference method generally exists ‘empty’ and noise, in order to select the two higher‐quality differential images to perform the logic AND operation, they add image quality assessment to the three‐frame‐difference method to obtain more accurate moving objects. Moreover, the edge detection algorithm and morphological processing are integrated together to further boost the overall detecting performance. The extensive empirical evaluations on aerial videos demonstrate that the proposed detector is very promising for the various challenging scenarios. Yumin Tian, Chenhui Peng, Di Wang 0011, Bo Wan 0002 |
IET Image Process. | 4 |
| 2019 | Robust joint learning network: improved deep representation learning for person re-identification
Yumin Tian, Di Wang 0011, Bo Wan 0002 |
Multim. Tools Appl. | 4 |
| 2019 | Semi-paired and semi-supervised multimodal hashing via cross-modality label propagation
Di Wang 0011, Bin Shang, Quan Wang 0006, Bo Wan 0002 |
Multim. Tools Appl. | 4 |