Sidan Du

dblp:22/6701 · DBLP profile ↗
← Back
56ranked-venue papers
1as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 21 since 2021Artificial intelligence and machine learning · 14 · 11 since 2021Systems, architecture and hardware · 4 · 1 since 2021Computer networks · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-View Weakly-Supervised 3D Human Pose Estimation via Human Body Segmentation
abstract
We propose a multi-view weakly-supervised 3D human pose estimation network. In this network, we fuse multi-view features and generate the final 3D pose, supervising the 3D pose using human body segmentation generated by a human parsing network. Using human body segmentation provides powerful supervision for the network. We further propose the Pose2Seg algorithm to transform 3D pose into simple simulated segmentation for all views, which allows the network to utilize the human body segmentation to supervise the 3D pose generated by the network, without the need for complex rendering processes and estimation of human body shape. We demonstrate the effectiveness of our method on various datasets and compare our method to other state-of-the-art algorithms, showing its advantages in terms of both quality and ability to generalize.
Siyuan Bei, Yu Zhou 0007, Yao Yu 0001, Sidan Du
Comput. Vis. Media4
2026 MotionLifter: Unsupervised text-driven 3D human motion generation with semantic and layout self-supervision
Jinghao Cao, Sidan Du
Neurocomputing6
2026 URPose: The model with unbiased rectified projection and reconstruction error for monocular unsupervised 3D human pose estimation
Sheng Liu 0013, Yang Li 0063, Sidan Du
Neurocomputing3
2026 Enhancing Multi-View Omnidirectional Depth Estimation With Semantic-Aware Cost Aggregation and Spatial Propagation
abstract
Omnidirectional depth estimation predicts 360-degree depth information using multiple fisheye cameras arranged in a surround-view configuration. However, due to the lack of reference panorama and differences between the predicted depth viewpoint and input cameras, it is challenging to construct and utilize semantic information to improve depth accuracy, resulting in limited accurate in complex regions such as non-overlapping, weak textures, object boundaries and occlusions. This paper proposes a novel model architecture that effectively extracts and leverages semantic information to enhance the accuracy of omnidirectional depth estimation. Specifically, the proposed algorithm combines the variance and mean of multi-view image features to construct the fused matching cost and utilize both geometry and semantic constraints. The model extracts 360-degree semantic context during matching cost aggregation, and predict the corresponding panoramas jointly with omnidirectional depth maps. A semantic-aware spatial propagation module is then employed to further refine the depth estimation. We leverage a multi-scale multi-task learning strategy to supervise the prediction of omnidirectional depth maps and panoramas jointly. The proposed approach achieves state-of-the-art performance on public datasets, and also demonstrates high-precision results on real-world data. The experiments with varying camera configurations validate the generalization ability and flexibility of the algorithm.
Ming Li 0069, Xuejiao Hu, Zihang Gao, Sidan Du, Yang Li 0063
IEEE Trans. Circuits Syst. Video Technol.4
2026 A qubit as a Kernel: an efficient quantum-classical network for image classification with Quantum Independent Convolution-like Kernel Layer
Wencong Cai, Yuanfeng Wang, Yongzhen Xu, Sidan Du
J. Supercomput.4
2025 Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
abstract
The target of video moment retrieval (VMR) is predicting temporal spans within a video that semantically match a given linguistic query. Existing VMR methods based on multimodal large language models (MLLMs) overly rely on expensive high-quality datasets and time-consuming fine-tuning. Although some recent studies introduce a zero-shot setting to avoid fine-tuning, they overlook inherent language bias in the query, leading to erroneous localization. To tackle the aforementioned challenges, this paper proposes Moment-GPT, a tuning-free pipeline for zero-shot VMR utilizing frozen MLLMs. Specifically, we first employ LLaMA-3 to correct and rephrase the query to mitigate language bias. Subsequently, we design a span generator combined with MiniGPT-v2 to produce candidate spans adaptively. Finally, to leverage the video comprehension capabilities of MLLMs, we apply Video-ChatGPT and span scorer to select the most appropriate spans. Our proposed method substantially outperforms the state-of-the-art MLLM-based and zero-shot models on several public datasets, including QVHighlights, ActivityNet-Captions, and Charades-STA.
Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Ming Li 0069, Wenxin Liang, Yang Li 0063, Sidan Du
AAAI7
2025 HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion
abstract
Recent advancements in diffusion-based technologies have made significant strides, particularly in identity-preserved portrait generation (IPG). However, when using multiple reference images from the same ID, existing methods typically produce lower-fidelity portraits and struggle to customize face attributes precisely. To address these issues, this paper presents HiFi-Portrait, a high-fidelity method for zero-shot portrait generation. Specifically, we first introduce the face refiner and landmark generator to obtain fine-grained multi-face features and 3D-aware face landmarks. The landmarks include the reference ID and the target attributes. Then, we design HiFi-Net to fuse multi-face features and align them with landmarks, which improves ID fidelity and face control. In addition, we devise an automated pipeline to construct an ID-based dataset for training HiFi-Portrait. Extensive experimental results demonstrate that our method surpasses the SOTA approaches in face similarity and controllability. Furthermore, our method is also compatible with previous SDXL-based works.
Yifang Xu, Benxiang Zhai, Yunzhuo Sun, Ming Li 0069, Yang Li 0063, Sidan Du
CVPR6
2025 FaceSnap: Enhanced ID-Fidelity Network for Tuning-Free Portrait Customization
Benxiang Zhai, Yifang Xu, Guofeng Zhang 0026, Yang Li 0063, Sidan Du
ICANN (2)5
2025 Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts
abstract
We present Uni-Inter, a unified framework for human motion generation that supports a wide range of interaction scenarios: including human-human, human-object, and human-scene—within a single, task-agnostic architecture. In contrast to existing methods that rely on task-specific designs and exhibit limited generalization, Uni-Inter introduces the Unified Interactive Volume (UIV), a volumetric representation that encodes heterogeneous interactive entities into a shared spatial field. This enables consistent relational reasoning and compound interaction modeling. Motion generation is formulated as joint-wise probabilistic prediction over the UIV, allowing the model to capture fine-grained spatial dependencies and produce coherent, context-aware behaviors. Experiments across three representative interaction tasks demonstrate that Uni-Inter achieves competitive performance and generalizes well to novel combinations of entities. These results suggest that unified modeling of compound interactions offers a promising direction for scalable motion synthesis in complex environments.
Sheng Liu 0013, Yuanzhi Liang, Jiepeng Wang 0005, Sidan Du, Chi Zhang 0067, Xuelong Li 0001
SIGGRAPH Asia4
2025 Authentic 3D Structure Preserved Surround View System for Automobile Driving Assistance
abstract
The automation and intellectualization of vehicles have become a development trend in recent years. Compared to Autonomous Driving systems, Advanced Driver Assistance Systems (ADAS) offer a more developed and practical choice. As an ADAS, the 3D surround view system can provide drivers with panoramic environmental information around the vehicle. However, traditional methods face issues such as distortion of close-range obstacles and ghosting in texture stitching. In this article, we propose an authentic 3D structure preserved surround view system based on depth information surface reconstruction. We introduce a novel dual-layer surface of foreground and background to restore structural information of close-range obstacles. In addition, we propose an innovative texture acquisition method to address the ghosting problem in stitching. Finally, we evaluate our method on both simulated and real datasets collected by ourselves, concluding that our approach can better provide drivers with information about the surroundings of their vehicles. More detailed material is available at https://nju-ee.github.io/Autonomous_Driving_Research_Group.page/surround/.
Chin Sheng Low, Jinghao Cao, Sidan Du
SMC4
2025 The Patch-Based Multi-Task Multi-Scale Reborn Network for Global Gaze Following in 360-Degree Images
abstract
ABSTRACT In this paper, we propose a global gaze following method using the patched‐based multi‐task multi‐scale reborn network (MMRGaze360) specifically designed for panorama images. Unlike existing approaches that rely on spherical networks or process only local regions, our architecture thoroughly accounts for the distortions introduced by the sphere‐to‐plane projection, enabling gaze following in comprehensive 360‐degree images. MMRGaze360 incorporates field‐of‐view (360‐FoV) and sight line (360‐Gaze) generators to model gaze behaviours and scene information in 360‐degree images. A multi‐task multi‐scale module is introduced to capture features from multiple patches centred around the estimated points located in the 360‐Gaze, using multi‐scale attention maps. These features, along with the 360‐FoV, are fused to produce a final heatmap. Additionally, we employ multi‐layer perceptions and convolutional networks using the reborn mechanism to enhance information usage and feature representation. Moreover, we establish a novel dataset, SRGaze360, which contains more conditions of the sphere‐to‐plane distortion. Experimental results on the GazeFollow360 and SRGaze360 datasets demonstrate the superiority of our method over previous works. It can be validated that our approach effectively addresses the limitations of 2D gaze following in handling out‐of‐frame gaze positions and distortions in 360‐degree images.
Jingzhao Dai, Sidan Du
IET Image Process.3
2025 Robust and Flexible Omnidirectional Depth Estimation With Multiple 360-Degree Cameras
Ming Li 0069, Xueqian Jin, Xuejiao Hu, Jinghao Cao, Sidan Du, Yang Li 0063
IET Image Process.5
2025 An efficient action proposal processing approach for temporal action detection
Xuejiao Hu, Jingzhao Dai, Ming Li 0069, Yang Li 0063, Sidan Du
Neurocomputing5
2025 ATHENA - Autonomous Vehicle Trajectory Planning Considered Human Action Awareness
abstract
Large language models have brought revolutionary changes to autonomous driving algorithms, ushering them into the era of multimodality. However, existing vehicle trajectory planning methods primarily focus on obstacle avoidance in autonomous driving scenarios, overlooking interactions with entities within the scene, such as humans. In this letter, we propose a new research direction: vehicle trajectory planning that takes into account human actions. We establish ATHENA, the first autonomous driving dataset that integrates multimodal human actions, comprising 33,855 scenarios. Each scenario contains status information about ego vehicle, as well as pedestrian actions that actively or passively interact with the vehicle, such as signaling the vehicle to proceed by waving and the unexpected falls by pedestrians that force the vehicle to stop. Based on each type of interaction, ATHENA also provides the corresponding driving suggestions and the reasons behind them. Moreover, we present an LLM-based baseline that consists of two agents: the Action Understanding Agent and the Vehicle Control Agent. Our baseline implements the generation of driving recommendations and vehicle control functions, which are guided by pedestrian actions. Experiments demonstrate the effectiveness and strong performance of our method. Our dataset and code will be publicly available at https://github.com/dogooooo/ATHENA.
Jinghao Cao, Sheng Liu 0013, Chaofan Wu, Yang Li 0063, Sidan Du
IEEE Signal Process. Lett.5
2024 Multi-Modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection
abstract
Given a video and a linguistic query, video moment retrieval and highlight detection (MR&HD) aim to locate all the relevant spans, while simultaneously predicting saliency scores. Most existing methods utilize RGB images as input, overlooking the inherent multi-modal visual signals like optical flow and depth. In this paper, we propose a Multi-modal Fusion and Query Refinement Network (MRNet) to learn complementary information from multi-modal cues. Specifically, we design a multi-modal fusion module to dynamically combine RGB, optical flow, and depth map. Furthermore, to simulate human understanding of sentences, we introduce a query refinement module that merges text at different granularities, containing word-, phrase-, and sentence-wise levels. Comprehensive experiments on QVHighlights and Charades datasets indicate that MRNet outperforms current SOTA methods, achieving notable improvements in MR-mAP@Avg (+3.41) and HD-HIT@1 (+3.46) on QVHighlights.
Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Zien Xie, Youyao Jia, Sidan Du
ICME6
2024 MH-DETR: Video Moment and Highlight Detection with Cross-modal Transformer
abstract
With the increasing demand for video understanding, video moment and highlight detection (MHD) has emerged as a critical research topic. MHD aims to localize all moments and predict clip-wise saliency scores simultaneously. Despite progress made by existing DETR-based methods, we observe that these methods coarsely fuse features from different modalities, which weakens the temporal intra-modal context and results in insufficient cross-modal interaction. To address this issue, we propose MH-DETR (Moment and Highlight DEtection TRansformer) tailored for MHD. Specifically, we introduce a simple yet efficient pooling operator within the uni-modal encoder to capture global intra-modal context. Moreover, to obtain temporally aligned cross-modal features, we design a plug-and-play cross-modal interaction module between the encoder and decoder, seamlessly integrating visual and textual features. Comprehensive experiments on QVHighlights, Charades-STA, Activity-Net, and TVSum datasets show that MH-DETR outperforms existing state-of-the-art methods, demonstrating its effectiveness and superiority. Our code is available at https://github.com/YoucanBaby/MH-DETR.
Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Youyao Jia, Sidan Du
IJCNN5
2024 A Stereo Matching Method for Specular Objects via Cascaded Network and Joint Supervision
Yongkang Feng, Jianghai Shuai, Pinzhi Wang, Yang Li 0063, Sidan Du
PRCV (3)5
2024 CASSC: Context-aware method for depth guided semantic scene completion
abstract
Abstract Semantic scene completion is a crucial end‐to‐end 3D perception task, and the 3D information perception subjects is vital for autonomous driving. This paper presents CASSC, a novel adaptive context‐aware method based on Transformer networks, aimed at realizing camera‐based semantic scene completion algorithms. The key idea is to leverage rich context information from images to obtain pixel‐level label proposals, followed by designing a multiscale fusion mechanism to merge this information and match it with voxel space. A weakly supervised training strategy is proposed to obtain semantic label distribution features from images and introduce an adaptive multiscale fusion module to fuse and adaptively match these features with voxel space. Here, CASSC achieves state‐of‐the‐art performance on the SemanticKITTI dataset and demonstrates excellent performance on the SSC‐Bench dataset. Ablation experiments validate the rationality and effectiveness of our design, and the model and code of CASSC will be open‐sourced on https://github.com/dogooooo/CASSC .
Jinghao Cao, Ming Li 0069, Sheng Liu 0013, Yang Li 0063, Sidan Du
IET Image Process.5
2024 Time-attentive fusion network: An efficient model for online detection of action start
abstract
Abstract Online detection of action start is a significant and challenging task that requires prompt identification of action start positions and corresponding categories within streaming videos. This task presents challenges due to data imbalance, similarity in boundary content, and real‐time detection requirements. Here, a novel Time‐Attentive Fusion Network is introduced to address the requirements of improved action detection accuracy and operational efficiency. The time‐attentive fusion module is proposed, which consists of long‐term memory attention and the fusion feature learning mechanism, to improve spatial‐temporal feature learning. The temporal memory attention mechanism captures more effective temporal dependencies by employing weighted linear attention. The fusion feature learning mechanism facilitates the incorporation of current moment action information with historical data, thus enhancing the representation. The proposed method exhibits linear complexity and parallelism, enabling rapid training and inference speed. This method is evaluated on two challenging datasets: THUMOS’14 and ActivityNet v1.3. The experimental results demonstrate that the proposed method significantly outperforms existing state‐of‐the‐art methods in terms of both detection accuracy and inference speed.
Xuejiao Hu, Ming Li 0069, Yang Li 0063, Sidan Du
IET Image Process.5
2024 ARES: Text-Driven Automatic Realistic Simulator for Autonomous Traffic
abstract
The large-scale generation of real-world scenario datasets is a pivotal task in the field of autonomous driving. Existing methods emphasize solely on single-frame rendering, which need complex inputs for continuous scenario rendering. In this letter, ARES: a text-driven automatic realistic simulator is proposed, which can generate extensive realistic datasets with just a single text input. Its core idea is to generate vehicle trajectories based on the textual description, and then render the scenario by vehicle attributes associated with these trajectories. For learning trajectories generating, supervisory signal temporal logic is proposed to assist conditional diffusion model, which incorporates prior physical information. We annotate textual descriptions for KITTI-MOT dataset and establish an objective quantitative evaluation system. The superiority of our method is demonstrated by its high performance, which is reflected in a matching score of 3.54 and an FID of 8.93in the trajectory reconstruction task, along with a speed accuracy of 0.99 and a direction accuracy of 0.93in the trajectory editing task. The scenarios rendered by the proposed method exhibit high quality and realism, which indicates its great potential in testing of autonomous driving algorithms with vehicle-in-the-loop simulations.
Jinghao Cao, Sheng Liu 0013, Yang Li 0063, Sidan Du
IEEE Signal Process. Lett.5
2024 Distribution-Aware Activity Boundary Representation for Online Detection of Action Start in Untrimmed Videos
abstract
The Online Detection of Action Start (ODAS) has attracted the attention of researchers because of its practical applications in areas such as security and emergency response. However, online detection of activity boundaries remains a challenging task due to the inherent ambiguity of boundary definition and the significant imbalance in the number of boundaries and nonboundary points. To address this issue, this study proposes a novel Distribution-aware Activity Boundary Representation (DABR) method that utilizes a continuous probability density function to smooth the probability of moments near activity boundaries. The proposed DABR reduces the penalty for detecting moments near ground-truth boundary points, while increasing the number of samples related to boundary points. Additionally, we introduce a two-stage framework that incorporates class-informed information in temporal localization for more efficient activity boundary localization. Extensive experiments demonstrate that our method achieves state-of-the-art results on two standard datasets, particularly exhibiting a significant improvement of 11.5% at average p-mAP on the THUMOS'14 dataset.
Xuejiao Hu, Ming Li 0069, Yang Li 0063, Sidan Du
IEEE Signal Process. Lett.5
2024 GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity Features
abstract
Moment retrieval (MR) and highlight detection (HD) aim to identify relevant moments and highlights in video from corresponding natural language query. Large language models (LLMs) have demonstrated proficiency in various computer vision tasks. However, existing methods for MR&HD have not yet been integrated with LLMs. In this letter, we propose a novel two-stage model that takes the output of LLMs as the input to the second-stage transformer encoder-decoder. First, MiniGPT-4 is employed to generate the detailed description of the video frame and rewrite the query statement, fed into the encoder as new features. Then, semantic similarity is computed between the generated description and the rewritten queries. Finally, continuous high-similarity video frames are converted into span anchors, serving as prior position information for the decoder. Experiments demonstrate that our approach achieves a state-of-the-art result, and by using only span anchors and similarity scores as outputs, positioning accuracy outperforms traditional methods, like Moment-DETR.
Yunzhuo Sun, Yifang Xu, Zien Xie, Yukun Shu, Sidan Du
IEEE Signal Process. Lett.5
2023 MMDA: Multi-person marginal distribution awareness for monocular 3D pose estimation
abstract
Abstract Most existing 3D pose representations cannot completely decouple the overlapping two or more human joints of the same type. In this paper, the authors propose a novel 2.5 D representation of the human pose by projecting human joints in 3D space onto the three orthogonal planes. The authors apply for the first time the permutation module to a multi‐person 3D human pose estimation task and use Geometric Constraints Loss (GCL) to guide the learning of the model. The authors overcome the negative effects of the inductive bias of convolutional neural networks (CNNs) by aligning the intermediate feature space with the output feature space. The effectiveness of the authors’ approach is validated on the carnegie mellon university (CMU) panoptic dataset and MuPoTS‐3D dataset. The authors’ proposed representations can effectively decouple the human joints in their selected data from overlapping human joints.
Sheng Liu 0013, Jianghai Shuai, Yang Li 0063, Sidan Du
IET Image Process.4
2023 Query-Guided Refinement and Dynamic Spans Network for Video Highlight Detection and Temporal Grounding in Online Information Systems
abstract
With the surge in online video content, finding highlights and key video segments have garnered widespread attention. Given a textual query, video highlight detection (HD) and temporal grounding (TG) aim to predict frame-wise saliency scores from a video while concurrently locating all relevant spans. Despite recent progress in DETR-based works, these methods crudely fuse different inputs in the encoder, which limits effective cross-modal interaction. To solve this challenge, the authors design QD-Net (query-guided refinement and dynamic spans network) tailored for HD&TG. Specifically, they propose a query-guided refinement module to decouple the feature encoding from the interaction process. Furthermore, they present a dynamic span decoder that leverages learnable 2D spans as decoder queries, which accelerates training convergence for TG. On QVHighlights dataset, the proposed QD-Net achieves 61.87 HD-HIT@1 and 61.88 [email protected], yielding a significant improvement of +1.88 and +8.05, respectively, compared to the state-of-the-art method.
Yifang Xu, Yunzhuo Sun, Zien Xie, Benxiang Zhai, Youyao Jia, Sidan Du
Int. J. Semantic Web Inf. Syst.6
2023 The multi-learning for food analyses in computer vision: a survey
Jingzhao Dai, Xuejiao Hu, Ming Li 0069, Yang Li 0063, Sidan Du
Multim. Tools Appl.5
2022 MODE: Multi-view Omnidirectional Depth Estimation with 360$^\circ $ Cameras
Ming Li 0069, Xueqian Jin, Xuejiao Hu, Jingzhao Dai, Sidan Du, Yang Li 0063
ECCV (33)5
2022 Online human action detection and anticipation in videos: A survey
Xuejiao Hu, Jingzhao Dai, Ming Li 0069, Chenglei Peng, Yang Li 0063, Sidan Du
Neurocomputing6
2022 The study of stereo matching optimization based on multi-baseline trinocular model
Chenglei Peng, Ming Li 0069, Yang Li 0063, Sidan Du
Multim. Tools Appl.5
2021 A Study of General Data Improvement for Large-Angle Head Pose Estimation
Jue Bai, Chenglei Peng, Zhaoxu Li, Sidan Du, Yang Li 0063
CAIP (2)4
2021 Pyramid Feature Attention Network for Monocular Depth Prediction
abstract
Deep convolutional neural networks (DCNNs) have achieved great success in monocular depth estimation (MDE). However, few existing works take the contributions for MDE of different levels feature maps into account, leading to inaccurate spatial layout, ambiguous boundaries and discontinuous object surface in the prediction. To better tackle these problems, we propose a Pyramid Feature Attention Network (PFANet) to improve the high-level context features and lowlevel spatial features. In the proposed PFANet, we design a Dual-scale Channel Attention Module (DCAM) to employ channel attention in different scales, which aggregate global context and local information from the high-level feature maps. To exploit the spatial relationship of visual features, we design a Spatial Pyramid Attention Module (SPAM) which can guide the network attention to multi-scale detailed information in the low-level feature maps. Finally, we introduce scale-invariant gradient loss to increase the penalty on errors in depth-wise discontinuous regions. Experimental results show that our method outperforms state-of-the-art methods on the KITTI dataset.
Yifang Xu, Chenglei Peng, Ming Li 0069, Yang Li 0063, Sidan Du
ICME5
2021 Jitter: Random Jittering Loss Function
abstract
Regularization plays a vital role in machine learning optimization. One novel regularization method called flooding makes the training loss fluctuate around the flooding level. It intends to make the model continue to “random walk” until it comes to a flat loss landscape to enhance generalization. However, the hyper-parameter flooding level of the flooding method fails to be selected properly and uniformly. We propose a novel method called Jitter to improve it. Jitter is essentially a kind of random loss function. Before training, we randomly sample the “Jitter Point” from a specific probability distribution. The flooding level should be replaced by Jitter point to obtain a new target function and train the model accordingly. As Jitter point acting as a random factor, we actually add some randomness to the loss function, which is consistent with the fact that there exists innumerable random behaviors in the learning process of the machine learning model and is supposed to make the model more robust. In addition, Jitter performs “random walk” randomly which divides the loss curve into small intervals and then flipping them over, ideally making the loss curve much flatter and enhancing generalization ability. Moreover, Jitter can be a domain-, task-, and model-independent regularization method and train the model effectively after the training error reduces to zero. Our experimental results show that Jitter method can improve model performance more significantly than the previous flooding method and make the test loss curve descend twice.
Zhicheng Cai, Chenglei Peng, Sidan Du
IJCNN3
2021 Exploiting Edge Computing in Internet of Space Things Networks: Dynamic and Static Server Placement
abstract
Internet of Space Things (IoST), which extends the concept of Internet of Things (IoT) to space, has emerged as a new paradigm for offering monitoring/reconnaissance, in-space backhaul, and cyber-physical integration services. As Low Earth Orbit (LEO) satellites are increasingly deployed for global Internet services, Mobile Edge Computing (MEC) is being introduced into satellite networks to provision computing services by placing edge servers on satellites. Nevertheless, it is a nontrivial and unexplored task to efficiently choose edge server deployment locations from a large number of satellites. In this paper, we address this issue in detail towards average response delay minimization considering propagation delay, forwarding delay, and service delay. In particular, we formulate the dynamic server placement problem as well as the static server placement problem, and devise a genetic algorithm-based heuristic approach to solve them. Simulation results compare the two placement strategies with two benchmarks and demonstrate the performance of our genetic algorithm-based approach. Furthermore, a comparison between the dynamic placement and the static placement is investigated.
Zhibo Yan, Tomaso de Cola, Kanglian Zhao, Wenfeng Li 0003, Sidan Du
VTC Fall5
2021 Omnidirectional stereo depth estimation based on spherical deep network
Ming Li 0069, Xuejiao Hu, Jingzhao Dai, Yang Li 0063, Sidan Du
Image Vis. Comput.5
2020 Learn a Global Appearance Semi-Supervisedly for Synthesizing Person Images
abstract
We present a novel approach for person images synthesis in this paper, that can generate person images in arbitrary poses, shapes and views. Unlike existing methods just using keypoints' locations in heatmaps format, we propose to render SMPL model to UV maps, which can provide human structural information about poses and shapes. Thus, by varying the parameters of poses, shapes and camera in SMPL model, we can generate different person images with various attributions in a simple way, while in most cases we can only obtain new shapes of people by computer graphics methods. We train an end to end generative adversarial network with unlabeled data. As our SMPL parameters come from a pretrained model, we call our overall network semi- supervised. Our network keeps a global appearance during the fine-tuning stage of the target person, thus we can get a complete appearance of the target person, rather than the inaccurate appearance caused by inferencing without enough information. Experiments on Human3.6M Dataset and a self-collected dataset demonstrate the excellent effectiveness of our approach on person images synthesis for different applications.
Zhipeng Ge, Sidan Du, Yao Yu 0001, Yu Zhou 0007
WACV3
2019 Image based fruit category classification by 13-layer deep convolutional neural network and data augmentation
Yudong Zhang 0001, Zhengchao Dong, Xianqing Chen, Wen-Juan Jia 0001, Sidan Du, Khan Muhammad 0001, Shuihua Wang
Multim. Tools Appl.5
2018 Application of stationary wavelet entropy in pathological brain detection
Shuihua Wang, Sidan Du, Abdon Atangana, Zeyuan Lu
Multim. Tools Appl.2
2018 Host load prediction with long short-term memory in cloud computing
Binbin Song, Yao Yu 0001, Yu Zhou 0007, Sidan Du
J. Supercomput.5
2017 Real-time walkthrough of outdoor scenes using TRI-view morphing
abstract
In this paper, an image-based walkthrough system is presented for navigating real-world outdoor scenes based on only three uncalibrated images without reconstructing 3D model. An image-based rendering operation is performed on these sample images and generates photorealistic in-between views in real time. We extend the traditional two-step view morphing to real-time tri-view morphing based on epipolar constraint. Compared with other tri-view morphing methods, our scheme copes well with both complex outdoor scenes and wide baseline scenes, which also poses a qualified alternative to state-of-the-art dense 3D reconstruction. Putting on a head mount display (HMD) like Oculus Rift, user can experience realistic and immersive exploration of real-world scenes.
Yu Zhou 0007, Yao Yu 0001, Sidan Du
ICIP4
2017 Hearing Loss Detection in Medical Multimedia Data by Discrete Wavelet Packet Entropy and Single-Hidden Layer Neural Network Trained by Adaptive Learning-Rate Back Propagation
Shuihua Wang, Sidan Du, Yang Li 0063, Huimin Lu 0001, Ming Yang 0011, Bin Liu 0043, Yudong Zhang 0001
ISNN (2)2
2017 Human Pose Estimation Using Deep Structure Guided Learning
abstract
In this paper, we propose a novel approach to incorporate structure knowledge into Convolutional Neural Networks (CNNs) for articulated human pose estimation from a single still image. Recent research on pose estimation adopt CNNs as base blocks to combine with other graphical models. Different from existing methods using features from CNNs to model the tree structure, we directly use the structure pose prior to guide the learning of CNN. First, we introduce a deep CNN with effective receptive fields which capture the holistic context of the whole image. Second, limb loss is used as intermediate supervision of CNN to learn the correlations of joints. Both parts and joints features are extracted in the middle of neural network and then are used to guide the following network learning. The proposed framework can exploit an implicit structure model of human body. Only using one stage and without any complex post processing, our method achieves state-of-art results on both FLIC and LSP benchmarks.
Baole Ai, Yu Zhou 0007, Yao Yu 0001, Sidan Du
WACV4
2017 Pathological Brain Detection via Wavelet Packet Tsallis Entropy and Real-Coded Biogeography-based Optimization
abstract
(Aim) In order to detect pathological brains in a more efficient way, (Method) we proposed a novel system of pathological brain detection (PBD) that combined wavelet packet Tsallis entropy (WPTE), feedforward neural network (FNN), and real-coded biogeography-based optimization (RCBBO). (Results) Th e experiments showed the proposed WPTE + FNN + RCBBO approach yielded an average accuracy of 99.49% over a 255-image dataset. (Conclusions) The WPTE + FNN + RCBBO performed better than 10 state-of-the-art approaches.
Shuihua Wang, Peng Li 0002, Peng Chen 0018, Preetha Phillips, Sidan Du, Yudong Zhang 0001
Fundam. Informaticae6
2017 Tea Category Identification using Computer Vision and Generalized Eigenvalue Proximal SVM
abstract
(Objective) In order to increase classification accuracy of tea-category identification (TCI) system, this paper proposed a novel approach. (Method) The proposed methods first extracted 64 color histogram to obtain color information, and 16 wavelet packet entropy to obtain the texture information. With the aim of reducing the 80 features, principal component analysis was harnessed. The reduced features were used as input to generalized eigenvalue proximal support vector machine (GEPSVM). Winner-takes-all (WTA) was used to handle the multiclass problem. Two kernels were tested, linear kernel and Radial basis function (RBF) kernel. Ten repetitions of 10-fold stratified cross validation technique were used to estimate the out-of-sample errors. We named our method as GEPSVM + RBF + WTA and GEPSVM + WTA. (Result) The results showed that PCA reduced the 80 features to merely five with explaining 99.90% of total variance. The recall rate of GEPSVM + RBF + WTA achieved the highest overall recall rate of 97.9%. (Conclusion) This was higher than the result of GEPSVM + WTA and other five state-of-the-art algorithms: back propagation neural network, RBF support vector machine, genetic neural-network, linear discriminant analysis, and fitness-scaling chaotic artificial bee colony artificial neural network.
Shuihua Wang, Preetha Phillips, Sidan Du
Fundam. Informaticae4
2017 Robust multi-view stereo synthesized by various parameters model
Yongming Nie, Tao Yue 0003, Hao Zhu 0004, Sidan Du, Xun Cao
J. Vis. Commun. Image Represent.4
2017 Fine-Grained Vehicle Model Recognition Using A Coarse-to-Fine Convolutional Neural Network Architecture
abstract
Fine-grained vehicle model recognition is a challenging problem in intelligent transportation systems due to the subtle intra-category appearance variation. In this paper, we demonstrate that this problem can be addressed by locating discriminative parts, where the most significant appearance variation appears, based on the large-scale training set. We also propose a corresponding coarse-to-fine method to achieve this, in which these discriminative regions are detected automatically based on feature maps extracted by convolutional neural network. A mapping from feature maps to the input image is established to locate the regions, and these regions are repeatedly refined until there are no more qualified ones. The global and local features are then extracted from the whole vehicle images and the detected regions, respectively. Based upon the holistic cues and the subordinate-level variation within these global and local features, an one-versus-all support vector machine classifier is applied for classification. The experimental results show that our framework outperforms most of the state-of-the-art approaches, achieving 98.29% accuracy over 281 vehicle makes and models.
Yu Zhou 0007, Yao Yu 0001, Sidan Du
IEEE Trans. Intell. Transp. Syst.4
2016 Higher-order class-specific priors for semantic segmentation of 3D outdoor scenes
abstract
Given 3D outdoor scenes acquired by a LIDAR sensor, we address the problem of semantic segmentation of 3D point clouds involving simultaneously segmenting and classifying the data. The capability of semantic segmentation is essential for several applications, such as autonomous robot navigation and 3D reconstruction of point clouds. In this paper, we present a higher-order class-specific CRF model according to the discriminative priors of object classes to capture the rich statistics of natural scenes. Consequently we cast this model in an energy minimization framework and propose the detailed energy potentials based on the class-specific priors of higher-order cliques. Then the SOSPD algorithm is adapted to optimize the energy function. To evaluate the performance of our method, we provide both quantitative and qualitative results on a challenging dataset. The results show an average Fl-score of 0.82 compared to the state-of-the-art Fl-score of 0.73.
Bingxiao Tang, Yu Zhou 0007, Yao Yu 0001, Sidan Du
WACV4
2016 Efficient inter-carrier interference cancellation transmissions for cooperative networks with frequency offsets
abstract
This study discusses new frequency reversal orthogonal frequency‐division multiplexing schemes for cooperative networks in order to minimise the effect of multiple carrier frequency offsets. Theorem 1 proposed and proved in this study tells that zero‐forcing (ZF) decoding is equivalent to maximum‐likelihood (ML) decoding if the equivalent channel matrix is complex orthogonal. Frequency reversal Alamouti code for two‐relay scenario and new space‐frequency codes for more relay nodes scenario are designed under the guidance of Theorem 1. The reversal operations of these codes make the equivalent channel matrix between the relays and the destination complex orthogonalised. Thus, ZF decoding achieves the advantages of ML decoding, and simulation results also confirm that. Benefiting from the simplified ZF decoding method in Theorem 1, the decoding complexity becomes really low.
Yao Yu 0001, Sidan Du
IET Commun.4
2015 Fusion of Local Manifold Learning Methods
abstract
Different local manifold learning methods are developed based on different geometric intuitions and each method only learns partial information of the true geometric structure of the underlying manifold. In this letter, we introduce a novel method to fuse the geometric information learned from local manifold learning algorithms to discover the underlying manifold structure more faithfully. We first use local tangent coordinates to compute the local objects from different local algorithms, then utilize the selection matrix to connect the local objects with a global functional and finally develop an alternating optimization-based algorithm to discover the low-dimensional embedding. Experiments on synthetic as well as real datasets demonstrate the effectiveness of our proposed method.
Xianglei Xing, Zhuowen Lv, Yu Zhou 0007, Sidan Du
IEEE Signal Process. Lett.5
2015 Multi-step-ahead host load prediction using autoencoder and echo state networks in cloud computing
Qiangpeng Yang, Yu Zhou 0007, Yao Yu 0001, Xianglei Xing, Sidan Du
J. Supercomput.6
2015 Learning human shape model from multiple databases with correspondence considering kinematic consensus
Yao Yu 0001, Yu Zhou 0007, Sidan Du, Zhengyu Cai
Vis. Comput.3
2014 Power allocation for orthogonal frequency division multiplexing-based cognitive radio networks with cooperative relays
abstract
The power allocation problem in orthogonal frequency division multiplexing‐based cognitive radio (CR) networks with cooperative relays has been investigated here, where both the interference to primary users (PUs) and the power budget of the CR network are considered. The authors try to maximise the overall throughput of the CR network within the given constraints. The coupling variables in the formulated problem make it hard to solve, so an iterative optimisation scheme to find out the optimal solution with a controllable complexity is developed. First, the original problem is decomposed into two subproblems that can be solved independently. A fast barrier method has been employed to work out the optimal solution to one of the subproblems with a complexity of O ( L 2 N ), where L and N are the number of PUs and subcarriers, respectively. Then, an iterative procedure is developed to solve the other subproblem. Numerical results show that the proposed method can significantly increase the throughput of the CR system, comparing with other representative ones. Furthermore, the proposed algorithm gives a general power optimisation framework for CR networks with cooperative relays.
Sidan Du, Fangjiang Huang, Shaowei Wang 0001
IET Commun.1
2014 A new method based on PSR and EA-GMDH for host load prediction in cloud computing system
Qiangpeng Yang, Chenglei Peng, He Zhao 0008, Yao Yu 0001, Yu Zhou 0007, Sidan Du
J. Supercomput.7
2013 Distributed optimal cyclotomic space-time coding for full-duplex cooperative relay networks
abstract
In this paper a diagonal cyclotomic space-time coding (DCSTC) transmission scheme for wireless cooperative relay networks is proposed. We consider the full-duplex cooperation which is more spectrally efficient than the half-duplex cooperation. By the combination of multidimensional complex constellation symbols rotation and Hadamard transform, the proposed scheme is capable of achieving full cooperative diversity. A large number of relay nodes may be employed in the cooperative system. The pairwise error probability (PEP) and the coding gain with the optimal total transmit power allocation is analyzed. It is not required with decoding and demodulation at the relay node, resulting in lower complexity. Besides, there is no delay with regards to re-transmission from delay nodes to destinations. At the destination node, the DCSTC structure can be exploited on the received signals, which is easily decoded by the fast sphere decoder. Simulation results show that the proposed scheme achieves full cooperative diversity and largely improves the network performance.
Xianglei Xing, Sidan Du
IWCMC3
2013 A multi-manifold semi-supervised Gaussian mixture model for pattern classification
Xianglei Xing, Yao Yu 0001, Sidan Du
Pattern Recognit. Lett.4
2013 Efficient early direct mode decision for multi-view video coding
Fengsui Wang, Huanqiang Zeng, Qinghong Shen, Sidan Du
Signal Process. Image Commun.4
2012 A low complexity fast lattice reduction algorithm for MIMO detection
abstract
Based on the well known Lenstra Lenstra Lovász (LLL) algorithm, we propose a possible swap LLL algorithm (P-SLLL) for lattice reduction aided (LRA) MIMO detection in this paper. The reduction process of the new algorithm is modified by searching for the next column swap through the whole basis, instead of the sequential implementation in the original LLL algorithm. Two different searching criteria are proposed, i.e. the random selection criterion and the optimal swap selection criterion. Comparing to the LLL algorithm, the PSLLL algorithm enjoys fast termination property and lower computational complexity, which can benefit practical hardware implementation. Simulation results prove our analysis and show that PSLLL aided linear MIMO detectors achieve the same performance as the LLL aided methods.
Kanglian Zhao, Yang Li 0063, Sidan Du
PIMRC4
2008 Fast super-resolution for license plate image reconstruction
abstract
A fast super-resolution reconstruction algorithm designed for license plate recognition is proposed in this paper. It uses a new reduced cost function to produce images of higher resolution from low resolution frame sequences. Computational cost required in this algorithm is much lower compared with other methods. The effectiveness of the proposed algorithm is demonstrated through blind reconstruction experiments with real videos, whose result images are nearly equivalent to those yielded by classical MAP-based approaches. The presented algorithm can be applied in real-time recognition systems to improve their performances, and to reduce the requirement of imaging hardware.
Sidan Du
ICPR2