Bowen Du 0002

dblp:94/4213-2 · DBLP profile ↗
← Back
40ranked-venue papers
0as first author
24since 2021 · last 2026
0000-0002-3755-4870ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 8 since 2021Software engineering, systems software and programming languages · 9 · 5 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021Computer networks · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SART: Sign-Absolute Reformulation Theory for Binary Variable Reduction in Neural Network Verification
abstract
Complete formal verification of neural networks is crucial for their deployment in safety-critical domains. A key bottleneck stems from encoding complexity: traditional methods assign one binary variable per unstable ReLU neuron. We propose the Sign-Absolute Reformulation Theory (SART), which fundamentally breaks the conventional one-to-one mapping between unstable neurons and binary variables by establishing formal reducibility criteria. This allows for finer-grained modeling, where each unstable neuron corresponds on average to fewer than one binary variable, thereby reducing verification complexity at its source. Based on SART, we derive a theoretical lower bound on the number of binary variables required for complete verification and, under the assumption that 𝑃 ≠ 𝑁𝑃, prove that variables in the final layer can be compressed by 50%, while the number of variables in intermediate layers cannot be further reduced. To overcome the apparent “last-layer-only” limitation, we recast verification as a sequential process and, crucially, show that the gain lifts to the entire network: LayerABS, a SART-based progressive tightening verifier, iteratively treats intermediate layers as temporary final layers and propagates tight bounds that shrink the global search space and binaryvariable counts. Furthermore, we reveal a structural law influencing verification complexity: when the signs of weights of unstable neurons satisfy numerical symmetry, with positive and negative weights equal or differing by at most one, the worst-case verification complexity achieves the theoretical optimum, offering theoretical guidance for the design of verification-friendly architectures. As a general-purpose underlying encoding, the value of SART is independent of specific algorithms. To comprehensively evaluate its effectiveness, we first evaluate the abstraction-free SART encoding, and then integrate it with abstraction techniques to construct the complete verifier LayerABS and its incomplete variant Incomplete-LayerABS. Across benchmarks, our methods surpass state-of-the-art baselines, validating SART’s practical impact.
Jin Xu 0002, Miaomiao Zhang 0003, Bowen Du 0002
Proc. ACM Program. Lang.3
2025 ODCCMamba-Unet: A Mamba-Unet Based Model with Omnidirectional Divide-and-Conquer Scanning Mechanism for Remote Sensing Image Change Detection
Hongming Zhu, Qin Liu 0004, Qiulu Dai, Bowen Du 0002
PRCV (2)5
2025 Robust spatio-temporal demand prediction for bike-sharing systems with dynamic hierarchical structure
abstract
Abstract Bike-sharing systems have gained popularity as a solution for short-distance urban travel, highlighting the need for precise demand prediction. Most existing methods predict bike-sharing demand across all regions without adequately accounting for spatial heterogeneity, meaning they overlook that different regions may have distinct and skewed demand distributions. Moreover, these models frequently miss the temporal heterogeneity introduced by varying demand patterns, as they tend to model temporal correlations with a uniform parameter set across all time periods, failing to account for the dynamic fluctuations in demand over time. To address these shortcomings, we introduce a novel Robust Spatio-Temporal Demand Prediction (RST) model that enhances bike-sharing demand prediction. The proposed approach employs a convolutional spatio-temporal graph, which integrates the features from multiple sources, including Points of Interest (POIs), road networks, and weather. Upon that, it proposes a scheme to classify the stations based on their traffic irregularity. The proposed scheme captures the traffic dependencies between stations by generating an embedding sequence from the historical data and ignoring correlations to the low-relevant stations, ensuring accurate predictions. Furthermore, the model introduces a dynamic hierarchical structure, where an additional retraining layer is activated only for stations with underperforming predictions, leveraging data from similar stations to refine and improve accuracy. This two-layered hierarchical structure ensures robustness and adaptability in complex scenarios such as different weather conditions or events. Our methodology sets a new standard in demand prediction for bike-sharing, demonstrating significant improvements in prediction accuracy on real-world datasets from New York City’s Citi-Bike and Beijing’s BJ-Bike. The results validate our model’s superior performance, offering a scalable and effective strategy for enhancing urban mobility.
Bowen Du 0002, Man Luo 0001, Hongkai Wen 0001
CCF Trans. Pervasive Comput. Interact.2
2025 STGAN-CR: A Semantics-Aware Cloud Removal Network Integrating Swin Transformer and GANs for Remote Sensing Applications
Hongming Zhu, Zeju Wang, Manxin Xu, Jinfeng Jiang, Hongfei Fan, Qin Liu 0004, Bowen Du 0002
Int. J. Softw. Eng. Knowl. Eng.8
2025 SHBC-VDRM: Space-Hard Block Cipher for Video Digital Rights Management
abstract
Video has become the primary channel for people to access information and engage in leisure and entertainment activities. Video service platforms have emerged to meet these demands and create significant profits. These platforms aggregate diverse videos and offer subscription services, aiming to provide convenient and personalized viewing experiences. To safeguard the copyright of video content against unauthorized access, these platforms employ digital rights management (DRM) technologies. However, the use of DRM faces significant security challenges due to the presence of malware or malicious users, which can compromise the decryption program running on user devices and lead to illegal video content distribution. This paper proposes a high-performance encryption scheme for video DRM protection. We establish the security of our proposed scheme through rigorous mathematical proof and analyze its effectiveness in the context of the video DRM scenario. Furthermore, the experimental results demonstrate that our scheme outperforms other schemes. Overall, our proposed scheme offers a promising solution to the security challenges faced by DRM technologies in protecting video content. Its efficiency, security, and resilience to various attacks make it a viable option for video service platforms seeking to enhance the security of their digital content.
Yang Shi 0002, Qiaoliang Ouyang, Jinkun Wang, Guodong Ye, Bowen Du 0002, Jiayao Gao
IEEE Trans. Netw. Serv. Manag.6
2024 RSTIE-KGC: A Relation Sensitive Textual Information Enhanced Knowledge Graph Completion Model
abstract
Nowadays, many knowledge graph completion models are proposed to assist the automatic construction of large knowledge graphs. While knowledge graph embedding models and textual information enhanced models are two tendencies in this area, they both suffer from some shortcomings, for example, relying too much on one single modal information may cause the limitation of performance. However, only a few works attempt to integrate multiple modalities. Besides, we found that the semantic similarity between head and tail entities correlates to the enhancing effect of textual information, and the semantic similarity is also related to relation types. We call this phenomenon relational sensitivity. To address these issues, we propose a relation sensitive textual information enhanced knowledge graph completion model (RSTIE-KGC). In our work, by integrating pre-trained language models (PLM) with knowledge graph embedding models, we fuse structural information and textual information, taking advantage of both topological and semantic features. Instead of finetuning the whole PLMs as existing models do, we choose a more efficient way, using frozen PLMs followed by an adapter to fit the text embeddings to our tasks, so that we can reduce training costs and enlarge the scale of negative samples, which maintains accuracy and effectiveness of our model. Based on the relational sensitivity, we propose our RelRank block, which selectively enhances textual information by calculating and filtering mean rank proportion (MRP) score for each relation type, thereby making better use of beneficial semantic information and reducing the noise and redundancy caused by textual information. The prediction process makes more refined use of textual information, as a result improving the accuracy of the model in the link prediction task. We conducted link prediction experiments on two real-world datasets. On FB15k-237, our model outperformed the current state-of-art textual information enhanced models in both MRR and Hit@k metrics. On WN18RR, our model also showed a stable prediction performance, and the results of both MRR and Hits@1 metrics were better than the most textual information enhanced models.
Sapae Phyu, Qin Liu 0004, Bowen Du 0002, Hongming Zhu
CSCWD4
2024 STGAN-CR: A Swin Transformer-Enhanced GAN Framework for Effective Cloud Removal in Satellite Imagery
abstract
Advancements in satellite remote sensing have enhanced Earth observation, enabling the acquisition of images crucial for various remote sensing applications.However, cloud cover degrades image quality and obstructs surface information.Traditional cloud removal techniques struggle to restore both low-level and high-level features, especially under complex conditions.To address these challenges, we propose STGAN-CR, a novel framework integrating Swin Transformer with Generative Adversarial Networks (GANs) to optimize image detail recovery.Leveraging the Swin Transformer's global modeling capabilities, our method enhances feature extraction and restoration, overcoming limitations of conventional models.We introduce a new evaluation metric focused on scene classification accuracy postde-clouding to better assess practical utility.Extensive experiments and ablation studies show that STGAN-CR outperforms existing models in visual quality and classification performance.These advancements offer an effective solution for enhancing the quality and utility of remote sensing images, balancing the restoration of both low-level and high-level features, and providing more meaningful de-clouded images for downstream applications.
Hongming Zhu, Zeju Wang, Manxin Xu, Qin Liu 0004, Bowen Du 0002
SEKE6
2024 Space-Hard Obfuscation Against Shared Cache Attacks and its Application in Securing ECDSA for Cloud-Based Blockchains
abstract
In cloud computing environments, virtual machines (VMs) running on cloud servers are vulnerable to shared cache attacks, such as Spectre and Foreshadow. By exploiting memory sharing among VMs, these attacks can compromise cryptographic keys in software modules. Program obfuscation serves as a promising countermeasure against key compromises by transforming a program into an unintelligent form while preserving its functionality. Unfortunately, for certain cryptographic algorithms such as the digital signature schemes, it is extremely difficult to construct provably secure obfuscators using traditional obfuscation approaches. To address such a challenge, this study proposes a novel approach to construct obfuscators for cryptographic algorithms named space-hard obfuscation, which can mitigate the threats from adversaries with the capability of acquiring a limited size of memory in shared cache attacks. Considering the extensive use of the Elliptic Curve Digital Signature Algorithm (ECDSA) in cloud-based Blockchain-as-a-Service (BaaS) and its potential vulnerability to shared cache attacks, we construct an exemplary scheme with provable security using space-hard obfuscation for ECDSA. Experimental results have demonstrated the scheme's high efficiency on cloud servers, as well as its successful integration with Hyperledger Fabric and Ethereum, two widely used blockchain systems.
Yang Shi 0002, Tianyuan Luo, Xiong Jiang, Bowen Du 0002, Hongfei Fan
IEEE Trans. Cloud Comput.5
2024 D3C2-Net: Dual-Domain Deep Convolutional Coding Network for Compressive Sensing
abstract
By mapping iterative optimization algorithms into neural networks (NNs), deep unfolding networks (DUNs) exhibit well-defined and interpretable structures and achieve remarkable success in the field of compressive sensing (CS). However, most existing DUNs solely rely on the image-domain unfolding, which restricts the information transmission capacity and reconstruction flexibility, leading to their loss of image details and unsatisfactory performance. To overcome these limitations, this paper develops a dual-domain optimization framework that combines the priors of (1) image- and (2) convolutional-coding-domains and offers generality to CS and other inverse imaging tasks. By converting this optimization framework into deep NN structures, we present a Dual-Domain Deep Convolutional Coding Network (D3C2-Net), which enjoys the ability to efficiently transmit high-capacity self-adaptive convolutional features across all its unfolded stages. Our theoretical analyses and experiments on simulated and real captured data, covering 2D and 3D natural, medical, and scientific signals, demonstrate the effectiveness, practicality, superior performance, and generalization ability of our method over other competing approaches and its significant potential in achieving a balance among accuracy, complexity, and interpretability. Code is available at https://github.com/lwq20020127/D3C2-Net.
Bin Chen 0006, Shijie Zhao 0001, Bowen Du 0002, Yongbing Zhang 0002, Jian Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.5
2023 MTN: A Multi-Scale Transformer Network for Different Resolution Remote Sensing Images Change Detection
abstract
In the field of change detection, detecting changes in images with different resolutions is crucial for both long-term interval scenes and scenarios that require rapid detection. However, existing methods face two main issues. Firstly, they require more stringent prior knowledge. The SPM-based methods require pixel-level class labeling of high-resolution(HR) images, while the approaches based on image super-resolution require HR images corresponding to low-resolutionr(LR) images. Such prior knowledge is either costly for labeling or difficult to meet in realistic scenarios. The second is that redundant error accumulation affects detection accuracy. Whether in traditional sub-pixel mapping(SPM) methods or deep learning methods based on image super-resolution, the final detection results are obtained after generating HR images from LR images. The redundant error produced in this step will be accumulated and affect the final detection results. Although the unsupervised methods do not have this problem, the detection accuracy is not as good as that of the supervised. To address these issues, we propose a multi-scale Transformer network(MTN). This model first uses a multi-scale feature extractor(MFE) to extract multi-scale features and perform scale matching at the feature level. Then, the Transformer is used to extract long-range relationships of ground objects on the multi-scale features to enhance the features. Finally, the multi-scale features are fused, and a classifier composed of a convolutional network is used to obtain binary change detection results. In addition, we consider that the edges of objects may be affected during the scale matching process, and introduce a CEBoundary Loss to better detect object edges. The results on the LEVIR and Google datasets demonstrate the effectiveness of our proposed method. The source code of MTN is available at https://github.com/Gavin-debug/MultiResolutionCD.
Hongming Zhu, Guodong Wu, Zeju Wang, Manxin Xu, Qin Liu 0004, Sicong Liu 0001, Bowen Du 0002
SMC7
2023 Towards a model of human-cyber-physical automata and a synthesis framework for control policies
Xiaochen Tang, Miaomiao Zhang 0003, Wanwei Liu, Bowen Du 0002, Zhiming Liu 0001
J. Syst. Archit.4
2023 Automatic modelling and verification of Autosar architectures
Miaomiao Zhang 0003, Yu Teng, Hui Kong 0007, John W. Baugh Jr., Junri Mi, Bowen Du 0002
J. Syst. Softw.7
2023 Landslide-Induced Seismic Signal Clustering With Outlier Removal
abstract
Seismic clustering is a vital technique in seismology for identifying patterns in seismic events and providing insights into geological processes. However, its application to ongoing landslide-induced signals and impact of outliers have not received much research attention. This letter presents a novel consensus clustering strategy with outlier removal for landslide-induced seismic signals. The proposed approach incorporates a parameter setting method to improve clustering accuracy and robustness. Experimental results demonstrate that the proposed approach outperforms the state-of-the-art clustering methods.
Bowen Du 0002
IEEE Geosci. Remote. Sens. Lett.3
2023 Integrating Real-Time and Non-Real-Time Collaborative Programming: Workflow, Techniques, and Prototypes
abstract
Real-time collaborative programming enables a group of programmers to edit shared source code at the same time, which significantly complements the traditional non-real-time collaborative programming supported by version control systems. However, one critical issue with this emerging technique is the lack of integration with non-real-time collaboration. Specifically, contributions from multiple programmers in a real-time collaboration session cannot be distinguished and accurately recorded in the version control system. In this study, we propose a scheme that integrates real-time and non-real-time collaborative programming with a novel workflow, and contribute enabling techniques to realize such integration. As a proof-of-concept, we have successfully implemented two prototype systems named CoEclipse and CoIDEA, which allow programmers to closely collaborate in a real-time fashion while preserving the work's compatibility with traditional non-real-time collaboration. User evaluation and performance experiments have confirmed the feasibility of the approach and techniques, demonstrated the good system performance, and presented the satisfactory usability of the prototypes.
Batu Qi, Wenhua Xu, Bowen Du 0002, Hongfei Fan
Proc. ACM Hum. Comput. Interact.5
2023 Fleet Rebalancing for Expanding Shared e-Mobility Systems: A Multi-Agent Deep Reinforcement Learning Approach
abstract
The electrification of shared mobility has become popular across the globe. Many cities have their new shared e-mobility systems deployed, with continuously expanding coverage from central areas to the city edges. A key challenge in the operation of these systems is fleet rebalancing, i.e., how EVs should be repositioned to better satisfy future demand. This is particularly challenging in the context of expanding systems, because i) the range of the EVs is limited while charging time is typically long, which constrain the viable rebalancing operations; and ii) the EV stations in the system are dynamically changing, i.e., the legitimate targets for rebalancing operations can vary over time. We tackle these challenges by first investigating rich sets of data collected from a real-world shared e-mobility system for one year, analyzing the operation model, usage patterns and expansion dynamics of this new mobility mode. With the learned knowledge we design a high-fidelity simulator, which is able to abstract key operation details of EV sharing at fine granularity. Then we model the rebalancing task for shared e-mobility systems under continuous expansion as a Multi-Agent Reinforcement Learning (MARL) problem, which directly takes the range and charging properties of the EVs into account. We further propose a novel policy optimization approach with action cascading, which is able to cope with the expansion dynamics and solve the formulated MARL. We evaluate the proposed approach extensively, and experimental results show that our approach outperforms the state-of-the-art, offering significant performance gain in both satisfied demand and net revenue.
Man Luo 0001, Bowen Du 0002, Tianyou Song, Kun Li 0030, Hongming Zhu, Mark H. Birkin, Hongkai Wen 0001
IEEE Trans. Intell. Transp. Syst.2
2022 A Multiple Locking Group Scheme for Flexible Semantic Conflict Prevention in Real-Time Collaborative Programming
abstract
Real-time collaborative programming has attracted increasing attention and interest in recent years. To resolve the semantic conflict problem in real-time collaborative programming, a Dependency-based Automatic Locking (DAL) scheme was proposed in prior work. The DAL scheme prevents other collaborators from editing semantically related regions by automatically detecting depended regions and locking them. However, the DAL scheme lacks flexibility and is not well suited to the needs of programmers in a real-world development scenario. When the programmer switches to a new region, all previous locks are automatically released. For this reason, we propose the Multiple Locking Group (MLG) scheme, where each programmer can hold multiple locking groups and switch freely between multiple working regions. Accordingly, three release modes for releasing locking groups are proposed. Each programmer can customize the release modes in a fine-grained manner. In supporting the scheme, we have devised techniques and solutions, implemented a prototype system and conducted a preliminary user evaluation to validate the feasibility, effectiveness and usability of the MLG scheme.
Wenhua Xu, Hongguang Zhou, Bowen Du 0002, Hongfei Fan
CSCWD5
2022 Context-based Operation Merging in Real-Time Collaborative Programming Environments
abstract
Real-time collaborative programming environments support a team of programmers to edit shared source code at the same time, where each local editing operation is captured and immediately transmitted to remote sites in a fine-grained manner. However, under real-world network conditions, collaborators are usually plagued by data congestion and transmitted errors. In this study, we define and analyze the content relationships among operations, and propose a Context-based Operation Merging Algorithm (COMA). Technically, the COMA examines the content relationships among a series of editing operations and merges content-related operations. Powered by the COMA, real-time collaborative programming environments can significantly compress editing operations to cope with complex network situations (such as network fluctuation and interruption) and improve the user experience of real-time collaboration. The proposed COMA has been implemented in a real-time collaborative programming environment prototype, namely CoEclipse. Preliminary user evaluations, correctness analysis and performance evaluations have demonstrated the effectiveness, correctness, and efficiency of the algorithm.
Hongguang Zhou, Wenhua Xu, Bowen Du 0002, Hongfei Fan
CSCWD5
2022 Human-Cyber-Physical Automata and Their Synthesis
Miaomiao Zhang 0003, Wanwei Liu, Xiaochen Tang, Bowen Du 0002, Zhiming Liu 0001
ICTAC4
2022 Event-Stream Representation for Human Gaits Identification Using Deep Neural Networks
abstract
Dynamic vision sensors (event cameras) have recently been introduced to solve a number of different vision tasks such as object recognition, activities recognition, tracking, etc. Compared with the traditional RGB sensors, the event cameras have many unique advantages such as ultra low resources consumption, high temporal resolution and much larger dynamic range. However, these cameras only produce noisy and asynchronous events of intensity changes, i.e., event-streams rather than frames, where conventional computer vision algorithms can't be directly applied. In our opinion the key challenge for improving the performance of event cameras in vision tasks is finding the appropriate representations of the event-streams so that cutting-edge learning approaches can be applied to fully uncover the spatio-temporal information contained in the event-streams. In this paper, we focus on the event-based human gait identification task and investigate the possible representations of the event-streams when deep neural networks are applied as the classifier. We propose new event-based gait recognition approaches basing on two different representations of the event-stream, i.e., graph and image-like representations, and use graph-based convolutional network (GCN) and convolutional neural networks (CNN) respectively to recognize gait from the event-streams. The two approaches are termed as EV-Gait-3DGraph and EV-Gait-IMG. To evaluate the performance of the proposed approaches, we collect two event-based gait datasets, one from real-world experiments and the other by converting the publicly available RGB gait recognition benchmark CASIA-B. Extensive experiments show that EV-Gait-3DGraph achieves significantly higher recognition accuracy than other competing methods when sufficient training samples are available. However, EV-Gait-IMG converges more quickly than graph-based approaches while training and shows good accuracy with only few number of training samples (less than ten). So image-like presentation is preferable when the amount of training data is limited.
Yiran Shen 0001, Bowen Du 0002, Guangrong Zhao, Li-Zhen Cui 0001, Hongkai Wen 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Deployment Optimization for Shared e-Mobility Systems With Multi-Agent Deep Neural Search
abstract
Shared e-mobility services have been widely tested and piloted in cities across the globe, and already woven into the fabric of modern urban planning. This paper studies a practical yet important problem in those systems: how to deploy and manage their infrastructure across space and time, so that the services areubiquitousto the users whilesustainablein profitability. However, in real-world systems evaluating the performance of different deployment strategies and then finding the optimal plan is prohibitively expensive, as it is often infeasible to conduct many iterations of trial-and-error. We tackle this by designing a high-fidelity simulation environment, which abstracts the key operation details of the shared e-mobility systems at fine-granularity, and is calibrated using data collected from the real-world. This allows us to try out arbitrary deployment plans to learn the optimal given specific context, before actually implementing any in the real-world systems. In particular, we propose a novel multi-agent neural search approach, in which we design a hierarchical controller to produce tentative deployment plans. The generated deployment plans are then tested using a multi-simulation paradigm, i.e., evaluated in parallel, where the results are used to train the controller with deep reinforcement learning. With this closed loop, the controller can be steered to have higher probability of generating better deployment plans in future iterations. The proposed approach has been evaluated extensively in our simulation environment, and experimental results show that it outperforms baselines e.g., human knowledge, and state-of-the-art heuristic-based optimization approaches in both service coverage and net revenue.
Man Luo 0001, Bowen Du 0002, Konstantin Klemmer, Hongming Zhu, Hongkai Wen 0001
IEEE Trans. Intell. Transp. Syst.2
2021 Supporting Cross-Platform Real-Time Collaborative Programming: Architecture, Techniques, and Prototype System
Brian Chiu, Jinfeng Jiang, Bowen Du 0002, Hongfei Fan
CollaborateCom (2)6
2021 Hybrid Semantic Conflict Prevention in Real-Time Collaborative Programming
Wenhua Xu, Brian Chiu, Jinfeng Jiang, Bowen Du 0002, Hongfei Fan
CollaborateCom (2)6
2021 Conv-Reluplex : A Verification Framework For Convolution Neural Networks (S)
abstract
In recent years, machine learning has demonstrated impressive performance in many real-world tasks, especially in computer vision and natural language processing.However, to apply them in safety-critical systems one needs formal guarantees on the neural network outputs.The Reluplex tool is proposed to verify the safety of deep neural networks (DNNs), and in case the DNN fails to give a correct output, can generate adversarial examples.Since the tool can only handle DNNs, it is necessary to extend the tool to process image data.Therefore, in this paper, we propose the Conv-Reluplex framework, which is designed to verify the convolutional layer and pooling layer in convolutional neural networks(CNNs), and generate adversarial examples when classification is misguided.We conduct several experiments on MNIST to evaluate our approaches.The results show that the original CNN is improved using the adversarial examples generated by our tool, and the precision of classification can be increased significantly.
Jin Xu 0002, Zishan Li, Miaomiao Zhang 0003, Bowen Du 0002
SEKE4
2021 LeaD: Learn to Decode Vibration-based Communication for Intelligent Internet of Things
abstract
In this article, we propose, LeaD , a new vibration-based communication protocol to Lea rn the unique patterns of vibration to D ecode the short messages transmitted to smart IoT devices. Unlike the existing vibration-based communication protocols that decode the short messages symbol-wise, either in binary or multi-ary, the message recipient in LeaD receives vibration signals corresponding to bits-groups. Each group consists of multiple symbols sent in a burst and the receiver decodes the group of symbols as a whole via machine learning-based approach. The fundamental behind LeaD is different combinations of symbols (1 s or 0 s) in a group will produce unique and reproducible patterns of vibration. Therefore, decoding in vibration-based communication can be modeled as a pattern classification problem. We design and implement a number of different machine learning models as the core engine of the decoding algorithm of LeaD to learn and recognize the vibration patterns. Through the intensive evaluations on large amount of datasets collected, the Convolutional Neural Network (CNN)-based model achieves the highest accuracy of decoding (i.e., lowest error rate), which is up to 97% at relatively high bits rate of 40 bits/s. While its competing vibration-based communication protocols can only achieve transmission rate of 10 bits/s and 20 bits/s with similar decoding accuracy. Furthermore, we evaluate its performance under different challenging practical settings and the results show that LeaD with CNN engine is robust to poses, distances (within valid range), and types of devices, therefore, a CNN model can be generally trained beforehand and widely applicable for different IoT devices under different circumstances. Finally, we implement LeaD on both off-the-shelf smartphone and smart watch to measure the detailed resources consumption on smart devices. The computation time and energy consumption of its different components show that LeaD is lightweight and can run in situ on low-cost smart IoT devices, e.g., smartwatches, without accumulated delay and introduces only marginal system overhead.
Guangrong Zhao, Bowen Du 0002, Yiran Shen 0001, Zhenyu Lao, Li-Zhen Cui 0001, Hongkai Wen 0001
ACM Trans. Sens. Networks2
2020 Rebalancing Expanding EV Sharing Systems with Deep Reinforcement Learning
abstract
Electric Vehicle (EV) sharing systems have recently experienced unprecedented growth across the world. One of the key challenges in their operation is vehicle rebalancing, i.e., repositioning the EVs across stations to better satisfy future user demand. This is particularly challenging in the shared EV context, because i) the range of EVs is limited while charging time is substantial, which constrains the rebalancing options; and ii) as a new mobility trend, most of the current EV sharing systems are still continuously expanding their station networks, i.e., the targets for rebalancing can change over time. To tackle these challenges, in this paper we model the rebalancing task as a Multi-Agent Reinforcement Learning (MARL) problem, which directly takes the range and charging properties of the EVs into account. We propose a novel approach of policy optimization with action cascading, which isolates the non-stationarity locally, and use two connected networks to solve the formulated MARL. We evaluate the proposed approach using a simulator calibrated with 1-year operation data from a real EV sharing system. Results show that our approach significantly outperforms the state-of-the-art, offering up to 14% gain in order satisfied rate and 12% increase in net revenue.
Man Luo 0001, Tianyou Song, Kun Li 0030, Hongming Zhu, Bowen Du 0002, Hongkai Wen 0001
IJCAI6
2020 Reluplex made more practical: Leaky ReLU
abstract
In recent years, Deep Neural Networks (DNNs) have been experiencing rapid development and have been widely used in various fields. However, while DNNs have shown strong capabilities, their security problems have gradually been exposed. Therefore, the formal guarantee of neural network output is needed. Prior to the appearance of the Reluplex algorithm, the verification of DNNs was always a difficult problem. Reluplex algorithm is specially used to verify DNNs with ReLU activation function. This is an excellent and effective algorithm, but it cannot verify more activation functions. ReLU activation function will bring about "Dead Neuron" problem, and Leaky ReLU activation function can solve this problem, so it is necessary to verify DNNs based on Leaky ReLU activation function. Therefore, we propose the Leaky-Reluplex algorithm, which is based on the Reluplex algorithm. Leaky-Reluplex algorithm can verify DNNs based on Leaky ReLU activation function.
Jin Xu 0002, Zishan Li, Bowen Du 0002, Miaomiao Zhang 0003, Jing Liu 0012
ISCC3
2020 Modeling Relation Path for Knowledge Graph via Dynamic Projection
Hongming Zhu, Yizhi Jiang, Xiaowen Wang 0003, Hongfei Fan, Qin Liu 0004, Bowen Du 0002
SEKE6
2020 Securing Cyber-Physical Social Interactions on Wrist-Worn Devices
abstract
Since ancient Greece, handshaking has been commonly practiced between two people as a friendly gesture to express trust and respect, or form a mutual agreement. In this article, we show that such physical contact can be used to bootstrap secure cyber contact between the smart devices worn by users. The key observation is that during handshaking, although belonged to two different users, the two hands involved in the shaking events are often rigidly connected, and therefore exhibit very similar motion patterns. We propose a novel key generation system, which harvests motion data during user handshaking from the wrist-worn smart devices such as smartwatches or fitness bands, and exploits the matching motion patterns to generate symmetric keys on both parties. The generated keys can be then used to establish a secure communication channel for exchanging data between devices. This provides a much more natural and user-friendly alternative for many applications, e.g., exchanging/sharing contact details, friending on social networks, or even making payments, since it doesn’t involve extra bespoke hardware, nor require the users to perform pre-defined gestures. We implement the proposed key generation system on off-the-shelf smartwatches, and extensive evaluation shows that it can reliably generate 128-bit symmetric keys just after around 1s of handshaking (with success rate >99%), and is resilient to different types of attacks including impersonate mimicking attacks, impersonate passive attacks, or eavesdropping attacks. Specifically, for real-time impersonate mimicking attacks, in our experiments, the Equal Error Rate (EER) is only 1.6% on average. We also show that the proposed key generation system can be extremely lightweight and is able to run in-situ on the resource-constrained smartwatches without incurring excessive resource consumption.
Yiran Shen 0001, Bowen Du 0002, Weitao Xu, Chengwen Luo 0001, Bo Wei 0003, Li-Zhen Cui 0001, Hongkai Wen 0001
ACM Trans. Sens. Networks2
2019 EV-Gait: Event-Based Robust Gait Recognition Using Dynamic Vision Sensors
abstract
In this paper, we introduce a new type of sensing modality, the Dynamic Vision Sensors (Event Cameras), for the task of gait recognition. Compared with the traditional RGB sensors, the event cameras have many unique advantages such as ultra low resources consumption, high temporal resolution and much larger dynamic range. However, those cameras only produce noisy and asynchronous events of intensity changes rather than frames, where conventional vision-based gait recognition algorithms can’t be directly applied. To address this, we propose a new Event-based Gait Recognition (EV-Gait) approach, which exploits motion consistency to effectively remove noise, and uses a deep neural network to recognise gait from the event streams. To evaluate the performance of EV-Gait, we collect two event-based gait datasets, one from real-world experiments and the other by converting the publicly available RGB gait recognition benchmark CASIA-B. Extensive experiments show that EV-Gait can get nearly 96% recognition accuracy in the real-world settings, while on the CASIA-B benchmark it achieves comparable performance with state-of-the-art RGB-based gait recognition approaches.
Bowen Du 0002, Yiran Shen 0001, Guangrong Zhao, Hongkai Wen 0001
CVPR2
2019 High-Speed Rail Operating Environment Recognition Based on Neural Network and Adversarial Training
abstract
Neural network is one of the key technologies for deep learning. Experiments on some standard test datasets show that their recognition ability has reached the level of human beings. However, they are extremely vulnerable to adversarial examples, that is, adding some subtle perturbations to the input example can cause the model to give a wrong output with high confidence. In this paper, we propose a non-contact approach based on neural network and adversarial training to recognize the high-speed rail operating environment. We first built the environment dataset and trained neural network models to do the recognition. We found that our model had high prediction accuracy, but with poor security since it was easy to attack our model using Basic Iterative Methods (BIM). To improve its security, we performed adversarial training based on the adversarial training dataset we built. The evaluation experiments indicated that this approach could improve the security of our model at the same time ensuring the prediction accuracy on the original test dataset.
Xiaoxue Hou, Jie An 0001, Miaomiao Zhang 0003, Bowen Du 0002, Jing Liu 0012
ICTAI4
2019 Weakly Supervised Brain Lesion Segmentation via Attentional Representation Learning
Bowen Du 0002, Man Luo 0001, Hongkai Wen 0001, Yiran Shen 0001, Jianfeng Feng
MICCAI (3)2
2019 Generating commit messages from diffs using pointer-generator network
abstract
The commit messages in source code repositories are valuable but not easy to be generated manually in time for tracking issues, reporting bugs, and understanding codes. Recently published works indicated that the deep neural machine translation approaches have drawn considerable attentions on automatic generation of commit messages. However, they could not deal with out-of-vocabulary (OOV) words, which are essential context-specific identifiers such as class names and method names in code diffs. In this paper, we propose PtrGNCMsg, a novel approach which is based on an improved sequence-to-sequence model with the pointer-generator network to translate code diffs into commit messages. By searching the smallest identifier set with the highest probability, PtrGNCMsg outperforms recent approaches based on neural machine translation, and first enables the prediction of OOV words. The experimental results based on the corpus of diffs and manual commit messages from the top 2,000 Java projects in GitHub show that PtrGNCMsg outperforms the state-of-the-art approach with improved BLEU by 1.02, ROUGE-1 by 4.00 and ROUGE-L by 3.78, respectively.
Qin Liu 0004, Hongming Zhu, Hongfei Fan, Bowen Du 0002
MSR5
2019 Multistep Flow Prediction on Car-Sharing Systems: A Multi-Graph Convolutional Neural Network with Attention Mechanism
abstract
Multistep flow prediction is an essential task for the car-sharing systems.An accurate flow prediction model can help system operators to pre-allocate the cars to meet the demand of users.However, this task is challenging due to the complex spatial and temporal relations among stations.Existing works only considered temporal relations (e.g., using LSTM) or spatial relations (e.g., using CNN) independently.In this paper, we propose an attention multi-graph convolutional sequenceto-sequence model (AMGC-Seq2Seq), which is a novel deep learning model for multistep flow prediction.The proposed model uses the encoder-decoder architecture, wherein the encoder part, spatial and temporal relations are encoded simultaneously.Then the encoded information is passed to the decoder to generate multistep outputs.In this work, specific multiple graphs are constructed to reflect spatial relations from different aspects, and we model them by using the proposed multi-graph convolution.Attention mechanism is also used to capture the important relations from previous information.Experiments on a large-scale real-world car-sharing dataset demonstrate the effectiveness of our approach over state-of-the-art methods.
Qin Liu 0004, Hongming Zhu, Hongfei Fan, Tianyou Song, Bowen Du 0002
SEKE7
2019 Autonomous Learning for Face Recognition in the Wild via Ambient Wireless Cues
abstract
Facial recognition is a key enabling component for emerging Internet of Things (IoT) services such as smart homes or responsive offices. Through the use of deep neural networks, facial recognition has achieved excellent performance. However, this is only possibly when trained with hundreds of images of each user in different viewing and lighting conditions. Clearly, this level of effort in enrolment and labelling is impossible for wide-spread deployment and adoption. Inspired by the fact that most people carry smart wireless devices with them, e.g. smartphones, we propose to use this wireless identifier as a supervisory label. This allows us to curate a dataset of facial images that are unique to a certain domain e.g. a set of people in a particular office. This custom corpus can then be used to finetune existing pre-trained models e.g. FaceNet. However, due to the vagaries of wireless propagation in buildings, the supervisory labels are noisy and weak. We propose a novel technique, AutoTune, which learns and refines the association between a face and wireless identifier over time, by increasing the inter-cluster separation and minimizing the intra-cluster distance. Through extensive experiments with multiple users on two sites, we demonstrate the ability of AutoTune to design an environment-specific, continually evolving facial recognition system with entirely no user effort.
Xiaoxuan Lu 0001, Xuan Kan, Bowen Du 0002, Changhao Chen, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, John A. Stankovic
WWW3
2019 Multistep Flow Prediction on Car-Sharing Systems: A Multi-Graph Convolutional Neural Network with Attention Mechanism
abstract
Multistep flow prediction is an essential task for the car-sharing systems. An accurate flow prediction model can help system operators to pre-allocate the cars to meet the demand of users. However, this task is challenging due to the complex spatial and temporal relations among stations. Existing works only considered temporal relations (e.g. using LSTM) or spatial relations (e.g. using CNN) independently. In this paper, we propose an attention to multi-graph convolutional sequence-to-sequence model (AMGC-Seq2Seq), which is a novel deep learning model for multistep flow prediction. The proposed model uses the encoder–decoder architecture, wherein the encoder part, spatial and temporal relations are encoded simultaneously. Then the encoded information is passed to the decoder to generate multistep outputs. In this work, specific multiple graphs are constructed to reflect spatial relations from different aspects, and we model them by using the proposed multi-graph convolution. Attention mechanism is also used to capture the important relations from previous information. Experiments on a large-scale real-world car-sharing dataset demonstrate the effectiveness of our approach over state-of-the-art methods.
Hongming Zhu, Qin Liu 0004, Hongfei Fan, Tianyou Song, Bowen Du 0002
Int. J. Softw. Eng. Knowl. Eng.7
2019 Efficient Indoor Positioning with Visual Experiences via Lifelong Learning
abstract
Positioning with visual sensors in indoor environments has many advantages: it doesn’t require infrastructure or accurate maps, and is more robust and accurate than other modalities such as WiFi. However, one of the biggest hurdles that prevents its practical application on mobile devices is the time-consuming visual processing pipeline. To overcome this problem, this paper proposes a novel lifelong learning approach to enable efficient and real-time visual positioning. We explore the fact that when following a previous visual experience for multiple times, one could gradually discover clues on how to traverse it with much less effort, e.g., which parts of the scene are more informative, and what kind of visual elements we should expect. Such second-order information is recorded as parameters, which provide key insights of the context and empower our system to dynamically optimise itself to stay localised with minimum cost. We implement the proposed approach on an array of mobile and wearable devices, and evaluate its performance in two indoor settings. Experimental results show our approach can reduce the visual processing time up to two orders of magnitude, while achieving sub-metre positioning accuracy.
Hongkai Wen 0001, Ronald Clark, Sen Wang 0002, Xiaoxuan Lu 0001, Bowen Du 0002, Wen Hu 0001, Agathoniki Trigoni
IEEE Trans. Mob. Comput.5
2018 Deepauth: in-situ authentication for smartwatches via deeply learned behavioural biometrics
abstract
This paper proposes DeepAuth, an in-situ authentication framework that leverages the unique motion patterns when users entering passwords as behavioural biometrics. It uses a deep recurrent neural network to capture the subtle motion signatures during password input, and employs a novel loss function to learn deep feature representations that are robust to noise, unseen passwords, and malicious imposters even with limited training data. DeepAuth is by design optimised for resource constrained platforms, and uses a novel split-RNN architecture to slim inference down to run in real-time on off-the-shelf smartwatches. Extensive experiments with real-world data show that DeepAuth outperforms the state-of-the-art significantly in both authentication performance and cost, offering real-time authentication on a variety of smartwatches.
Xiaoxuan Lu 0001, Bowen Du 0002, Peijun Zhao, Hongkai Wen 0001, Yiran Shen 0001, Andrew Markham, Agathoniki Trigoni
UbiComp2
2018 Shake-n-Shack: Enabling Secure Data Exchange Between Smart Wearables via Handshakes
abstract
Since ancient Greece, handshaking has been commonly practiced between two people as a friendly gesture to express trust and respect, or form a mutual agreement. In this paper, we show that such physical contact can be used to bootstrap secure cyber contact between the smart devices worn by users. The key observation is that during handshaking, although belonged to two different users, the two hands involved in the shaking events are often rigidly connected, and therefore exhibit very similar motion patterns. We propose a novel Shake-n-Shack system, which harvests motion data during user handshaking from the wrist worn smart devices such as smartwatches or fitness bands, and exploits the matching motion patterns to generate symmetric keys on both parties. The generated keys can be then used to establish a secure communication channel for exchanging data between devices. This provides a much more natural and user-friendly alternative for many applications, e.g., exchanging/sharing contact details, friending on social networks, or even making payments, since it doesn't involve extra bespoke hardware, nor require the users to perform pre-defined gestures. We implement the proposed Shake-n-Shack1system on off-the-shelf smartwatches, and extensive evaluation shows that it can reliably generate 128-bit symmetric keys just after around 1s of handshaking (with success rate >99%), and is resilient to real-time mimicking attacks: in our experiments the Equal Error Rate (EER) is only 1.6% on average. We also show that the proposed Shake-n-Shack system can be extremely lightweight, and is able to run in-situ on the resource-constrained smartwatches without incurring excessive resource consumption.
Yiran Shen 0001, Fengyuan Yang 0001, Bowen Du 0002, Weitao Xu, Chengwen Luo 0001, Hongkai Wen 0001
PerCom3
2018 Automatic Face Recognition Adaptation via Ambient Wireless Identifiers
abstract
Face recognition is a key enabling service for smart-spaces, allowing building management agents to easily monitor 'who is where', anticipating user needs and tailoring their local environment and experiences. Although facial recognition, especially through the use of deep neural networks, has achieved stellar performance over large datasets, the majority of approaches require supervised learning, that is, to be trained with tens or hundreds of images of users in different poses and lighting conditions. In this paper, we motivate that this enrollment effort is unnecessary if the smart-space has access to a wireless identifier e.g., through a smart-phone's MAC address. By learning and refining the noisy and weak association between a user's smart-phone and facial images, AutoTune can fine-tune a deep neural network to tailor it to the environment, users and conditions of a particular camera or set of cameras.
Xiaoxuan Lu 0001, Peijun Zhao, Bowen Du 0002, Hongkai Wen 0001, Andrew Markham, Stefano Rosa, Agathoniki Trigoni
SenSys3
2017 Towards Self-supervised Face Labeling via Cross-modality Association
abstract
Face recognition has become the de facto authentication solution in a broad spectrum of applications, from smart buildings, to industrial monitoring and security services. However, in many of those real-world scenarios, tracking or identifying people with facial recognition is extremely challenging due to the variations in the environment such as lighting conditions, camera viewing angles and subject motion. For most of the state-of-the-art face recognition systems, they need to be trained on a large dataset containing a good variety of labelled face images to work well. However, collecting and manually labelling such datasets is difficult and time consuming, probably more so than developing the algorithms. In this paper, we propose a novel framework to automatically label user identities with their face images in smart spaces, exploiting the fact that the users tend to carry their smart devices while seen by the surveillance cameras. We evaluate our method on 10 users in a smart building setting, and the experimental results show that our method can achieve > 0.9 f1 score on average.
Xiaoxuan Lu 0001, Xuan Kan, Stefano Rosa, Bowen Du 0002, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni
SenSys4