Mu Yang

dblp:32/4027 · DBLP profile ↗
← Back
35ranked-venue papers
12as first author
20since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Systems, architecture and hardware · 3 · 2 since 2021Security and privacy · 3 · 2 first-authorTheory of computation · 2Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
abstract
One challenge of integrating speech input with large language models (LLMs) stems from the discrepancy between the continuous nature of audio data and the discrete tokenbased paradigm of LLMs. To mitigate this gap, we propose a method for integrating vector quantization (VQ) into LLM-based automatic speech recognition (ASR). Using the LLM embedding table as the VQ codebook, the VQ module aligns the continuous representations from the audio encoder with the discrete LLM inputs, enabling the LLM to operate on a discretized audio representation that better reflects the linguistic structure. We further create a “soft discretization” of the audio representation by updating the codebook and performing a weighted sum over the codebook embeddings. Empirical results demonstrate that our proposed method significantly improves upon the LLMbased ASR baseline, particularly in out-of-domain conditions. This work highlights the potential of soft discretization as a modality bridge in LLM-based ASR.
Mu Yang, Szu-Jui Chen, Jiamin Xie, John H. L. Hansen
ASRU1
2025 UniScene: Unified Occupancy-centric Driving Scene Generation
abstract
Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles to model the direct layout-to-data distribution. In this paper, we introduce UniScene, the first unified framework for generating three key data forms — semantic occupancy, video, and LiDAR — in driving scenes. UniScene employs a progressive generation process that decomposes the complex task of scene generation into two hierarchical steps: (a) first generating semantic occupancy from a customized scene layout as a meta scene representation rich in both semantic and geometric information, and then (b) conditioned on occupancy, generating video and LiDAR data, respectively, with two novel transfer strategies of Gaussian-based Joint Rendering and Prior-guided Sparse Modeling. This occupancy-centric approach reduces the generation burden, especially for intricate scenes, while providing detailed intermediate representations for the subsequent generation stages. Extensive experiments demonstrate that UniScene outperforms previous SOTAs in the occupancy, video, and LiDAR generation, which also indeed benefits downstream driving tasks. The Project is available at https://arlo0o.github.io/uniscene/.
Bohan Li 0015, Jiazhe Guo, Hongsi Liu, Yingshuang Zou, Yikang Ding, Xiwu Chen, Hu Zhu, Feiyang Tan, Tiancai Wang, Shuchang Zhou 0001, Li Zhang 0040, Xiaojuan Qi 0001, Hao Zhao 0002, Mu Yang, Wenjun Zeng 0001, Xin Jin 0014
CVPR15
2025 A Neural Codec Approach for Noise-Robust Bandwidth Expansion
Mu Yang, Szu-Jui Chen, John H. L. Hansen
INTERSPEECH2
2025 Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
Mu Yang, Bowen Shi 0002, Matt Le 0001, Wei-Ning Hsu, Andros Tjandra
INTERSPEECH1
2025 MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving
abstract
This paper introduces MCTrack, a new 3D multi-object tracking method that achieves performance across KITTI, nuScenes, and Waymo datasets. Addressing the gap in existing tracking paradigms, which often perform well on specific datasets but lack generalizability, MCTrack offers a unified solution. Additionally, we have standardized the format of perceptual results across various datasets, termed BaseVersion, facilitating researchers in the field of MOT) to concentrate on the core algorithmic development without the undue burden of data preprocessing. Finally, recognizing the limitations of current evaluation metrics, we introduce a novel set of metrics designed to evaluate the output of motion information, including velocity and acceleration, which are essential for subsequent tasks. The source codes of the proposed method are available at this link: https://github.com/megvii-research/MCTrack
Xiyang Wang 0002, Shouzheng Qi, Jieyou Zhao, Hangning Zhou, Siyu Zhang 0002, Guoan Wang, Kai Tu, Songlin Guo, Jianbo Zhao 0001, Hailong Qin, Mu Yang
IROS12
2024 Diarist: Streaming Speech Translation with Speaker Diarization
abstract
End-to-end speech translation (ST) for conversation recordings involves several under-explored challenges such as speaker diarization (SD) without accurate word time stamps and handling of overlapping speech in a streaming fashion. In this work, we propose DiariST, the first streaming ST and SD solution. It is built upon a neural transducer-based streaming ST system and integrates tokenlevel serialized output training and t-vector, which were originally developed for multi-talker speech recognition. Due to the absence of evaluation benchmarks in this area, we develop a new evaluation dataset, DiariST-AliMeeting, by translating the reference Chinese transcriptions of the AliMeeting corpus into English. We also propose new metrics, called speaker-agnostic BLEU and speaker-attributed BLEU, to measure the ST quality while taking SD accuracy into account. Our system achieves a strong ST and SD capability compared to offline systems based on Whisper, while performing streaming inference for overlapping speech. To facilitate the research in this new direction, we release the evaluation data, the offline baseline systems, and the evaluation code.
Mu Yang, Naoyuki Kanda, Xiaofei Wang 0009, Jun-Kun Chen, Jinyu Li 0001, Takuya Yoshioka
ICASSP1
2024 Prioritizing user requirements for digital products using explainable artificial intelligence: A data-driven analysis on video conferencing apps
abstract
The advent of Industry 5.0 has brought a wealth of digital information to mobile app stores. With the help of emerging technologies such as machine learning and explainable artificial intelligence (XAI), these large amounts of user-generated data can be efficiently captured and analyzed. In this study, we propose an app store analysis framework and demonstrate the utility of the framework by mining and prioritizing user requirements in three popular video conferencing apps. We used the Sentistrength sentiment analysis tool, structural topic modeling, the Gephi web analysis tool, machine learning, and XAI techniques to conduct an in-depth analysis of user requirements in Microsoft Teams, ZOOM Cloud Meetings, and Google Meet. The findings indicated that Steal data, Audio and video quality, Customer service, Hacker issues, Meeting and account passwords, Mute and unmute, Features, and Office platform were the web conferencing system's key areas for improvement. The study demonstrated the usability of app store analysis frameworks and the great potential of XAI to provide insights about requirements prioritization by interpreting machine learning models. Additionally, it offered valuable suggestions for app developers on using the massive data in app stores to improve their apps.
Shizhen Bai, Songlin Shi, Chunjia Han, Mu Yang, Brij B. Gupta, Varsha Arya
Future Gener. Comput. Syst.4
2024 Risk disclosure and entrepreneurial resource acquisition in crowdfunding digital platforms: Evidence from digital technology ventures
abstract
The widespread development of digital technology facilitates the emergence of new entrepreneurial modes, of which crowdfunding digital platforms are one. In the digital environment of crowdfunding platforms, digital entrepreneurs can obtain the essential resources necessary for their startups' rapid and cost-effcient development. However, the information asymmetry derived from the digital nature of crowdfunding platforms leads to a lower chance of success for entrepreneurial ventures in this market, especially those in digital technology, limiting the important role that crowdfunding platforms can play in digital entrepreneurship. To this end, we focus on the risk disclosure section introduced by crowdfunding platforms to alleviate information asymmetry and explore the influence mechanism of the content of risk disclosure on entrepreneurial resource acquisition in crowdfunding digital platforms. By employing a novel text mining technique structural topic modelling, we analyse the risk disclosure texts of 4,284 digital technology crowdfunding projects and successfully identify various factors that constrain the development of digital technology ventures in crowdfunding platforms. Furthermore, we find that the risk topics digital entrepreneurs disclose negatively affect entrepreneurial resource acquisition. However, this relationship is moderated by the reward structure setting in the context of reward-based crowdfunding. The findings of this study not only enrich the literature on crowdfunding and digital entrepreneurship but also provide valuable practical implications on how crowdfunding digital platforms can be used to promote the development of digital entrepreneurship.
Chunjia Han, Mu Yang, Wen-Long Shang
Inf. Process. Manag.4
2024 Unveiling the Evolution of Enterprise Digital Innovation Strategies: Insights From U.S.-Listed Companies' Annual Reports
abstract
This article introduces a new metric for evaluating digital innovation in enterprise transformation using textual analysis of annual reports from U.S.-listed companies. Through network analysis and topic modeling, we identified 12 topics categorized into three main areas: digital technology innovation, customer-oriented digital strategy, and digital transformation in traditional business operations. Our research indicates that digital innovation strategies are critical for maintaining competitiveness and have shifted to a more innovation management-oriented approach. We also found differences in digital innovation strategies between companies and industries. Our study contributes to the theoretical significance of enterprise management and sustainable development.
Shizhen Bai, Yongbo Tan, Chunjia Han, Mu Yang, Brij B. Gupta, Varsha Arya, Neeraj Kumar 0001
IEEE Trans. Comput. Soc. Syst.4
2024 Learning What and Where to Learn: A New Perspective on Self-Supervised Learning
abstract
Self-supervised learning (SSL) has demonstrated its power in generalized model acquisition by leveraging the discriminative semantic and explicit positional information of unlabeled datasets. Unfortunately, mainstream contrastive learning-based methods excessive focus on semantic information and ignore the position is also the carrier of image content, resulting in inadequate data utilization and extensive computational consumption. To address these issues, we present an efficient SSL framework, learning What and Where to learn (W2SSL), to aggregate semantic and position features. Concretely, we devise a spatially-coupled sampling manner to process images through pre-defined rules, which integrates the advantage of semantic (What) and positional (Where) features into framework to enrich the diversity of feature representation capabilities and improve data utilization. Besides, a spectrum of latent vectors is obtained by mapping the positional features, which implicitly explores the relationship between these vectors. Whereafter, the corresponding discriminative and contrastive optimization objectives are seamlessly embedded in the framework via a cascade paradigm to explore semantic and positional features. The proposed W2SSL is verified on different types of datasets, which demonstrates that it still outperforms state-of-the-art SSL methods even with half the computational consumption. Code will be available at https://github.com/WilyZhao8/W2SSL.
Wenyi Zhao, Lu Yang 0006, Weidong Zhang 0007, Yongqin Tian, Wenhe Jia, Wei Li 0243, Mu Yang, Xipeng Pan
IEEE Trans. Circuits Syst. Video Technol.7
2023 Learning ASR Pathways: A Sparse Multilingual ASR Model
abstract
Neural network pruning compresses automatic speech recognition (ASR) models effectively. However, in multilingual ASR, language-agnostic pruning may lead to severe performance drops on some languages because language-agnostic pruning masks may not fit all languages and discard important language-specific parameters. In this work, we present ASR pathways, a sparse multilingual ASR model that activates language-specific sub-networks ("pathways"), such that the parameters for each language are learned explicitly. With the overlapping sub-networks, the shared parameters can also enable knowledge transfer for lower-resource languages via joint multilingual training. We propose a novel algorithm to learn ASR pathways, and evaluate the proposed method on 4 languages with a streaming RNN-T model. Our proposed ASR pathways outperform both dense models and a language-agnostically pruned model, and provide better performance on low-resource languages compared to the monolingual sparse models.
Mu Yang, Andros Tjandra, Chunxi Liu, Ozlem Kalinli
ICASSP1
2023 Assessment of Non-Native Speech Intelligibility using Wav2vec2-based Mispronunciation Detection and Multi-level Goodness of Pronunciation Transformer
Ram C. M. C. Shekar, Mu Yang, Kevin Hirschi, Stephen D. Looney, Okim Kang, John H. L. Hansen
INTERSPEECH2
2023 What Can an Accent Identifier Learn? Probing Phonetic and Prosodic Information in a Wav2vec2-based Accent Identification Model
Mu Yang, Ram C. M. C. Shekar, Okim Kang, John H. L. Hansen
INTERSPEECH1
2023 A semisupervised classification algorithm combining noise learning theory and a disagreement cotraining framework
Zaoli Yang, Chunjia Han, Yuchen Li 0002, Mu Yang, Petros Ieromonachou
Inf. Sci.5
2023 Exploring the Over-Time Variation in Customer Concerns on Sharing Economy Services
abstract
The sharing economy represented by Airbnb has evolved rapidly. It is particularly important to identify and understand how consumer concerns change over time. As a result, this study employs structural topic modelling using room type and time as covariates to extract topics from 896,658 Airbnb reviews in London and to observe the variation in the prevalence of topics over time. The findings show that the topic proportion changed relatively sharply in the early years of Airbnb (2010-2013) and during the COVID-19 pandemic (2020-2022), but relatively smoothly in the middle period (2014-2019). This research also discovered that the proportion of topics on customers' special experiences has been decreasing while the proportion of topics on their overall experience has been increasing. This shift could be attributed to an increase in the number of professional hosts, which has accelerated the standardisation of the Airbnb service.
Shizhen Bai, Xinrui Bi, Chunjia Han, Mu Yang, Hao He 0013
J. Glob. Inf. Manag.4
2022 Towards Lifelong Learning of Multilingual Text-to-Speech Synthesis
abstract
This work presents a lifelong learning approach to train a multilingual Text-To-Speech (TTS) system, where each language was seen as an individual task and was learned sequentially and continually. It does not require pooled data from all languages altogether, and thus alleviates the storage and computation burden. One of the challenges of lifelong learning methods is "catastrophic forgetting": in TTS scenario it means that model performance quickly degrades on previous languages when adapted to a new language. We approach this problem via a data-replay-based lifelong learning method. We formulate the replay process as a supervised learning problem, and propose a simple yet effective dual-sampler framework to tackle the heavily language-imbalanced training samples. Through objective and subjective evaluations, we show that this supervised learning formulation outperforms other gradient-based and regularization-based lifelong learning methods, achieving 43% Mel-Cepstral Distortion reduction compared to a fine-tuning baseline.
Mu Yang, Shaojin Ding, Tianlong Chen 0001, Zhangyang Wang
ICASSP1
2022 Joint Hypoglycemia Prediction and Glucose Forecasting via Deep Multi-Task Learning
abstract
We present a multitask learning approach to the problem of hypoglycemia (HG) prediction in diabetes. The approach is based on a state-of-the-art time series forecasting model, N-BEATS, and extends it by adding a classification task so that the model performs both glucose forecasting (i.e., predicting future glucose values) and HG prediction (i.e., probability of future HG events sometime within the prediction horizon). We also propose an alternative loss function that penalizes forecasting errors in the HG range. We evaluate the approach on a dataset containing over 1.6M recordings from 112 patients with type 1 diabetes who wore a continuous glucose monitor (CGM) for 90 days. Our results show that the classification branch significantly outperforms the forecasting branch on the problem of HG prediction, and that the new loss function is more effective at reducing forecasting errors in the HG range than multi-task learning.
Mu Yang, Darpit Dave, Madhav Erraguntla, Gerard L. Coté, Ricardo Gutierrez-Osuna
ICASSP1
2022 Improving Mispronunciation Detection with Wav2vec2-based Momentum Pseudo-Labeling for Accentedness and Intelligibility Assessment
abstract
Current leading mispronunciation detection and diagnosis (MDD) systems achieve promising performance via end-to-end phoneme recognition.One challenge of such end-to-end solutions is the scarcity of human-annotated phonemes on natural L2 speech.In this work, we leverage unlabeled L2 speech via a pseudo-labeling (PL) procedure and extend the fine-tuning approach based on pre-trained self-supervised learning (SSL) models.Specifically, we use Wav2vec 2.0 as our SSL model, and fine-tune it using original labeled L2 speech samples plus the created pseudo-labeled L2 speech samples.Our pseudo labels are dynamic and are produced by an ensemble of the online model on-the-fly, which ensures that our model is robust to pseudo label noise.We show that fine-tuning with pseudo labels achieves a 5.35% phoneme error rate reduction and 2.48% MDD F1 score improvement over a labeled-samples-only finetuning baseline.The proposed PL method is also shown to outperform conventional offline PL methods.Compared to the state-of-the-art MDD systems, our MDD solution produces a more accurate and consistent phonetic error diagnosis.In addition, we conduct an open test on a separate UTD-4Accents dataset, where our system recognition outputs show a strong correlation with human perception, based on accentedness and intelligibility.
Mu Yang, Kevin Hirschi, Stephen D. Looney, Okim Kang, John H. L. Hansen
INTERSPEECH1
2022 Research on Dual Channel Supply Chain Decision Making of New Retailing Enterprises Considering Service Behavior in the Era of Big Data
abstract
Drawing from extant retailing and supply chain research, this paper studies the dual channel supply chain decision-making of member channel, and obtains the optimal price strategy, maximum demand and maximum total revenue of the supply chain of network channel and retailing channel under the centralized decision-making and decentralized decision-making respectively. The contributions of this study identify that investing in big data within a certain threshold can improve the channel service level, reduce the channel price and improve the income of the supply chain. Supply chain members improve the channel service level and increase the corresponding channel price. The supply chain can get the most advantages when manufacturers and retailers make centralized decisions. This paper provides a starting point for new retailing academic and practical research in a domain that is deficient in empirical research, provides the theoretical framework to new retailing enterprises and decision-making model for their sustainable competitive advantage.
Di Rong, Chunjia Han, Mu Yang, Fengtao Liu
J. Glob. Inf. Manag.4
2021 Proximal Policy Optimization with Elo-based Opponent Selection and Combination with Enhanced Rolling Horizon Evolution Algorithm
abstract
Two-player zero-sum video game is a basic and important problem in game artificial intelligence. In 2020, enhanced rolling horizon evolution algorithm with policy gradient (ERHEAPI) beat heuristics, Monte-Carlo tree search and other methods to win the championship of Fighting Game Artificial Intelligence Competition (FTGAIC). However, the performance of ERHEAPI in the first round was not good. In this paper, we present an effective method noted as ERHEAPPO that combines proximal policy optimization (PPO) and enhanced rolling horizon evolution algorithm (ERHEA) with opponent model learning to further improve performance. We train the PPO agent and find that the Elo-based opponent selection can improve the sample efficiency. We compare the performance of the proposed ERHEAPPO with ERHEAPI. The experimental results demonstrate the effectiveness of ERHEAPPO.
Rongqin Liang, Yuanheng Zhu, Zhentao Tang, Mu Yang
CoG4
2020 Headword-Oriented Entity Linking: A Special Entity Linking Task with Dataset and Baseline
abstract
In this paper, we design headword-oriented entity linking (HEL), a specialized entity linking problem in which only the headwords of the entities are to be linked to knowledge bases; mention scopes of the entities do not need to be identified in the problem setting. This special task is motivated by the fact that in many articles referring to specific products, the complete full product names are rarely written; instead, they are often abbreviated to shorter, irregular versions or even just to their headwords, which are usually their product types, such as “stick” or “mask” in a cosmetic context. To fully design the special task, we construct a labeled cosmetic corpus as a public benchmark for this problem, and propose a product embedding model to address the task, where each product corresponds to a dense representation to encode the different information on products and their context jointly. Besides, to increase training data, we propose a special transfer learning framework in which distant supervision with heuristic patterns is first utilized, followed by supervised learning using a small amount of manually labeled data. The experimental results show that our model provides a strong benchmark performance on the special task.
Mu Yang, Chi-Yen Chen, Yi-Hui Lee, Qian-hui Zeng, Wei-Yun Ma, Chen-Yang Shih, Wei-Jhih Chen
LREC1
2020 An improved clear cell renal cell carcinoma stage prediction model based on gene sets
abstract
BACKGROUND: Clear cell renal cell carcinoma (ccRCC) is the most common subtype of renal cell carcinoma and accounts for cancer-related deaths. Survival rates are very low when the tumor is discovered in the late-stage. Thus, developing an efficient strategy to stratify patients by the stage of the cancer and inner mechanisms that drive the development and progression of cancers is critical in early prevention and treatment. RESULTS: In this study, we developed new strategies to extract important gene features and trained machine learning-based classifiers to predict stages of ccRCC samples. The novelty of our approach is that (i) We improved the feature preprocessing procedure by binning and coding, and increased the stability of data and robustness of the classification model. (ii) We proposed a joint gene selection algorithm by combining the Fast-Correlation-Based Filter (FCBF) search with the information value, the linear correlation coefficient, and variance inflation factor, and removed irrelevant/redundant features. Then the logistic regression-based feature selection method was used to determine influencing factors. (iii) Classification models were developed using machine learning algorithms. This method is evaluated on RNA expression value of clear cell renal cell carcinoma derived from The Cancer Genome Atlas (TCGA). The results showed that the result on the testing set (accuracy of 81.15% and AUC 0.86) outperformed state-of-the-art models (accuracy of 72.64% and AUC 0.81) and a gene set FJL-set was developed, which contained 23 genes, far less than 64. Furthermore, a gene function analysis was used to explore molecular mechanisms that might affect cancer development. CONCLUSIONS: The results suggested that our model can extract more prognostic information, and is worthy of further investigation and validation in order to understand the progression mechanism.
Fangjun Li, Mu Yang, Mingqiang Zhang, Dongfeng Yuan, Dongqi Tang
BMC Bioinform.2
2019 Deep Structured Neural Network for Event Temporal Relation Extraction
abstract
We propose a novel deep structured learning framework for event temporal relation extraction.The model consists of 1) a recurrent neural network (RNN) to learn scoring functions for pair-wise relations, and 2) a structured support vector machine (SSVM) to make joint predictions.The neural network automatically learns representations that account for long-term contexts to provide robust features for the structured model, while the SSVM incorporates domain knowledge such as transitive closure of temporal relations as constraints to make better globally consistent decisions.By jointly training the two components, our model combines the benefits of both data-driven learning and knowledge exploitation.Experimental results on three highquality event temporal relation datasets (TCR, MATRES, and TB-Dense) demonstrate that incorporated with pre-trained contextualized embeddings, the proposed model achieves significantly better performances than the stateof-the-art methods on all three datasets.We also provide thorough ablation studies to investigate our model.
Rujun Han, I-Hung Hsu, Mu Yang, Aram Galstyan, Ralph M. Weischedel, Nanyun Peng 0001
CoNLL3
2019 Spoken Language Intent Detection Using Confusion2Vec
abstract
Decoding speaker's intent is a crucial part of spoken language understanding (SLU). The presence of noise or errors in the text transcriptions, in real life scenarios make the task more challenging. In this paper, we address the spoken language intent detection under noisy conditions imposed by automatic speech recognition (ASR) systems. We propose to employ confusion2vec word feature representation to compensate for the errors made by ASR and to increase the robustness of the SLU system. The confusion2vec, motivated from human speech production and perception, models acoustic relationships between words in addition to the semantic and syntactic relations of words in human language. We hypothesize that ASR often makes errors relating to acoustically similar words, and the confusion2vec with inherent model of acoustic relationships between words is able to compensate for the errors. We demonstrate through experiments on the ATIS benchmark dataset, the robustness of the proposed model to achieve state-of-the-art results under noisy ASR conditions. Our system reduces classification error rate (CER) by 20.84% and improves robustness by 37.48% (lower CER degradation) relative to the previous state-of-the-art going from clean to noisy transcripts. Improvements are also demonstrated when training the intent detection models on noisy transcripts.
Prashanth Gurunath Shivakumar, Mu Yang, Panayiotis G. Georgiou
INTERSPEECH2
2017 A Distributed Infrastructure for Democratic Cloud Federations
abstract
Cloud federation is a novel concept that has been drawing attention from research and industry. However, there is a lack of solid proposal that can be widely adopted in practice to guarantee adequate governance of federations, especially in the Public Sector contexts due to legal requirements. In this paper, we propose an innovative governance approach that ensures distributed and democratic control in cloud federations. Starting from FaaS, a recent cloud federation proposal, we propose a blockchain infrastructure for the federation registry that implements the proposed governance approach.
Andrea Margheri, Md Sadek Ferdous, Mu Yang, Vladimiro Sassone
CLOUD3
2017 Enabling End-Users to Protect their Privacy
abstract
In this paper we present our ongoing work to build an approach to empower users of IoT-based cyber physical systems to protect their privacy by themselves. Our approach allows users to identify the privacy risks involved in sharing private data with a data consumer, assess the value of their private data based on identified risks and take a pragmatic data sharing decision balancing the risks with the benefits generated by the sharing. Our approach features a knowledgebase, called the Privacy Oracle, that exploits the power of the Semantic Web to determine how raw metadata can be combined by data consumers to infer privacy-sensitive information as well as the privacy risks associated with the disclosure of inferred information.
Mahmoud Barhamgi, Mu Yang, Chia-Mu Yu, Yijun Yu 0001, Arosha K. Bandara, Djamal Benslimane, Bashar Nuseibeh
AsiaCCS2
2017 Decentralised Runtime Monitoring for Access Control Systems in Cloud Federations
abstract
Cloud federation is an emergent cloud-computing paradigm where partner organisations share data and services hosted on their own cloud platforms. In this context, it is crucial to enforce access control policies that satisfy data protection and privacy requirements of partner organisations. However, due to the distributed nature of cloud federations, the access control system alone does not guarantee that its deployed components cannot be circumvented while processing access requests. In order to promote accountability and reliability of a distributed access control system, we present a decentralised runtime monitoring architecture based on blockchain technology.
Md Sadek Ferdous, Andrea Margheri, Federica Paci, Mu Yang, Vladimiro Sassone
ICDCS4
2017 Learning to share: engineering adaptive decision-support for online social networks
abstract
Some online social networks (OSNs) allow users to define friendship-groups as reusable shortcuts for sharing information with multiple contacts. Posting exclusively to a friendship-group gives some privacy control, while supporting communication with (and within) this group. However, recipients of such posts may want to reuse content for their own social advantage, and can bypass existing controls by copy-pasting into a new post; this cross-posting poses privacy risks. This paper presents a learning to share approach that enables the incorporation of more nuanced privacy controls into OSNs. Specifically, we propose a reusable, adaptive software architecture that uses rigorous runtime analysis to help OSN users to make informed decisions about suitable audiences for their posts. This is achieved by supporting dynamic formation of recipient-groups that benefit social interactions while reducing privacy risks. We exemplify the use of our approach in the context of Facebook.
Yasmin Rafiq, Luke Dickens, Alessandra Russo, Arosha K. Bandara, Mu Yang, Avelie Stuart, Mark Levine, Gül Çalikli, Blaine A. Price, Bashar Nuseibeh
ASE5
2015 Randomized spatial pooling in deep convolutional networks for scene recognition
abstract
The spatial layout of scene images is essential to recognizing them. Without considering spatial layout information, the deep convolutional neural networks could not achieve satisfied performance on scene recognition. In this paper a novel network architecture, namely the randomized spatial pooling (RS-Pooling) layer, is proposed to incorporate the spatial layout information into the deep model. By partitioning the feature maps via randomized patterns, the RS-Pooling layer offers the probability to handle various image layouts. Moreover, a maxout objective function is adopted to adaptively choose the optimal partition pattern to characterize the image layout. The experimental results on the scene recognition benchmarks demonstrate the effectiveness of the proposed RS-Pooling architecture.
Mu Yang, Brian Li, Haoqiang Fan, Yuning Jiang 0001
ICIP1
2015 An analysis of trust in anonymity networks in the presence of adaptive attackers
abstract
Anonymity is a security property of paramount importance, as we move steadily towards a wired, online community. Its import touches upon subjects as different as eGovernance, eBusiness and eLeisure, as well as personal freedom of speech in authoritarian societies. Trust metrics are used in anonymity networks to support and enhance reliability in the absence of verifiable identities, and a variety of security attacks currently focus on degrading a user's trustworthiness in the eyes of the other users.
Sardaouna Hamadou, Vladimiro Sassone, Mu Yang
Math. Struct. Comput. Sci.3
2014 Learning Compact Face Representation: Packing a Face into an int32
abstract
This paper addresses the problem of producing very compact representation of a face image for large-scale face search and analysis tasks. In tradition, the compactness of face representation is achieved by a dimension reduction step after representation extraction. However, the dimension reduction usually degrades the discriminative ability of the original representation drastically. In this paper, we present a deep learning framework which optimizes the compactness and discriminative ability jointly. The learnt representation can be as compact as 32 bit (same as the int32) and still produce highly discriminative performance (91.4% on LFW benchmark). Based on the extreme compactness, we show that traditional face analysis tasks (e.g. gender analysis) can be effectively solved by a Look-Up-Table approach given a large-scale face data set.
Haoqiang Fan, Mu Yang, Zhimin Cao, Yuning Jiang 0001, Qi Yin
ACM Multimedia2
2014 Adaptive Sharing for Online Social Networks: A Trade-off Between Privacy Risk and Social Benefit
abstract
Online social networks such as Facebook allow users to control which friend sees what information, but it can be a laborious process for users to specify every receiver for each piece of information they share. Therefore, users usually group their friends into social circles, and select the most appropriate social circle to share particular information with. However, social circles are not formed for setting privacy policies, and even the most appropriate social circle still cannot adapt to the changes of users' privacy requirements influenced by the changes in context. This problem drives the need for better privacy control which can adaptively filter the members in a selected social circle to satisfy users' requirements while maintaining users' social needs. To enable such adaptive sharing, this paper proposes a utility-based trade-off framework that models users' concerns (i.e. Potential privacy risks) and incentives of sharing (i.e. Potential social benefits), and quantifies users' requirements as a trade-off between these two types of utilities. By balancing these two metrics, our framework suggests a subset of a selected circle that aims to maximise users' overall utility of sharing. Numerical simulation results compare the outcome of three sharing strategies in randomly changing contexts.
Mu Yang, Yijun Yu 0001, Arosha K. Bandara, Bashar Nuseibeh
TrustCom1
2012 Design and implementation of the highly-reliable, low-cost housekeeping system in the ZDPS-1A pico-satellite
abstract
The ZDPS-1A pico-satellite designed in Zhejiang University with a mass of 3.5 kg and a power consumption of less than 3.5 W is the smallest satellite in China up to now. The housekeeping system (HKS) is the core part of ZDPS-1A. The reliability of HKS has an important influence on the safety of the satellite. Traditional fault-tolerant methods do not apply to ZDPS-1A due to such pico-satellite characteristics as light weight, compactness in size, energy saving, and high integration. This paper deals with a highly-reliable, low-cost design for HKS using industrial devices. The reliable strategies of HKS include a dual modular redundancy scheme, CPU warm backup, a static triple modular redundancy scheme, and two-level watchdogs. Recursive experiments, special tests, and environmental tests show that this system meets the design target. This design has already been applied to ZDPS-1A, which was launched to execute in-orbit tasks on Sept. 22, 2010. To date, the satellite has been in a proper state for more than 15 months.
Yang-Ming Zheng, Mu Yang
J. Zhejiang Univ. Sci. C3
2011 Minimising Anonymity Loss in Anonymity Networks under DoS Attacks
Mu Yang, Vladimiro Sassone
ICICS1
2010 Trust in Anonymity Networks
Vladimiro Sassone, Sardaouna Hamadou, Mu Yang
CONCUR3