VLDB 2026 Research / reviewers in the wild / expert
Yike Guo
dblp:g/YikeGuo · also Yi-Ke Guo
· DBLP profile ↗
216ranked-venue papers
6as first author
106since 2021 · last 2026
0000-0002-3075-2161ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 97 · 75 since 2021Applied, interdisciplinary, general and emerging computing · 44 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 27 since 2021Systems, architecture and hardware · 24 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 24 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 13 · 2 since 2021Security and privacy · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Theory of computation · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inference-time Scaling for Diffusion-based Audio Super-resolutionabstractDiffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album mastering, substantial computational budgets are available for achieving superior audio quality. However, while existing diffusion approaches typically increase sampling steps to improve quality, the performance remains fundamentally limited by the stochastic nature of the sampling process, leading to high-variance and quality-limited outputs. Here, rather than simply increasing the number of sampling steps, we propose a different paradigm through inference-time scaling for SR, which explores multiple solution trajectories during the sampling process. Different task-specific verifiers are developed, and two search algorithms, including the random search and zero-order search for SR, are introduced. By actively guiding the exploration of the high-dimensional solution space through verifier-algorithm combinations, we enable more robust and higher-quality outputs. Through extensive validation across diverse audio domains (speech, music, sound effects) and frequency ranges, we demonstrate consistent performance gains, achieving improvements of up to 9.70% in aesthetics, 5.88% in speaker similarity, 15.20% in word error rate, and 46.98% in spectral distance for speech SR from 4 kHz to 24 kHz, showcasing the effectiveness of our approach. Yizhu Jin, Zhen Ye 0006, Zeyue Tian, Haohe Liu, Qiuqiang Kong, Yike Guo, Wei Xue 0002 |
AAAI | 6 |
| 2026 | Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert MergingabstractMixture of Experts (MoE) LLMs face significant obstacles due to their massive parameter scale, which imposes memory, storage, and deployment challenges. Although recent expert merging methods aim to achieve greater efficiency by consolidating several experts, they are fundamentally hindered by parameter conflicts arising from expert specialization. In this paper, we present Sub-MoE, a novel MoE compression framework via Subspace Expert Merging. Our key insight is to perform joint Singular Value Decomposition (SVD) on concatenated expert weights, reducing conflicting parameters by extracting shared U-matrices while enabling effective merging of the expert-specific V components. Specifically, Sub-MoE consists of two innovative stages: (1) Adaptive Expert Clustering, which groups functionally coherent experts via K-means clustering based on cosine similarity of expert outputs; and (2) Subspace Expert Merging, which first performs Experts Union Decomposition to derive the shared U-matrix across experts in the same group, then applies frequency-based merging for individual V-matrices, and completes expert reconstruction using the merged V-matrix. In this way, we align and fuse experts in a shared subspace. Additionally, the framework can be extended with intra-expert compression for further inference optimization. Extensive experiments on Mixtral, DeepSeek, and Qwen-1.5/3 MoE LLMs demonstrate that our Sub-MoE significantly outperforms existing expert pruning and merging methods. Notably, our Sub-MoE maintains 96%/86% of original performance with 25%/50% expert reduction on Mixtral-8×7B in zero-shot benchmarks. Lujun Li 0001, Qiyuan Zhu, Xiaoyu Qin 0001, Wei Li 0286, Hao Gu 0001, Sirui Han, Yike Guo |
AAAI | 8 |
| 2026 | Outlier Matters: Efficient Long-to-Short Reasoning via Outlier-Guided Model MergingabstractLarge Reasoning Language Models (LRMs) have recently shown remarkable performance in complex reasoning tasks, but their extensive reasoning chains incur substantial computational overhead. To address this challenge, we propose Outlier-aware Reasoning Conciseness Adaptive Merge (ORCA), a novel plug-and-play model merging framework that leverages outlier activation patterns to fuse base models with reasoning models. Our ORCA introduces three key innovations: (1) adaptive alignment that reduces conflicts between disparate activation patterns during merging, (2) outlier-guided allocation that assigns merging coefficients proportional to each layer's reasoning importance as indicated by outlier concentrations, and (3) dynamic probe-based adjustment that adapts merging coefficients during inference based on input-specific activation characteristics. These strategies allow seamless integration into existing merging pipelines while creating unified models that maintain reasoning accuracy with significantly reduced response verbosity. Comprehensive evaluation across six benchmarks using Qwen and LLaMA models shows ORCA reduces average response length by 55% while improving accuracy by 2.4∼5.7% over existing methods. Qiyuan Zhu, Lujun Li 0001, Xiaoyu Qin 0001, Wei Li 0286, Hao Gu 0001, Sirui Han, Yike Guo |
AAAI | 9 |
| 2026 | Benchmarking Fine-Grained Error Detection in Multimodal ReasoningabstractChi-Min Chan, Han Zhu, Chunyang Jiang, Jiaming Ji, Juntao Dai, Wei Xue, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chi-Min Chan, Jiaming Ji, Juntao Dai, Wei Xue 0002, Sirui Han, Yike Guo |
ACL (1) | 8 |
| 2026 | Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across ModalitiesabstractChi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin, Jiaming Ji, Juntao Dai, Wei Xue, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin, Jiaming Ji, Juntao Dai, Wei Xue 0002, Sirui Han, Yike Guo |
ACL (1) | 9 |
| 2026 | BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary CodebookabstractHao Gu, Lujun Li, Hao Wang, Lei Wang, Zheyu Wang, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Gu 0001, Lujun Li 0001, Hao Wang 0097, Jiacheng Liu 0001, Qiyuan Zhu, Sirui Han, Yike Guo |
ACL (1) | 10 |
| 2026 | Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning ModelsabstractHao Wang, Hao Gu, Hongming Piao, Kaixiong Gong, Yuxiao Ye, Xiangyu Yue, Sirui Han, Yike Guo, Dapeng Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Wang 0193, Hao Gu 0001, Hongming Piao, Kaixiong Gong, Yuxiao Ye, Xiangyu Yue 0001, Sirui Han, Yike Guo, Dapeng Oliver Wu |
ACL (1) | 8 |
| 2026 | Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMsabstractBinxing Xu, Hao Gu, Lujun Li, Hao Wang, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Xintong Yang, Chao Li, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Binxing Xu, Hao Gu 0001, Lujun Li 0001, Hao Wang 0097, Jiacheng Liu 0001, Qiyuan Zhu, Xintong Yang, Chao Li 0009, Sirui Han, Yike Guo |
ACL (1) | 11 |
| 2026 | SafeMT: Multi-turn Safety for Multimodal Language ModelsabstractHan Zhu, Juntao Dai, Jiaming Ji, Haoran Li, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen, Yaodong Yang, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Juntao Dai, Jiaming Ji, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen 0008, Yaodong Yang 0001, Sirui Han, Yike Guo |
ACL (1) | 11 |
| 2026 | Reimagining Legal Fact Verification with GenAI: Toward Effective Human-AI CollaborationabstractFact verification is a critical yet underexplored component of non-litigation legal practice. While existing research has examined automation in legal workflow and human-AI collaboration in high-stakes domains, little is known about how GenAI can support fact verification, a task that demands prudent judgment and strict accountability. To address this, we conducted semi-structured interviews with 18 lawyers to understand their current verification practices, attitudes toward GenAI adoption, and expectations for future systems. We found that while lawyers use GenAI for low-risk tasks like drafting and language optimization, concerns over accuracy, confidentiality, and liability are currently limiting its adoption for fact verification. These concerns translate into core design requirements for AI systems that are trustworthy and accountable. Based on these, we contribute design insights for human-AI collaboration in legal fact verification, emphasizing the development of auditable systems that balance efficiency with professional judgment and uphold ethical and legal accountability in high-stakes practice. Sirui Han, Yuyao Zhang 0006, Yidan Huang, Chengzhong Liu, Yike Guo |
CHI | 6 |
| 2026 | CareerCraft: Supporting New Graduates on Job Hunting with LLM-Assisted Self-Construction of Career ProfileabstractStarting the job hunt is often challenging for new graduates, who face barriers in translating experiences into actionable career profiles due to limited self-awareness and unclear skill mapping. Through formative study with new graduates and early-career professionals, we concluded specific challenges in experience extraction, skill organization, and expressive confidence. Drawing on these insights, we designed CareerCraft, an interactive system that scaffolds the construction of coherent career stories and supports tailored job searching via experience card extraction, guided profile building, and LLM-powered recommendations. In a within-subject evaluation (N=16), participants rated the efficacy of CareerCraft against the baseline condition without the tool in improving profile structuring, clarifying their self-awareness and competencies, and supporting informed job direction choices. Based on the findings, we concluded that CareerCraft offered a promising pathway to career readiness among new graduates to the workforce. We further summarized the design considerations for LLM products emphasizing on users’ self-exploration. Xinyue Qi, Chengzhong Liu, Xiangyu Long, Zhizhuo Kou, Sirui Han, Yike Guo |
CHI | 6 |
| 2026 | Spatiotemporal Graph Learning with Direct Volumetric Information Passing and Feature EnhancementabstractData-driven learning of physical systems has attracted significant attention, where many neural models have been developed. In particular, mesh-based graph neural networks (GNNs) have demonstrated considerable potential in modeling spatiotemporal dynamics across arbitrary geometric domains. However, the existing node-edge message-passing and aggregation mechanism in GNNs limits the representation learning capability. In this paper, we propose a dual-module framework, Cell-embedded and Feature-enhanced Graph Neural Network (CeFeGNN), for learning spatiotemporal dynamics. Specifically, we embed learnable cell attributions to the common node-edge message passing process, thereby better capturing the spatial dependency of regional features. Such a strategy essentially upgrades the local aggregation scheme from first order (e.g., from edge to node) to a higher order (e.g., from volume and edge to node), which takes advantage of volumetric information in message passing. Meanwhile, a novel feature-enhanced block is designed to further improve the model's performance and alleviate the over-smoothing problem. Extensive experiments on various PDE systems and a real-world dataset demonstrate that CeFeGNN achieves superior performance compared with other baselines. Yuan Mi, Qi Wang 0123, Xueqin Hu, Yike Guo, Ji-Rong Wen, Yang Liu 0130, Hao Sun 0002 |
KDD (1) | 4 |
| 2026 | Adaptive weighted disentangling variational autoencoder with fine-grained feedback
Zhenyao Yu, Zitu Liu, Yike Guo, Qun Liu 0005, Guoyin Wang 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM PromptsabstractAbstract The potential for higher-resolution image generation using pretrained diffusion models is immense. However, these models often struggle with object repetition and structural artifacts especially when scaling to 4K resolution and beyond. Our analysis reveals that causes the problem, a single prompt for the generation of multiple scales provides insufficient efficacy. To address this, we propose HiPrompt, a new tuning-free solution that tackles the above problems by introducing hierarchical prompts. The hierarchical prompts provide both global and local semantic guidance. Specifically, the global prompt captures overall scene semantics from user input, while local guidance comes from patch-wise descriptions generated by MLLMs to refine regional structures and textures. Furthermore, during inverse denoising, noise is decomposed into low- and high-frequency components, each conditioned on different prompt levels, facilitating prompt-guided denoising under hierarchical semantic guidance. It further allows the generation to focus more on local spatial regions and ensures the generated images maintain coherent local and global semantics, structures, and textures with high definition. Extensive experiments demonstrate that HiPrompt outperforms state-of-the-art works in higher-resolution image generation, significantly reducing object repetition and enhancing structural quality. The demo and code can be found on the project website: https://liuxinyv.github.io/HiPrompt/ . Yingqing He, Lanqing Guo, Bu Jin, Chi-Min Chan, Wei Xue 0002, Wenhan Luo, Yike Guo |
Int. J. Comput. Vis. | 11 |
| 2026 | Improving Model Fusion by Training-Time Neuron Alignment With Fixed Neuron AnchorsabstractModel fusion aims to integrate several deep neural network (DNN) models' knowledge into one by fusing parameters, and it has promising applications, such as improving the generalization of foundation models and parameter averaging in federated learning. However, models under different settings (data, hyperparameter, etc.) have diverse neuron permutations; in other words, from the perspective of loss landscape, they reside in different loss basins, thus hindering model fusion performances. To alleviate this issue, previous studies highlighted the role of permutation invariance and have developed methods to find correct network permutations for neuron alignment after training. Orthogonal to previous attempts, this paper studies training-time neuron alignment, improving model fusion without the need for post-matching. Training-time alignment is cheaper than post-alignment and is applicable in various model fusion scenarios. Starting from fundamental hypotheses and theorems, a simple yet lossless algorithm called TNA-PFN is introduced. TNA-PFN utilizes partially fixed neuron weights as anchors to reduce the potential of training-time permutations, and it is empirically validated in reducing the barriers of linear mode connectivity and multi-model fusion. It is also validated that TNA-PFN can improve the fusion of pretrained models under the setting of model soup (vision transformers) and ColD fusion (pretrained language models). Based on TNA-PFN, two federated learning methods, FedPFN and FedPNU, are proposed, showing the prospects of training-time neuron alignment. FedPFN and FedPNU reach state-of-the-art performances in federated learning under heterogeneous settings and can be compatible with the server-side algorithm. Zexi Li 0001, Zhiqi Li 0004, Tao Shen 0002, Jun Xiao 0001, Yike Guo, Tao Lin 0004, Chao Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Importance Weighting Can Help Large Language Models Self-ImproveabstractLarge language models (LLMs) have shown remarkable capability in numerous tasks and applications. However, fine-tuning LLMs using high-quality datasets under external supervision remains prohibitively expensive. In response, LLM self-improvement approaches have been vibrantly developed recently. The typical paradigm of LLM self-improvement involves training LLM on self-generated data, part of which may be detrimental and should be filtered out due to the unstable data quality. While current works primarily employs filtering strategies based on answer correctness, in this paper, we demonstrate that filtering out correct but with high distribution shift extent (DSE) samples could also benefit the results of self-improvement. Given that the actual sample distribution is usually inaccessible, we propose a new metric called DS weight to approximate DSE, inspired by the Importance Weighting methods. Consequently, we integrate DS weight with self-consistency to comprehensively filter the self-generated samples and fine-tune the language model. Experiments show that with only a tiny valid set (up to 5% size of the training set) to compute DS weight, our approach can notably promote the reasoning ability of current LLM self-improvement methods. The resulting performance is on par with methods that rely on external supervision from pre-trained reward models. Chi-Min Chan, Wei Xue 0002, Yike Guo |
AAAI | 5 |
| 2025 | Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language ModelabstractRecent advancements in audio generation have been significantly propelled by the capabilities of Large Language Models (LLMs). The existing research on audio LLM has primarily focused on enhancing the architecture and scale of audio language models, as well as leveraging larger datasets, and generally, acoustic codecs, such as EnCodec, are used for audio tokenization. However, these codecs were originally designed for audio compression, which may lead to suboptimal performance in the context of audio LLM. Our research aims to address the shortcomings of current audio LLM codecs, particularly their challenges in maintaining semantic integrity in generated audio. For instance, existing methods like VALL-E, which condition acoustic token generation on text transcriptions, often suffer from content inaccuracies and elevated word error rates (WER) due to semantic misinterpretations of acoustic tokens, resulting in word skipping and errors. To overcome these issues, we propose a straightforward yet effective approach called X-Codec. X-Codec incorporates semantic features from a pre-trained semantic encoder before the Residual Vector Quantization (RVQ) stage and introduces a semantic reconstruction loss after RVQ. By enhancing the semantic ability of the codec, X-Codec significantly reduces WER in speech synthesis tasks and extends these benefits to non-speech applications, including music and sound generation. Our experiments in text-to-speech, music continuation, and text-to-sound tasks demonstrate that integrating semantic information substantially improves the overall performance of language models in audio generation. Zhen Ye 0006, Peiwen Sun, Jiahe Lei, Hongzhan Lin 0001, Xu Tan 0003, Zheqi Dai, Qiuqiang Kong, Jianyi Chen, Yike Guo, Wei Xue 0002 |
AAAI | 11 |
| 2025 | FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning EvaluationabstractJunyu Luo, Zhizhuo Kou, Liming Yang, Xiao Luo, Jinsheng Huang, Zhiping Xiao, Jingshu Peng, Chengzhong Liu, Jiaming Ji, Xuanzhe Liu, Sirui Han, Ming Zhang, Yike Guo. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Junyu Luo 0002, Zhizhuo Kou, Xiao Luo 0001, Jinsheng Huang, Zhiping Xiao 0001, Jingshu Peng, Chengzhong Liu, Jiaming Ji, Xuanzhe Liu, Sirui Han, Ming Zhang 0004, Yike Guo |
ACL (1) | 13 |
| 2025 | PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human PreferenceabstractJiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Alex Qiu, Jiayi Zhou, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen 0008, Josef Dai, Boren Zheng, Tianyi Qiu, Kaile Wang, Boxun Li, Sirui Han, Yike Guo, Yaodong Yang 0001 |
ACL (1) | 12 |
| 2025 | LegalReasoner: Step-wised Verification-Correction for Legal Judgment ReasoningabstractWeijie Shi, Han Zhu, Jiaming Ji, Mengze Li, Jipeng Zhang, Ruiyuan Zhang, Jia Zhu, Jiajie Xu, Sirui Han, Yike Guo. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiaming Ji, Mengze Li 0001, Ruiyuan Zhang, Jia Zhu 0003, Jiajie Xu 0001, Sirui Han, Yike Guo |
ACL (1) | 10 |
| 2025 | PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit RemeshingabstractPhotorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, existing methods for monocular full-body reconstruction, typically relying on front and/or predicted back view, still struggle with satisfactory performance due to the ill-posed nature of the problem and sophisticated self-occlusions. In this paper, we propose PSHuman, a novel framework that explicitly reconstructs human meshes utilizing priors from the multiview diffusion model. It is found that directly applying multiview diffusion on single-view human images leads to severe geometric distortions, especially on generated faces. To address it, we propose a cross-scale diffusion that models the joint probability distribution of global full-body shape and local facial characteristics, enabling identity-preserved novel-view generation without geometric distortion. Moreover, to enhance cross-view body shape consistency of varied human poses, we condition the generative model on parametric models (SMPL-X), which provide body priors and prevent unnatural views inconsistent with human anatomy. Leveraging the generated multiview normal and color images, we present SMPLX-initialized explicit human carving to recover realistic textured human meshes efficiently. Extensive experiments on CAPE and THuman2.1 demonstrate PSHuman’s superiority in geometry details, texture fidelity, and generalization capability. Wangguandong Zheng, Yuan Liu 0025, Tao Yu 0007, Yangguang Li 0001, Xingqun Qi, Xiaowei Chi, Si-Yu Xia, Yan-Pei Cao 0001, Wei Xue 0002, Wenhan Luo, Yike Guo |
CVPR | 12 |
| 2025 | VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term ModelingabstractIn this work, we systematically study music generation conditioned solely on the video. First, we present a large-scale dataset by collecting 360K video-music pairs, including various genres such as movie trailers, advertisements, and documentaries. Furthermore, we propose VidMuse, a simple framework for generating music aligned with video inputs. VidMuse stands out by producing high-fidelity music that is both acoustically and semantically aligned with the video. By incorporating local and global visual cues, VidMuse enables the creation of coherent music tracks that consistently match the video content through Long-Short-Term modeling. Through extensive experiments, VidMuse outperforms existing models in terms of audio quality, diversity, and audio-visual alignment. The code and datasets are available at https://vidmuse.github.io/ Zeyue Tian, Zhaoyang Liu 0001, Ruibin Yuan, Xu Tan 0003, Qifeng Chen 0001, Wei Xue 0002, Yike Guo |
CVPR | 9 |
| 2025 | Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem ProvingabstractChuxue Cao, Mengze Li, Juntao Dai, Jinluan Yang, Zijian Zhao, Shengyu Zhang, Weijie Shi, Chengzhong Liu, Sirui Han, Yike Guo. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chuxue Cao, Mengze Li 0001, Juntao Dai, Jinluan Yang, Zijian Zhao 0002, Shengyu Zhang 0001, Chengzhong Liu, Sirui Han, Yike Guo |
EMNLP | 10 |
| 2025 | Graceful Forgetting in Generative Language ModelsabstractRecently, the pretrain-finetune paradigm has become a cornerstone in various deep learning areas.While in general the pre-trained model would promote both effectiveness and efficiency of downstream tasks fine-tuning, studies have shown that not all knowledge acquired during pre-training is beneficial.Some of the knowledge may actually bring detrimental effects to the fine-tuning tasks, which is also known as negative transfer.To address this problem, graceful forgetting has emerged as a promising approach.The core principle of graceful forgetting is to enhance the learning plasticity of the target task by selectively discarding irrelevant knowledge.However, this approach remains underexplored in the context of generative language models, and it is often challenging to migrate existing forgetting algorithms to these models due to architecture incompatibility.To bridge this gap, in this paper we propose a novel framework, Learning With Forgetting (LWF), to achieve graceful forgetting in generative language models.With Fisher Information Matrix weighting the intended parameter updates, LWF computes forgetting confidence to evaluate selfgenerated knowledge regarding the forgetting task, and consequently, knowledge with high confidence is periodically unlearned during fine-tuning.Our experiments demonstrate that, although thoroughly uncovering the mechanisms of knowledge interaction remains challenging in pre-trained language models, applying graceful forgetting can contribute to enhanced fine-tuning performance. Chi-Min Chan, Yiyang Cai, Wei Xue 0002, Yike Guo |
EMNLP | 6 |
| 2025 | Efficient Fine-Tuning of Large Models Via Nested Low-Rank Adaptation
Lujun Li 0001, Cheng Lin 0001, You-Liang Huang, Wei Li 0286, Jie Zou 0001, Wei Xue 0002, Sirui Han, Yike Guo |
ICCV | 10 |
| 2025 | AIRA: Activation-Informed Low-Rank Adaptation for Large ModelsabstractLow-Rank Adaptation (LoRA) is a widely used method for efficiently fine-tuning large models by introducing lowrank matrices into weight updates. However, existing LoRA techniques fail to account for activation information, such as outliers, which significantly impact model performance. This omission leads to suboptimal adaptation and slower convergence. To address this limitation, we present Activation-Informed Low-Rank Adaptation (AIRA), a novel approach that integrates activation information into initialization, training, and rank assignment to enhance model performance. Specifically, AIRA introduces: (1) Outlierweighted SVD decomposition to reduce approximation errors in low-rank weight initialization, (2) Outlier-driven dynamic rank assignment using offline optimization for better layer-wise adaptation, and (3) Activation-informed training to amplify updates on significant weights. This cascaded activation-informed paradigm enables faster convergence and fewer fine-tuned parameters while maintaining high performance. Extensive experiments on multiple large models demonstrate that AIRA outperforms state-of-the-art LoRA variants, achieving superior performance-efficiency trade-offs in vision-language instruction tuning, few-shot learning, and image generation. Codes are available at https://github.com/lliai/LoRA-Zoo. Lujun Li 0001, Cheng Lin 0001, Wei Li 0286, Wei Xue 0002, Sirui Han, Yike Guo |
ICCV | 7 |
| 2025 | STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMsabstractIn this paper, we present the first structural binarization method for LLM compression to less than 1-bit precision. Although LLMs have achieved remarkable performance, their memory-bound nature during the inference stage hinders the adoption of resource-constrained devices. Reducing weights to 1-bit precision through binarization substantially enhances computational efficiency. We observe that randomly flipping some weights in binarized LLMs does not significantly degrade the model's performance, suggesting the potential for further compression. To exploit this, our STBLLM employs an N:M sparsity technique to achieve structural binarization of the weights. Specifically, we introduce a novel Standardized Importance (SI) metric, which considers weight magnitude and input feature norm to more accurately assess weight significance. Then, we propose a layer-wise approach, allowing different layers of the LLM to be sparsified with varying N:M ratios, thereby balancing compression and accuracy. Furthermore, we implement a fine-grained grouping strategy for less important weights, applying distinct quantization schemes to sparse, intermediate, and dense regions. Finally, we design a specialized CUDA kernel to support structural binarization. We conduct extensive experiments on LLaMA, OPT, and Mistral family. STBLLM achieves a perplexity of 11.07 at 0.55 bits per weight, outperforming the BiLLM by 3×. The results demonstrate that our approach performs better than other compressed binarization LLM methods while significantly reducing memory requirements. Code is released at https://github.com/pprp/STBLLM. Peijie Dong, Lujun Li 0001, Yuedong Zhong, Dayou Du, Ruibo Fan, Yuhan Chen 0008, Zhenheng Tang, Qiang Wang 0022, Wei Xue 0002, Yike Guo, Xiaowen Chu 0001 |
ICLR | 10 |
| 2025 | Co3Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion
Xingqun Qi, Yatian Wang, Wei Xue 0002, Shanghang Zhang, Wenhan Luo, Yike Guo |
ICLR | 9 |
| 2025 | Both Ears Wide Open: Towards Language-Driven Spatial Audio GenerationabstractRecently, diffusion models have achieved great success in mono-channel audio generation.
However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions.
Controlling stereo audio with spatial contexts remains challenging due to high data costs and unstable generative models.
To the best of our knowledge, this work represents the first attempt to address these issues.
We first construct a large-scale, simulation-based, and GPT-assisted dataset, BEWO-1M, with abundant soundscapes and descriptions even including moving and multiple sources.
Beyond text modality, we have also acquired a set of images and rationally paired stereo audios through retrieval to advance multimodal generation.
Existing audio generation models tend to generate rather random and indistinct spatial audio.
To provide accurate guidance for Latent Diffusion Models, we introduce the SpatialSonic model utilizing spatial-aware encoders and azimuth state matrices to reveal reasonable spatial guidance.
By leveraging spatial guidance, our model not only achieves the objective of generating immersive and controllable spatial audio from text but also extends to other modalities as the pioneer attempt.
Finally, under fair settings, we conduct subjective and objective evaluations on simulated and real-world data to compare our approach with prevailing methods.
The results demonstrate the effectiveness of our method, highlighting its capability to generate spatial audio that adheres to physical rules. Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye 0006, Huadai Liu, Honggang Zhang 0002, Wei Xue 0002, Yike Guo |
ICLR | 8 |
| 2025 | Empowering World Models with Reflection for Embodied Video PredictionabstractVideo generation models have made significant progress in simulating future states, showcasing their potential as world simulators in embodied scenarios. However, existing models often lack robust understanding, limiting their ability to perform multi-step predictions or handle Out-of-Distribution (OOD) scenarios. To address this challenge, we propose the Reflection of Generation (RoG), a set of intermediate reasoning strategies designed to enhance video prediction. It leverages the complementary strengths of pre-trained vision-language and video generation models, enabling them to function as a world model in embodied scenarios. To support RoG, we introduce Embodied Video Anticipation Benchmark(EVA-Bench), a comprehensive benchmark that evaluates embodied world models across diverse tasks and scenarios, utilizing both in-domain and OOD datasets. Building on this foundation, we devise a world model, Embodied Video Anticipator (EVA), that follows a multistage training paradigm to generate high-fidelity video frames and apply an autoregressive strategy to enable adaptive generalization for longer video sequences. Extensive experiments demonstrate the efficacy of EVA in various downstream tasks like video generation and robotics, thereby paving the way for large-scale pre-trained models in real-world video prediction applications. The video demos are available at https://sites.google.com/view/icml-eva. Xiaowei Chi, Chun-Kai Fan, Xingqun Qi, Rongyu Zhang, Anthony Chen, Chi-Min Chan, Wei Xue 0002, Shanghang Zhang, Yike Guo |
ICML | 11 |
| 2025 | Delta Decompression for MoE-based LLMs CompressionabstractMixture-of-Experts (MoE) architectures in large language models (LLMs) achieve exceptional performance, but face prohibitive storage and memory requirements. To address these challenges, we present $D^2$-MoE, a new delta decompression compressor for reducing the parameters of MoE LLMs. Based on observations of expert diversity, we decompose their weights into a shared base weight and unique delta weights. Specifically, our method first merges each expert's weight into the base weight using the Fisher information matrix to capture shared components. Then, we compress delta weights through Singular Value Decomposition (SVD) by exploiting their low-rank properties.
Finally, we introduce a semi-dynamical structured pruning strategy for the base weights, combining static and dynamic redundancy analysis to achieve further parameter reduction while maintaining input adaptivity. In this way, our $D^2$-MoE successfully compacts MoE LLMs to high compression ratios without additional training. Extensive experiments highlight the superiority of our approach, with over 13\% performance gains than other compressors on Mixtral|Phi-3.5|DeepSeek|Qwen2 MoE LLMs at 40$\sim$60\% compression rates. Codes are available in https://github.com/lliai/D2MoE. Hao Gu 0001, Wei Li 0286, Lujun Li 0001, Qiyuan Zhu, Mark Lee 0001, Wei Xue 0002, Yike Guo |
ICML | 8 |
| 2025 | MoE-SVD: Structured Mixture-of-Experts LLMs Compression via Singular Value DecompositionabstractMixture of Experts (MoE) architecture improves Large Language Models (LLMs) with better scaling, but its higher parameter counts and memory demands create challenges for deployment. In this paper, we present MoE-SVD, a new decomposition-based compression framework tailored for MoE LLMs without any extra training. By harnessing the power of Singular Value Decomposition (SVD), MoE-SVD addresses the critical issues of decomposition collapse and matrix redundancy in MoE architectures. Specifically, we first decompose experts into compact low-rank matrices, resulting in accelerated inference and memory optimization. In particular, we propose selective decomposition strategy by measuring sensitivity metrics based on weight singular values and activation statistics to automatically identify decomposable expert layers. Then, we share a single V-matrix across all experts and employ a top-k selection for U-matrices. This low-rank matrix sharing and trimming scheme allows for significant parameter reduction while preserving diversity among experts. Comprehensive experiments on Mixtral, Phi-3.5, DeepSeek, and Qwen2 MoE LLMs show MoE-SVD outperforms other compression methods, achieving a 60% compression ratio and 1.5$\times$ faster inference with minimal performance loss. Wei Li 0286, Lujun Li 0001, Hao Gu 0001, You-Liang Huang, Mark Lee 0001, Wei Xue 0002, Yike Guo |
ICML | 8 |
| 2025 | Conservation-informed Graph Learning for Spatiotemporal Dynamics PredictionabstractData-centric methods have shown great potential in understanding and predicting spatiotemporal dynamics, enabling better design and control of the object system. However, deep learning models often lack interpretability, fail to obey intrinsic physics, and struggle to cope with the various domains. While geometry-based methods, e.g., graph neural networks (GNNs), have been proposed to further tackle these challenges, they still need to find the implicit physical laws from large datasets and rely excessively on rich labeled data. In this paper, we herein introduce the conservation-informed GNN (CiGNN), an end-to-end explainable learning framework, to learn spatiotemporal dynamics based on limited training data. The network is designed to conform to the general conservation law via symmetry, where conservative and non-conservative information passes over a multiscale space enhanced by a latent temporal marching strategy. The efficacy of our model has been verified in various spatiotemporal systems based on synthetic and real-world datasets, showing superiority over baseline models. Results demonstrate that CiGNN exhibits remarkable accuracy and generalizability, and is readily applicable to learning for prediction of various spatiotemporal dynamics in a spatial domain with complex geometry. Yuan Mi, Pu Ren, Hongteng Xu, Hongsheng Liu 0002, Zidong Wang 0010, Yike Guo, Ji-Rong Wen, Hao Sun 0002, Yang Liu 0005 |
KDD (1) | 6 |
| 2025 | Outlier-Aware Model Merging for Efficient Multitask InferenceabstractModel merging techniques aim to consolidate multiple fine-tuned models into a single unified model, reducing both storage and computational overhead while retaining task-specific performance. However, existing methods face several limitations: monotonous compression techniques that fail to account for task-specific weight distribution characteristics, weight-magnitude-based compression that fails to consider functional importance revealed by activation patterns, and non-adaptive allocation strategies that ignores task-specific layer importance. To overcome these challenges, we propose OA-Merge, a novel Outlier-Aware Model Merging framework that leverages task activation outliers to enable adaptive compression and resource allocation across tasks. OA-Merge comprises three key components: (1) dynamic hybrid decomposition technique that formulates task vectors as tailored combinations of low-rank and sparse components adapted to task-specific statistical distributions, (2) activation-informed compression methodology that incorporates task-specific activation statistics to prioritize functionally important weights, and (3) task-related allocation that optimizes the distribution of compression resources according to layer-specific importance metrics derived from activation outlier analysis. These hybrid outlier-aware strategies adapt dynamically to each task's intrinsic characteristics, avoiding the pitfalls of one-size-fits-all ways. Extensive experiments on both vision models (e.g., ViT) and language models (e.g., RoBERTa, Qwen) demonstrate that OA-Merge outperforms state-of-the-art baselines, achieving average performance gains of 3.2% on vision tasks and 2.8% on language tasks. Qiyuan Zhu, Lujun Li 0001, Jiacheng Liu 0001, Pengyu Cheng, Sirui Han, Yike Guo |
ACM Multimedia | 8 |
| 2025 | Foundation Cures Personalization: Improving Personalized Models' Prompt Consistency via Hidden Foundation KnowledgeabstractFacial personalization faces challenges to maintain identity fidelity without disrupting the foundation model's prompt consistency. The mainstream personalization models employ identity embedding to integrate identity information within the attention mechanisms. However, our preliminary findings reveal that identity embeddings compromise the effectiveness of other tokens in the prompt, thereby limiting high prompt consistency and attribute-level controllability. Moreover, by deactivating identity embedding, personalization models still demonstrate the underlying foundation models' ability to control facial attributes precisely. It suggests that such foundation models' knowledge can be leveraged to cure the ill-aligned prompt consistency of personalization models. Building upon these insights, we propose FreeCure, a framework that improves the prompt consistency of personalization models with their latent foundation models' knowledge. First, by setting a dual inference paradigm with/without identity embedding, we identify attributes (e.g., hair, accessories, etc.) for enhancements. Second, we introduce a novel foundation-aware self-attention module, coupled with an inversion-based process to bring well-aligned attribute information to the personalization process. Our approach is training-free, and can effectively enhance a wide array of facial attributes; and it can be seamlessly integrated into existing popular personalization models based on both Stable Diffusion and FLUX. FreeCure has consistently shown significant improvements in prompt consistency across these facial personalization models while maintaining the integrity of their original identity fidelity. Yiyang Cai, Zhengkai Jiang 0001, Wei Xue 0002, Yike Guo, Wenhan Luo |
NeurIPS | 6 |
| 2025 | InterMT: Multi-Turn Interleaved Preference Alignment with Human FeedbackabstractAs multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: \textbf{\textit{What essential capabilities are still missing? }}A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation.To move closer to human-level intelligence, models must similarly support \textbf{multi-turn}, \textbf{multimodal interaction}. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges.In this work, we present \textbf{an initial exploration} through the \textsc{InterMT} -- \textbf{the first preference dataset for \textit{multi-turn} multimodal interaction}, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. \textsc{InterMT} captures human preferences at both global and local levels into nine sub-dimensions, consists of 15.6k prompts, 52.6k multi-turn dialogue instances, and 32.4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances.To further this goal, we introduce \textsc{InterMT-Bench} to assess the ability ofMLLMs in assisting judges with multi-turn, multimodal tasks.We demonstrate the utility of \textsc{InterMT} through applications such as judge moderation and further reveal the \textit{multi-turn scaling law} of judge model.We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step. Boyuan Chen 0008, Donghai Hong, Jiaming Ji, Jiacheng Zheng, Kaile Wang, Juntao Dai, Xuyao Wang, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 14 |
| 2025 | Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human FeedbackabstractMultimodal large language models (MLLMs) are essential for building general-purpose AI assistants; however, they pose increasing safety risks. How can we ensure safety alignment of MLLMs to prevent undesired behaviors? Going further, it is critical to explore how to fine-tune MLLMs to preserve capabilities while meeting safety constraints. Fundamentally, this challenge can be formulated as a min-max optimization problem. However, existing datasets have not yet disentangled single preference signals into explicit safety constraints, hindering systematic investigation in this direction. Moreover, it remains an open question whether such constraints can be effectively incorporated into the optimization process for multi-modal models. In this work, we present the first exploration of the Safe RLHF-V -- the first multimodal safety alignment framework. The framework consists of: (I) BeaverTails-V, the first open-source dataset featuring dual preference annotations for helpfulness and safety, supplemented with multi-level safety labels (minor, moderate, severe); (II) Beaver-Guard-V, a multi-level guardrail system to proactively defend against unsafe queries and adversarial attacks. Applying the guard model over five rounds of filtering and regeneration significantly enhances the precursor model’s overall safety by an average of 40.9%. (II) Based on dual preference, we initiate the first exploration of multi-modal safety alignment within a constrained optimization. Experimental results demonstrate that Safe RLHF effectively improves both model helpfulness and safety. Specifically, Safe RLHF-V enhances model safety by 34.2% and helpfulness by 34.3%. Jiaming Ji, Donghai Hong, Boyuan Chen 0008, Kaile Wang, Juntao Dai, Chi-Min Chan, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 13 |
| 2025 | Generative RLHF-V: Learning Principles from Multi-modal Human PreferenceabstractTraining multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, \textit{e.g.,} reinforcement learning from human feedback (RLHF). Generative reward models (GRMs) leverage MLLMs' intrinsic reasoning capabilities to discriminate pair-wise responses, but their pair-wise paradigm makes it hard to generalize to learnable rewards. We introduce Generative RLHF-V, a novel alignment framework that integrates GRMs with multi-modal RLHF. We propose a two-stage pipeline: \textbf{multi-modal generative reward modeling from RL}, where RL guides GRMs to actively capture human intention, then predict the correct pair-wise scores; and \textbf{RL optimization from grouped comparison}, which enhances multi-modal RL scoring precision by grouped responses comparison. Experimental results demonstrate that, besides out-of-distribution generalization of RM discrimination, our framework improves 4 MLLMs' performance across 7 benchmarks by 18.1\%, while the baseline RLHF is only 5.3\%. We further validate that Generative RLHF-V achieves a near-linear improvement with an increasing number of candidate responses. Jiaming Ji, Boyuan Chen 0008, Jiapeng Sun, Donghai Hong, Sirui Han, Yike Guo, Yaodong Yang 0001 |
NeurIPS | 8 |
| 2025 | MultiGranDTI: an explainable multi-granularity representation framework for drug-target interaction prediction
Qun Liu 0005, Yike Guo, Guoyin Wang 0001 |
Appl. Intell. | 4 |
| 2025 | Dynamic link prediction: Using language models and graph structures for temporal knowledge graph completion with emerging entities and relations
Ryan Ong, Yike Guo, Ovidiu Serban |
Expert Syst. Appl. | 3 |
| 2025 | PT-VAE: Variational autoencoder with prior concept transformation
Zitu Liu, Zhenyao Yu, Qingshan Fu, Yike Guo, Qun Liu 0005, Guoyin Wang 0001 |
Neurocomputing | 6 |
| 2025 | Intuitively interpreting GANs latent space using semantic distribution
Ruqi Wang, Guoyin Wang 0001, Lihua Gu, Qun Liu 0005, Yike Guo |
Knowl. Based Syst. | 6 |
| 2025 | OVST: online video stabilization with two-stage training transformer
Xing Wu 0001, Junfeng Yao, Quan Qian, Yike Guo |
Neural Comput. Appl. | 8 |
| 2025 | MIFS: An adaptive multipath information fused self-supervised framework for drug discovery
Qun Liu 0005, Rui Han 0001, Yike Guo, Guoyin Wang 0001 |
Neural Networks | 4 |
| 2025 | MDFCL: Multimodal data fusion-based graph contrastive learning framework for molecular property prediction
Maotao Liu, Qun Liu 0005, Yike Guo, Guoyin Wang 0001 |
Pattern Recognit. | 4 |
| 2025 | $ \tt {zkFL}$zkFL: Zero-Knowledge Proof-Based Gradient Aggregation for Federated LearningabstractFederated learning (FL) is a machine learning paradigm, which enables multiple and decentralized clients to collaboratively train a model under the orchestration of a central aggregator. FL can be a scalable machine learning solution inbig datascenarios. Traditional FL relies on the trust assumption of the central aggregator, which forms cohorts of clients honestly. However, a malicious aggregator, in reality, could abandon and replace the client's training models, or insert fake clients, to manipulate the final training results. In this work, we introducezkFL, which leverages zero-knowledge proofs to tackle the issue of a malicious aggregator during the training model aggregation process. To guarantee the correct aggregation results, the aggregator provides a proof per round, demonstrating to the clients that the aggregator executes the intended behavior faithfully. To further reduce the verification cost of clients, we use blockchain to handle the proof in a zero-knowledge way, where miners (i.e., the participants validating and maintaining the blockchain data) can verify the proof without knowing the clients' local and aggregated models. The theoretical analysis and empirical results show thatzkFLachieves better security and privacy than traditional FL, without modifying the underlying FL network structure or heavily compromising the training speed. Zhipeng Wang 0009, Nanqing Dong, William J. Knottenbelt, Yike Guo |
IEEE Trans. Big Data | 5 |
| 2025 | Every Angle is Worth a Second Glance: Mining Kinematic Skeletal Structures From Multi-View Joint CloudabstractMulti-person motion capture over sparse angular observations is a challenging problem under interference from both self- and mutual-occlusions. Existing works produce accurate 2D joint detection, however, when these are triangulated and lifted into 3D, available solutions all struggle in selecting the most accurate candidates and associating them to the correct joint type and target identity. As such, in order to fully utilize all accurate 2D joint location information, we propose to independently triangulate between all same-typed 2D joints from all camera views regardless of their target ID, forming the Joint Cloud. Joint Cloud consist of both valid joints lifted from the same joint type and target ID, as well as falsely constructed ones that are from different 2D sources. These redundant and inaccurate candidates are processed over the proposed Joint Cloud Selection and Aggregation Transformer (JCSAT) involving three cascaded encoders which deeply explore the trajectile, skeletal structural, and view-dependent correlations among all 3D point candidates in the cross-embedding space. An Optimal Token Attention Path (OTAP) module is proposed which subsequently selects and aggregates informative features from these redundant observations for the final prediction of human motion. To demonstrate the effectiveness of JCSAT, we build and publish a new multi-person motion capture dataset BUMocap-X with complex interactions and severe occlusions. Comprehensive experiments over the newly presented as well as benchmark datasets validate the effectiveness of the proposed framework, which outperforms all existing state-of-the-art methods, especially under challenging occlusion scenarios. Junkun Jiang, Jie Chen 0026, Ho Yin Au, Wei Xue 0002, Yike Guo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Multi-View Large Reconstruction Model via Geometry-Aware Positional Encoding and AttentionabstractDespite recent advancements in the Large Reconstruction Model (LRM) demonstrating impressive results, when extending its input from single image to multiple images, it exhibits inefficiencies, subpar geometric and texture quality, as well as slower convergence speed than expected. It is attributed to that, LRM formulates 3D reconstruction as a naive images-to-3D translation problem, ignoring the strong 3D coherence among the input images. In this article, we propose a Multi-view Large Reconstruction Model (M-LRM) designed to reconstruct high-quality 3D shapes from multi-views in a 3D-aware manner. Specifically, we introduce a multi-view consistent cross-attention scheme to enable M-LRM to accurately query information from the input images. Moreover, we employ the 3D priors of the input multi-view images to initialize the triplane tokens. Compared to previous methods, the proposed M-LRM can generate 3D shapes of high fidelity. Experimental studies demonstrate that our model achieves a significant performance gain and faster training convergence. Xiaoxiao Long, Yixun Liang, Yuan Liu 0025, Wenhan Luo, Wenping Wang 0001, Yike Guo |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2024 | APK-MRL: An Adaptive Pre-training Framework with Knowledge-enhanced for Molecular Representation LearningabstractAs a dominant pre-training paradigm, molecular contrastive learning (MCL) has been proven effective in learning molecular representations with unlabeled data. However, high data dependency and scarce domain knowledge caused by data augmentation in MCL limit the model’s generalization and performance. To address these issues, we propose an adaptive pre-training framework with knowledge-enhanced for molecular representation learning, named APK-MRL. It seamlessly integrates diverse prior information on hierarchical skeletons and the chemical semantics of molecules, aiming to obtain stronger stability and generalization. Extensive computational experiments demonstrate that APK-MRL can achieve competitive performances over state-of-the-art baselines on both drug-target interaction and molecular properties prediction tasks. All code is released at https://github.com/lukcats/APK-MRL. Qun Liu 0005, Rui Han 0001, Li Liu 0030, Yike Guo, Guoyin Wang 0001 |
BIBM | 5 |
| 2024 | Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-Speech Gesture GenerationabstractGenerating vivid and emotional 3D co-speech gestures is crucial for virtual avatar animation in human-machine interaction applications. While the existing methods enable generating the gestures to follow a single emotion label, they overlook that long gesture sequence modeling with emotion transition is more practical in real scenes. In addition, the lack of large-scale available datasets with emotional transition speech and corresponding 3D human gestures also limits the addressing of this task. To fulfill this goal, we first incorporate the ChatGPT-4 and an audio inpainting approach to construct the high-fidelity emotion transition human speeches. Considering obtaining the realistic 3D pose annotations corresponding to the dynamically inpainted emotion transition audio is extremely difficult, we propose a novel weakly supervised training strategy to encourage authority gesture transitions. Specifically, to enhance the coordination of transition gestures w. r. t. different emotional ones, we model the temporal association representation between two different emotional gesture sequences as style guidance and infuse it into the transition generation. We further devise an emotion mixture mechanism that provides weak supervision based on a learnable mixed emotion label for transition gestures. Last, we present a keyframe sampler to supply effective initial posture cues in long sequences, enabling us to generate diverse gestures. Extensive experiments demonstrate that our method outperforms the state-of-the-art models constructed by adapting single emotion-conditioned counterparts on our newly defined emotion transition task and datasets. Our code and dataset will be released on the project page: https://xingqunqi-lab.github.io/Emo-Transition-Gesture/. Xingqun Qi, Ruibin Yuan, Xiaowei Chi, Wenhan Luo, Wei Xue 0002, Shanghang Zhang, Yike Guo |
CVPR | 11 |
| 2024 | Auto-GAS: Automated Proxy Discovery for Training-Free Generative Architecture Search
Lujun Li 0001, Haosen Sun, Shiwen Li, Peijie Dong, Wenhan Luo, Wei Xue 0002, Yike Guo |
ECCV (5) | 8 |
| 2024 | AttnZero: Efficient Attention Discovery for Vision Transformers
Lujun Li 0001, Zimian Wei, Peijie Dong, Wenhan Luo, Wei Xue 0002, Yike Guo |
ECCV (5) | 7 |
| 2024 | Combined Global and Local Information Diffusion of Neural Processes
Jinyang Tai, Yike Guo |
ICANN (1) | 2 |
| 2024 | Topology of Neural Processes
Jinyang Tai, Yike Guo |
ICANN (1) | 2 |
| 2024 | Freeze the Backbones: a Parameter-Efficient Contrastive Approach to Robust Medical Vision-Language Pre-TrainingabstractModern healthcare often utilises radiographic images alongside textual reports for diagnostics, encouraging the use of Vision-Language Self-Supervised Learning (VL-SSL) with large pre-trained models to learn versatile medical vision representations. However, most existing VL-SSL frameworks are trained end-to-end, which is computation-heavy and can lose vital prior information embedded in pre-trained encoders. To address both issues, we introduce the backbone-agnostic Adaptor framework, which preserves medical knowledge in pre-trained image and text encoders by keeping them frozen, and employs a lightweight Adaptor module for cross-modal learning. Experiments on medical image classification and segmentation tasks across three datasets reveal that our framework delivers competitive performance while cutting trainable parameters by over 90% compared to current pre-training approaches. Notably, when fine-tuned with just 1% of data, Adaptor outperforms several Transformer-based methods trained on full datasets in medical image segmentation. Jiuming Qin, Che Liu 0002, Sibo Cheng, Yike Guo, Rossella Arcucci |
ICASSP | 4 |
| 2024 | EDM: Synthetic Data from Exemplar Diffusion Model Improves Non-Communicable Diseases DetectionabstractThere have been researches revealing obvious associations between facial phenotypes and non-communicable diseases (NCDs), which enables effective health assessment with the integration of model-based learning methods. However, the paucity and poor quality of available datasets hinder the development of potent algorithms to detect NCDs. To meet this challenge, we propose a method called Exemplar Diffusion Model (EDM), the objective of proposed EDM is to generate facial images that illustrate simulated non-communicable diseases, utilizing a normal facial image as input. Extensive experimental results show that the proposed EDM method outperforms the state-of-the-art methods in terms of Frechet Inception Distance (FID) and Quality Score (QS), with improvements of 0.11 and 0.74, respectively. Furthermore, comprehensive ablation studies and comparative experiments prove the value of proposed EDM method in large-scale facial image dataset generation and non-communicable disease detection. Xing Wu 0001, Junfeng Yao, Quan Qian, Yike Guo |
ICASSP | 7 |
| 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingabstractSelf-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech.
Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music.
To address this research gap, we propose an acoustic **M**usic und**ER**standing model with large-scale self-supervised **T**raining (**MERT**), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training.
In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance.
This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT).
Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters.
Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores. Ruibin Yuan, Ge Zhang 0009, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin 0002, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi 0001, Wenhao Huang 0001, Yike Guo, Jie Fu 0001 |
ICLR | 19 |
| 2024 | DetKDS: Knowledge Distillation Search for Object DetectorsabstractIn this paper, we present DetKDS, the first framework that searches for optimal detection distillation policies. Manual design of detection distillers becomes challenging and time-consuming due to significant disparities in distillation behaviors between detectors with different backbones, paradigms, and label assignments. To tackle these challenges, we leverage search algorithms to discover optimal distillers for homogeneous and heterogeneous student-teacher pairs. Firstly, our search space encompasses global features, foreground-background features, instance features, logits response, and localization response as inputs. Then, we construct omni-directional cascaded transformations and obtain the distiller by selecting the advanced distance function and common weight value options. Finally, we present a divide-and-conquer evolutionary algorithm to handle the explosion of the search space. In this strategy, we first evolve the best distiller formulations of individual knowledge inputs and then optimize the combined weights of these multiple distillation losses. DetKDS automates the distillation process without requiring expert design or additional tuning, effectively reducing the teacher-student gap in various scenarios. Based on the analysis of our search results, we provide valuable guidance that contributes to detection distillation designs. Comprehensive experiments on different detectors demonstrate that DetKDS outperforms state-of-the-art methods in detection and instance segmentation tasks. For instance, DetKDS achieves significant gains than baseline detectors: $+3.7$, $+4.1$, $+4.0$, $+3.7$, and $+3.5$ AP on RetinaNet, Faster-RCNN, FCOS, RepPoints, and GFL, respectively. Code at: https://github.com/lliai/DetKDS. Lujun Li 0001, Yufan Bao, Peijie Dong, Chuanguang Yang, Anggeng Li, Wenhan Luo, Wei Xue 0002, Yike Guo |
ICML | 9 |
| 2024 | FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
Jianyi Chen, Wei Xue 0002, Xu Tan 0003, Zhen Ye 0006, Yike Guo |
IJCAI | 6 |
| 2024 | Hierarchical Linear Symbolized Tree-Structured Neural ProcessesabstractTraditional Neural Processes (NPs) and their variants aim to learn relationships between context sample points but do not consider multi-level information, resulting in a limited ability to learn complex distributions.This paper draws inspiration from features such as the hierarchical nature and interpretability of tree-like structures.This paper proposes a Hierarchical Linear Symbolized Treestructured Neural Processes (HLNPs) architecture.This framework utilizes variables to build a top-down hierarchical linear symbolized tree-structured network architecture, enhancing positional representation information in a hierarchical manner along the deterministic path.In the latent distribution, the hierarchical linear symbolized tree-structured network approximates functions discretely through a layered approach.By decomposing the latent complex distribution into several simpler sub-problems using sum and product symbols, the upper bound of optimization is thereby increased.The tree structure discretizes variables to capture model uncertainty in the form of entropy.This approach also imparts a causal effect to the HLNPs model.Finally, we demonstrate the effectiveness of the HLNPs models for 1D data, Bayesian optimization, and 2D data. Jinyang Tai, Yike Guo |
KDD | 2 |
| 2024 | FlashSpeech: Efficient Zero-Shot Speech SynthesisabstractRecent progress in large-scale zero-shot speech synthesis has been significantly advanced by language models and diffusion models. However, the generation process of both methods is slow and computationally intensive. Efficient speech synthesis using a lower computing budget to achieve quality on par with previous work remains a significant challenge. In this paper, we present FlashSpeech, a large-scale zero-shot speech synthesis system with approximately 5% of the inference time compared with previous work. FlashSpeech is built on the latent consistency model and applies a novel adversarial consistency training approach that can train from scratch without the need for a pre-trained diffusion model as the teacher. Furthermore, a new prosody generator module enhances the diversity of prosody, making the rhythm of the speech sound more natural. The generation processes of FlashSpeech can be achieved efficiently with one or two sampling steps while maintaining high audio quality and high similarity to the audio prompt for zero-shot speech generation. Our experimental results demonstrate the superior performance of FlashSpeech. Notably, FlashSpeech can be about 20 times faster than other zero-shot speech synthesis systems while maintaining comparable performance in terms of voice quality and similarity. Furthermore, FlashSpeech demonstrates its versatility by efficiently performing tasks like voice conversion, speech editing, and diverse speech sampling. Audio samples can be found in https://flashspeech.github.io/ Zhen Ye 0006, Zeqian Ju, Haohe Liu, Xu Tan 0003, Jianyi Chen, Peiwen Sun, Weizhen Bian, Shulin He, Wei Xue 0002, Yike Guo |
ACM Multimedia | 13 |
| 2024 | Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingabstractLLMs are computationally expensive to pre-train due to their large scale.
Model growth emerges as a promising approach by leveraging smaller models to accelerate the training of larger ones.
However, the viability of these model growth methods in efficient LLM pre-training remains underexplored.
This work identifies three critical $\underline{\textit{O}}$bstacles: ($\textit{O}$1) lack of comprehensive evaluation, ($\textit{O}$2) untested viability for scaling, and ($\textit{O}$3) lack of empirical guidelines.
To tackle $\textit{O}$1, we summarize existing approaches into four atomic growth operators and systematically evaluate them in a standardized LLM pre-training setting.
Our findings reveal that a depthwise stacking operator, called $G_{\text{stack}}$, exhibits remarkable acceleration in training, leading to decreased loss and improved overall performance on eight standard NLP benchmarks compared to strong baselines.
Motivated by these promising results, we conduct extensive experiments to delve deeper into $G_{\text{stack}}$ to address $\textit{O}$2 and $\textit{O}$3.
For $\textit{O}$2 (untested scalability), our study shows that $G_{\text{stack}}$ is scalable and consistently performs well, with experiments up to 7B LLMs after growth and pre-training LLMs with 750B tokens.
For example, compared to a conventionally trained 7B model using 300B tokens, our $G_{\text{stack}}$ model converges to the same loss with 194B tokens, resulting in a 54.6\% speedup.
We further address $\textit{O}$3 (lack of empirical guidelines) by formalizing guidelines to determine growth timing and growth factor for $G_{\text{stack}}$, making it practical in general LLM pre-training.
We also provide in-depth discussions and comprehensive ablation studies of $G_{\text{stack}}$.
Our code and pre-trained model are available at https://llm-stacking.github.io/. Wenyu Du, Tongxu Luo, Zihan Qiu, Yikang Shen, Reynold Cheng, Yike Guo, Jie Fu 0001 |
NeurIPS | 7 |
| 2024 | Discovering Sparsity Allocation for Layer-wise Pruning of Large Language ModelsabstractIn this paper, we present DSA, the first automated framework for discovering sparsity allocation schemes for layer-wise pruning in Large Language Models (LLMs). LLMs have become increasingly powerful, but their large parameter counts make them computationally expensive. Existing pruning methods for compressing LLMs primarily focus on evaluating redundancies and removing element-wise weights. However, these methods fail to allocate adaptive layer-wise sparsities, leading to performance degradation in challenging tasks. We observe that per-layer importance statistics can serve as allocation indications, but their effectiveness depends on the allocation function between layers. To address this issue, we develop an expression discovery framework to explore potential allocation strategies. Our allocation functions involve two steps: reducing element-wise metrics to per-layer importance scores, and modelling layer importance to sparsity ratios. To search for the most effective allocation function, we construct a search space consisting of pre-process, reduction, transform, and post-process operations. We leverage an evolutionary algorithm to perform crossover and mutation on superior candidates within the population, guided by performance evaluation. Finally, we seamlessly integrate our discovered functions into various uniform methods, resulting in significant performance improvements. We conduct extensive experiments on multiple challenging tasks such as arithmetic, knowledge reasoning, and multimodal benchmarks spanning GSM8K, MMLU, SQA, and VQA, demonstrating that our DSA method achieves significant performance gains on the LLaMA-1|2|3, Mistral, and OPT models. Notably, the LLaMA-1|2|3 model pruned by our DSA reaches 4.73\%|6.18\%|10.65\% gain over the state-of-the-art techniques (e.g., Wanda and SparseGPT). Lujun Li 0001, Peijie Dong, Zhenheng Tang, Xiang Liu 0001, Qiang Wang 0022, Wenhan Luo, Wei Xue 0002, Xiaowen Chu 0001, Yike Guo |
NeurIPS | 10 |
| 2024 | Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise AttentionabstractIn this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resulting in poor-quality multiview images. Specifically, these methods assume that the input images should comply with a predefined camera type, e.g. a perspective camera with a fixed focal length, leading to distorted shapes when the assumption fails. Moreover, the full-image or dense multiview attention they employ leads to a dramatic explosion of computational complexity as image resolution increases, resulting in prohibitively expensive training costs. To bridge the gap between assumption and reality, Era3D first proposes a diffusion-based camera prediction module to estimate the focal length and elevation of the input image, which allows our method to generate images without shape distortions. Furthermore, a simple but efficient attention layer, named row-wise attention, is used to enforce epipolar priors in the multiview diffusion, facilitating efficient cross-view information fusion. Consequently, compared with state-of-the-art methods, Era3D generates high-quality multiview images with up to a 512×512 resolution while reducing computation complexity of multiview attention by 12x times. Comprehensive experiments demonstrate the superior generation power of Era3D- it can reconstruct high-quality and detailed 3D meshes from diverse single-view input images, significantly outperforming baseline multiview diffusion methods. Yuan Liu 0025, Xiaoxiao Long, Feihu Zhang, Cheng Lin 0001, Xingqun Qi, Shanghang Zhang, Wei Xue 0002, Wenhan Luo, Ping Tan 0002, Wenping Wang 0001, Yike Guo |
NeurIPS | 14 |
| 2024 | Dirichlet Continual Learning: Tackling Catastrophic Forgetting in NLPabstractCatastrophic forgetting poses a significant challenge in continual learning (CL). In the context of Natural Language Processing, generative-based rehearsal CL methods have made progress in avoiding expensive retraining. However, generating pseudo samples that accurately capture the task-specific distribution remains a daunting task. In this paper, we propose Dirichlet Continual Learning (DCL), a novel generative-based rehearsal strategy designed specifically for CL. Different from the conventional use of Gaussian latent variable in Conditional Variational Autoencoder, DCL employs the flexibility of the Dirichlet distribution to model the latent variable. This allows DCL to effectively capture sentence-level features from previous tasks and guide the generation of pseudo samples. Additionally, we introduce Jensen-Shannon Knowledge Distillation, a robust logit-based knowledge distillation method that enhances knowledge transfer during pseudo-sample generation. Our extensive experiments show that DCL outperforms state-of-the-art methods in two typical tasks of task-oriented dialogue systems, demonstrating its efficacy. Haiqin Yang, Wei Xue 0002, Yike Guo |
UAI | 5 |
| 2024 | Causal inference in the medical domain: a survey
Xing Wu 0001, Shaoqi Peng, Weimin Li 0001, Quan Qian, Yike Guo |
Appl. Intell. | 9 |
| 2024 | FedEL: Federated ensemble learning for non-iid data
Xing Wu 0001, Jie Pei, Xianhua Han, Yen-Wei Chen 0001, Junfeng Yao, Yang Liu 0005, Quan Qian, Yike Guo |
Expert Syst. Appl. | 8 |
| 2024 | GNN-MgrPool: Enhanced graph neural networks with multi-granularity pooling for graph classification
Haichao Sun, Guoyin Wang 0001, Qun Liu 0005, Yike Guo |
Inf. Sci. | 4 |
| 2024 | The potential and pitfalls of using a large language model such as ChatGPT, GPT-4, or LLaMA as a clinical assistantabstractOBJECTIVES: This study aims to evaluate the utility of large language models (LLMs) in healthcare, focusing on their applications in enhancing patient care through improved diagnostic, decision-making processes, and as ancillary tools for healthcare professionals. MATERIALS AND METHODS: We evaluated ChatGPT, GPT-4, and LLaMA in identifying patients with specific diseases using gold-labeled Electronic Health Records (EHRs) from the MIMIC-III database, covering three prevalent diseases-Chronic Obstructive Pulmonary Disease (COPD), Chronic Kidney Disease (CKD)-along with the rare condition, Primary Biliary Cirrhosis (PBC), and the hard-to-diagnose condition Cancer Cachexia. RESULTS: In patient identification, GPT-4 had near similar or better performance compared to the corresponding disease-specific Machine Learning models (F1-score ≥ 85%) on COPD, CKD, and PBC. GPT-4 excelled in the PBC use case, achieving a 4.23% higher F1-score compared to disease-specific "Traditional Machine Learning" models. ChatGPT and LLaMA3 demonstrated lower performance than GPT-4 across all diseases and almost all metrics. Few-shot prompts also help ChatGPT, GPT-4, and LLaMA3 achieve higher precision and specificity but lower sensitivity and Negative Predictive Value. DISCUSSION: The study highlights the potential and limitations of LLMs in healthcare. Issues with errors, explanatory limitations and ethical concerns like data privacy and model transparency suggest that these models would be supplementary tools in clinical settings. Future studies should improve training datasets and model designs for LLMs to gain better utility in healthcare. CONCLUSION: The study shows that LLMs have the potential to assist clinicians for tasks such as patient identification but false positives and false negatives must be mitigated before LLMs are adequate for real-world clinical assistance. Jingqing Zhang, Kai Sun 0005, Akshay Jagadeesh, Parastoo Falakaflaki, Elena Kayayan, Guanyu Tao, Mahta Haghighat Ghahfarokhi, Ashok Gupta, Vibhor Gupta, Yike Guo |
J. Am. Medical Informatics Assoc. | 11 |
| 2024 | Improving disentanglement in variational auto-encoders via feature imbalance-informed dimension weighting
Zhenyao Yu, Zitu Liu, Ziyi Yu, Xinyan Yang, Xingyue Li, Yike Guo, Qun Liu 0005, Guoyin Wang 0001 |
Knowl. Based Syst. | 7 |
| 2024 | Labelling with dynamics: A data-efficient learning paradigm for medical image segmentationabstractThe success of deep learning on image classification and recognition tasks has led to new applications in diverse contexts, including the field of medical imaging. However, two properties of deep neural networks (DNNs) may limit their future use in medical applications. The first is that DNNs require a large amount of labeled training data, and the second is that the deep learning-based models lack interpretability. In this paper, we propose and investigate a data-efficient framework for the task of general medical image segmentation. We address the two aforementioned challenges by introducing domain knowledge in the form of a strong prior into a deep learning framework. This prior is expressed by a customized dynamical system. We performed experiments on two different datasets, namely JSRT and ISIC2016 (heart and lungs segmentation on chest X-ray images and skin lesion segmentation on dermoscopy images). We have achieved competitive results using the same amount of training data compared to the state-of-the-art methods. More importantly, we demonstrate that our framework is extremely data-efficient, and it can achieve reliable results using extremely limited training data. Furthermore, the proposed method is rotationally invariant and insensitive to initialization. Yuanhan Mo, Fangde Liu, Guang Yang 0006, Shuo Wang 0011, Jian-Qing Zheng, Fuping Wu, Bartlomiej Wladyslaw Papiez, Douglas McIlwraith, Taigang He, Yike Guo |
Medical Image Anal. | 10 |
| 2024 | Video-Instrument Synergistic Network for Referring Video Instrument Segmentation in Robotic SurgeryabstractSurgical instrument segmentation is fundamentally important for facilitating cognitive intelligence in robot-assisted surgery. Although existing methods have achieved accurate instrument segmentation results, they simultaneously generate segmentation masks of all instruments, which lack the capability to specify a target object and allow an interactive experience. This paper focuses on a novel and essential task in robotic surgery, i.e., Referring Surgical Video Instrument Segmentation (RSVIS), which aims to automatically identify and segment the target surgical instruments from each video frame, referred by a given language expression. This interactive feature offers enhanced user engagement and customized experiences, greatly benefiting the development of the next generation of surgical education systems. To achieve this, this paper constructs two surgery video datasets to promote the RSVIS research. Then, we devise a novel Video-Instrument Synergistic Network (VIS-Net) to learn both video-level and instrument-level knowledge to boost performance, while previous work only utilized video-level information. Meanwhile, we design a Graph-based Relation-aware Module (GRM) to model the correlation between multi-modal information (i.e., textual description and video frame) to facilitate the extraction of instrument-level information. Extensive experimental results on two RSVIS datasets exhibit that the VIS-Net can significantly outperform existing state-of-the-art referring segmentation methods. We will release our code and dataset for future research (https://github.com/whq-xxh/RSVIS). Hongqiu Wang, Guang Yang 0006, Harry Qin, Yike Guo, Yueming Jin, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Unsupervised Zero-Shot Learning for Achieve Cross-Modal Alignment with CounterfactualsabstractZero-Shot Learning (ZSL) is a technique that transfers knowledge from seen classes to unseen classes by establishing cross-modal mapping relationships. However, traditional ZSL methods heavily rely on a large number of expensive labeled data, which may not be readily available in practical applications. In practical applications there is often a lack of labels, and the approach implies that the lack of effective supervised information in the transfer process of seen classes can lead to ’negative causality’ problem between different modalities. Therefore, we propose an unsupervised counterfactual approach to solve the above problem. Therefore, we propose an unsupervised counterfactual approach to solve the above problem. In this paper, we propose an unsupervised learning model and use a Counterfactual Causal Inference framework to cross-modal mapping relationship adjustment (CMRA). Specifically, we aim to regard images as cause and Wikipedia text as effect form a causal relationship diagram. First, it uses multiple attributes attention to learn the visual semantic attributes of images and the corresponding Wikipedia text description words to form cross-modal alignment. Then, we combine contrastive learning and stop-gradient techniques to create a novel cross-modal mapping relationship. Finally, we conducted an investigation to assess the consistency of the multiple attribute attention in the distribution of visual semantic attributes before and after image transformation. To tackle this issue, we implemented a deactivation strategy specifically designed to eliminate the multiple attribute attention of visual semantic attributes that displayed noticeable distribution gaps at different stages. This approach also involves eliminating the mapping relationships between their corresponding Wikipedia text description words. This model evaluates the classification accuracy in AWA, CUB, APY, SUN. The experimental results show that the algorithm outperforms the state-of-the-art algorithms technology approaches. Jinyang Tai, Yike Guo |
ECAI | 2 |
| 2023 | NAS-FM: Neural Architecture Search for Tunable and Interpretable Sound Synthesis Based on Frequency ModulationabstractDeveloping digital sound synthesizers is crucial to the music industry as it provides a low-cost way to produce high-quality sounds with rich timbres. Existing traditional synthesizers often require substantial expertise to determine the overall framework of a synthesizer and the parameters of submodules. Since expert knowledge is hard to acquire, it hinders the flexibility to quickly design and tune digital synthesizers for diverse sounds. In this paper, we propose ``NAS-FM'', which adopts neural architecture search (NAS) to build a differentiable frequency modulation (FM) synthesizer. Tunable synthesizers with interpretable controls can be developed automatically from sounds without any prior expert knowledge and manual operating costs. In detail, we train a supernet with a specifically designed search space, including predicting the envelopes of carriers and modulators with different frequency ratios. An evolutionary search algorithm with adaptive oscillator size is then developed to find the optimal relationship between oscillators and the frequency ratio of FM. Extensive experiments on recordings of different instrument sounds show that our algorithm can build a synthesizer fully automatically, achieving better results than handcrafted synthesizers. Audio samples are available at https://nas-fm.github.io/ Zhen Ye 0006, Wei Xue 0002, Xu Tan 0003, Yike Guo |
IJCAI | 5 |
| 2023 | CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency ModelabstractDenoising diffusion probabilistic models (DDPMs) have shown promising performance for speech synthesis. However, a large number of iterative steps are required to achieve high sample quality, which restricts the inference speed. Maintaining sample quality while increasing sampling speed has become a challenging task. In this paper, we propose a Consistency Model-based Speech synthesis method, CoMoSpeech, which achieve speech synthesis through a single diffusion sampling step while achieving high audio quality. The consistency constraint is applied to distill a consistency model from a well-designed diffusion-based teacher model, which ultimately yields superior performances in the distilled CoMoSpeech. Our experiments show that by generating audio recordings by a single sampling step, the CoMoSpeech achieves an inference speed more than 150 times faster than real-time on a single NVIDIA A100 GPU, which is comparable to FastSpeech2, making diffusion-sampling based speech synthesis truly practical. Meanwhile, objective and subjective evaluations on text-to-speech and singing voice synthesis show that the proposed teacher models yield the best audio quality, and the one-step sampling based CoMoSpeech achieves the best inference speed with better or comparable audio quality to other conventional multi-step diffusion model baselines. Audio samples and codes are available at https://comospeech.github. https://comospeech.github.io/. Zhen Ye 0006, Wei Xue 0002, Xu Tan 0003, Jie Chen 0026, Yike Guo |
ACM Multimedia | 6 |
| 2023 | MARBLE: Music Audio Representation Benchmark for Universal EvaluationabstractIn the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the limited work on deep music representations, the scarcity of large-scale datasets, and the absence of a universal and community-driven benchmark. To address this issue, we introduce the Music Audio Representation Benchmark for universaL Evaluation, termed MARBLE. It aims to provide a benchmark for various Music Information Retrieval (MIR) tasks by defining a comprehensive taxonomy with four hierarchy levels, including acoustic, performance, score, and high-level description. We then establish a unified protocol based on 18 tasks on 12 public-available datasets, providing a fair and standard assessment of representations of all open-sourced pre-trained models developed on music recordings as baselines. Besides, MARBLE offers an easy-to-use, extendable, and reproducible suite for the community, with a clear statement on copyright issues on datasets. Results suggest recently proposed large-scale pre-trained musical language models perform the best in most tasks, with room for further improvement. The leaderboard and toolkit repository are published to promote future music AI research. Ruibin Yuan, Yinghao Ma, Ge Zhang 0009, Xingran Chen, Hanzhi Yin, Le Zhuo, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin 0002, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger B. Dannenberg, Wenhu Chen, Gus Xia, Wei Xue 0002, Shi Wang 0002, Ruibo Liu, Yike Guo, Jie Fu 0001 |
NeurIPS | 24 |
| 2023 | Space or time for video classification transformers
Xing Wu 0001, Chenjie Tao, Jianjia Wang, Weimin Li 0001, Yike Guo |
Appl. Intell. | 8 |
| 2023 | STR Transformer: A Cross-domain Transformer for Scene Text Recognition
Xing Wu 0001, Bin Tang 0010, Jianjia Wang, Yike Guo |
Appl. Intell. | 5 |
| 2023 | Accelerating Multi-Exit BERT Inference via Curriculum Learning and Knowledge DistillationabstractThe real-time deployment of bidirectional encoder representations from transformers (BERT) is limited by its slow inference caused by its large number of parameters. Recently, multi-exit architecture has garnered scholarly attention for its ability to achieve a trade-off between performance and efficiency. However, its early exits suffer from a considerable performance reduction compared to the final classifier. To accelerate inference with minimal compensation of performance, we propose a novel training paradigm for multi-exit BERT performing at two levels: training samples and intermediate features. Specifically, for the training samples level, we leverage curriculum learning to guide the training process and improve the generalization capacity of the model. For the intermediate features level, we employ layer-wise distillation learning from shallow to deep layers to resolve the performance deterioration of early exits. The experimental results obtained on the benchmark datasets of textual entailment and answer selection demonstrate that the proposed training paradigm is effective and achieves state-of-the-art results. Furthermore, the layer-wise distillation can completely replace vanilla distillation and deliver superior performance on text entailment datasets. Shengwei Gu, Xiangfeng Luo, Xinzhi Wang 0001, Yike Guo |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2023 | Cloud-Cluster: An uncertainty clustering algorithm based on cloud model
Zitu Liu, Yike Guo, Qun Liu 0005, Guoyin Wang 0001 |
Knowl. Based Syst. | 4 |
| 2023 | Conflict-aware multilingual knowledge graph completionabstractKnowledge graph completion (KGC), a task that aims at predicting missing links with existing information inside a knowledge graph (KG), has emerged as a popular research area in recent years. While many existing works have demonstrated effectiveness on a single knowledge graph KGC, limited effort has been devoted to exploring the potentially complementary nature of multiple KGs. In this work, we proposed a novel method called CA-MKGC (Conflict-aware Multilingual Knowledge Graph Completion) for multiple knowledge graph completion (MKGC), aiming to alleviate the sparseness of a single knowledge graph by leveraging information from other knowledge graphs. We designed an intra-KG graph convolutional network encoder that regards the seed alignments between KGs as edges for intra-KG message propagation to model all KGs in a unified model while also adopting an iterative mechanism to progressively incorporate newly predicted alignments along with the newly inferred facts into the learning process. Additionally, we employed an active learning mechanism and a greedy approximation to a semi-constrained optimization problem to focus on learning the structural prior knowledge that is difficult to learn in semantic space, limiting the propagation of error in the iterative training process. Experimental results on multilingual KG datasets demonstrated that our method achieved state-of-the-art results. Ovidiu Serban, Yike Guo |
Knowl. Based Syst. | 4 |
| 2023 | Cloud-VAE: Variational autoencoder with concepts embedded
Zitu Liu, Zhenyao Yu, Yike Guo, Qun Liu 0005, Guoyin Wang 0001 |
Pattern Recognit. | 5 |
| 2023 | Federated Active Learning for Multicenter Collaborative Disease DiagnosisabstractCurrent computer-aided diagnosis system with deep learning method plays an important role in the field of medical imaging. The collaborative diagnosis of diseases by multiple medical institutions has become a popular trend. However, large scale annotations put heavy burdens on medical experts. Furthermore, the centralized learning system has defects in privacy protection and model generalization. To meet these challenges, we propose two federated active learning methods for multicenter collaborative diagnosis of diseases: the Labeling Efficient Federated Active Learning (LEFAL) and the Training Efficient Federated Active Learning (TEFAL). The proposed LEFAL applies a task-agnostic hybrid sampling strategy considering data uncertainty and diversity simultaneously to improve data efficiency. The proposed TEFAL evaluates the client informativeness with a discriminator to improve client efficiency. On the Hyper-Kvasir dataset for gastrointestinal disease diagnosis, with only 65% of labeled data, the LEFAL achieves 95% performance on the segmentation task with whole labeled data. Moreover, on the CC-CCII dataset for COVID-19 diagnosis, with only 50 iterations, the accuracy and F1-score of TEFAL are 0.90 and 0.95, respectively on the classification task. Extensive experimental results demonstrate that the proposed federated active learning methods outperform state-of-the-art methods on segmentation and classification tasks for multicenter collaborative disease diagnosis. Xing Wu 0001, Jie Pei, Cheng Chen 0075, Jianjia Wang, Quan Qian, Yike Guo |
IEEE Trans. Medical Imaging | 9 |
| 2022 | A continuous glucose monitoring measurements forecasting approach via sporadic blood glucose monitoringabstractContinuous glucose monitoring prediction is a crucial yet challenging task in precision medicine. This paper presents a novel neural ODE based approach for predicting continuous glucose monitoring (CGM) levels purely based on sporadic self-monitoring signals. We integrate the expert knowledge from physiological model into our model to improve the accuracy. Experiments on the real-world data demonstrate that our method outperforms other state-of-the-art methods on NRMSE metrics. Yuting Xing, Hangting Ye, Xiaoyu Zhang 0008, Wei Cao 0007, Shun Zheng 0001, Jiang Bian 0002, Yike Guo |
BIBM | 7 |
| 2022 | OmiTrans: Generative Adversarial Networks Based Omics-to-omics Translation FrameworkabstractWith the significant development of high-throughput experimental technologies, different types of omics (e.g., genomics, epigenomics, transcriptomics, proteomics, and metabolomics) data are produced from clinical samples at an unprecedented speed. The correlations between different types of omics data attract enormous research interest. In contrast, the study on genome-wide omics-to-omics translation (i.e., generation and prediction of one type of omics data from another type of omics data) is almost blank. Generative adversarial networks and variants are among the most state-of-the-art deep learning technologies, which have shown great success in image-to-image translation, text-to-image translation, etc. Here we propose OmiTrans, a deep learning framework that adopted the idea of generative adversarial networks to achieve omics-to-omics translation with promising results. OmiTrans can faithfully reconstruct genome-wide gene expression profiles from DNA methylation data with high accuracy and excellent model generalisation, as demonstrated in the experiments. Xiaoyu Zhang 0008, Yike Guo |
BIBM | 2 |
| 2022 | Improving Deep Embedded Clustering via Learning Cluster-level RepresentationsabstractDriven by recent advances in neural networks, various Deep Embedding Clustering (DEC) based short text clustering models are being developed. In these works, latent representation learning and text clustering are performed simultaneously. Although these methods are becoming increasingly popular, they use pure cluster-oriented objectives, which can produce meaningless representations. To alleviate this problem, several improvements have been developed to introduce additional learning objectives in the clustering process, such as models based on contrastive learning. However, existing efforts rely heavily on learning meaningful representations at the instance level. They have limited focus on learning global representations, which are necessary to capture the overall data structure at the cluster level. In this paper, we propose a novel DEC model, which we named the deep embedded clustering model with cluster-level representation learning (DECCRL) to jointly learn cluster and instance level representations. Here, we extend the embedded topic modelling approach to introduce reconstruction constraints to help learn cluster-level representations. Experimental results on real-world short text datasets demonstrate that our model produces meaningful clusters. Qing Yin, Zhihua Wang 0008, Yunya Song, Liang Bai 0001, Yike Guo, Xian Yang 0001 |
COLING | 7 |
| 2022 | Causal Reasoning Methods in Medical Domain: A Review
Xing Wu 0001, Quan Qian, Yike Guo |
IEA/AIE | 5 |
| 2022 | ChoreoGraph: Music-conditioned Automatic Dance Choreography over a Style and Tempo Consistent Dynamic GraphabstractTo generate dance that temporally and aesthetically matches the music is a challenging problem, as the following factors need to be considered. First, the aesthetic styles and messages conveyed by the motion and music should be consistent. Second, the beats of the generated motion should be locally aligned to the musical features. And finally, basic choreomusical rules should be observed, and the motion generated should be diverse. To address these challenges, we propose ChoreoGraph, which choreographs high-quality dance motion for a given piece of music over a Dynamic Graph. A data-driven learning strategy is proposed to evaluate the aesthetic style and rhythmic connections between music and motion in a progressively learned cross-modality embedding space. The motion sequences will be beats-aligned based on the music segments and then incorporated as nodes of a Dynamic Motion Graph. Compatibility factors such as the style and tempo consistency, motion context connection, action completeness, and transition smoothness are comprehensively evaluated to determine the node transition in the graph. We demonstrate that our repertoire-based framework can generate motions with aesthetic consistency and robustly extensible in diversity. Both quantitative and qualitative experiment results show that our proposed model outperforms other baseline models. Ho Yin Au, Jie Chen 0026, Junkun Jiang, Yike Guo |
ACM Multimedia | 4 |
| 2022 | A Dual-Masked Auto-Encoder for Robust Motion Capture with Spatial-Temporal Skeletal Token CompletionabstractMulti-person motion capture can be challenging due to ambiguities caused by severe occlusion, fast body movement, and complex interactions. Existing frameworks build on 2D pose estimations and triangulate to 3D coordinates via reasoning the appearance, trajectory, and geometric consistencies among multi-camera observations. However, 2D joint detection is usually incomplete and with wrong identity assignments due to limited observation angle, which leads to noisy 3D triangulation results. To overcome this issue, we propose to explore the short-range autoregressive characteristics of skeletal motion using transformer. First, we propose an adaptive, identity-aware triangulation module to reconstruct 3D joints and identify the missing joints for each identity. To generate complete 3D skeletal motion, we then propose a Dual-Masked Auto-Encoder (D-MAE) which encodes the joint status with both skeletal-structural and temporal position encoding for trajectory completion. D-MAE's flexible masking and encoding mechanism enable arbitrary skeleton definitions to be conveniently deployed under the same framework. In order to demonstrate the proposed model's capability in dealing with severe data loss scenarios, we contribute a high-accuracy and challenging motion capture dataset of multi-person interactions with severe occlusion. Evaluations on both benchmark and our new dataset demonstrate the efficiency of our proposed model, as well as its advantage against the other state-of-the-art methods. Junkun Jiang, Jie Chen 0026, Yike Guo |
ACM Multimedia | 3 |
| 2022 | TRCA: Text Restoration for Chinese ASR with BERTabstractText restoration plays a vital role in Chinese automatic speech recognition (ASR), which includes punctuation prediction and error correction. However, there are two inevitable challenges for this task. On the one hand, there are no public dataset and model for Chinese punctuation prediction. On the other hand, current text restoration methods for automatic speech recognition only focus on Chinese error correction instead of combining with Chinese punctuation prediction task. To address these problems, a BERT-based text restoration method called TRCA is proposed for Chinese ASR consisting of a Chinese punctuation prediction model and a Chinese error correction model. Experiments demonstrate that the proposed TRCA method outperforms state-of-the-art methods for both punctuation prediction and error correction tasks, among which the proposed TRCA improves the average accuracy to 98% in Chinese punctuation prediction. Xing Wu 0001, Jianjia Wang, Yike Guo |
SoMeT | 4 |
| 2022 | Speech synthesis with face embeddings
Xing Wu 0001, Sihui Ji, Jianjia Wang, Yike Guo |
Appl. Intell. | 4 |
| 2022 | Face aging with pixel-level alignment GAN
Xing Wu 0001, Qing Li 0011, Yangyang Qi, Jianjia Wang, Yike Guo |
Appl. Intell. | 6 |
| 2022 | Digital twins based on bidirectional LSTM and GAN for modelling the COVID-19 pandemic
César Quilodrán Casas, Vinicius L. S. Silva, Rossella Arcucci, Claire E. Heaney, Yike Guo, Christopher C. Pain |
Neurocomputing | 5 |
| 2022 | FTAP: Feature transferring autonomous machine learning pipeline
Xing Wu 0001, Cheng Chen 0075, Mingyu Zhong, Jianjia Wang, Quan Qian, Junfeng Yao, Yike Guo |
Inf. Sci. | 9 |
| 2022 | Weather-degraded image semantic segmentation with multi-task knowledge distillation
Xing Wu 0001, Jianjia Wang, Yike Guo |
Image Vis. Comput. | 4 |
| 2022 | UBAR: User Behavior-Aware Recommendation with knowledge graph
Xing Wu 0001, Yisong Li, Jianjia Wang, Quan Qian, Yike Guo |
Knowl. Based Syst. | 5 |
| 2022 | Suggestive annotation of brain MR images with gradient-guided sampling
Chengliang Dai, Shuo Wang 0011, Yuanhan Mo, Elsa D. Angelini, Yike Guo, Wenjia Bai |
Medical Image Anal. | 5 |
| 2022 | Bayesian data assimilation for estimating instantaneous reproduction numbers during epidemics: Applications to COVID-19abstractEstimating the changes of epidemiological parameters, such as instantaneous reproduction number, Rt, is important for understanding the transmission dynamics of infectious diseases. Current estimates of time-varying epidemiological parameters often face problems such as lagging observations, averaging inference, and improper quantification of uncertainties. To address these problems, we propose a Bayesian data assimilation framework for time-varying parameter estimation. Specifically, this framework is applied to estimate the instantaneous reproduction number Rt during emerging epidemics, resulting in the state-of-the-art 'DARt' system. With DARt, time misalignment caused by lagging observations is tackled by incorporating observation delays into the joint inference of infections and Rt; the drawback of averaging is overcome by instantaneously updating upon new observations and developing a model selection mechanism that captures abrupt changes; the uncertainty is quantified and reduced by employing Bayesian smoothing. We validate the performance of DARt and demonstrate its power in describing the transmission dynamics of COVID-19. The proposed approach provides a promising solution for making accurate and timely estimation for transmission dynamics based on reported data. Xian Yang 0001, Shuo Wang 0011, Yuting Xing, Ling Li 0010, Karl J. Friston, Yike Guo |
PLoS Comput. Biol. | 7 |
| 2022 | Scale-Consistent Fusion: From Heterogeneous Local Sampling to Global Immersive RenderingabstractImage-based geometric modeling and novel view synthesis based on sparse large-baseline samplings are challenging but important tasks for emerging multimedia applications such as virtual reality and immersive telepresence. Existing methods fail to produce satisfactory results due to the limitation on inferring reliable depth information over such challenging reference conditions. With the popularization of commercial light field (LF) cameras, capturing LF images (LFIs) is as convenient as taking regular photos, and geometry information can be reliably inferred. This inspires us to use a sparse set of LF captures to render high-quality novel views globally. However, the fusion of LF captures from multiple angles is challenging due to the scale inconsistency caused by various capture settings. To overcome this challenge, we propose a novel scale-consistent volume rescaling algorithm that robustly aligns the disparity probability volumes (DPV) among different captures for scale-consistent global geometry fusion. Based on the fused DPV projected to the target camera frustum, novel learning-based modules (i.e., the attention-guided multi-scale residual fusion module, and the disparity field-guided deep re-regularization module), which comprehensively regularize noisy observations from heterogeneous captures for high-quality rendering of novel LFIs, have been proposed. Both quantitative and qualitative experiments over the Stanford Lytro Multi-view LF dataset show that the proposed method outperforms state-of-the-art methods significantly under different experiment settings for disparity inference and LF synthesis. Wenpeng Xing, Jie Chen 0026, Zaifeng Yang, Qiang Wang 0022, Yike Guo |
IEEE Trans. Image Process. | 5 |
| 2021 | Label-dependent and event-guided interpretable disease risk prediction using EHRsabstractElectronic health records (EHRs) contain patients’ heterogeneous data that are collected from medical providers involved in the patient’s care, including medical notes, clinical events, laboratory test results, symptoms, and diagnoses. In the field of modern healthcare, predicting whether patients would experience any risks based on their EHRs has emerged as a promising research area, in which artificial intelligence (AI) plays a key role. To make AI models practically applicable, it is required that the prediction results should be both accurate and interpretable. To achieve this goal, this paper proposed a label-dependent and event-guided risk prediction model (LERP) to predict the presence of multiple disease risks by mainly extracting information from unstructured medical notes. Our model is featured in the following aspects. First, we adopt a label-dependent mechanism that gives greater attention to words from medical notes that are semantically similar to the names of risk labels. Secondly, as the clinical events (e.g., treatments and drugs) can also indicate the health status of patients, our model utilizes the information from events and uses them to generate an event-guided representation of medical notes. Thirdly, both label-dependent and event-guided representations are integrated to make a robust prediction, in which the interpretability is enabled by the attention weights over words from medical notes. To demonstrate the applicability of the proposed method, we apply it to the MIMIC-III dataset, which contains real-world EHRs collected from hospitals. Our method is evaluated in both quantitative and qualitative ways. Yunya Song, Qing Yin, Yike Guo, Xian Yang 0001 |
BIBM | 4 |
| 2021 | Self-Supervised Detection of Contextual Synonyms in a Multi-Class Setting: Phenotype Annotation Use CaseabstractJingqing Zhang, Luis Bolanos Trujillo, Tong Li, Ashwani Tanwar, Guilherme Freire, Xian Yang, Julia Ive, Vibhor Gupta, Yike Guo. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Jingqing Zhang, Luis Bolanos, Ashwani Tanwar, Guilherme Freire, Xian Yang 0001, Julia Ive, Vibhor Gupta, Yike Guo |
EMNLP (1) | 9 |
| 2021 | Label Dependent Attention Model for Disease Risk Prediction Using Multimodal Electronic Health RecordsabstractDisease risk prediction has attracted increasing attention in the field of modern healthcare, especially with the latest advances in artificial intelligence (AI). Electronic health records (EHRs), which contain heterogeneous patient information, are widely used in disease risk prediction tasks. One challenge of applying AI models for risk prediction lies in generating interpretable evidence to support the prediction results while retaining the prediction ability. In order to address this problem, we propose the method of jointly embedding words and labels whereby attention modules learn the weights of words from medical notes according to their relevance to the names of risk prediction labels. This approach boosts interpretability by employing an attention mechanism and including the names of prediction tasks in the model. However, its application is only limited to the handling of textual inputs such as medical notes. In this paper, we propose a label dependent attention model LDAM to 1) improve the interpretability by exploiting Clinical-BERT (a biomedical language model pre-trained on a large clinical corpus) to encode biomedically meaningful features and labels jointly; 2) extend the idea of joint embedding to the processing of timeseries data, and develop a multi-modal learning framework for integrating heterogeneous information from medical notes and time-series health status indicators. To demonstrate our method, we apply LDAM to the MIMIC-III dataset to predict different disease risks. We evaluate our method both quantitatively and qualitatively. Specifically, the predictive power of LDAM will be shown, and case studies will be carried out to illustrate its interpretability. Qing Yin, Yunya Song, Yike Guo, Xian Yang 0001 |
ICDM | 4 |
| 2021 | Joint Motion Correction and Super Resolution for Cardiac Segmentation via Latent Optimisation
Shuo Wang 0011, Chen Qin, Nicolò Savioli, Chen Chen 0042, Declan P. O'Regan, Stuart A. Cook, Yike Guo, Daniel Rueckert, Wenjia Bai |
MICCAI (3) | 7 |
| 2021 | XOmiVAE: an interpretable deep learning model for cancer classification using high-dimensional omics dataabstractThe lack of explainability is one of the most prominent disadvantages of deep learning applications in omics. This 'black box' problem can undermine the credibility and limit the practical implementation of biomedical deep learning models. Here we present XOmiVAE, a variational autoencoder (VAE)-based interpretable deep learning model for cancer classification using high-dimensional omics data. XOmiVAE is capable of revealing the contribution of each gene and latent dimension for each classification prediction and the correlation between each gene and each latent dimension. It is also demonstrated that XOmiVAE can explain not only the supervised classification but also the unsupervised clustering results from the deep learning network. To the best of our knowledge, XOmiVAE is one of the first activation level-based interpretable deep learning models explaining novel clusters generated by VAE. The explainable results generated by XOmiVAE were validated by both the performance of downstream tasks and the biomedical knowledge. In our experiments, XOmiVAE explanations of deep learning-based cancer classification and clustering aligned with current domain knowledge including biological annotation and academic literature, which shows great potential for novel biomedical knowledge discovery from deep learning models. Eloise Withnell, Xiaoyu Zhang 0008, Kai Sun 0005, Yike Guo |
Briefings Bioinform. | 4 |
| 2021 | Privacy preservation in federated learning: An insightful survey from the GDPR perspectiveabstractIn recent years, along with the blooming of Machine Learning (ML)-based applications and services, ensuring data privacy and security have become a critical obligation. ML-based service providers not only confront with difficulties in collecting and managing data across heterogeneous sources but also challenges of complying with rigorous data protection regulations such as EU/UK General Data Protection Regulation (GDPR). Furthermore, conventional centralised ML approaches have always come with long-standing privacy risks to personal data leakage, misuse, and abuse. Federated learning (FL) has emerged as a prospective solution that facilitates distributed collaborative learning without disclosing original training data. Unfortunately, retaining data and computation on-device as in FL are not sufficient for privacy-guarantee because model parameters exchanged among participants conceal sensitive information that can be exploited in privacy attacks. Consequently, FL-based systems are not naturally compliant with the GDPR. This article is dedicated to surveying of state-of-the-art privacy-preservation techniques in FL in relations with GDPR requirements. Furthermore, insights into the existing challenges are examined along with the prospective approaches following the GDPR regulatory guidelines that FL-based systems shall implement to fully comply with the GDPR. © 2021 Nguyen Binh Truong, Kai Sun 0005, Siyao Wang, Florian Guitton, Yike Guo |
Comput. Secur. | 5 |
| 2021 | A blockchain-based trust system for decentralised applications: When trustless needs trustabstractBlockchain technology has been envisaged to commence an era of decentralised applications and services (DApps) without the need for a trusted intermediary. Such DApps open a marketplace in which services are delivered to end-users by contributors which are then incentivised by cryptocurrencies in an automated, peer-to-peer, and trustless fashion. However, blockchain, consolidated by smart contracts, only ensures on-chain data security, autonomy and integrity of the business logic execution defined in smart contracts. It cannot guarantee the quality of service of DApps, which entirely depends on the services’ performance. Thus, there is a critical need for a trust system to reduce the risk of dealing with fraudulent counterparts in a blockchain network. These reasons motivate us to develop a fully decentralised trust framework deployed on top of a blockchain platform, operating along with DApps in the marketplace to demoralise deceptive entities while encouraging trustworthy ones. The trust system works as an underlying decentralised service providing a feedback mechanism for end-users and maintaining trust relationships among them in the ecosystem accordingly. We believe this research fortifies the DApps ecosystem by introducing an universal trust middleware for DApps as well as shedding light on the implementation of a decentralised trust system. Nguyen Binh Truong, Gyu Myoung Lee, Kai Sun 0005, Florian Guitton, Yike Guo |
Future Gener. Comput. Syst. | 5 |
| 2020 | Suggestive Annotation of Brain Tumour Images with Gradient-Guided Sampling
Chengliang Dai, Shuo Wang 0011, Yuanhan Mo, Kaichen Zhou, Elsa D. Angelini, Yike Guo, Wenjia Bai |
MICCAI (4) | 6 |
| 2020 | Deep Generative Model-Based Quality Control for Cardiac MRI Segmentation
Shuo Wang 0011, Giacomo Tarroni, Chen Qin, Yuanhan Mo, Chengliang Dai, Chen Chen 0042, Ben Glocker, Yike Guo, Daniel Rueckert, Wenjia Bai |
MICCAI (4) | 8 |
| 2020 | Improving taxonomic relation learning via incorporating relation descriptions into word embeddingsabstractSummary Taxonomic relations play an important role in various Natural Language Processing (NLP) tasks (eg, information extraction, question answering and knowledge inference). Existing approaches on embedding‐based taxonomic relation learning mainly rely on the word embeddings trained using co‐occurrence‐based similarity learning. However, the performance of these approaches is not quite satisfactory due to the lack of sufficient taxonomic semantic knowledge within word embeddings. To solve this problem, we propose an improved embedding‐based approach to learn taxonomic relations via incorporating relation descriptions into word embeddings. First, to capture additional taxonomic semantic knowledge, we train special word embeddings using not only co‐occurrence information of words but also relation descriptions (eg, taxonomic seed relations and their contextual triples). Then, using the trained word embeddings as features, we employ two learning models to identify and predict taxonomic relations, namely, offset‐based classification model and offset‐based similarity model. Experimental results on four real‐world domain datasets demonstrate that our proposed approach can capture additional taxonomic semantic knowledge and reduce dependence on the training dataset, outperforming the state‐of‐the‐art compared approaches on the taxonomic relation learning task. Subin Huang, Xiangfeng Luo, Hao Wang 0097, Shengwei Gu, Yike Guo |
Concurr. Comput. Pract. Exp. | 6 |
| 2020 | Multiple premises entailment recognition based on attention and gate mechanism
Pin Wu, Zhidan Lei, Rukang Zhu, Xuting Chang, Junwu Sun, Wenjie Zhang 0005, Yike Guo |
Expert Syst. Appl. | 8 |
| 2020 | Towards a large-scale twitter observatory for political events
Senaka Fernando, Julio Amador Díaz López, Ovidiu Serban, Juan Gómez-Romero, Miguel Molina-Solana, Yike Guo |
Future Gener. Comput. Syst. | 6 |
| 2020 | Open Visualization Environment (OVE): A web framework for scalable rendering of data visualizations
Senaka Fernando, James Scott-Brown, Ovidiu Serban, David Birch, David Akroyd, Miguel Molina-Solana, Thomas Heinis, Yike Guo |
Future Gener. Comput. Syst. | 8 |
| 2020 | Tuoris: A middleware for visualizing dynamic graphics in scalable resolution display environments
Víctor Martínez, Senaka Fernando, Miguel Molina-Solana, Yike Guo |
Future Gener. Comput. Syst. | 4 |
| 2020 | The assessment of small bowel motility with attentive deformable neural network
Xing Wu 0001, Mingyu Zhong, Yike Guo, Hamido Fujita |
Inf. Sci. | 3 |
| 2020 | The autonomous navigation and obstacle avoidance for USVs with ANOA deep reinforcement learning method
Xing Wu 0001, Haolei Chen, Changgu Chen, Mingyu Zhong, Shaorong Xie, Yike Guo, Hamido Fujita |
Knowl. Based Syst. | 6 |
| 2020 | GDPR-Compliant Personal Data Management: A Blockchain-Based SolutionabstractThe General Data Protection Regulation (GDPR) gives control of personal data back to the owners by appointing higher requirements and obligations on service providers who manage and process personal data. As the verification of GDPR-compliance, handled by a supervisory authority, is irregularly conducted; it is challenging to be certified that a service provider has been continuously adhering to the GDPR. Furthermore, it is beyond the data owner's capability to perceive whether a service provider complies with the GDPR and effectively protects her personal data. This motivates us to envision a design concept for developing a GDPR-compliant personal data management platform leveraging the emerging blockchain and smart contract technologies. The goals of the platform are to provide decentralised mechanisms to both service providers and data owners for processing personal data; meanwhile, empower data provenance and transparency by leveraging advanced features of the blockchain technology. The platform enables data owners to impose data usage consent, ensures only designated parties can process personal data, and logs all data activities in an immutable distributed ledger using smart contract and cryptography techniques. By honestly participating in the platform, a service provider can be endorsed by the blockchain network that it is fully GDPR-compliant; otherwise, any violation is immutably recorded and is easily figured out by associated parties. We then demonstrate the feasibility and efficiency of the proposed design concept by developing a profile management platform implemented on top of the Hyperledger Fabric permissioned blockchain framework, following by valuable analysis and discussion. Nguyen Binh Truong, Kai Sun 0005, Gyu Myoung Lee, Yike Guo |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | Unsupervised Annotation of Phenotypic Abnormalities via Semantic Latent Representations on Electronic Health RecordsabstractThe extraction of phenotype information which is naturally contained in electronic health records (EHRs) has been found to be useful in various clinical informatics applications such as disease diagnosis. However, due to imprecise descriptions, lack of gold standards and the demand for efficiency, annotating phenotypic abnormalities on millions of EHR narratives is still challenging. In this work, we propose a novel unsupervised deep learning framework to annotate the phenotypic abnormalities from EHRs via semantic latent representations. The proposed framework takes the advantage of Human Phenotype Ontology (HPO), which is a knowledge base of phenotypic abnormalities, to standardize the annotation results. Experiments have been conducted on 52,722 EHRs from MIMIC-III dataset. Quantitative and qualitative analysis have shown the proposed framework achieves state-of-the-art annotation performance and computational efficiency compared with other methods. Jingqing Zhang, Xiaoyu Zhang 0008, Kai Sun 0005, Xian Yang 0001, Chengliang Dai, Yike Guo |
BIBM | 6 |
| 2019 | Integrated Multi-omics Analysis Using Variational Autoencoders: Application to Pan-cancer ClassificationabstractOmics data are normally high dimensional with large number of molecular features and relatively small number of available samples with clinical labels. The “curse of dimensionality” makes it challenging to train a machine learning model using high dimensional omics data like DNA methylation and gene expression profiles. Here we propose an end-to-end deep learning model called OmiVAE to extract low dimensional features and classify samples from multi-omics data. OmiVAE combines the basic structure of variational autoencoders with a classifier to achieve task-oriented feature extraction and multi-class classification. The training procedure of OmiVAE is comprised of an unsupervised phase and a supervised phase. During the unsupervised phase, a hierarchical cluster structure of samples can be automatically formed without the need for labels. And in the supervised phase, OmiVAE achieved an average accuracy of 97.49% after 10-fold cross-validation among 33 tumour types and normal samples, which shows better performance than existing methods. The integrated model learned from multi-omics datasets outperformed those using only one type of omics data, which indicates that the complementary information from different omics datatypes provides useful insights for biomedical tasks like cancer classification. Xiaoyu Zhang 0008, Jingqing Zhang, Kai Sun 0005, Xian Yang 0001, Chengliang Dai, Yike Guo |
BIBM | 6 |
| 2019 | Generative Creativity: Adversarial Learning for Bionic Design
Simiao Yu, Hao Dong 0003, Pan Wang 0005, Chao Wu 0001, Yike Guo |
ICANN (3) | 5 |
| 2019 | SIMGAN: Photo-Realistic Semantic Image Manipulation Using Generative Adversarial NetworksabstractSemantic image manipulation (SIM) aims to generate realistic images from an input source image and a target text description, such that the generated images not only match the content of the description, but also maintain text-irrelevant features of the source image. It requires to learn a good mapping between visual features and linguistic features. Previous works on SIM can only generate images of limited resolution that typically lack of fine and clear details. In this work, we aim to generate high-resolution photo-realistic images for SIM. Specifically, we propose SIMGAN, a generative adversarial networks (GAN) based architecture that is capable of generating images of size 256 × 256 for SIM. We demonstrate the effectiveness of SIMGAN and its superiority over existing methods via qualitative and quantitative evaluation on Caltech-200 and Oxford-102 datasets. Simiao Yu, Hao Dong 0003, Felix Liang, Yuanhan Mo, Chao Wu 0001, Yike Guo |
ICIP | 6 |
| 2019 | Compositional Microservices for Immersive Social Visual AnalyticsabstractAs humans, we have developed to process highly complex visual data from our surroundings. This is why data visualization and interaction is one of the quickest ways to facilitate investigation and communicate understanding. To perform visual analytics effectively at the big data scale it is crucial that we develop an integrated processing and visualization ecosystem. However, to date, in Large High-Resolution Display (LHRD) environments the worlds of data processing and visualization remain largely disconnected. In this paper, we propose a common architectural approach to enable integrated data processing and distributed visualization via the composition of discrete microservices. Each of these microservices provides a very specific clearly-defined function, such as analyzing data, creating a visualization, sharding data or providing a synchronization source. By defining common transport, data and API formats we enable the composition of these microservices from processing raw data through to analytics, visualization and rendering. This compositionality, inspired by successful data-driven visualization frameworks provides a common platform for immersive social visual analytics. Senaka Fernando, David Birch, Miguel Molina-Solana, Douglas McIlwraith, Yike Guo |
IV (1) | 5 |
| 2019 | Self-Supervised Learning for Cardiac MR Image Segmentation by Anatomical Position Prediction
Wenjia Bai, Chen Chen 0042, Giacomo Tarroni, Jinming Duan 0001, Florian Guitton, Steffen E. Petersen, Yike Guo, Paul M. Matthews, Daniel Rueckert |
MICCAI (2) | 7 |
| 2019 | Blockchain-based Personal Data Management: From Fiction to SolutionabstractThe emerging blockchain technology has enabled various decentralised applications in a trustless environment without relying on a trusted intermediary. It is expected as a promising solution to tackle sophisticated challenges on personal data management, thanks to its advanced features such as immutability, decentralisation and transparency. Although certain approaches have been proposed to address technical difficulties in personal data management; most of them only provided preliminary methodological exploration. Alarmingly, when utilising Blockchain for developing a personal data management system, fictions have occurred in existing approaches and been promulgated in the literature. Such fictions are theoretically doable; however, by thoroughly breaking down consensus protocols and transaction validation processes, we clarify that such existing approaches are either impractical or highly inefficient due to the natural limitations of the blockchain and Smart Contracts technologies. This encourages us to propose a feasible solution in which such fictions are reduced by designing a novel system architecture with a blockchain-based “proof of permission” protocol. We demonstrate the feasibility and efficiency of the proposed models by implementing a clinical data sharing service built on top of a public blockchain platform. We believe that our research resolves existing ambiguity and take a step further on providing a practically feasible solution for decentralised personal data management. Nguyen Binh Truong, Kai Sun 0005, Yike Guo |
NCA | 3 |
| 2019 | The Collaborative Strategy of Multiple USVs with Deep Reinforcement Learning MethodabstractThe unmanned surface vehicle (USV) has been widely used to accomplish tasks that cannot be completed by ships with human drivers on certain sea areas. It is not only necessary but essential to obtain a robust strategy in order to ensure multiple USVs accomplish collaborative tasks successfully and efficiently. To meet the challenge, a deep reinforcement learning method is proposed, which is combined with an improved A star algorithm. A statistically promising collaborative strategy is achieved by the proposed method under the guidance from the unmanned aerial vehicles (UAVs). After the collaborative strategy is generated, the improved A star algorithm is used to navigate the USVs. To verify the proposed algorithm, several tasks are tested on a simulation platform. Experimental results demonstrate that the proposed method outperforms state-of-the-art reinforcement learning methods such as DQN and DeepSarsa. © 2019 The authors and IOS Press. All rights reserved. Xing Wu 0001, Mingyu Zhong, Guofei Feng, Shaorong Xie, Yike Guo |
SoMeT | 5 |
| 2019 | Hierarchical attention based long short-term memory for Chinese lyric generation
Xing Wu 0001, Zhikang Du, Yike Guo, Hamido Fujita |
Appl. Intell. | 3 |
| 2019 | A machine learning attack against variable-length Chinese character CAPTCHAs
Xing Wu 0001, Shuji Dai, Yike Guo, Hamido Fujita |
Appl. Intell. | 3 |
| 2019 | Navigating the disease landscape: knowledge representations for contextualizing molecular signaturesabstractLarge amounts of data emerging from experiments in molecular medicine are leading to the identification of molecular signatures associated with disease subtypes. The contextualization of these patterns is important for obtaining mechanistic insight into the aberrant processes associated with a disease, and this typically involves the integration of multiple heterogeneous types of data. In this review, we discuss knowledge representations that can be useful to explore the biological context of molecular signatures, in particular three main approaches, namely, pathway mapping approaches, molecular network centric approaches and approaches that represent biological statements as knowledge graphs. We discuss the utility of each of these paradigms, illustrate how they can be leveraged with selected practical examples and identify ongoing challenges for this field of research. Mansoor A. S. Saqi, Artem Lysenko, Yike Guo, Tatsuhiko Tsunoda, Charles Auffray |
Briefings Bioinform. | 3 |
| 2019 | Data and knowledge management in translational research: implementation of the eTRIKS platform for the IMI OncoTrack consortiumabstractBACKGROUND: For large international research consortia, such as those funded by the European Union's Horizon 2020 programme or the Innovative Medicines Initiative, good data coordination practices and tools are essential for the successful collection, organization and analysis of the resulting data. Research consortia are attempting ever more ambitious science to better understand disease, by leveraging technologies such as whole genome sequencing, proteomics, patient-derived biological models and computer-based systems biology simulations. RESULTS: The IMI eTRIKS consortium is charged with the task of developing an integrated knowledge management platform capable of supporting the complexity of the data generated by such research programmes. In this paper, using the example of the OncoTrack consortium, we describe a typical use case in translational medicine. The tranSMART knowledge management platform was implemented to support data from observational clinical cohorts, drug response data from cell culture models and drug response data from mouse xenograft tumour models. The high dimensional (omics) data from the molecular analyses of the corresponding biological materials were linked to these collections, so that users could browse and analyse these to derive candidate biomarkers. CONCLUSIONS: In all these steps, data mapping, linking and preparation are handled automatically by the tranSMART integration platform. Therefore, researchers without specialist data handling skills can focus directly on the scientific questions, without spending undue effort on processing the data and data integration, which are otherwise a burden and the most time-consuming part of translational research data analysis. Reha Yildirimman, Emmanuel Van der Stuyft, Denny Verbeeck, Sascha Herzinger, Venkata P. Satagopam, Adriano Barbosa-Silva, Reinhard Schneider 0002, Bodo M. H. Lange, Hans Lehrach, Yike Guo, David Henderson, Anthony Rowe 0002 |
BMC Bioinform. | 11 |
| 2019 | Entity emotion mining in social media environmentabstractSummary With the thriving of the social media (eg, Sina Microblog and Twitter), the public is keen on expressing their opinions or views on entities as celebrities and products. Emotion mining on social media can be applied in diverse areas, such as helping government or organizations understand people's attitudes so as to make right decisions for business services or political campaigns. One kind of the mainstream approaches for emotion mining are based on topic models, while most of these existing approaches aim at analyzing entire sentiments for the whole Microblog, rather than explicitly assigning the relevant sentiments to the specific entities. In addition, topic model is mainly based on the bag‐of‐words model but ignores the semantic relations of entity‐word in the modeling process, which brings low accuracy and poor interpretability to the sentiment analysis. To overcome the aforementioned difficulties, an Entity Sentiment Topic Model (ESTM) is proposed, which carries out entity‐dependent sentiment analysis. To improve the accuracy of sentiment analysis and enhance the interpretability of the results, ESTM is integrated with relations of entity‐word and a six‐dimensional emotion lexicon as weakly supervised information. Experiments have shown promising results on sentiment classification accuracy, interpretability, and quality of coherent topics for entities. Xuefeng Fu, Xiangfeng Luo, Yike Guo |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Explanations by arbitrated argumentative dispute
Kristijonas Cyras, David Birch, Yike Guo, Francesca Toni, Rajvinder Dulay, Sally Turvey, Daniel Greenberg, Tharindi Hapuarachchi |
Expert Syst. Appl. | 3 |
| 2019 | An artificial intelligence based data-driven approach for design ideation
Liuqing Chen 0002, Pan Wang 0005, Hao Dong 0003, Feng Shi 0007, Yike Guo, Peter R. N. Childs, Jun Xiao 0001, Chao Wu 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2019 | An unsupervised approach for learning a Chinese IS-A taxonomy from an unstructured corpus
Subin Huang, Xiangfeng Luo, Yike Guo, Shengwei Gu |
Knowl. Based Syst. | 4 |
| 2019 | I_MDS: an inflammatory bowel disease molecular activity score to classify patients with differing disease-driving pathways and therapeutic response to anti-TNF treatmentabstractCrohn's disease and ulcerative colitis are driven by both common and distinct underlying mechanisms of pathobiology. Both diseases, exhibit heterogeneity underscored by the variable clinical responses to therapeutic interventions. We aimed to identify disease-driving pathways and classify individuals into subpopulations that differ in their pathobiology and response to treatment. We applied hierarchical clustering of enrichment scores derived from gene set variation analysis of signatures representative of various immunological processes and activated cell types, to a colonic biopsy dataset that included healthy volunteers, Crohn's disease and ulcerative colitis patients. Patient stratification at baseline or after anti-TNF treatment in clinical responders and non-responders was queried. Signatures with significantly different enrichment scores were identified using a general linear model. Comparisons to healthy controls were made at baseline in all participants and then separately in responders and non-responders. Fifty-nine percent of the signatures were commonly enriched in both conditions at baseline, supporting the notion of a disease continuum within ulcerative colitis and Crohn's disease. Signatures included T cells, macrophages, neutrophil activation and poly:IC signatures, representing acute inflammation and a complex mix of potential disease-driving biology. Collectively, identification of significantly enriched signatures allowed establishment of an inflammatory bowel disease molecular activity score which uses biopsy transcriptomics as a surrogate marker to accurately track disease severity. This score separated diseased from healthy samples, enabled discrimination of clinical responders and non-responders at baseline with 100% specificity and 78.8% sensitivity, and was validated in an independent data set that showed comparable classification. Comparing responders and non-responders separately at baseline to controls, 43% and 70% of signatures were enriched, respectively, suggesting greater molecular dysregulation in TNF non-responders at baseline. This methodological approach could facilitate better targeted design of clinical studies to test therapeutics, concentrating on patient subsets sharing similar underlying pathobiology, therefore increasing the likelihood of clinical response. Stelios Pavlidis, Calixte Monast, Matthew J. Loza, Patrick Branigan, Kiang F. Chung, Ian M. Adcock, Yike Guo, Anthony Rowe 0002, Frédéric Baribaud |
PLoS Comput. Biol. | 7 |
| 2019 | An Information-Theoretical Framework for Cluster EnsembleabstractCluster ensemble is a very important tool that aggregates several base clusterings to generate a single output clustering with improved robustness and stability. However, the quality of the final clustering is often affected by uncertainties on the generation and integration of base clusterings. In this paper, we develop an information-theoretical framework which makes an effort to obtain a final clustering with high consensus on both the original data set and the base clustering set by minimizing the two uncertainties of cluster ensemble. In this framework, we provide a weighted consensus measure based on information entropy to evaluate the quality of a clustering, the similarity between clusters and the similarity between objects. Based on the measure, we propose three weighted cluster ensemble algorithms with different ensemble strategies in the framework, including the weighted feature consensus algorithm, the weighted relabeling consensus algorithm and the weighted pairwise-similarity consensus algorithm. In the experimental analysis, we compare the proposed algorithms with other existing clustering ensemble algorithms on several data sets. The comparison results illustrate the proposed algorithms are very effective and robust. Liang Bai 0001, Jiye Liang, Hangyuan Du, Yike Guo |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2018 | Optimising Toward Completed Videos in an Online Video Advertising ExchangeabstractPredictive modelling is used extensively within online programmatic advertising [1], [2]. By predicting online behaviour [3], advertisers and intermediaries can ensure that they are targeting the most profitable users. This has typically been used to predict user clicks in ‘display advertising’, where static images are shown, however in recent years we’ve seen a rise in more complex advert types such as video. Due to the maturity of programmatic video advertising, similar modelling techniques for video events have yet to be demonstrated. In this paper, we outline a deployed system that automatically optimises toward video adverts that people complete (i.e. reach the end of). We discuss the specifics of a cloud based pipeline for data processing which handles 240M records a day in less than 1 hour aggregate wall clock time and provides a 4.4% point increase in completion rate over a two week trial. We provide results demonstrating the success of our system in deployment. Douglas McIlwraith, Andrea Catalucci, Sam Boyd, Raouf Aghrout, Yike Guo |
COMPSAC (1) | 5 |
| 2018 | Multiple Feature Fusion for Automatic Emotion Recognition Using EEG SignalsabstractAutomatic emotion recognition based on electroencephalo-graphic (EEG) signals has received increasing attention in recent years. The Deep Residual Networks (ResNets) can solve vanishing gradient problem and exploding gradient problem well in computer vision and can learn more profound semantic information. And for traditional methods, frequency features often play important role in signal processing area. Thus, in this paper, we use the pre-trained ResNets to extract deep semantic information and the linear-frequency cepstral coefficients (LFCC) as features from raw EEG signals. Then the two features are fused to improve the emotion classification performance of our approach. Moreover, several classifiers are used for our fused features to evaluate the performance and it shows that the proposed approach is effective for emotion classification. We find that the best performance is achieved when use k-nearst neighbor (KNN) as classifier, and we provide a detailed discussion for the reason. Ningjie Liu, Yuchun Fang, Ling Li 0010, Limin Hou, Fenglei Yang, Yike Guo |
ICASSP | 6 |
| 2018 | A Multi Tenant Computational Platform for Translational MedicineabstractTranslational biomedical research has become a science driven by big data. Improving patient care by developing personalized therapies and new drugs depends increasingly on an organization's ability to rapidly and intelligently leverage complex molecular and clinical data from a variety of large-scale internal and external, partner and public, data sources. As analysing these large-scale and complex datasets has become increasingly computationally expensive, it is of paramount importance to enable researchers to seamlessly scale up their computation platform while being able to manage complex yet flexible scenario that biomedical scientists are asking for. We developed a new platform as an answer to those needs of analysing and exploring massive amounts of medical data with the constrain of enabling the broadest audience, ranging from the medical doctor to the advanced coders, to easily and intuitively exploit this new resource. The platform consists of three main components: Borderline UI, the eTRIKS Analytical Environment (eAE) and the eTRIKS Data Platform (eDP). Each component has been developed independently to address specific sets of problems, then, loosely connected to each other components as to form a coherent platform for large scale medical data analysis. Axel Oehmichen, Florian Guitton, Paul Agapow, Ibrahim Emam, Yike Guo |
ICDCS | 5 |
| 2018 | Deep Sequence Learning with Auxiliary Information for Traffic PredictionabstractPredicting traffic conditions from online route queries is a challenging task as there are many complicated interactions over the roads and crowds involved. In this paper, we intend to improve traffic prediction by appropriate integration of three kinds of implicit but essential factors encoded in auxiliary information. We do this within an encoder-decoder sequence learning framework that integrates the following data: 1) offline geographical and social attributes. For example, the geographical structure of roads or public social events such as national celebrations; 2) road intersection information. In general, traffic congestion occurs at major junctions; 3) online crowd queries. For example, when many online queries issued for the same destination due to a public performance, the traffic around the destination will potentially become heavier at this location after a while. Qualitative and quantitative experiments on a real-world dataset from Baidu have demonstrated the effectiveness of our framework. Binbing Liao, Jingqing Zhang, Chao Wu 0001, Douglas McIlwraith, Tong Chen 0006, Shengwen Yang, Yike Guo, Fei Wu 0001 |
KDD | 7 |
| 2018 | The Deep Poincaré Map: A Novel Approach for Left Ventricle Segmentation
Yuanhan Mo, Fangde Liu, Douglas McIlwraith, Guang Yang 0006, Jingqing Zhang, Taigang He, Yike Guo |
MICCAI (4) | 7 |
| 2018 | Dest-ResNet: A Deep Spatiotemporal Residual Network for Hotspot Traffic Speed PredictionabstractWith the ever-increasing urbanization process, the traffic jam has become a common problem in the metropolises around the world, making the traffic speed prediction a crucial and fundamental task. This task is difficult due to the dynamic and intrinsic complexity of the traffic environment in urban cities, yet the emergence of crowd map query data sheds new light on it. In general, a burst of crowd map queries for the same destination in a short duration (called "hotspot'') could lead to traffic congestion. For example, queries of the Capital Gym burst on weekend evenings lead to traffic jams around the gym. However, unleashing the power of crowd map queries is challenging due to the innate spatiotemporal characteristics of the crowd queries. To bridge the gap, this paper firstly discovers hotspots underlying crowd map queries. These discovered hotspots address the spatiotemporal variations. Then Dest-ResNet (Deep spatiotemporal Residual Network) is proposed for hotspot traffic speed prediction. Dest-ResNet is a sequence learning framework that jointly deals with two sequences in different modalities, i.e., the traffic speed sequence and the query sequence. The main idea of Dest-ResNet is to learn to explain and amend the errors caused when the unimodal information is applied individually. In this way, Dest-ResNet addresses the temporal causal correlation between queries and the traffic speed. As a result, Dest-ResNet shows a 30% relative boost over the state-of-the-art methods on real-world datasets from Baidu Map. Binbing Liao, Jingqing Zhang, Siliang Tang, Chao Wu 0001, Shengwen Yang, Wenwu Zhu 0001, Yike Guo, Fei Wu 0001 |
ACM Multimedia | 9 |
| 2018 | The Multiple Unmanned Surface Vehicles Cooperative Defense Based on PM-PSO and GA-PSO in the Sophisticated Sea EnvironmentabstractThe unmanned surface vehicles (USVs) have become a major trend in the construction of naval equipment and its flexibility and intelligence making it widely used in real-scenes. For cooperative defense with multiple USVs to intercept intruders, it is proposed that planning the path with obstacle avoidance and protecting the target by task allocation actions. The particle swarm optimization based on probe mechanism (PM-PSO) is proposed for pathing planning with obstacle avoidance. With the consideration of the constraints of different defense schemes such as the path cost, the interception loss, the defense income and so on, it is proposed that the dispersed particle swarm optimization based on genetic algorithm (GA-PSO) for the interception task allocation. Furthermore, the fitness function is proposed to evaluate the feasibility of the interception path and the quality of the allocation scheme. Extensive simulation experiments are conducted and demonstrated the effectiveness, rationality and superiority of the proposed methods. Yuan Liu 0025, Xing Wu 0001, Yike Guo, Shaorong Xie, Huayan Pu, Yan Peng 0001 |
SoMeT | 3 |
| 2018 | Crowdsourcing with online quantitative design analysisabstractDesign is a balancing act between people’s competing concerns, design options and design performance. Recently collecting data on such concerns such as sustainability or aesthetics has become possible through online crowdsourcing, particularly in 3d. However, such systems rarely present more than a single design alternative or allow users to change the design and seldom provide quantitative design analysis to gauge design performance. This precludes a more participatory approach including a wider audience and their insight in the design process. To improve the design process we propose a system to assist the design team in exploring the balance of concerns, design options and their performance. We augment a 3d visualisation crowdsourcing environment with quantitative on-demand assessment of design variants run in the cloud. This enables crowdsourced exploration of the design space and its performance. Automated participant tracking and explicit submitted feedback on design options are collated and presented to aid the design team in balancing the demands of urban master planning. We report application of this system to an urban masterplan with Arup. David Birch, Alvise Simondetti, Yike Guo |
Adv. Eng. Informatics | 3 |
| 2018 | Visualizing large knowledge graphs: A performance analysisabstractKnowledge graphs are an increasingly important source of data and context information in Data Science. A first step in data analysis is data exploration, in which visualization plays a key role. Currently, Semantic Web technologies are prevalent for modeling and querying knowledge graphs; however, most visualization approaches in this area tend to be overly simplified and targeted to small-sized representations. In this work, we describe and evaluate the performance of a Big Data architecture applied to large-scale knowledge graph visualization. To do so, we have implemented a graph processing pipeline in the Apache Spark framework and carried out several experiments with real-world and synthetic graphs. We show that distributed implementations of the graph building, metric calculation and layout stages can efficiently manage very large graphs, even without applying partitioning or incremental processing strategies. Juan Gómez-Romero, Miguel Molina-Solana, Axel Oehmichen, Yike Guo |
Future Gener. Comput. Syst. | 4 |
| 2018 | A novel community detection algorithm based on simplification of complex networks
Liang Bai 0001, Jiye Liang, Hangyuan Du, Yike Guo |
Knowl. Based Syst. | 4 |
| 2018 | A visual attention-based keyword extraction for document classification
Xing Wu 0001, Zhikang Du, Yike Guo |
Multim. Tools Appl. | 3 |
| 2018 | An Ensemble Clusterer of Multiple Fuzzy k-Means Clusterings to Recognize Arbitrarily Shaped ClustersabstractFuzzy cluster ensemble is an important research component of ensemble learning, which is used to aggregate several fuzzy base clusterings to generate a single output clustering with improved robustness and quality. However, since clustering is unsupervised, where “accuracy” does not have a clear meaning, it is difficult for existing ensemble methods to integrate multiple fuzzy k-means clusterings to find arbitrarily shaped clusters. To overcome the deficiency, we propose a new ensemble clusterer (algorithm) of multiple fuzzy k-means clusterings based on a local hypothesis. In the new algorithm, we study the extraction of local-credible memberships from a base clustering, the production of multiple base clusterings with different local-credible spaces, and the construction of cluster relation based on indirect overlap of local-credible spaces. The proposed ensemble clusterer not only inherits the scalability of fuzzy k-means but also overcomes the inability to find arbitrarily shaped clusters. We compare the proposed algorithm with other cluster ensemble algorithms on several synthetical and real datasets. The experimental results illustrate the effectiveness and efficiency of the proposed algorithm. Liang Bai 0001, Jiye Liang, Yike Guo |
IEEE Trans. Fuzzy Syst. | 3 |
| 2018 | Dropping Activation Outputs With Localized First-Layer Deep Network for Enhancing User Privacy and Data SecurityabstractDeep learning methods can play a crucial role in anomaly detection, prediction, and supporting decision making for applications like personal health-care, pervasive body sensing, and so on. However, current architecture of deep networks suffers the privacy issue that users need to give out their data to the model (typically hosted in a server or a cluster on Cloud) for training or prediction. This problem is getting more severe for those sensitive health-care or medical data (e.g., fMRI or body sensors measures like EEG signals). In addition to this, there is also a security risk of leaking these data during the data transmission from user to the model (especially when it is through the Internet). Targeting at these issues, in this paper, we proposed a new architecture for deep network in which users do not reveal their original data to the model. In our method, feed-forward propagation and data encryption are combined into one process: we migrate the first layer of deep network to users' local devices and apply the activation functions locally, and then use the “dropping activation output” method to make the output non-invertible. The resulting approach is able to make model prediction without accessing users' sensitive raw data. The experiment conducted in this paper showed that our approach achieves the desirable privacy protection requirement and demonstrated several advantages over the traditional approach with encryption/decryption. Hao Dong 0003, Chao Wu 0001, Yike Guo |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | DAGAN: Deep De-Aliasing Generative Adversarial Networks for Fast Compressed Sensing MRI ReconstructionabstractCompressed sensing magnetic resonance imaging (CS-MRI) enables fast acquisition, which is highly desirable for numerous clinical applications. This can not only reduce the scanning cost and ease patient burden, but also potentially reduce motion artefacts and the effect of contrast washout, thus yielding better image quality. Different from parallel imaging-based fast MRI, which utilizes multiple coils to simultaneously receive MR signals, CS-MRI breaks the Nyquist-Shannon sampling barrier to reconstruct MRI images with much less required raw data. This paper provides a deep learning-based strategy for reconstruction of CS-MRI, and bridges a substantial gap between conventional non-learning methods working only on data from a single image, and prior knowledge from large training data sets. In particular, a novel conditional Generative Adversarial Networks-based model (DAGAN)-based model is proposed to reconstruct CS-MRI. In our DAGAN architecture, we have designed a refinement learning method to stabilize our U-Net based generator, which provides an end-to-end network to reduce aliasing artefacts. To better preserve texture and edges in the reconstruction, we have coupled the adversarial loss with an innovative content loss. In addition, we incorporate frequency-domain information to enforce similarity in both the image and frequency domains. We have performed comprehensive comparison studies with both conventional CS-MRI reconstruction methods and newly investigated deep learning approaches. Compared with these methods, our DAGAN method provides superior reconstruction with preserved perceptual image details. Furthermore, each image is reconstructed in about 5 ms, which is suitable for real-time processing. Guang Yang 0006, Simiao Yu, Hao Dong 0003, Gregory Slabaugh, Pier Luigi Dragotti, Xujiong Ye, Fangde Liu, Simon R. Arridge, Jennifer Keegan, Yike Guo, David N. Firmin |
IEEE Trans. Medical Imaging | 10 |
| 2017 | eTRIKS analytical environment: A modular high performance framework for medical data analysisabstractTranslational research is quickly becoming a science driven by big data. Improving patient care, developing personalized therapies and new drugs depend increasingly on an organization's ability to rapidly and intelligently leverage complex molecular and clinical data from a variety of large-scale partner and public sources. As analysing these large-scale datasets becomes computationally increasingly expensive, traditional analytical engines are struggling to provide a timely answer to the questions that biomedical scientists are asking. Designing such a framework is developing for a moving target as the very nature of biomedical research based on big data requires an environment capable of adapting quickly and efficiently in response to evolving questions. The resulting framework consequently must be scalable in face of large amounts of data, flexible, efficient and resilient to failure. In this paper we design the eTRIKS Analytical Environment (eAE), a scalable and modular framework for the efficient management and analysis of large scale medical data, in particular the massive amounts of data produced by high-throughput technologies. We particularly discuss how we design the eAE as a modular and efficient framework enabling us to add new components or replace old ones easily. We further elaborate on its use for a set of challenging big data use cases in medicine and drug discovery. Axel Oehmichen, Florian Guitton, Kai Sun 0005, Jean Grizet, Thomas Heinis, Yike Guo |
IEEE BigData | 6 |
| 2017 | Semantic Image Synthesis via Adversarial LearningabstractIn this paper, we propose a way of synthesizing realistic images directly with natural language description, which has many useful applications, e.g. intelligent image manipulation. We attempt to accomplish such synthesis: given a source image and a target text description, our model synthesizes images to meet two requirements: 1) being realistic while matching the target text description; 2) maintaining other image features that are irrelevant to the text description. The model should be able to disentangle the semantic information from the two modalities (image and text), and generate new images from the combined semantics. To achieve this, we proposed an end-to-end neural architecture that leverages adversarial learning to automatically learn implicit loss functions, which are optimized to fulfill the aforementioned two requirements. We have evaluated our model by conducting experiments on Caltech-200 bird dataset and Oxford-102 flower dataset, and have demonstrated that our model is capable of synthesizing realistic images that match the given descriptions, while still maintain other features of original images. Hao Dong 0003, Simiao Yu, Chao Wu 0001, Yike Guo |
ICCV | 4 |
| 2017 | I2T2I: Learning text to image synthesis with textual data augmentationabstractTranslating information between text and image is a fundamental problem in artificial intelligence that connects natural language processing and computer vision. In the past few years, performance in image caption generation has seen significant improvement through the adoption of recurrent neural networks (RNN). Meanwhile, text-to-image generation begun to generate plausible images using datasets of specific categories like birds and flowers. We've even seen image generation from multi-category datasets such as the Microsoft Common Objects in Context (MSCOCO) through the use of generative adversarial networks (GANs). Synthesizing objects with a complex shape, however, is still challenging. For example, animals and humans have many degrees of freedom, which means that they can take on many complex shapes. We propose a new training method called Image-Text-Image (I2T2I) which integrates text-to-image and image-to-text (image captioning) synthesis to improve the performance of text-to-image synthesis. We demonstrate that I2T2I can generate better multi-categories images using MSCOCO than the state-of-the-art. We also demonstrate that I2T2I can achieve transfer learning by using a pre-trained image captioning module to generate human images on the MPII Human Pose dataset (MHP) without using sentence annotation. Hao Dong 0003, Jingqing Zhang, Douglas McIlwraith, Yike Guo |
ICIP | 4 |
| 2017 | Assimilated Learning - Bridging the Gap between Big Data and Smart Data
Yike Guo |
IoTBDS | 1 |
| 2017 | TensorLayer: A Versatile Library for Efficient Deep Learning DevelopmentabstractRecently we have observed emerging uses of deep learning techniques in multimedia systems. Developing a practical deep learning system is arduous and complex. It involves labor-intensive tasks for constructing sophisticated neural networks, coordinating multiple network models, and managing a large amount of training-related data. To facilitate such a development process, we propose TensorLayer which is a Python-based versatile deep learning library. TensorLayer provides high-level modules that abstract sophisticated operations towards neuron layers, network models, training data and dependent training jobs. In spite of offering simplicity, it has transparent module interfaces that allows developers to flexibly embed low-level controls within a backend engine, with the aim of supporting fine-grain tuning towards training. Real-world cluster experiment results show that TensorLayeris able to achieve competitive performance and scalability in critical deep learning tasks. TensorLayer was released in September 2016 on GitHub. Since after, it soon become one of the most popular open-sourced deep learning library used by researchers and practitioners. Hao Dong 0003, Akara Supratak, Luo Mai, Fangde Liu, Axel Oehmichen, Simiao Yu, Yike Guo |
ACM Multimedia | 7 |
| 2017 | The Cooperative Defense Strategy by Multi-USVsabstractBased on the multi-agents system control theory and technology, this paper presents the cooperative defense process of multiple unmanned surface vehicles (USVs) operating in complicated sea environment, and explains the quantification of the battle effectiveness, cooperative strategy, task allocation and finally describes in detail the cooperative strategies on random graph. Then we point out the problems in the current cooperative defense process and the future development direction. The cooperative defense research of USVs in the sea environment has a great significance to the effective promotion of social and military efficiency. Yuan Liu 0025, Xing Wu 0001, Yike Guo, Shaorong Xie, Huayan Pu, Yan Peng 0001 |
SoMeT | 3 |
| 2017 | The Combat of Unmanned Surface Vehicles Based on Wolves AttackabstractUnmanned combat system is one of the development trend of modern weapons and equipment and has applied to military affairs. The major goal of the Unmanned Surface Vehicles (USVs in short) is to destroy protected targets in the shortest time. This paper originates from biology, putting forward a new attack strategy—The USV combat Based on Wolves Attack with Weight. Namely, using the characteristics of the wolves attack to study the process of attacking. It also discusses the attack measures from weights, velocity and firepower of USVs through weight distribution and summarizes the advantages of wolves attack in USV combat. Juan Pu, Xing Wu 0001, Yike Guo, Shaorong Xie, Huayan Pu, Yan Peng 0001 |
SoMeT | 3 |
| 2017 | The Cooperative Defense System by Team of USVs in Complicated Sea EnvironmentabstractBased on the multi-agents system control theory and technology, this paper presents the construction of cooperative defense system of multiple unmanned surface vehicles (USVs) operating in complicated sea environment, develops the mathematical formula describing the trajectory equation when the USV intercepts the intruder, and explains the proposed coordination control method used in the corresponding defense system. Then we point out the future development requirement of the cooperative defense system. The cooperative defense system research of USVs in the sea environment has a great significance to the effective promotion of social and military efficiency, and it is the basis of the implement about cooperative strategies. Xing Wu 0001, Yuan Liu 0025, Yike Guo, Shaorong Xie, Huayan Pu, Yan Peng 0001 |
SoMeT | 3 |
| 2017 | tranSMART-XNAT Connector tranSMART-XNAT connector - image selection based on clinical phenotypes and genetic profilesabstractMotivation: TranSMART has a wide range of functionalities for translational research and a large user community, but it does not support imaging data. In this context, imaging data typically includes 2D or 3D sets of magnitude data and metadata information. Imaging data may summarise complex feature descriptions in a less biased fashion than user defined plain texts and numeric numbers. Imaging data also is contextualised by other data sets and may be analysed jointly with other data that can explain features or their variation. Results: Here we describe the tranSMART-XNAT Connector we have developed. This connector consists of components for data capture, organisation and analysis. Data capture is responsible for imaging capture either from PACS system or directly from an MRI scanner, or from raw data files. Data are organised in a similar fashion as tranSMART and are stored in a format that allows direct analysis within tranSMART. The connector enables selection and download of DICOM images and associated resources using subjects' clinical phenotypic and genotypic criteria. Availability and Implementation: tranSMART-XNAT connector is written in Java/Groovy/Grails. It is maintained and available for download at https://github.com/sh107/transmart-xnat-connector.git. Contact: [email protected]. Sijin He, May Yong, Paul M. Matthews, Yike Guo |
Bioinform. | 4 |
| 2017 | Fast graph clustering with a new description model for community detection
Liang Bai 0001, Xueqi Cheng 0001, Jiye Liang, Yike Guo |
Inf. Sci. | 4 |
| 2017 | Small bowel motility assessment based on fully convolutional networks and long short-term memory
Mengqi Pei, Xing Wu 0001, Yike Guo, Hamido Fujita |
Knowl. Based Syst. | 3 |
| 2017 | Fast density clustering strategies based on the k-means algorithm
Liang Bai 0001, Xueqi Cheng 0001, Jiye Liang, Huawei Shen, Yike Guo |
Pattern Recognit. | 5 |
| 2017 | ACM TIST Special Issue on Data-Driven Intelligence for Wireless NetworkingabstractNo abstract available. Wenwu Zhu 0001, Jean C. Walrand, Yike Guo, Zhi Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2016 | CGDM: collaborative genomic data model for molecular profiling data using NoSQLabstractMOTIVATION: High-throughput molecular profiling has greatly improved patient stratification and mechanistic understanding of diseases. With the increasing amount of data used in translational medicine studies in recent years, there is a need to improve the performance of data warehouses in terms of data retrieval and statistical processing. Both relational and Key Value models have been used for managing molecular profiling data. Key Value models such as SeqWare have been shown to be particularly advantageous in terms of query processing speed for large datasets. However, more improvement can be achieved, particularly through better indexing techniques of the Key Value models, taking advantage of the types of queries which are specific for the high-throughput molecular profiling data. RESULTS: In this article, we introduce a Collaborative Genomic Data Model (CGDM), aimed at significantly increasing the query processing speed for the main classes of queries on genomic databases. CGDM creates three Collaborative Global Clustering Index Tables (CGCITs) to solve the velocity and variety issues at the cost of limited extra volume. Several benchmarking experiments were carried out, comparing CGDM implemented on HBase to the traditional SQL data model (TDM) implemented on both HBase and MySQL Cluster, using large publicly available molecular profiling datasets taken from NCBI and HapMap. In the microarray case, CGDM on HBase performed up to 246 times faster than TDM on HBase and 7 times faster than TDM on MySQL Cluster. In single nucleotide polymorphism case, CGDM on HBase outperformed TDM on HBase by up to 351 times and TDM on MySQL Cluster by up to 9 times. AVAILABILITY AND IMPLEMENTATION: The CGDM source code is available at https://github.com/evanswang/CGDM. CONTACT: [email protected]. Shicai Wang, Mihaela A. Mares, Yike Guo |
Bioinform. | 3 |
| 2016 | Wiki-Health: From Quantified Self to Self-Understanding
Yang Li 0003, Yike Guo |
Future Gener. Comput. Syst. | 2 |
| 2015 | Joint affine transformation and loop pipelining for mapping nested loop on CGRAs
Shouyi Yin, Dajiang Liu, Leibo Liu, Shaojun Wei, Yike Guo |
DATE | 5 |
| 2015 | Cooperatively managing dynamic writeback and insertion policies in a last-level DRAM cache
Shouyi Yin, Leibo Liu, Shaojun Wei, Yike Guo |
DATE | 5 |
| 2015 | Resampling-Based Variable Selection with Lasso for p >> n and Partially Linear ModelsabstractThe linear model of the regression function is a widely used and perhaps, in most cases, highly unrealistic simplifying assumption, when proposing consistent variable selection methods for large and highly-dimensional datasets. In this paper, we study what happens from theoretical point of view, when a variable selection method assumes a linear regression function and the underlying ground-truth model is composed of a linear and a non-linear term, that is at most partially linear. We demonstrate consistency of the Lasso method when the model is partially linear. However, we note that the algorithm tends to increase even more the number of selected false positives on partially linear models when given few training samples. That is usually because the values of small groups of samples happen to explain variation coming from the non-linear part of the response function and the noise, using a linear combination of wrong predictors. We demonstrate theoretically that false positives are likely to be selected by the Lasso method due to a small proportion of samples, which happen to explain some variation in the response variable. We show that this property implies that if we run the Lasso on several slightly smaller size data replications, sampled without replacement, and intersect the results, we are likely to reduce the number of false positives without losing already selected true positives. We propose a novel consistent variable selection algorithm based on this property and we show it can outperform other variable selection methods on synthetic datasets of linear and partially linear models and datasets from the UCI machine learning repository. Mihaela A. Mares, Yike Guo |
ICMLA | 2 |
| 2014 | Enabling Performance as a Service for a Cloud Storage SystemabstractOne of the main contributions of the paper is that we introduce "performance as a service" as a key component for future cloud storage environments. This is achieved through demonstration of the design and implementation of a multi-tier cloud storage system (CACSS), and the illustration of a linear programming model that helps to predict future data access patterns for efficient data caching management. The proposed caching algorithm aims to leverage the cloud economy by incorporating both potential performance improvement and revenue-gain into the storage systems. Yang Li 0003, Li Guo 0002, Akara Supratak, Yike Guo |
IEEE CLOUD | 4 |
| 2014 | Optimising parallel R correlation matrix calculations on gene expression data using MapReduceabstractBACKGROUND: High-throughput molecular profiling data has been used to improve clinical decision making by stratifying subjects based on their molecular profiles. Unsupervised clustering algorithms can be used for stratification purposes. However, the current speed of the clustering algorithms cannot meet the requirement of large-scale molecular data due to poor performance of the correlation matrix calculation. With high-throughput sequencing technologies promising to produce even larger datasets per subject, we expect the performance of the state-of-the-art statistical algorithms to be further impacted unless efforts towards optimisation are carried out. MapReduce is a widely used high performance parallel framework that can solve the problem. RESULTS: In this paper, we evaluate the current parallel modes for correlation calculation methods and introduce an efficient data distribution and parallel calculation algorithm based on MapReduce to optimise the correlation calculation. We studied the performance of our algorithm using two gene expression benchmarks. In the micro-benchmark, our implementation using MapReduce, based on the R package RHIPE, demonstrates a 3.26-5.83 fold increase compared to the default Snowfall and 1.56-1.64 fold increase compared to the basic RHIPE in the Euclidean, Pearson and Spearman correlations. Though vanilla R and the optimised Snowfall outperforms our optimised RHIPE in the micro-benchmark, they do not scale well with the macro-benchmark. In the macro-benchmark the optimised RHIPE performs 2.03-16.56 times faster than vanilla R. Benefiting from the 3.30-5.13 times faster data preparation, the optimised RHIPE performs 1.22-1.71 times faster than the optimised Snowfall. Both the optimised RHIPE and the optimised Snowfall successfully performs the Kendall correlation with TCGA dataset within 7 hours. Both of them conduct more than 30 times faster than the estimated vanilla R. CONCLUSIONS: The performance evaluation found that the new MapReduce algorithm and its implementation in RHIPE outperforms vanilla R and the conventional parallel algorithms implemented in R Snowfall. We propose that MapReduce framework holds great promise for large molecular data analysis, in particular for high-dimensional genomic data such as that demonstrated in the performance evaluation described in this paper. We aim to use this new algorithm as a basis for optimising high-throughput molecular data correlation calculation for Big Data. Shicai Wang, Ioannis Pandis, David Johnson 0006, Ibrahim Emam, Florian Guitton, Axel Oehmichen, Yike Guo |
BMC Bioinform. | 7 |
| 2014 | An iterative parameter estimation method for biological systems and its parallel implementationabstractSUMMARY One difficulty in building a mechanistic model of biological systems lies in determining correct parameter values. This paper proposes a novel parameter estimation method to infer unknown parameters, such as kinetic rates, from noisy experimental observations. Derived from the approximate Bayesian computation sequential Monte Carlo algorithm, our method predicts the distribution of each parameter rather than a single value via several intermediate distributions. Motivated by the computational intensity of the method, we improve the approximate Bayesian computation sequential Monte Carlo method in two aspects. First, to increase the efficiency, a windowing method is developed to reduce the parameter‐searching space, and an adaptive sampling weight mechanism is introduced to make the intermediate distributions converge to the target distributions in a much quicker manner. Second, to speed up the estimation process, we implement our method in a parallel computing environment to speed up the sampling process. Copyright © 2013 John Wiley & Sons, Ltd. Xian Yang 0001, Yike Guo, Li Guo 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Enabling cost-aware and adaptive elasticity of multi-tier cloud applications
Rui Han 0001, Moustafa Ghanem, Li Guo 0002, Yike Guo, Michelle Osmond |
Future Gener. Comput. Syst. | 4 |
| 2014 | Cloud computing in e-Science: research challenges and opportunities
David Wallom, Simon Waddington, Jianwu Wang 0001, Arif Shaon, Brian Matthews, Michael D. Wilson, Yike Guo, Li Guo 0002, Jonathan D. Blower, Athanasios V. Vasilakos, Philip Kershaw |
J. Supercomput. | 8 |
| 2013 | Elastic algorithms for guaranteeing quality monotonicity in big data miningabstractWhen mining large data volumes in big data applications users are typically willing to use algorithms that produce acceptable approximate results satisfying the given resource and time constraints. Two key challenges arise when designing such algorithms. The first relates to reasoning about tradeoffs between the quality of data mining output, e.g. prediction accuracy for classification tasks and available resource and time budgets. The second is organizing the computation of the algorithm to guarantee producing better quality of results as more budget is used. Little work has addressed these two challenges together in a generic way. In this paper, we propose a novel framework for developing elastic big data mining algorithms. Based on Shannon's entropy, an information-theoretic approach is introduced to reason about how result quality is affected by the allocated budget. This is then used to guide the development of algorithms that adapt to the available time budgets while guaranteeing producing better quality results as more budgets are used. We demonstrate the application of the framework by developing elastic k-Nearest Neighbour (kNN) classification and collaborative filtering (CF) recommendation algorithms as two examples. The core of both elastic algorithms is to use a naïve kNN classification or CF algorithm over R-tree data structures that successively approximate the entire datasets. Experimental evaluation was performed using prediction accuracy as quality metric on real datasets. The results show that elastic mining algorithms indeed produce results with consistent increase in observable qualities, i.e., prediction accuracy, in practice. Rui Han 0001, Lei Nie 0008, Moustafa Ghanem, Yike Guo |
IEEE BigData | 4 |
| 2013 | Building a generic platform for big sensor data applicationabstractThe drive toward smart cities alongside the rising adoption of personal sensors is leading to a torrent of sensor data. While systems exist for storing and managing sensor data, the real value of such data is the insight which can be generated from it. However there is currently no platform which enables sensor data to be taken from collection, through use in models to produce useful data products. The architecture of such a platform is a current research question in the field of Big Data and Smart Cities. In this paper we explore five key challenges in this field and provide a response through a sensor data platform “Concinnity” which can take sensor data from collection to final product via a data repository and workflow system. This will enable rapid development of applications built on sensor data using data fusion and the integration and composition of models to form novel workflows. We summarize the key features of our approach, exploring how it enables value to be derived from sensor data efficiently. Chun-Hsiang Lee, David Birch, Chao Wu 0001, Dilshan Silva, Orestis Tsinalis, Yang Li 0003, Shulin Yan, Moustafa Ghanem, Yike Guo |
IEEE BigData | 9 |
| 2013 | Enhanced user data privacy with pay-by-data modelabstractPersonal data collection is becoming pervasive these days, these data has the risk of being abused by current application and application marketplace model, because only the price of application is explicitly indicated without clear agreement on usage of data, and the granularity of data access authentication is not enough to protect users privacy. In this short paper, we propose a new model of user data privacy. Data usage of the application is explicitly shown, and controlled by an authentication service, to protect users from the abuse of their data, especially in mobile application. Chao Wu 0001, Yike Guo |
IEEE BigData | 2 |
| 2013 | Cloud Resource Monitoring for Intrusion DetectionabstractWe present a novel security monitoring framework for intrusion detection in IaaS cloud infrastructures. The framework uses statistical anomaly detection techniques over data monitored both inside and outside each Virtual Machine instance. We present the architecture of our monitoring framework and describe the implementation of the real-time monitors and detectors. We also describe how the framework is used in three different attack scenarios. For each of the three attack scenarios, we describe how the attack itself works and how it could be detected. We describe what data is monitored in our framework and how the detection is conducted using anomaly detection methods. We also present evaluation of the detection using synthetic and real data sets. Our experimental evaluation across all three scenarios shows that our tools perform well in practical situations and provide a promising direction for future research. Sijin He, Moustafa Ghanem, Li Guo 0002, Yike Guo |
CloudCom (2) | 4 |
| 2013 | Sensor Deployment in Bayesian Compressive Sensing Based Environmental Monitoring
Chao Wu 0001, Di Wu 0002, Shulin Yan, Yike Guo |
MobiQuitous | 4 |
| 2012 | Improving Resource Utilisation in the Cloud Environment Using Multivariate Probabilistic ModelsabstractResource provisioning based on virtual machine (VM) has been widely accepted and adopted in cloud computing environments. A key problem resulting from using static scheduling approaches for allocating VMs on different physical machines (PMs) is that resources tend to be not fully utilised. Although some existing cloud reconfiguration algorithms have been developed to address the problem, they normally result in high migration costs and low resource utilisation due to ignoring the multi-dimensional characteristics of VMs and PMs. In this paper we present and evaluate a new algorithm for improving resource utilisation for cloud providers. By using a multivariate probabilistic model, our algorithm selects suitable PMs for VM re-allocation which are then used to generate a reconfiguration plan. We also describe two heuristics metrics which can be used in the algorithm to capture the multi-dimensional characteristics of VMs and PMs. By combining these two heuristics metrics in our experiments, we observed that our approach improves the resource utilisation level by around 8% for cloud providers, such as IC Cloud, which accept user-defined VM configurations and 14% for providers, such as Amazon EC2, which only provide limited types of VM configurations. Sijin He, Li Guo 0002, Moustafa Ghanem, Yike Guo |
IEEE CLOUD | 4 |
| 2012 | Elastic Application Container: A Lightweight Approach for Cloud Resource ProvisioningabstractVirtual machine (VM) based virtual infrastructure has been adopted widely in cloud computing environment for elastic resource provisioning. Performing resource management using VMs, however, is a heavyweight task. In practice, we have identified two scenarios where VM based resource management is less feasible and less resource-efficient. In this paper, we propose a lightweight resource management model that is called Elastic Application Container (EAC). EAC is a virtual resource unit for delivering better resource efficiency and more scalable cloud applications. We describe the EAC system architecture and components, and also present an algorithm for EAC resource provisioning. We also describe an implementation of the EAC-oriented platform to support multi-tenant cloud use. To evaluate our approach and implementation, we conducted experiments and collected performance data by comparing VM-based and EAC-based resource management with regards to their feasibility and resource-efficiency. The experiment results show that our proposed EAC-based resource management approach outperforms the VM-based approach in terms of feasibility and resource-efficiency. Sijin He, Li Guo 0002, Yike Guo, Chao Wu 0001, Moustafa Ghanem, Rui Han 0001 |
AINA | 3 |
| 2012 | Lightweight Resource Scaling for Cloud ApplicationsabstractElastic resource provisioning is a key feature of cloud computing, allowing users to scale up or down resource allocation for their applications at run-time. To date, most practical approaches to managing elasticity are based on allocation/de-allocation of the virtual machine (VM) instances to the application. This VM-level elasticity typically incurs both considerable overhead and extra costs, especially for applications with rapidly fluctuating demands. In this paper, we propose a lightweight approach to enable cost-effective elasticity for cloud applications. Our approach operates fine-grained scaling at the resource level itself (CPUs, memory, I/O, etc) in addition to VM-level scaling. We also present the design and implementation of an intelligent platform for light-weight resource management of cloud applications. We describe our algorithms for light-weight scaling and VM-level scaling and show their interaction. We then use an industry standard benchmark to evaluate the effectiveness of our approach and compare its performance against traditional approaches. Rui Han 0001, Li Guo 0002, Moustafa Ghanem, Yike Guo |
CCGRID | 4 |
| 2012 | Developing a novel integrated model of p38 MAPK and glucocorticoid signalling pathwaysabstractGlucocorticoid (GC) resistance is a key mechanism by which traditional asthma treatments become ineffective for patients, yet the molecular characteristics of the associated regulatory changes are largely unknown. Significant evidence suggests that crosstalk between p38 Mitogen Activated Protein Kinase (MAPK) and GC signalling pathways may contribute to this resistance. Based on a number of studies, a simplified GC signalling pathway model was developed and integrated with a pre-existing model of the p38 MAPK pathway. It is predicted that with experimental data, the validity and use of this model can be confirmed, corrections and updates can be made where necessary, and that through the two pathways' interface points the existence and scale of crosstalk can be examined. Alex Holehouse, Xian Yang 0001, Ian M. Adcock, Yike Guo |
CIBCB | 4 |
| 2012 | CACSS: Towards a Generic Cloud Storage Service
Yang Li 0003, Li Guo 0002, Yike Guo |
CLOSER | 3 |
| 2012 | Does the Cloud need new algorithms? An introduction to elastic algorithmsabstractCloud computing has emerged as a cost-effective way to deliver metered computing resources. Within a Cloud, elasticity of resource usage is typically realized through the “on-demand” provision principle supported by the “Pay-as-You-Go” business model. However, little, or no work, has investigated elasticity of algorithms for Cloud computing. In this paper, we introduce novel research on elastic algorithms (EA) where the computation itself is organized in a “Pay-as-You-Go” fashion. In contrast to conventional algorithms, where computation is a deterministic process that only produces an “ali-or-nothing” result, an EA generates a sequence of approximate results corresponding to its resource consumption. As more resources are consumed, better results will be derived. In this sense, the quality of the algorithm is elastic to its resource consumption. In the paper, we formalize the proeprties of elasticity and also formalize desirable properties for elastic algorithms themselves. We illustrate the design of an EA for kNN classification in the context of machine learning and discuss its properties. Finally we provide an ambitious agenda for future research in this area. Yike Guo, Moustafa Ghanem, Rui Han 0001 |
CloudCom | 1 |
| 2012 | A New Paradigm for Web App Development, Deployment, Distribution, and Collaboration
Chao Wu 0001, Yike Guo |
ICSOFT | 2 |
| 2012 | Modelling and performance analysis of clinical pathways using the stochastic process algebra PEPAabstractBACKGROUND: Hospitals nowadays have to serve numerous patients with limited medical staff and equipment while maintaining healthcare quality. Clinical pathway informatics is regarded as an efficient way to solve a series of hospital challenges. To date, conventional research lacks a mathematical model to describe clinical pathways. Existing vague descriptions cannot fully capture the complexities accurately in clinical pathways and hinders the effective management and further optimization of clinical pathways. METHOD: Given this motivation, this paper presents a clinical pathway management platform, the Imperial Clinical Pathway Analyzer (ICPA). By extending the stochastic model performance evaluation process algebra (PEPA), ICPA introduces a clinical-pathway-specific model: clinical pathway PEPA (CPP). ICPA can simulate stochastic behaviours of a clinical pathway by extracting information from public clinical databases and other related documents using CPP. Thus, the performance of this clinical pathway, including its throughput, resource utilisation and passage time can be quantitatively analysed. RESULTS: A typical clinical pathway on stroke extracted from a UK hospital is used to illustrate the effectiveness of ICPA. Three application scenarios are tested using ICPA: 1) redundant resources are identified and removed, thus the number of patients being served is maintained with less cost; 2) the patient passage time is estimated, providing the likelihood that patients can leave hospital within a specific period; 3) the maximum number of input patients are found, helping hospitals to decide whether they can serve more patients with the existing resource allocation. CONCLUSIONS: ICPA is an effective platform for clinical pathway management: 1) ICPA can describe a variety of components (state, activity, resource and constraints) in a clinical pathway, thus facilitating the proper understanding of complexities involved in it; 2) ICPA supports the performance analysis of clinical pathway, thereby assisting hospitals to effectively manage time and resources in clinical pathway. Xian Yang 0001, Rui Han 0001, Yike Guo, Jeremy T. Bradley, Benita Cox, Robert Dickinson, Richard Kitney |
BMC Bioinform. | 3 |
| 2011 | Real Time Elastic Cloud Management for Limited ResourcesabstractAn Infrastructure-as-a-Service (IaaS) provider is usually assumed to own a large data centre with significant computational resources. For a small or medium sized Internet Data Centre (IDC), offering cloud computing service is a nature of business model but there are technical barriers which need to be resolved. One of the key issues is ineffective resource management given such an IDC usually has only limited resource. In this paper, we propose an efficient resource management solution specially designed for helping small and medium sized IaaS cloud providers to better utilise their hardware resources with minimum operational cost. Such an optimised resource utilisation is achieved by a well-designed underlying hardware infrastructure, an efficient resource scheduling algorithm and a set of migrating operations of VMs. Sijin He, Li Guo 0002, Yike Guo |
IEEE CLOUD | 3 |
| 2011 | A Deployment Platform for Dynamically Scaling Applications in the CloudabstractSimplifying the process of deploying applications is almost essential in the cloud. However, existing techniques can automate applications' initial deployment but have not yet adequately addressed their scaling problem. In this paper, a deployment platform to enable a novel dynamic scaling technique is introduced. This platform employs: (i) an extensible specification that describes all aspects of applications, (ii) a flexible analytical model that determines how many servers to be deployed for an application in each scaling. The platform's ability to handle dynamic workloads and to scale applications quickly enough to maintain the response time target is demonstrated. Rui Han 0001, Li Guo 0002, Yike Guo, Sijin He |
CloudCom | 3 |
| 2011 | Finding consistent disease subnetworks across microarray datasetsabstractBACKGROUND: While contemporary methods of microarray analysis are excellent tools for studying individual microarray datasets, they have a tendency to produce different results from different datasets of the same disease. We aim to solve this reproducibility problem by introducing a technique (SNet). SNet provides both quantitative and descriptive analysis of microarray datasets by identifying specific connected portions of pathways that are significant. We term such portions within pathways as "subnetworks". RESULTS: We tested SNet on independent datasets of several diseases, including childhood ALL, DMD and lung cancer. For each of these diseases, we obtained two independent microarray datasets produced by distinct labs on distinct platforms. In each case, our technique consistently produced almost the same list of significant nontrivial subnetworks from two independent sets of microarray data. The gene-level agreement of these significant subnetworks was between 51.18% to 93.01%. In contrast, when the same pairs of microarray datasets were analysed using GSEA, t-test and SAM, this percentage fell between 2.38% to 28.90% for GSEA, 49.60% tp 73.01% for t-test, and 49.96% to 81.25% for SAM. Furthermore, the genes selected using these existing methods did not form subnetworks of substantial size. Thus it is more probable that the subnetworks selected by our technique can provide the researcher with more descriptive information on the portions of the pathway actually affected by the disease. CONCLUSIONS: These results clearly demonstrate that our technique generates significant subnetworks and genes that are more consistent and reproducible across datasets compared to the other popular methods available (GSEA, t-test and SAM). The large size of subnetworks which we generate indicates that they are generally more biologically significant (less likely to be spurious). In addition, we have chosen two sample subnetworks and validated them with references from biological literature. This shows that our algorithm is capable of generating descriptive biologically conclusions. Donny Soh, Difeng Dong, Yike Guo, Limsoon Wong |
BMC Bioinform. | 3 |
| 2010 | IC Cloud: A Design Space for Composable Cloud ComputingabstractCloud computing has attracted great interest from both academic and industrial communities. Different paradigms, architectures and applications have emerged. However, to the best of our knowledge, only few efforts have been devoted to study the architecture as well as implementation details for building up a cloud computing system. In this paper, we present our design and implementation of\textit{Imperial College Cloud (IC Cloud)}. The goal of IC Cloud is to provide a generic design space where various cloud computing architectures and implementation strategies can be systematically studied. The IC Cloud design strictly follows the SOA principle and incorporates a highly flexible system design approach. Li Guo 0002, Yike Guo, Xiangchuan Tian |
IEEE CLOUD | 2 |
| 2010 | Consistency, comprehensiveness, and compatibility of pathway databasesabstractBACKGROUND: It is necessary to analyze microarray experiments together with biological information to make better biological inferences. We investigate the adequacy of current biological databases to address this need. DESCRIPTION: Our results show a low level of consistency, comprehensiveness and compatibility among three popular pathway databases (KEGG, Ingenuity and Wikipathways). The level of consistency for genes in similar pathways across databases ranges from 0% to 88%. The corresponding level of consistency for interacting genes pairs is 0%-61%. These three original sources can be assumed to be reliable in the sense that the interacting gene pairs reported in them are correct because they are curated. However, the lack of concordance between these databases suggests each source has missed out many genes and interacting gene pairs. CONCLUSIONS: Researchers will hence find it challenging to obtain consistent pathway information out of these diverse data sources. It is therefore critical to enable them to access these sources via a consistent, comprehensive and unified pathway API. We accumulated sufficient data to create such an aggregated resource with the convenience of an API to access its information. This unified resource can be accessed at http://www.pathwayapi.com. Donny Soh, Difeng Dong, Yike Guo, Limsoon Wong |
BMC Bioinform. | 3 |
| 2009 | Real-Time Data Mining Methodology and a Supporting FrameworkabstractThe need for real-time data mining has long been recognized in various application domains. However existing methodologies are still limited to the optimization of single classical data mining algorithms. In this paper, we investigate the development of a general purpose methodology for real-time data mining and propose a novel supporting framework. In the methodology, definition, characteristics and principles of real-time data mining are finely studied. The framework is proposed based on the novel dynamic data mining process model. The model offers the ability to incrementally update data mining knowledge and synchronously execute data mining tasks; an implementation of the framework and a case study are also presented. Xiong Deng, Moustafa Ghanem, Yike Guo |
NSS | 3 |
| 2009 | Eight Times Acceleration of Geospatial Data Archiving and Distribution on the GridsabstractA grid-powered Web Geographical Information Science (GIS)/Web Processing Service (WPS) system has been developed for archiving and distributing large volumes of geospatial data. However, users, WPS servers, and data resources are always distributed across different locations, attempting to access and archive geospatial data from a GIS survey via conventional Hypertext Transport Protocol, Network File System Protocol, and File Transfer Protocol, which often encounters long waits and frustration in wide area network (WAN) environments. To provide a “local-like” performance, a WAN/grid-optimized protocol known as “GridJet” developed at our lab was used as the underlying engine between WPS servers and clients, which utilizes a wide range of technologies including the one of paralleling the remote file access. No change in the way of using software is required since the multistreamed GridJet protocol remains fully compatible with the existing IP infrastructures. Our recent progress includes a real-world test that PyWPS and Google Earth over the GridJet protocol beat those over the classic ones by a factor of two to eight, where the distribution/archiving distance is over 10 000 km. Frank Wang, Na Helian, Sining Wu, Yike Guo, Yuhui Deng 0001, Lingkui Meng, Wen Zhang 0011, Jon Crowcroft, Jean Bacon, Michael Andrew Parker |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2009 | Correction to "Eight Times Acceleration of Geospatial Data Archiving and Distribution on the Grids"abstractIn the above-named work the name of one of the authors is incorrectly given. Included here also is the biography with the missing IEEE membership information. Frank Wang, Na Helian, Sining Wu, Yike Guo, Yuhui Deng 0001, Lingkui Meng, Wen Zhang 0011, Jon Crowcroft, Jean Bacon, Michael Andrew Parker |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2007 | Supporting scientific discovery processes in Discovery NetabstractAbstract The activity of e‐Science involves making discoveries by analysing data to find new knowledge. Discoveries of value cannot be made by simply performing a pre‐defined set of steps to produce a result. Rather, there is an original, creative aspect to the activity that by its nature cannot be automated. In addition to finding new knowledge, discovery therefore also concerns finding a process to find new knowledge. How discovery processes are modelled is therefore key to effectively practicing e‐Science. We argue that since a discovery process instance serves a similar purpose to a mathematical proof it should have similar properties, namely it allows results to be deterministically reproduced when re‐executed and that intermediate results can be viewed to aid examination and comprehension. We examine the issues involved for software environments used to make discoveries to preserve these properties, and show how they are tackled in the Discovery Net system. Copyright © 2006 John Wiley & Sons, Ltd. Jameel Syed, Moustafa Ghanem, Yike Guo |
Concurr. Comput. Pract. Exp. | 3 |
| 2007 | Grid-Oriented Storage: A Single-Image, Cross-Domain, High-Bandwidth ArchitectureabstractThis paper describes the grid-oriented storage (GOS) architecture and its implementations. A GOS-specific file system (GOS-FS), the single-purpose intent of a GOS OS, and secure interfaces via grid security infrastructure (GSI) motivate and enable this new architecture. As an FTP server, GOS with a slimmed OS, with a total volume of around 150 MB, outperforms the standard GridFTP by 20-40 percent. As a file server, GOS-FS acts as a network/grid interface, enabling a user to perform searches and access resources without downloading them locally. In the real-world tests between Cambridge and Beijing, where the transfer distance is 10,000 km, the multistreamed GOS-FS file opening/saving resulted in a remarkable performance increase of about 2-25 times, compared to the single-streamed network file system (NFSv4). GOS is expected to be a variant of or successor to the well-used network-attached storage (NAS) and/or storage area network (SAN) products in the grid era Frank Wang, Sining Wu, Na Helian, Michael Andrew Parker, Yike Guo, Yuhui Deng 0001, Vineet R. Khare |
IEEE Trans. Computers | 5 |
| 2006 | Achievements and Experiences from a Grid-Based Earthquake Analysis and Modelling StudyabstractWe have developed and used a grid-based geoinformatics infrastructure and analytical methods for investigating the relationship between macro and microscale earthquake deformational processes by linking geographically distributed and computationally intensive earthquake monitoring and modelling tools. Using this infrastructure, measurement of lateral co-seismic deformation is carried out with imageodesy algorithms running on servers at the London eScience Centre. The resultant deformation field is used to initialise geomechanical simulations of the earthquake deformation running on supercomputers based at the University of Oklahoma. This paper describes the details of our work, summarizes our scientific results and details our experiences from implementing and testing the distributed infrastructure and analysis workflow. Jian Guo Liu 0005, Moustafa Ghanem, Vasa Curcin, Christian Haselwimmer, Yike Guo, Gareth Morgan, Kyran Mish |
e-Science | 5 |
| 2006 | KDE Bioscience: Platform for bioinformatics analysis workflows
Pei Hao, Vasa Curcin, Wei-Zhong He, Qingming Luo, Yike Guo |
J. Biomed. Informatics | 7 |
| 2005 | A grid infrastructure for mixed bioinformatics data and text miningabstractSummary form only given. We present a distributed infrastructure for mixed data and text mining. Our approach is based on extending the discovery net infrastructure, a grid-computing environment for knowledge discovery, to allow end users to construct complex distributed text and data mining workflows. We describe our architecture, data model and visual programming approach and present a number of mixed data text mining examples over biological data. Moustafa Ghanem, Alexandros Chortaras, Yike Guo, Anthony Rowe 0002, Jon Ratcliffe |
AICCSA | 3 |
| 2005 | Bridging the Macro and Micro: A Computing Intensive Earthquake Study Using Discovery NetabstractWe present the development and use of a novel distributed geohazard modeling environment for the analysis and interpretation of large scale earthquake data sets. Our work demonstrates, for the first time, how earthquake-related surface deformation measured from satellite images using imageodesy algorithms is coupled with analysis and simulation using finite-element numerical models. Our work realises a real time distributed analytical environment where analysis and simulation are closely coupled; integrating high performance implementations of image mining components executing on dedicated Discovery Net servers at Imperial College London, UK and high performance implementations of finiteelement models executing at specialised servers at the University of Oklahoma, USA. Novel scientific results produced using our data sets provide a valuable insight into earthquake analysis. In addition, our informatics work provides a novel high performance computing framework and methods for the application of complex knowledge discovery methods to understanding earthquake dynamics. Furthermore, the realisation of our distributed computing platform is based on the implementation of a set of open standards, making its results accessible over the Grid to the wider scientific community. Yike Guo, Jian Guo Liu 0005, Moustafa Ghanem, Kyran Mish, Vasa Curcin, Christian Haselwimmer, D. Sotiriou, K. K. Muraleetharan, L. Taylor |
SC | 1 |
| 2005 | Using dragpushing to refine centroid text classifiersabstractWe present a novel algorithm, DragPushing, for automatic text classification. Using a training data set, the algorithm first calculates the prototype vectors, or centroids, for each of the available document classes. Using misclassified examples, it then iteratively refines these centroids; by dragging the centroid of a correct class towards a misclassified example and in the same time pushing the centroid of an incorrect class away from the misclassified example. The algorithm is simple to implement and is computationally very efficient. Evaluation experiments conducted on two benchmark collections show that its classification accuracy is comparable to that of more complex methods, such as support vector machines (SVM). Songbo Tan, Xueqi Cheng 0001, Bin Wang 0004, Moustafa Ghanem, Yike Guo |
SIGIR | 6 |
| 2003 | InfoGrid: providing information integration for knowledge discovery
Nikolaos Giannadakis, Anthony Rowe 0002, Moustafa Ghanem, Yike Guo |
Inf. Sci. | 4 |
| 2002 | Grid-Based Knowledge Discovery Services for High Throughput InformaticsabstractDiscovery Net is an application layer for providing grid-based knowledge discovery services. These services allow scientists to create and manage complex knowledge discovery workflows that integrate data and analysis routines provided as remote services. They also allow scientists to store, share and execute these workflows as well as publish them as new services. Discovery Net provides a higher level of abstraction of the Grid for knowledge discovery activities, thus separating the end-users from resource management issues already handled by existing and emerging standards. Moustafa Ghanem, Yike Guo, Anthony Rowe 0002, Patrick Wendel |
HPDC | 2 |
| 2002 | Discovery net: towards a grid of knowledge discoveryabstractThis paper provides a blueprint for constructing collaborative and distributed knowledge discovery systems within Grid-based computing environments. The need for such systems is driven by the quest for sharing knowledge, information and computing resources within the boundaries of single large distributed organisations or within complex Virtual Organisations (VO) created to tackle specific projects. The proposed architecture is built on top of a resource federation management layer and is composed of a set of different resources. We show how this architecture will behave during a typical KDD process design and deployment, how it enables the execution of complex and distributed data mining tasks with high performance and how it provides a community of e-scientists with means to collaborate, retrieve and reuse both KDD algorithms, discovery processes and knowledge in a visual analytical environment. Vasa Curcin, Moustafa Ghanem, Yike Guo, Anthony Rowe 0002, Jameel Syed, Patrick Wendel |
KDD | 3 |
| 2001 | Design of Problem-Solving Environment for Contingent Claim Valuation
Francis Oliver Bunnin, Yike Guo, John Darlington |
Euro-Par | 2 |
| 2001 | Developing a distributed scalable Java component server
Yike Guo, Patrick Wendel |
Future Gener. Comput. Syst. | 1 |
| 2000 | New paradigms in information visualizationabstractWe present three new visualization front-ends that aid navigation through the set of documents returned by a search engine (hit documents). We cluster the hit documents to visually group these documents and label the groups with related words. The different front-ends cater for different user needs, but all can browse cluster information as well as drilling up or down in one or more clusters and refining the search using one or more of the suggested related keywords. Peter Au, Matthew Carey, Shalini Sewraz, Yike Guo, Stefan M. Rüger |
SIGIR | 4 |
| 2000 | Design of high performance financial modelling environment
Francis Oliver Bunnin, Yike Guo, Yuhe Ren, John Darlington |
Parallel Comput. | 2 |
| 1999 | Probing Knowledge in Distributed Data Mining
Yike Guo, Janjao Sutiwaraphun |
PAKDD | 1 |
| 1999 | Editorial
Yike Guo, Robert L. Grossman |
Data Min. Knowl. Discov. | 1 |
| 1998 | GOFFIN: Higher-Order Functions Meet Concurrent Constraints
Manuel M. T. Chakravarty, Yike Guo, Hendrik C. R. Lock |
Sci. Comput. Program. | 2 |
| 1997 | Parallel Induction Algorithms for Data Mining
John Darlington, Yike Guo, Janjao Sutiwaraphun, Hing Wing To |
IDA | 2 |
| 1997 | The Minimised Geometric Buchberger Algorithm: An Optimal Algebraic Algorithm for Integer ProgrammingabstractIP problems characterize combinatorial optimization problems where conventional numerical methods based on the hill-climbing technique can not be directly applied. Conventional methods for solving integer programming are based on searching algorithms where heuristics such as branch and bound are applied to reduce the search space. Recently, various algebraic IP solvers have been proposed based on the theory of Grobner bases. The key idea is to encode an IP problem IPA,C into a special ideal associated with the constraint matrix A and the cost (object) function C. An important property of such an encoding is that its Grobner basis corresponds directly to the test set of the IP problem. The main difficulty of these new methods is the size of the Grobner bases generated. In the proposed algorithms, large Grobner bases are caused by either introducing additional variables or by considering the generic IP problem IPA,C. Some improvements have been proposed such as the Hosten and Sturmfels method (GRIN) designed to avoid additional variables and the truncated Grobner basis method of Thomas which computes the Grobner basis for a specific IP problem IPA,C(b) (rather than its generalization IPA,C). In this paper we propose a new algebraic algorithm for solving integer programming problems. The new algorithm, called the Minimized Geometric Buchberger Algorithm (MGBA), combines the Hosten and Sturmfels method (GRIN) and Thomas's truncated GBA to compute the fundamental segments of a IP problem IPA,C directly in its original space and also the truncated Grobner basis for a specific IP problem IPA,C(b). We have carried out experiments to compare this algorithm with others such as the geometric Buchberger algorithm, the truncated geometric Buchberger algorithm, and the algorithm in GRIN. These experiments shows that the new algorithm offers significant performance improvement. Yike Guo, Tetsuo Ida, John Darlington |
ISSAC | 2 |
| 1997 | Large Scale Data Mining: Challenges and Responses
Jaturon Chattratichat, John Darlington, Moustafa Ghanem, Yike Guo, Harald Frank Hüning, Janjao Sutiwaraphun, Hing Wing To |
KDD | 4 |
| 1995 | Functional Skeletons for Parallel Coordination
John Darlington, Yike Guo, Hing Wing To, Jin Yang 0002 |
Euro-Par | 2 |
| 1995 | Parallel Skeletons for Structured CompositionabstractIn this paper, we propose a straightforward solution to the problems of compositional parallel programming by using skeletons as the uniform mechanism for structured composition. In our approach parallel programs are constructed by composing procedures in a conventional base language using a set of high-level, pre-defined, functional, parallel computational forms known as skeletons. The ability to compose skeletons provides us with the essential tools for building further and more complex application-oriented skeletons specifying important aspects of parallel computation. Compared with the process network based composition approach, such as PCN, the skeleton approach abstracts away the fine details of connecting communication ports to the higher level mechanism of making data distributions conform, thus avoiding the complexity of using lower level ports as the means of interaction. Thus, the framework provides a natural integration of the compositional programming approach with the data parallel programming paradigm. John Darlington, Yike Guo, Hing Wing To, Jin Yang 0002 |
PPoPP | 2 |
| 1994 | Constraint Logic Programming in the Sequent Calculus
John Darlington, Yike Guo |
LPAR | 2 |
| 1989 | Narrowing and Unification in Functional Programming - An Evaluation Mechanism for Absolute Set Abstraction
John Darlington, Yike Guo |
RTA | 2 |