EDBT 2026 Demo / reviewers in the wild / expert
Chengru Song
dblp:144/1365
· DBLP profile ↗
25ranked-venue papers
0as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 9 since 2021Systems, architecture and hardware · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Nappa: NNA-Compatible and Privacy-Preserving DNN Training Framework via Vector DecompositionabstractHow to preserve the data privacy during the training of deep neural network (DNN) is a key security concern in the artificial intelligence era. However, most existing solutions based on homomorphic encryption and Trusted Execution Environment (TEE) are incompatible with heterogeneous Neural Network Accelerators (NNAs), leading to significant performance loss. We propose a novel method based on vector decomposition to allocate operators across different NNAs, ensuring both throughput and privacy simultaneously. Furthermore, based on this approach, we have designed a compiler that automatically converts front-end model descriptions into backend encrypted computation graphs, which is running securely over trusted and untrusted hardware. This compiler heuristically determines the allocation scheme based on hardware affinity and cross-hardware communication costs, significantly reducing additional overhead. Experimental results demonstrate that our method does not incur extra accuracy costs and achieves a throughput significantly higher than existing methods. Deploying our approach at scale on a platform with a billion users, we have verified its negligible impact on real-world operations while ensuring the privacy protection capability for cross-domain data. Yan Zhang 0002, Qiushi Li 0002, Ju Ren 0001, Yiqiao Liao, Jin Ouyang, Chengru Song, Honghuan Wu, Kaiqiao Zhan, Ben Wang 0006, Xu Chen 0004, Yaoxue Zhang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | Cauchy: A Cost-Efficient LLM Serving System through Adaptive Heterogeneous DeploymentabstractRecent advances in large language models (LLMs) have intensified the need for serving LLMs that are cost-efficient and QoS-guaranteed. Existing frameworks often co-locate computationally distinct prefill and decode instances on homogeneous GPUs, overlooking their unique resource demands and under-utilizing heterogeneous GPUs. This leads to suboptimal resource utilization and increased capital expenditure. We present Cauchy, a LLM serving framework that adaptively deploys prefill and decode computation to the most suitable heterogeneous GPUs and dynamically schedules user requests. At the core of Cauchy is choosing proper GPU Combo, a conceptual GPU combination encompassing diverse GPU configurations, for their cost efficiency in running prefill-decode pairs. Cauchy deploys a set of combos to satisfy QoS requirements (e.g., goodput) of LLM inference. Cauchy further employs hierarchical scheduling to handle user requests, using opportunistic scheduling within the allocated GPU Combos and a goodput-weighted round-robin policy across GPU Combos. Dynamic autoscaling is used to stabilize the cost-efficiency in the face of surging requests. Experiments show that Cauchy achieves up to a 38.3% improvement in Tokens/USD efficiency over the state-of-the-art baselines, while maintaining strict Service Level Objectives (SLOs). Our work highlights the importance of leveraging workload and GPU heterogeneity to achieve superior cost-efficient LLM serving. Renyu Yang, Yuxi Luo, Menghao Zhang 0001, Li Li 0029, Chunming Hu, Tianyu Wo, Chengru Song, Jin Ouyang |
SoCC | 10 |
| 2025 | Cuckoo: Deadline-Aware Job Packing on Heterogeneous GPUs for DL Model TrainingabstractThe growing scale and heterogeneity of GPU clusters pose new challenges to deep learning (DL) job scheduling. While existing schedulers primarily focus on GPU utilization, they often ignore multi-dimensional resource demands of DL workloads and lack precise execution time estimation for co-located jobs. While Muri pioneered the use of interleaving execution to improve resource efficiency, it simplified interference when jobs using one resource simultaneously and is agnostic to the deadline constraints. The job grouping also comes to suboptimal when heterogeneous GPU devices are taken into account. In this paper, we propose Cuckoo, a scheduling system that packs deep learning jobs with stringent deadline requirements over a set of heterogeneous GPU devices where multi-dimensional resources are interleaved and shared by a group of jobs. Specifically, the interleaving execution of simultaneous jobs is characterized and modeled through stage-grained execution time estimation considering the runtime performance interference and the impact of GPU heterogeneity on the job performance. The job packing is formulated as a multi-objective optimization problem which is then solved by the maximum weight matching algorithm. Cuckoo then allocates heterogeneous resources to the packed job groups through a graph-based maximum flow and minimum cut algorithm. Experiments show that Cuckoo improves deadline satisfaction rate by 2.38x and reduces average job completion time (JCT) by 1.81x compared with the state-of-the-art approaches. Cuckoo is implemented based on Kubernetes and has been deployed in Kuaishou to serve thousands of model training jobs that can be interleaved on shared heterogeneous GPU clusters Yuzheng Zhang, Renyu Yang, Weihan Jiang, Tianyu Ye, Yiqiao Liao, Penghao Zhang, Tiezi Zhang, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang |
SoCC | 12 |
| 2025 | Enhancing LLM to Decompile Optimized PTX to Readable CUDA for Tensor ProgramsabstractThe growing demand for high-performance tensor programs on GPUs, especially for large language models (LLMs), necessitates advanced compilation and optimization techniques. However, the critical task of analyzing optimized, low-level PTX code for performance tuning or understanding poses significant challenges. While LLMs hold promise for PTX-to-CUDA de-compilation to improve code intelligibility, their effectiveness is severely limited by the scarcity of aligned training data and the inherent complexity of highly optimized, unrolled PTX code.In this work, we explore methodologies to significantly enhance LLM capabilities for accurate and readable PTX-to-CUDA decompilation and present PtxDec, a decompilation prototype implementing our approach. To overcome the critical barrier of data scarcity, we develop a compiler-based data augmentation framework coupled with rigorous post-processing, enabling the creation of a large-scale, high-quality dataset of 400K aligned CUDA-PTX kernel pairs for effective LLM training. Furthermore, to empower LLMs to handle the complexity of optimized PTX, we introduce Rolled-PTX—an intermediate representation generated through heuristic loop rerolling during preprocessing. Rolled-PTX condenses unrolled patterns, drastically simplifying the input structure presented to the LLM and aligning it better with higher-level loop constructs.Comprehensive evaluation demonstrates that PtxDec achieves substantial performance gains: our approach yields a 2.3×–3.1× improvement in functional accuracy over baseline methods, alongside significant enhancements in generated code readability and scheduling consistency with the original optimized kernels. Ablation studies further validate the contribution of each proposed component to the overall performance.To the best of our knowledge, this is the first work tackling PTX-to-CUDA decompilation, specifically focusing on and demonstrating effective strategies for augmenting LLMs to overcome the key challenges in this domain. Fugen Tang, Yu Zhang 0086, Chengru Song |
ASE | 5 |
| 2025 | KAIOPS: A Platform Solution of End-to-End Multi-Modal AIOps for AI Training at ScaleabstractThe resilience of large-scale AI training platforms are fundamental to enabling contemporary AI innovation and business development. However, with the rapid increase in the scale and complexity of AI model training tasks, anomalies become the norm rather than the exception at scale. Failing to handle them properly may lead to enormous resource waste and prolonged development cycles. Traditional anomaly detection methods struggle to tackle the complex temporal characteristics and extreme class imbalance inherently manifesting in training tasks, and fall short in automated solution to root cause analysis and the follow-up remediation. This paper proposes KAIOPS, an end-to-end automated platform solution for handling anomalies and engineering experience of daily operational maintenance for large-scale AI training clusters at Kuaishou. KAIOPS employs a Temporal Context Encoding mechanism to precisely capture and encode long-term trends and critical temporal context information within fault evolution. The detection model elaborates a dynamic class-weighted loss function for enhancing the detection performance. To deliver a complete end-to-end intelligent processing pipeline, KAIOPS further leverages knowledge graph and LLMs for automated root cause analysis and actionable solution generation. Extensive experiments, on the basis of data collected from Kuaishou’s production-grade training clusters, show the superior performance of our proposed approach. KAIOPS has been deployed in Kuaishou, in both testbed and production grade environments, consisting of with over 10,000 GPUs, and accelerate the reliability assurance for industry-scale model training and serving. Zeying Wang, Penghao Zhang, Xu Wang 0007, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang, Renyu Yang |
ASE | 8 |
| 2025 | Kair: A Statistical and Causal Approach to Pinpointing Stragglers in Distributed Model TrainingabstractThe distributed deep learning training process within large-scale clusters serves as the foundation of contemporary artificial intelligence. However, its inherent characteristics make it particularly sensitive to stragglers, specifically the presence of slow workers, which can significantly decelerate the entire procedure. Observability tools are essential for identifying stragglers within systems. However, the prevailing system profiling tools are either designed for single-node analysis, lacking visibility across multiple workers, or they recognize stragglers but only deliver high-level symptoms, providing engineers with insufficient insight into the underlying causes.We design Kair, a robust production-standard observability tool. Kair uses an innovative hierarchical approach, transitioning from statistical anomaly detection to causal inference. It employs Kolmogorov-Smirnov statistics for the identification of statistically anomalous workers and implements a causal path tracing algorithm to accurately determine the specific operations, such as computation or communication, that are responsible for the delay. Kair has been evaluated in a production cluster of 2,048 NVIDIA A800 GPUs and demonstrated high effectiveness in detecting latent stragglers at the framework level that are often overlooked by conventional tools. It offers precise suggestions that markedly reduce processing inefficiencies and engineering workload. Yitang Yang, Jiapeng Chen, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang, Renyu Yang |
ASE | 7 |
| 2025 | SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM TrainingabstractPipeline Parallelism serves as a crucial technique for training Large Language Models, as it alleviates memory pressure from model states with relatively low communication overhead. However, in long-context scenarios, existing pipeline parallelism methods fail to address the substantial activation memory pressure, primarily due to the peak memory consumption resulting from the accumulation of activations across multiple microbatches. Moreover, these approaches inevitably introduce considerable pipeline bubbles, further hindering efficiency. Zhouyang Li, Tailing Yuan, Chengru Song |
SC | 6 |
| 2024 | Kale: Elastic GPU Scheduling for Online DL Model TrainingabstractLarge-scale GPU clusters have been widely used for effectively training both online and offline deep learning (DL) jobs. However, elastic scheduling in most cases of resource schedulers is dedicated for offline model training where resource adjustment is planned ahead of time. The native autoscaling policy is on the basis of pre-defined threshold and, if applied directly in online model training, often suffers from belated resource adjustment, leading to diminished model accuracy. In this paper, we present Kale, a novel elastic GPU scheduling system to improve the performance of online DL model training. Through traffic forecasting and resource-throughput modeling, Kale automatically pinpoints the number of required GPUs that best accommodate the on-the-fly data samples before performing stabilized autoscaling. An advanced data shuffling strategy is further employed for balancing uneven samples among different training workers, thereby improving the runtime efficacy. Experiments show that Kale substantially outperforms the state-of-the-art solutions. Compared with the default HPA autoscaling strategy, Kale reduces the accumulated lag and downtime by 69.2% and 33.1%, respectively, whilst lowering the SLO violation rate from 19.57% to just 2.6%. Kale has been deployed at Kuaishou's production-level GPU clusters and successfully underpins real-time video recommendation and advertisement at scale. Renyu Yang, Jin Ouyang, Weihan Jiang, Tianyu Ye, Menghao Zhang 0001, Sui Huang, Chengru Song, Di Zhang 0026, Tianyu Wo, Chunming Hu |
SoCC | 9 |
| 2024 | Decoding at the Speed of Thought: Harnessing Parallel Decoding of Lexical Units for LLMsabstractLarge language models have demonstrated exceptional capability in natural language understanding and generation. However, their generation speed is limited by the inherently sequential nature of their decoding process, posing challenges for real-time applications. This paper introduces Lexical Unit Decoding (LUD), a novel decoding methodology implemented in a data-driven manner, accelerating the decoding process without sacrificing output quality. The core of our approach is the observation that a pre-trained language model can confidently predict multiple contiguous tokens, forming the basis for a lexical unit, in which these contiguous tokens could be decoded in parallel. Extensive experiments validate that our method substantially reduces decoding time while maintaining generation quality, i.e., 33% speed up on natural language generation with no quality loss, and 30% speed up on code generation with a negligible quality loss of 3%. Distinctively, LUD requires no auxiliary models and does not require changes to existing architectures. It can also be integrated with other decoding acceleration methods, thus achieving an even more pronounced inference efficiency boost. We posit that the foundational principles of LUD could define a new decoding paradigm for future language models, enhancing their applicability for a broader spectrum of applications. All codes are be publicly available at https://github.com/tjunlp-lab/Lexical-Unit-Decoding-LUD-. Zijia Lin, Zhongyuan Wang 0006, Chengru Song, Di Zhang 0026, Kun Gai, Deyi Xiong |
LREC/COLING | 8 |
| 2024 | Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual TokenizationabstractRecently, the remarkable advance of the Large Language Model (LLM) has inspired researchers to transfer its extraordinary reasoning capability to both vision and language data. However, the prevailing approaches primarily regard the visual input as a prompt and focus exclusively on optimizing the text generation process conditioned upon vision content by a frozen LLM. Such an inequitable treatment of vision and language heavily constrains the model's potential. In this paper, we break through this limitation by representing both vision and language in a unified form. Specifically, we introduce a well-designed visual tokenizer to translate the non-linguistic image into a sequence of discrete tokens like a foreign language that LLM can read. The resulting visual tokens encompass high-level semantics worthy of a word and also support dynamic sequence length varying from the image. Coped with this tokenizer, the presented foundation model called LaVIT can handle both image and text indiscriminately under the same generative learning paradigm. This unification empowers LaVIT to serve as an impressive generalist interface to understand and generate multi-modal content simultaneously. Extensive experiments further showcase that it outperforms the existing models by a large margin on massive vision-language tasks. Our code and models are available at https://github.com/jy0205/LaVIT. Kun Xu 0005, Chao Liao, Jianchao Tan, Quzhe Huang, Chengru Song, Dai Meng, Di Zhang 0026, Wenwu Ou, Kun Gai, Yadong Mu |
ICLR | 8 |
| 2024 | Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationabstractIn light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for effective large-scale pre-training due to the modeling of its spatiotemporal dynamics. In this paper, we address such limitations in video-language pre-training with an efficient video decomposition that represents each video as keyframes and temporal motions. These are then adapted to an LLM using well-designed tokenizers that discretize visual and temporal information as a few tokens, thus enabling unified generative pre-training of videos, images, and text. At inference, the generated tokens from the LLM are carefully recovered to the original continuous pixel space to create various video content. Our proposed framework is both capable of comprehending and generating image and video content, as demonstrated by its competitive performance across 13 multimodal benchmarks in image and video understanding and generation. Our code and models are available at https://video-lavit.github.io. Zhicheng Sun 0001, Kun Xu 0005, Hao Jiang 0032, Quzhe Huang, Chengru Song, Di Zhang 0026, Yang Song 0008, Kun Gai, Yadong Mu |
ICML | 8 |
| 2024 | Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid Parallelism
Tailing Yuan, Xucheng Ye, Shenglong Zhang, Jianchao Tan, Chengru Song |
USENIX ATC | 7 |
| 2024 | PHD-NAS: Preserving helpful data to promote Neural Architecture Search
Shun Lu 0001, Yu Hu 0001, Longxing Yang, Jilin Mei, Jianchao Tan, Chengru Song |
Neurocomputing | 7 |
| 2023 | SHARK: A Lightweight Model Compression Approach for Large-scale Recommender SystemsabstractIncreasing the size of embedding layers has shown to be effective in improving the performance of recommendation models, yet gradually causing their sizes to exceed terabytes in industrial recommender systems, and hence the increase of computing and storage costs. To save resources while maintaining model performances, we propose SHARK, the model compression practice we have summarized in the recommender system of industrial scenarios. SHARK consists of two main components. First, we use the novel first-order component of Taylor expansion as importance scores to prune the number of embedding tables (feature fields). Second, we introduce a new row-wise quantization method to apply different quantization strategies to each embedding. We conduct extensive experiments on both public and industrial datasets, demonstrating that each component of our proposed SHARK framework outperforms previous approaches. We conduct A/B tests in multiple models on Kuaishou, such as short video, e-commerce, and advertising recommendation models. The results of the online A/B test showed SHARK can effectively reduce the memory footprint of the embedded layer. For the short-video scenarios, the compressed model without any performance drop significantly saves 70% storage and thousands of machines, improves 30% queries per second (QPS), and has been deployed to serve hundreds of millions of users and process tens of billions of requests every day. Beichuan Zhang 0002, Chenggen Sun, Jianchao Tan, Xinjun Cai, Mengqi Miao, Chengru Song, Na Mou, Yang Song 0008 |
CIKM | 8 |
| 2023 | PA&DA: Jointly Sampling PAth and DAta for Consistent NASabstractBased on the weight-sharing mechanism, one-shot NAS methods train a supernet and then inherit the pre-trained weights to evaluate sub-models, largely reducing the search cost. However, several works have pointed out that the shared weights suffer from different gradient descent directions during training. And we further find that large gradient variance occurs during supernet training, which degrades the supernet ranking consistency. To mitigate this issue, we propose to explicitly minimize the gradient variance of the supernet training by jointly optimizing the sampling distributions of PAth and DAta (PA&DA). We theoretically derive the relationship between the gradient variance and the sampling distributions, and reveal that the optimal sampling probability is proportional to the normalized gradient norm of path and training data. Hence, we use the normalized gradient norm as the importance indicator for path and training data, and adopt an importance sampling strategy for the supernet training. Our method only requires negligible computation cost for optimizing the sampling distributions of path and data, but achieves lower gradient variance during supernet training and better generalization performance for the supernet, resulting in a more consistent NAS. We conduct comprehensive comparisons with other improved approaches in various search spaces. Results show that our method surpasses others with more reliable ranking performance and higher accuracy of searched architectures, showing the effectiveness of our method. Code is available at https://github.com/ShunLu91/PA-DA. Shun Lu 0001, Yu Hu 0001, Longxing Yang, Jilin Mei, Jianchao Tan, Chengru Song |
CVPR | 7 |
| 2023 | Privacy-Preserving DNN Training with Prefetched Meta-Keys on Heterogeneous Neural Network AcceleratorsabstractThe embedded software may migrate the collected data to the server for DNN computation acceleration, which may compromise privacy. We propose a DNN computation framework that combines TEE and NNA to address the privacy leakage problem. We design an NNA-friendly encryption method that enables NNA to correctly compute the encrypted linear input. Facing the overhead of TEE-NNA interaction, we design a pipeline-based prefetch mechanism that can reduce the TEE interaction overhead. Experimentally, our approach proves to be compatible with a wide range of NPUs and TPUs, and improves the performance by 8-19 times over the TEE scheme. Qiushi Li 0002, Ju Ren 0001, Yan Zhang 0104, Chengru Song, Yiqiao Liao, Yaoxue Zhang |
DAC | 4 |
| 2023 | Dynamic TF-TDNN: Dynamic Time Delay Neural Network Based on Temporal-Frequency Attention for Dialect RecognitionabstractDialect recognition aims to recognize dialect categories in utterances, which has been applied in many audio applications. Recently, various Time Delayed Neural Network (TDNN) based AI models are proposed to solve dialect recognition problems, such as D-TDNN, DMC-TDNN, and ECAPA-TDNN, however, most of them only perform temporal attention in the last statistical pooling layer of the TDNN network, which ignores the importance of simultaneously capturing both frequency and temporal key information in utterances under different receptive fields. In contrast, we introduce a hybrid attention mechanism in both the temporal and frequency domain, called the TF-attention module, which adaptively pays more attention to the indeed important frames and the frame-level important information under different receptive fields for dialect recognition. Moreover, we are the first to introduce a dynamic architecture mechanism in the field of dialect recognition to dynamically reduce the computational cost and the number of parameters of models. We evaluate the proposed dynamic TF-TDNN on the OLR challenge AP20-OLR-dialect task and achieve State-Of-The-Art (SOTA) performance with fewer model parameters. Chao Liao, Jinwen Huang, Huan Yuan, Jianchao Tan, Feng Deng, Chengru Song |
ICASSP | 9 |
| 2023 | NAS-DYMC: NAS-Based Dynamic Multi-Scale Convolutional Neural Network for Sound Event DetectionabstractCNN+RNN models have become the mainstream approach for semi-supervised sound event detection, and the CNN part is mainly a stack of several 2D convolutional layers to capture the representations of the time-frequency features. However, conventional 2D convolution is of limited ability in capturing detailed information about acoustic events. In this paper, to enhance the representation ability of CNN, we propose NAS-DYMC, a NAS-based dynamic multi-scale convolutional neural network to extract a more effective acoustic representation. Specifically, multi-scale convolution can capture the characteristics of sound events with different time-frequency distributions and dynamic convolution enhances the representation capability of conventional convolution by adapting attention weights onto basis kernels. Furthermore, a neural architecture search (NAS) method is adopted to find the optimal network architecture from the search space consisting of various dynamic multi-scale convolutions for the DCASE 2021 Task4 dataset. Experimental results demonstrate the superiority of our proposed method. Feng Deng, Jianchao Tan, Chengru Song |
ICASSP | 5 |
| 2023 | MaskFusion: Feature Augmentation for Click-Through Rate Prediction via Input-adaptive Mask Fusion
Chao Liao, Jianchao Tan, Jiyuan Jia, Chengru Song |
ICLR | 5 |
| 2023 | Resource Constrained Model Compression via Minimax Optimization for Spiking Neural NetworksabstractBrain-inspired Spiking Neural Networks (SNNs) have the characteristics of event-driven and high energy-efficient, which are different from traditional Artificial Neural Networks (ANNs) when deployed on edge devices such as neuromorphic chips. Most previous work focuses on SNNs training strategies to improve model performance and brings larger and deeper network architectures. It's difficult to deploy these complex networks on resource-limited edge devices directly. To meet such demand, people compress SNNs very cautiously to balance the performance and the computation efficiency. Existing compression methods either iteratively pruned SNNs using weights norm magnitude or formulated the problem as a sparse learning optimization. We propose an improved end-to-end Minimax optimization method for this sparse learning problem to better balance the model performance and the computation efficiency. We also demonstrate that jointly applying compression and finetuning on SNNs is better than sequentially, especially for extreme compression ratios. The compressed SNN models achieved state-of-the-art (SOTA) performance on various benchmark datasets and architectures. Our code is available athttps://github.com/chenjallen/Resource-Constrained-Compression-on-SNN . Huan Yuan, Jianchao Tan, Chengru Song, Di Zhang 0026 |
ACM Multimedia | 5 |
| 2022 | Conformer Space Neural Architecture Search for Multi-Task Audio Separation
Shun Lu 0001, Chenxing Li, Jianchao Tan, Feng Deng, Chengru Song |
INTERSPEECH | 8 |
| 2022 | WA-Transformer: Window Attention-based Transformer with Two-stage Strategy for Multi-task Audio Source Separation
Chenxing Li, Feng Deng, Shun Lu 0001, Jianchao Tan, Chengru Song |
INTERSPEECH | 7 |
| 2022 | FogChain: A Blockchain-Based Peer-to-Peer Solar Power Trading System Powered by Fog AIabstractMicrogrids, gaining traction from rising distributed generation for carbon reduction, demand novel solutions to regulate on- and off-grid operations, as well as both energy and monetary transfers between the microgrid and the central grid and among different microgrid participants. This research aims to develop and validate an intelligent microgrid management system to secure the competitiveness of Singapore’s energy market, by leveraging the inherent synergy between two emerging technologies, i.e., blockchain for Peer-to-Peer (P2P) solar power trading and fog computing for grid infrastructure management. For this vision, we have developed FogChain, an integrative, cost-effective, and scalable microgrid operating system (MGOS), consisting of three technical service layers: 1) a novel microgrid information infrastructure based on the fog-computing paradigm (i.e., intelligence on edge); 2) a blockchain-based microgrid service layer, providing smart contract and decentralized control capabilities for grid application development; and 3) a microgrid application layer (i.e., P2P energy trading) over the blockchain-based grid service. This MGOS would fundamentally transform how solar power is traded among participating electricity prosumers, leading to potentially new operational and business models. We have implemented the FogChain system and conducted extensive experiments to verify its performance advantages. Our results demonstrate that FogChain can efficiently process energy auction among 1000 participants with 1.1 s delay on average, reduce transmission cost up to 20% under the loss-aware trading mechanism, and reduce the solar yield prediction error to 0.11. Our system prototype suggests that FogChain provides a promising solution for efficient decentralized energy trading and intelligent distributed control for microgrids. Guanyu Gao, Chengru Song, T. G. Thusitha Asela Bandara, Meng Shen 0002, Fan Yang 0172, Wolf Posdorfer, Dacheng Tao, Yonggang Wen 0001 |
IEEE Internet Things J. | 2 |
| 2018 | Real-Time Bidding with Multi-Agent Reinforcement Learning in Display AdvertisingabstractReal-time advertising allows advertisers to bid for each impression for a visiting user. To optimize specific goals such as maximizing revenue and return on investment (ROI) led by ad placements, advertisers not only need to estimate the relevance between the ads and user's interests, but most importantly require a strategic response with respect to other advertisers bidding in the market. In this paper, we formulate bidding optimization with multi-agent reinforcement learning. To deal with a large number of advertisers, we propose a clustering method and assign each cluster with a strategic bidding agent. A practical Distributed Coordinated Multi-Agent Bidding (DCMAB) has been proposed and implemented to balance the tradeoff between the competition and cooperation among advertisers. The empirical study on our industry-scaled real-world data has demonstrated the effectiveness of our methods. Our results show cluster-based bidding would largely outperform single-agent and bandit approaches, and the coordinated bidding achieves better overall objectives than purely self-interested bidding agents. Junqi Jin, Chengru Song, Han Li 0005, Kun Gai, Jun Wang 0012, Weinan Zhang 0001 |
CIKM | 2 |
| 2018 | Deep Interest Network for Click-Through Rate PredictionabstractClick-through rate prediction is an essential task in industrial applications, such as online advertising. Recently deep learning based models have been proposed, which follow a similar Embedding&MLP paradigm. In these methods large scale sparse input features are first mapped into low dimensional embedding vectors, and then transformed into fixed-length vectors in a group-wise manner, finally concatenated together to fed into a multilayer perceptron (MLP) to learn the nonlinear relations among features. In this way, user features are compressed into a fixed-length representation vector, in regardless of what candidate ads are. The use of fixed-length vector will be a bottleneck, which brings difficulty for Embedding&MLP methods to capture user's diverse interests effectively from rich historical behaviors. In this paper, we propose a novel model: Deep Interest Network (DIN) which tackles this challenge by designing a local activation unit to adaptively learn the representation of user interests from historical behaviors with respect to a certain ad. This representation vector varies over different ads, improving the expressive ability of model greatly. Besides, we develop two techniques: mini-batch aware regularization and data adaptive activation function which can help training industrial deep networks with hundreds of millions of parameters. Experiments on two public datasets as well as an Alibaba real production dataset with over 2 billion samples demonstrate the effectiveness of proposed approaches, which achieve superior performance compared with state-of-the-art methods. DIN now has been successfully deployed in the online display advertising system in Alibaba, serving the main traffic. Guorui Zhou, Xiaoqiang Zhu, Chengru Song, Han Zhu 0001, Xiao Ma 0028, Yanghui Yan, Junqi Jin, Han Li 0005, Kun Gai |
KDD | 3 |