EDBT 2026 Demo / reviewers in the wild / expert
Sidi Lu
dblp:206/6156
· DBLP profile ↗
28ranked-venue papers
11as first author
20since 2021 · last 2025
0000-0001-9846-7570ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Context-Aware Perception: VLM-Augmented All-Weather Detection for Autonomous DrivingabstractAdverse weather, such as rain, fog, and snow, remains a major challenge for autonomous vehicles (AVs), degrading sensor reliability and compromising safety. To address this problem, we propose a novel perception pipeline augmented by a powerful vision-language model (VLM) to strengthen detection robustness and contextual understanding across diverse weather scenarios. The pipeline incorporates QwenVL to provide automatic weather labeling, semantically guided data augmentation, and adaptive sensor prioritization through weather-aware reasoning. We evaluate the approach on the BDD100K dataset using YOLOv10-M and RF-DETR, two strong detectors under adverse weather in our baseline comparisons. With VLM integration, both models achieve 8–12% gains in mean Average Precision in rainy and foggy conditions while incurring minimal additional latency. These results indicate that VLM-augmented perception can improve decision reliability and model explainability in autonomous driving. The findings underscore the value of context-aware, interpretable, and computationally efficient perception frameworks for achieving reliable all-weather autonomy. Johora Akter Polin, Sidi Lu |
SEC | 3 |
| 2025 | Training-Free Late Fusion across Geometry and BEV for Edge-Deployable LiDAR-Camera 3D PerceptionabstractAutonomous driving depends on accurate and reliable 3D perception, and LiDAR-camera fusion is central to that goal. However, most multisensor fusion pipelines limit portability and hinder real-world deployment. To understand the strengths of dominant LiDAR-camera fusion paradigms compared to single-sensor perception, we present a unified benchmark comparing three representative 3D perception pipelines: CenterPoint (LiDAR only), a geometry-level fusion model (MVP) that injects image cues into point space, and a bird's-eye view (BEV)-level fusion model (BEV-Fusion) that aggregates multimodal features in BEV. We further propose a training-free late fusion module that applies consensus and de-noising across the strongest models' predictions to improve detection quality across operating conditions. Our results show that (i) geometry-level and BEV-level fusion offer complementary strengths rather than a single winner: BEVFusion achieves the highest overall detection accuracy, while MVP provides more precise spatial localization; (ii) a late fusion stage can combine the advantages of both paradigms and is suitable for real-time deployment on in-vehicle edge hardware, without requiring retraining. Sidi Lu |
SEC | 2 |
| 2025 | EAA: Emotion-Aware Audio Large Language Models with Dual Cross-Attention and Context-Aware Instruction Tuning
Sidi Lu, Gang Zhou 0002, Ye Gao 0001 |
INTERSPEECH | 2 |
| 2025 | Mitigating In-Transit Vision Noise for Enhanced Vehicle SafetyabstractSoftware-defined vehicles (SDVs) rely on cameras for intelligent and safety-critical applications but face challenges from dynamic environmental noise, including weather and occlusions. Unlike static sensors, SDV cameras encounter noise patterns influenced by driving speed, a factor often overlooked in prior research. To address this gap, we conduct a quantitative analysis of the in-transit noise impact using data from public datasets, the CARLA simulator, a robotic vehicle, and a real vehicle. Our findings suggest that maintaining a speed below 40km/h may serve as a threshold for ensuring reliable camera-based applications under noisy urban conditions. In addition, we propose TransitNet, a novel model designed to mitigate in-transit camera noise and enhance driving safety, particularly at higher speeds. Compared to multiple baselines, experimental results show that TransitNet improves the F-measure by 5.1%, mAP@50 by 3.6%, and increases FPS by 56.7% across all datasets. We also provide detailed observations and insights from extensive testing. Junzhou Chen 0002, Sidi Lu |
SenSys | 4 |
| 2025 | Incremental Over-the-Air Update for Software-Defined Vehicles with C-V2XabstractSoftware-defined vehicles (SDVs) rely on intricate software systems distributed across multiple electronic control units (ECUs), making errors inevitable and costly to address via conventional recalls. Software over-the-air (OTA) update offers a real-time, cost-efficient solution while enabling software customization. This paper presents a comprehensive OTA framework for SDVs, utilizing a robotic vehicle and an industry-grade roadside unit (RSU). The framework leverages Docker to create isolated environments and Docker Compose to manage rolling updates, minimizing disruptions and enabling rollbacks. We also propose RSU- and vehicle-side mechanisms for incremental update and evaluate the performance of Cellular Vehicle-to-Everything (C-V2X), including V2X and 5G, alongside Wi-Fi for OTA transmission. Our study of model weight (ResNet) and software package (YOLO, UFAST) updates shows that C-V2X achieves over$30 \times$faster transmission than Wi-Fi. Using 5G within C-V2X further reduces transmission by up to$2339 \times$compared to V2X, establishing it as a highly effective OTA solution. Incremental updates enhance efficiency by cutting delays by$2 \times$versus full updates, with a 32 KB chunk size yielding the highest success rate and shortest times across distances. Incremental update also reduces transmission by up to$5000 \times$using C-V2X, enabling package downloads even while the SDV is in motion. Sidi Lu |
VTC2025-Spring | 2 |
| 2025 | Toward Real-Time and Efficient Perception Workflows in Software-Defined VehiclesabstractWith the growing demand for software-defined vehicles (SDVs), deep learning-based perception models have become increasingly important in intelligent transportation systems. However, these models face significant challenges in enabling real-time and efficient SDV solutions due to their substantial computational requirements, which are often unavailable in resource-constrained vehicles. As a result, these models typically suffer from low throughput, high latency, and excessive GPU/memory usage, making them impractical for real-time SDV applications. To address these challenges, our research focuses on optimizing model and workflow performance through the integration of pruning and quantization techniques across various computational environments, utilizing frameworks, such as PyTorch, open neural network exchange (ONNX), ONNX Runtime, and TensorRT. We systematically explore and evaluate three distinct pruning methods in combination with multiprecision quantization workflows (FP32, FP16, and INT8) and present the results based on four evaluation metrics: 1) inference throughput; 2) latency; 3) GPU/memory usage; and 4) accuracy. Our designed techniques, including pruning and quantization, along with optimized workflows, can achieve up to$18\times $faster inference speed and$16.5\times $higher throughput, while reducing GPU/memory usage by up to 30%, all with minimal impact on accuracy. Our work suggests using the Torch-ONNX-TensorRT workflow quantized with 16-bit floating point precision (FP16) precision and group pruning as the optimal strategy for maximizing inference performance. It demonstrates great potential in optimizing real-time, efficient perception workflows in SDVs, contributing to the enhanced application of deep learning models in resource-constrained environments. Sumaiya, Reza Jafarpourmarzouni, Sidi Lu, Zheng Dong 0002 |
IEEE Internet Things J. | 4 |
| 2024 | DiNADO: Norm-Disentangled Neurally-Decomposed Oracles for Controlling Language ModelsabstractNeurAlly-Decomposed Oracle (NADO) is a powerful approach for controllable generation with large language models. It is designed to avoid catastrophic forgetting while achieving guaranteed convergence to an entropy-maximized closed-form optimal solution with reasonable modeling capacity. Despite the success, several challenges arise when apply NADO to a wide range of scenarios. Vanilla NADO suffers from gradient vanishing for low-probability control signals and is highly reliant on a regularization to satisfy the stochastic version of Bellman equation. In addition, the vanilla implementation of NADO introduces a few additional transformer layers, suffering from a limited capacity especially compared to other finetune-based model adaptation methods like LoRA. In this paper, we propose a improved version of the NADO algorithm, namely DiNADO (norm-Disentangled NeurAlly-Decomposed Oracles), which improves the performance of the NADO algorithm through disentangling the step-wise global norm over the approximated oracle $R$-value for all potential next-tokens, allowing DiNADO to be combined with finetuning methods like LoRA. We discuss in depth how DiNADO achieves better capacity, stability and flexibility with both empirical and theoretical results. Experiments on formality control in machine translation and the lexically constrained generation task CommonGen demonstrates the significance of the improvements. Sidi Lu, Wenbo Zhao 0006, Chenyang Tao, Arpit Gupta, Shanchan Wu, Tagyoung Chung, Nanyun Peng 0001 |
ICML | 1 |
| 2024 | Open-Domain Text Evaluation via Contrastive Distribution MethodsabstractRecent advancements in open-domain text generation, driven by the power of large pre-trained language models (LLMs), have demonstrated remarkable performance. However, assessing these models’ generation quality remains a challenge. In this paper, we introduce a novel method for evaluating open-domain text generation called Contrastive Distribution Methods (CDM). Leveraging the connection between increasing model parameters and enhanced LLM performance, CDM creates a mapping from the contrast of two probabilistic distributions – one known to be superior to the other – to quality measures. We investigate CDM for open-domain text generation evaluation under two paradigms: 1) Generative CDM, which harnesses the contrast of two language models’ distributions to generate synthetic examples for training discriminator-based metrics; 2) Discriminative CDM, which directly uses distribution disparities between two language models for evaluation. Our experiments on coherence evaluation for multi-turn dialogue and commonsense evaluation for controllable generation demonstrate CDM’s superior correlate with human judgment than existing automatic evaluation metrics, highlighting the strong performance and generalizability of our approach. Sidi Lu, Asli Celikyilmaz, Nanyun Peng 0001 |
ICML | 1 |
| 2024 | Edge-Aware Dual Branch Network for Nucleus Instance SegmentationabstractIn mobile healthcare and remote diagnosis, nucleus segmentation is a critical step for pathological analysis, diagnosis, and classification, requiring real-time processing and high accuracy. However, variations in nucleus size, blurred contours, uneven staining, cell clustering, and overlapping cells hinder precise segmentation. Additionally, existing deep learning models often prioritize accuracy at the cost of increased complexity, making them unsuitable for resource-limited edge devices and real-world deployment. To address the aforementioned issues, we propose an edge-aware dual branch network for nucleus instance segmentation. The network simultaneously predicts target information and target contours. Within the network, we propose a context fusion block (CF-block) that effectively extracts and merges contextual information from the network. Additionally, we introduce a post-processing method that combines the target information and target contours to distinguish overlapping nuclei and generate an instance segmentation image. Extensive quantitative evaluations are conducted to assess the performance of our method. Experimental results demonstrate the superior performance of the proposed method compared to state-of-the-art approaches on the BNS, MoNuSeg, and CPM-17 datasets. Junzhou Chen 0002, Yanfu Zhang, Sidi Lu |
SEC | 3 |
| 2024 | An Efficient Data Transmission Framework for Connected VehiclesabstractConnected vehicles (CVs) face significant challenges in continuous big data transmission, resulting in high transmission bandwidth costs and impacting real-time decision-making. To address this, we propose two dynamic, driving-aware compression mechanisms based on reinforcement learning and temporal compressive sensing to intelligently compress video data. These mechanisms adapt to driving conditions, reducing bandwidth while preserving sufficient information for accurate applications such as object detection and ensuring high-quality reconstruction when needed. We also implement a Vehicle-EdgeServer-Cloud (VEC) closed-loop framework that integrates these mechanisms. Specifically, a lightweight vehicle model performs real-time detection on compressed data (measurements), while the EdgeServer receives measurements and reconstructs scenes if needed. The measurements, reconstructed video, and analysis results are then sent to the cloud for vehicle model updates. Unlike conventional methods, our framework seamlessly adapts across vehicles, Edge-Servers, and the cloud, supporting efficient data transmission and dynamic model updates. Extensive evaluations were conducted on our designed roadside unit platform and robotic vehicle, both equipped with industry-grade sensors and computing units. The results demonstrate an 18x reduction in bandwidth at 320KB/s while maintaining high detection accuracy and reconstruction quality compared to non-adaptive measurements, highlighting the framework's promising real-world applications for CVs. Yongtao Yao, Junzhou Chen 0002, Sidi Lu, Weisong Shi |
SEC | 4 |
| 2024 | Enabling Accurate and Timely Prognostics for Aircraft Turbofan EnginesabstractThe aviation industry relies on scheduled maintenance performed on aircraft engines, which ensures safety but incurs significant costs during routine inspections. Traditional preventative maintenance may introduce issues during inspections, leading to delays and cancellations. Predictive maintenance, powered by edge computing, offers a more efficient solution to predict engine failures, optimize schedules, and enhance safety. This paper explores the development and evaluation of a predictive model based on the Gradient-Boosting Regression Tree (GBRT) algorithm for turbofan engine prognostics. Our study uses a synthetic dataset to evaluate the performance of the proposed model through various external conditions and internal configurations. Through this analysis, we compare our model's performance to existing solutions and propose a new benchmarking metric, Margin-Adjusted Reliability Score (MARS), to better assess the applicability and effectiveness of predictive maintenance models in real-world scenarios. Philippa Scroggins, Sidi Lu |
SEC | 2 |
| 2023 | Poster: Edge-Assisted Over-the-Air Software UpdatesabstractThe exploration of software Over-the-Air (OTA) updates for automotive applications is currently very limited. Our work introduces an edge-assisted framework for automotive OTA updates that carefully accounts for various factors, including different software models in vehicles, communication distances, and cluster sizes. We present valuable insights using key evaluation metrics like update speed, data transmission efficiency, and success rate, accompanied by a thorough scalability analysis. Our research involves three distinct vehicle software models: ResNet-18 (46.8 MB), ResNet-50 (102.5 MB), and Faster R-CNN (175.2 MB). These models are used to evaluate update performance across eight distance categories ranging from 0 to 21 meters with a 3-meter interval. We also utilize diverse computing platforms to assess the success rate and conduct a comprehensive scalability analysis. This innovative approach significantly advances our understanding and practical implementation of OTA updates in the automotive field. Arpan Bhattacharjee, Hamza Mahmood, Sidi Lu, Nejib Ammar, Akila Ganlath, Weisong Shi |
SEC | 3 |
| 2023 | Reinforcement Learning for Adaptive Video Compressive SensingabstractWe apply reinforcement learning to video compressive sensing to adapt the compression ratio. Specifically, video snapshot compressive imaging (SCI), which captures high-speed video using a low-speed camera is considered in this work, in which multiple ( B ) video frames can be reconstructed from a snapshot measurement. One research gap in previous studies is how to adapt B in the video SCI system for different scenes. In this article, we fill this gap utilizing reinforcement learning (RL). An RL model, as well as various convolutional neural networks for reconstruction, are learned to achieve adaptive sensing of video SCI systems. Furthermore, the performance of an object detection network using directly the video SCI measurements without reconstruction is also used to perform RL-based adaptive video compressive sensing. Our proposed adaptive SCI method can thus be implemented in low cost and real time. Our work takes the technology one step further towards real applications of video SCI. Sidi Lu, Xin Yuan 0002, Aggelos K. Katsaggelos, Weisong Shi |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2022 | Poster: Towards Efficient Multilayer Collaboration for CAV ApplicationsabstractConnected and autonomous vehicles (CAV s) are facing increasing amounts of data and more complex data analysis, which creates challenges for them to make reliable decisions in real-time. To enable time-sensitive CAV applications, we design and implement a vehicle-edge-cloud framework that integrates compressed imaging (CI) and edge computing into CAV systems. Specifically, a lightweight model is used on the vehicle to perform real-time detection based on optical domain compressed data (called measurements). The edge is responsible for receiving the measurements and performing video reconstruction to support (more accurate) analysis based on the reconstructed video with a trigger. At the same time, the measurements, reconstructed videos, and analysis results are sent to the cloud to continuously update the vehicle model. In addition, we apply reinforcement learning to adapt the compression rate in different driving scenarios. The proposed framework is fully evaluated using our designed roadside platform and outdoor delivery vehicles. Sidi Lu, Weisong Shi |
SEC | 1 |
| 2022 | InsNet: An Efficient, Flexible, and Performant Insertion-based Text Generation ModelabstractWe propose InsNet, an expressive insertion-based text generator with efficient training and flexible decoding (parallel or sequential). Unlike most existing insertion-based text generation works that require re-encoding of the (decoding) context after each insertion operation and thus are inefficient to train, InsNet only requires one pass of context encoding for the entire insertion sequence during training by using a novel insertion-oriented position encoding to enable computation sharing. Furthermore, InsNet provides a controllable switch between parallel and sequential decoding, making it flexible to handle more parallelizable tasks such as machine translation to support efficient decoding, or less parallelizable tasks such as lexically constrained text generation to guarantee high-quality outputs. Experiments on two unsupervised lexically constrained text generation datasets and three machine translation datasets demonstrate InsNet’s advantages over previous insertion-based methods in terms of training speed, inference efficiency, and generation quality. Sidi Lu, Nanyun Peng 0001 |
NeurIPS | 1 |
| 2022 | Controllable Text Generation with Neurally-Decomposed OracleabstractWe propose a general and efficient framework to control auto-regressive generation models with NeurAlly-Decomposed Oracle (NADO). Given a pre-trained base language model and a sequence-level boolean oracle function, we aim to decompose the oracle function into token-level guidance to steer the base model in text generation. Specifically, the token-level guidance is provided by NADO, a neural model trained with examples sampled from the base model, demanding no additional auxiliary labeled data. Based on posterior regularization, we present the close-form optimal solution to incorporate the decomposed token-level guidance into the base model for controllable generation. We further discuss how the neural approximation affects the quality of the solution. These experiments conducted on two different applications: (1) text generation with lexical constraints and (2) machine translation with formality control demonstrate that our framework efficiently guides the base model towards the given oracle while keeping high generation quality. Sidi Lu, Nanyun Peng 0001, Kai-Wei Chang 0001 |
NeurIPS | 2 |
| 2022 | SC-UDA: Style and Content Gaps aware Unsupervised Domain Adaptation for Object DetectionabstractCurrent state-of-the-art object detectors can have significant performance drop when deployed in the wild due to domain gaps with training data. Unsupervised Domain Adaptation (UDA) is a promising approach to adapt detectors for new domains/environments without any expensive label cost. Previous mainstream UDA works for object detection usually focused on image-level and/or feature-level adaptation by using adversarial learning methods. In this work, we show that such adversarial-based methods can only reduce domain style gap, but cannot address the domain content gap that is also important for object detectors. To overcome this limitation, we propose the SC-UDA framework to concurrently reduce both gaps: We propose fine-grained domain style transfer to reduce the style gaps with finer image details preserved for detecting small objects; Then we leverage the pseudo label-based self-training to reduce content gaps; To address pseudo label error accumulation during self-training, novel optimizations are proposed, including uncertainty-based pseudo labeling and imbalanced mini-batch sampling strategy. Experiment results show that our approach consistently outperforms prior state-of-the-art methods (up to 8.6%, 2.7% and 2.5% mAP on three UDA benchmarks). Fuxun Yu, Di Wang 0003, Yinpeng Chen, Nikolaos Karianakis, Pei Yu, Dimitrios Lymberopoulos, Sidi Lu, Weisong Shi, Xiang Chen 0010 |
WACV | 8 |
| 2022 | EdgeWare: toward extensible and flexible middleware for connected vehicle services
Sidi Lu, Yongtao Yao, Zhifeng Yu, Weisong Shi |
CCF Trans. High Perform. Comput. | 1 |
| 2021 | Computing Systems for Autonomous Driving: State of the Art and ChallengesabstractThe recent proliferation of computing technologies (e.g., sensors, computer vision, machine learning, and hardware acceleration) and the broad deployment of communication mechanisms (e.g., dedicated short-range communication, cellular vehicle-to-everything, 5G) have pushed the horizon of autonomous driving, which automates the decision and control of vehicles by leveraging the perception results based on multiple sensors. The key to the success of these autonomous systems is making a reliable decision in real-time fashion. However, accidents and fatalities caused by early deployed autonomous vehicles arise from time to time. The real traffic environment is too complicated for current autonomous driving computing systems to understand and handle. In this article, we present state-of-the-art computing systems for autonomous driving, including seven performance metrics and nine key technologies, followed by 12 challenges to realize autonomous driving. We hope this article will gain attention from both the computing and automotive communities and inspire more research in this direction. Liangkai Liu, Sidi Lu, Ren Zhong, Baofu Wu, Yongtao Yao, Qingyang Zhang 0001, Weisong Shi |
IEEE Internet Things J. | 2 |
| 2021 | CLONE: Collaborative Learning on the EdgesabstractThe proliferation of edge computing technologies has boosted the development of new applications for a plethora of edge devices. However, many applications face privacy issues and bandwidth limitations. To solve these limitations, we propose a collaborative learning framework on the edges, named CLONE, which is steered by the real-world data sets collected from a large electric vehicle (EV) company and a grocery store of a shopping mall, respectively. We categorize two application scenarios for CLONE, i.e., CLONE in the training stage (CLONE_training) and CLONE in the inference stage (CLONE_inference). As to CLONE_training, we choose the failure prediction of EV battery and associated components as the first use case. While as for CLONE_inference, customer tracking in a grocery store is selected as another case study. In this work, the goal of the CLONE is to support real-time training and inference for connected vehicles and marketing intelligence services. Our experimental results on the EV data show that CLONE is able to reduce model training time without sacrificing algorithm performance. Furthermore, the experimental results on the video data from the grocery store reveal that CLONE is a useful approach to solve the multitarget multicamera tracking problem in a collaborative fashion. Sidi Lu, Yongtao Yao, Weisong Shi |
IEEE Internet Things J. | 1 |
| 2020 | Making Disk Failure Predictions SMARTer!
Sidi Lu, Tirthak Patel, Yongtao Yao, Devesh Tiwari, Weisong Shi |
FAST | 1 |
| 2020 | Edge Compression: An Integrated Framework for Compressive Imaging Processing on CAVsabstractMachine vision is the key to the successful deployment of many Advanced Driver Assistant System (ADAS) / Automated Driving System (ADS) functions, which require accurate high-resolution video processing in a real-time manner. Conventional approaches are either to reduce the frame rate or reduce the related frame size of the conventional camera videos, which lead to undesired consequences such as losing informative high-speed information and/or small objects in the video frames.Unlike conventional cameras, Compressive Imaging (CI) cameras are the promising implications of Compressive Sensing, which is an emerging field with the revelation that the optical domain compressed signal (a small number of linear projections of the original video image data) contains sufficient high-speed information for reconstruction and processing. Yet, CI cameras usually need complicated algorithms to retrieve the desired signal, leading to the corresponding high energy consumption. In this paper, we take a step further to the real applications of CI cameras in connected and autonomous vehicles (CAVs), with the primary goal of accelerating accurate video analysis and decreasing energy consumption. We propose a novel Vehicle Edge Server-Cloud closed-loop framework called Edge Compression for CI processing on CAVs. Our comprehensive experiments with four public datasets demonstrate that the detection accuracy of the compressed video images (named measurements) generated by the CI camera is close to the accuracy on reconstructed videos and comparable to the true value, which paves the way of applying CI in CAVs. Finally, six important observations with supporting evidence and analysis are presented to provide practical implications for researchers and domain experts. The code to reproduce our results is available at https://www.thecarlab.oryoutcomes/software. Sidi Lu, Xin Yuan 0002, Weisong Shi |
SEC | 1 |
| 2020 | A Cross-Layer Optimization Framework for Distributed Computing in IoT NetworksabstractIn Internet-of-Thing (IoT) networks, enormous low-power IoT devices execute latency-sensitive yet computation intensive machine learning tasks. However, the energy is usually scarce for IoT devices, especially for some without battery and relying on solar power or other renewables forms. In this paper, we introduce a cross-layer optimization framework for distributed computing among low-power IoT devices. Specifically, a programming layer design for distributed IoT networks is presented by addressing the problems of application partition, task scheduling, and communication overhead mitigation. Furthermore, the associated federated learning and local differential privacy schemes are developed in the communication layer to enable distributed machine learning with privacy preservation. In addition, we illustrate a three-dimensional network architecture with various network components to facilitate efficient and reliable information exchange among IoT devices. Moreover, a model quantization design for IoT devices is illustrated to reduce the cost of information exchange. Finally, a parallel and scalable neuromorphic computing system for IoT devices is established to achieve energy-efficient distributed computing platforms in the hardware layer. Based on the introduced cross-layer optimization framework, IoT devices can execute their machine learning tasks in an energy-efficient way while guaranteeing data privacy and reducing communication costs. Bodong Shang, Shiya Liu, Sidi Lu, Yang Yi 0002, Weisong Shi, Lingjia Liu 0001 |
SEC | 3 |
| 2019 | OpenEI: An Open Framework for Edge IntelligenceabstractIn the last five years, edge computing has attracted tremendous attention from industry and academia due to its promise to reduce latency, save bandwidth, improve availability, and protect data privacy to keep data secure. At the same time, we have witnessed the proliferation of AI algorithms and models which accelerate the successful deployment of intelligence mainly in cloud services. These two trends, combined together, have created a new horizon: Edge Intelligence (EI). The development of EI requires much attention from both the computer systems research community and the AI community to meet these demands. However, existing computing techniques used in the cloud are not applicable to edge computing directly due to the diversity of computing sources and the distribution of data sources. We envision that there missing a framework that can be rapidly deployed on edge and enable edge AI capabilities. To address this challenge, in this paper we first present the definition and a systematic review of EI. Then, we introduce an Open Framework for Edge Intelligence (OpenEI), which is a lightweight software platform to equip edges with intelligent processing and data sharing capability. We analyze four fundamental EI techniques which are used to build OpenEI and identify several open problems based on potential research directions. Finally, four typical application scenarios enabled by OpenEI are presented. Xingzhou Zhang, Yifan Wang 0005, Sidi Lu, Liangkai Liu, Lanyu Xu, Weisong Shi |
ICDCS | 3 |
| 2019 | Neurally-Guided Structure InferenceabstractMost structure inference methods either rely on exhaustive search or are purely data-driven. Exhaustive search robustly infers the structure of arbitrarily complex data, but it is slow. Data-driven methods allow efficient inference, but do not generalize when test data have more complex structures than training data. In this paper, we propose a hybrid inference algorithm, the Neurally-Guided Structure Inference (NG-SI), keeping the advantages of both search-based and data-driven methods. The key idea of NG-SI is to use a neural network to guide the hierarchical, layer-wise search over the compositional space of structures. We evaluate our algorithm on two representative structure inference tasks: probabilistic matrix decomposition and symbolic program parsing. It outperforms data-driven and search-based alternatives on both tasks. Sidi Lu, Jiayuan Mao, Josh Tenenbaum, Jiajun Wu 0001 |
ICML | 1 |
| 2019 | CoT: Cooperative Training for Generative Modeling of Discrete DataabstractIn this paper, we study the generative models of sequential discrete data. To tackle the exposure bias problem inherent in maximum likelihood estimation (MLE), generative adversarial networks (GANs) are introduced to penalize the unrealistic generated samples. To exploit the supervision signal from the discriminator, most previous models leverage REINFORCE to address the non-differentiable problem of sequential discrete data. However, because of the unstable property of the training signal during the dynamic process of adversarial training, the effectiveness of REINFORCE, in this case, is hardly guaranteed. To deal with such a problem, we propose a novel approach called Cooperative Training (CoT) to improve the training of sequence generative models. CoT transforms the min-max game of GANs into a joint maximization framework and manages to explicitly estimate and optimize Jensen-Shannon divergence. Moreover, CoT works without the necessity of pre-training via MLE, which is crucial to the success of previous methods. In the experiments, compared to existing state-of-the-art methods, CoT shows superior or at least competitive performance on sample quality, diversity, as well as training stability. Sidi Lu, Lantao Yu, Siyuan Feng 0007, Yaoming Zhu, Weinan Zhang 0001 |
ICML | 1 |
| 2018 | Long Text Generation via Adversarial Training with Leaked InformationabstractAutomatically generating coherent and semantically meaningful text has many applications in machine translation, dialogue systems, image captioning, etc. Recently, by combining with policy gradient, Generative Adversarial Nets(GAN) that use a discriminative model to guide the training of the generative model as a reinforcement learning policy has shown promising results in text generation. However, the scalar guiding signal is only available after the entire text has been generated and lacks intermediate information about text structure during the generative process. As such, it limits its success when the length of the generated text samples is long (more than 20 words). In this paper, we propose a new framework, called LeakGAN, to address the problem for long text generation. We allow the discriminative net to leak its own high-level extracted features to the generative net to further help the guidance. The generator incorporates such informative signals into all generation steps through an additional MANAGER module, which takes the extracted features of current generated words and outputs a latent vector to guide the WORKER module for next-word generation.Our extensive experiments on synthetic data and various real-world tasks with Turing test demonstrate that LeakGAN is highly effective in long text generation and also improves the performance in short text generation scenarios. More importantly, without any supervision, LeakGAN would be able to implicitly learn sentence structures only through the interaction between MANAGER and WORKER. Jiaxian Guo, Sidi Lu, Han Cai, Weinan Zhang 0001, Yong Yu 0001, Jun Wang 0012 |
AAAI | 2 |
| 2018 | Texygen: A Benchmarking Platform for Text Generation ModelsabstractWe introduce Texygen, a benchmarking platform to support research on open-domain text generation models. Texygen has not only implemented a majority of text generation models, but also covered a set of metrics that evaluate the diversity, the quality and the consistency of the generated texts. The Texygen platform could help standardize the research on text generation and improve the reproductivity and reliability of future research work in text generation. Yaoming Zhu, Sidi Lu, Lei Zheng 0004, Jiaxian Guo, Weinan Zhang 0001, Jun Wang 0012, Yong Yu 0001 |
SIGIR | 2 |