VLDB 2026 Research / reviewers in the wild / expert
Kaisheng Ma
dblp:133/4053
· DBLP profile ↗
89ranked-venue papers
8as first author
62since 2021 · last 2026
0000-0001-9226-3366ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 39 since 2021Systems, architecture and hardware · 37 · 7 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 20 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterizing Cloud-Native LLM Inference at Bytedance and Exposing Optimization Challenges and Opportunities for Future AI AcceleratorsabstractAs a major provider of LLM inference services, ByteDance has continuously explored diverse accelerator options to meet the rapidly growing inference demands of various heterogeneous LLM scenarios with higher cost-effectiveness, thereby enabling LLMs to serve more people worldwide. However, during this process, we have found that the complexity and opacity of cloud scenarios and corresponding cloud accelerators make it difficult for academia and many innovative chip startups to fully understand the real demands and challenges of these scenarios, which in turn severely restricts innovation and application potential in this field. To bridge this gap, we first present and analyze the data and characteristics of the ByteDance Doubao LLM app across multiple dimensions, helping the community understand real-world cloud scenarios, and detail the challenges and opportunities we have identified. Second, we propose and plan to open-source our multi-level evaluation framework, XPU-Perf, which includes benchmarks spanning instructions, operators, and models. This framework improves interpretability and trustworthiness, and helps promising new accelerator architectures gain wider adoption and development. Finally, we present comparative results of four typical accelerators, summarize their shortcomings and challenges, conduct in-depth analysis, and highlight numerous architectural and scheduling innovation opportunities we have observed. Jingwei Cai, Dehao Kong, Hantao Huang, Zishan Jiang, Zixuan Ma, Qingyu Guo, Guiming Shi, Mingyu Gao 0001, Kaisheng Ma, Minghui Yu |
HPCA | 10 |
| 2025 | MindPainter: Efficient Brain-Conditioned Painting of Natural Images via Cross-Modal Self-Supervised LearningabstractDespite significant advancements in image and text conditional image editing, the exploration of using brain signals, which are more direct and personalized to reflect user intentions, remains limited. An intuitive method is to convert implicit brain signals into explicit representations such as images, which can then serve as prompts for editing. However, such two-stage method suffers from low inference efficiency, inaccurate brain interpretation, and unnatural editing results. In this paper, we apply brain signals of visual perception as prompts and propose a cross-modal self-supervised learning for natural image painting (MindPainter). This method achieves efficient and natural brain-conditioned image editing in a straightforward manner. MindPainter is trained for reconstruction from masked images directly with pseudo-brain signals, which is simulated by the proposed Pseudo Brain Generator. It facilitates efficient cross-modal integration. The proposed Brain Adapter further eliminates the gap in implicit space between modalities, ensuring accurate semantic interpretation of brain signals and coherent consolidation. Besides, the designed Multi-Mask Generation Policy enhances the generalization, realizing high-quality editing in various painting scenarios, including inpainting and outpainting. To the best of our knowledge, MindPainter is the first work to achieve efficient brain-conditioned image painting, providing potential for direct brain control in creative AI. The code and the link to the extended version will be available on GitHub. Muzhou Yu, Shuyun Lin, Hongwei Yan, Kaisheng Ma |
AAAI | 4 |
| 2025 | WPC: Weight Plaintext Compression for CNN Inference based on RNS-CKKSabstractConvolutional neural network (CNN) inference based on RNS-CKKS enables secure processing on encrypted data but introduces significant weight size overhead. Weight plaintext, weight in RNS-CKKS format, can reach tens to hundreds of gigabytes. Existing compression methods either add high computational cost or yield low compression rates. In this work, we propose WPC, Weight Plaintext Compression, to compress weight plaintext for RNS-CKKS-based CNN inference. We observe that the transformation from the weight in CNN models to the weight plaintext in RNS-CKKS format involves an operation akin to the Discrete Fourier Transform, which shifts data between the time and frequency domains while retaining redundant information from periodic and discrete data. Based on this observation, we first introduce the Periodic Transmit Theorem, which states that periodic patterns can be preserved during the transformation process, thereby enabling compression. We then propose Channel Innermost Packing Scheme and Rotation Padding to rearrange the weight data into periodic patterns for compression. Results show that WPC achieves 1.25 to 2.18 times speedup on an A100 GPU and 46.08 to 139.11 times compression rate. Guiming Shi, Shengyu Fan, Xianglong Deng, Liang Kong 0005, Jingwei Cai, Shuwen Deng, Mingzhe Zhang 0005, Kaisheng Ma |
CCS | 10 |
| 2025 | Buffer Prospector: Discovering and Exploiting Untapped Buffer Resources in Many-Core DNN AcceleratorsabstractIn large-scale DNN inference accelerators, the many-core architecture has emerged as a predominant design, with layer-pipeline (LP) mapping being a mainstream mapping approach. However, our experimental findings and theoretical justifications uncover a hardware-independent and prevalent flaw in employing layer-pipeline mapping on many-core accelerators: a significant underutilization of buffer space across numerous cores, indicating substantial potential for optimization. Building on this discovery, we develop a universal and efficient buffer allocation strategy, BufferProspector, which includes a Buffer Requirement Calculator and Buffer Allocator, to capitalize on these unused buffers, addressing the timing mismatch challenge inherent in LP mapping. Compared to the state-of-the-art (SOTA) open-source LP mapping framework Tangram, BufferProspector averages a simultaneous increase in energy efficiency and performance by 1.44× and 2.26×, respectively. Moreover, we conduct some case studies on architecture and mapping. BufferProspector will be open-sourced. Jingwei Cai, Mingyu Gao 0001, Sen Peng, Zuotong Wu, Guiming Shi, Kaisheng Ma |
DAC | 7 |
| 2025 | SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN AcceleratorsabstractModern Deep Neural Network (DNN) accelerators are equipped with increasingly larger on-chip buffers to provide more opportunities to alleviate the increasingly severe DRAM bandwidth pressure. However, most existing research on buffer utilization still primarily focuses on single-layer dataflow scheduling optimization. As buffers grow large enough to accommodate most single-layer weights in most networks, the impact of single-layer dataflow optimization on DRAM communication diminishes significantly. Therefore, developing new paradigms that fuse multiple layers to fully leverage the increasingly abundant onchip buffer resources to reduce DRAM accesses has become particularly important, yet remains an open challenge.To address this challenge, we first identify the optimization opportunities in DRAM communication scheduling by analyzing the drawbacks of existing works on the layer fusion paradigm and recognizing the vast optimization potential in scheduling the timing of data prefetching from and storing to DRAM. To fully exploit these optimization opportunities, we develop a Tensor-centric Notation and its corresponding parsing method to represent different DRAM communication scheduling schemes and depict the overall space of DRAM communication scheduling. Then, to thoroughly and efficiently explore the space of DRAM communication scheduling for diverse accelerators and workloads, we develop an end-to-end scheduling framework, SoMa, which has already been developed into a compiler for our commercial accelerator product. Compared with the state-of-the-art (SOTA) Cocco framework, SoMa achieves, on average, a 2.11× performance improvement and a 37.3% reduction in energy cost simultaneously. Then, we leverage SoMa to study optimizations for LLM, perform design space exploration (DSE), and analyze the DRAM communication scheduling space through a practical example, yielding some interesting insights. Moreover, SoMa has been open-sourced at https://github.com/SET-Scheduling-Project/SoMa-HPCA2025. Jingwei Cai, Mingyu Gao 0001, Sen Peng, Zuotong Wu, Kaisheng Ma |
HPCA | 8 |
| 2025 | Enhancing Vision: Harmonizing Frequency for Imaging Quality and Perception AccuracyabstractIn low-level vision tasks, achieving harmony between visual quality and recognition accuracy is often challenging, as the two do not always align. Many existing approaches focus on optimizing downstream tasks by linking image quality to machine perception, typically incurring additional burdens such as extensive annotations and joint training. In this work, we demonstrate that independent low-level reconstruction algorithms can simultaneously enhance imaging quality and downstream perception accuracy. By conducting a comprehensive frequency-domain analysis, we identify high-frequency components as critical for both visual and perceptual tasks. To counteract the frequency information loss typically seen in ISP pipelines, we propose a Multi-Frequency Fusion Block (MFFB) for on-the-fly upsampling, alongside a Frequency-Aware Supervision (FAS) mechanism guided by discrete wavelet transform. Our method achieves a notable +0.32 dB improvement in smart ISP performance on the Zurich dataset. Moreover, without relying on assistance from downstream tasks, our approach demonstrates significant improvements in object detection and instance segmentation. Hongyang Chen 0003, Kaisheng Ma |
ICASSP | 2 |
| 2025 | MindCustomer: Multi-Context Image Generation Blended with Brain SignalabstractAdvancements in generative models have promoted text- and image-based multi-context image generation. Brain signals, offering a direct representation of user intent, present new opportunities for image customization. However, it faces challenges in brain interpretation, cross-modal context fusion and retention. In this paper, we present MindCustomer to explore the blending of visual brain signals in multi-context image generation. We first design shared neural data augmentation for stable cross-subject brain embedding by introducing the Image-Brain Translator (IBT) to generate brain responses from visual images. Then, we propose an effective cross-modal information fusion pipeline that mask-freely adapts distinct semantics from image and brain contexts within a diffusion model. It resolves semantic conflicts for context preservation and enables harmonious context integration. During the fusion pipeline, we further utilize the IBT to transfer image context to the brain representation to mitigate the cross-modal disparity. MindCustomer enables cross-subject generation, delivering unified, high-quality, and natural image outputs. Moreover, it exhibits strong generalization for new subjects via few-shot learning, indicating the potential for practical application. As the first work for multi-context blending with brain signal, MindCustomer lays a foundational exploration and inspiration for future brain-controlled generative technologies. Muzhou Yu, Shuyun Lin, Lei Ma 0008, Kaisheng Ma |
ICML | 5 |
| 2025 | Benchmarking Long-Horizon Mobile Manipulation in Multi-Room Dynamic EnvironmentsabstractLong-horizon reasoning and task execution are crucial for complex mobile manipulation tasks in household environments. Existing benchmarks and methods primarily focus on single-room or single-object mobile manipulation scenarios, limiting the scope of long-horizon planning and scene-level understanding. To address this gap, we introduce a novel benchmark for long-horizon mobile manipulation in multi-room household environments. Our task requires agents to follow a sequence of language instructions, each directing the movement of specific objects across receptacles and rooms. In this task, we investigate the role of long-term memory by constructing a hierarchical scene graph that captures the relationships between objects, furniture, and rooms. This scene graph-based memory is dynamically updated as the agent explores the environment, which effectively aligns the scene information with the targets and environmental context specified in the language instructions. Additionally, we benchmark the proposed task in dynamic environments where objects can be relocated during task execution, simulating real-world scenarios. Our results demonstrate that the scene graph-based memory significantly improves the agent’s performance in long-horizon mobile manipulation tasks. Moreover, dynamically updating the state of objects within the scene graph enables the agent to better adapt to dynamic conditions. Kaisheng Ma |
IROS | 2 |
| 2025 | SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationabstractWhile spatial reasoning has made progress in object localization relationships, it often overlooks object orientation—a key factor in 6-DoF fine-grained manipulation. Traditional pose representations rely on pre-defined frames or templates, limiting generalization and semantic grounding. In this paper, we introduce the concept of semantic orientation, which defines object orientations using natural language in a reference-frame-free manner (e.g., the ''plug-in'' direction of a USB or the ''handle'' direction of a cup). To support this, we construct OrienText300K, a large-scale dataset of 3D objects annotated with semantic orientations, and develop PointSO, a general model for zero-shot semantic orientation prediction. By integrating semantic orientation into VLM agents, our SoFar framework enables 6-DoF spatial reasoning and generates robotic actions. Extensive experiments demonstrated the effectiveness and generalization of our SoFar, e.g., zero-shot 48.7\% successful rate on Open6DOR and zero-shot 74.9\% successful rate on SIMPLER-Env. Zekun Qi, Yufei Ding 0002, Runpei Dong, Xinqiang Yu, Baoyu Li, Xialin He, Guofan Fan, Jiazhao Zhang, Jiawei He 0002, Jiayuan Gu, Xin Jin 0014, Kaisheng Ma, Zhizheng Zhang 0011, He Wang 0010, Li Yi 0001 |
NeurIPS | 15 |
| 2024 | CutFreq: Cut-and-Swap Frequency Components for Low-Level Vision AugmentationabstractLow-level vision plays a crucial role in a wide range of imaging quality and image recognition applications. However, the limited size, quality, and diversity of datasets often pose significant challenges for low-level tasks. Data augmentation is the most effective and practical way of sample expansion, but the commonly used augmentation methods in high-level tasks have limited improvement in the low-level due to the boundary effects or the non-realistic context information. In this paper, we propose the Cut-and-Swap Frequency Components (CutFreq) method for low-level vision, which aims to preserve high-level representations with directionality and improve image synthesis quality. Observing the significant frequency domain differences between reconstructed images and real ones, in CutFreq, we propose to transform the input and real images separately in the frequency domain, then define two stages for the model training process, and finally swap the specified frequency bands respectively and inversely transform to generate augmented samples. The experimental results show the superior performance of CutFreq on five low-level vision tasks. Moreover, we demonstrate the effectiveness of CutFreq in the low-data regime. Code is available at https://github.com/DreamerCCC/CutFreq. Hongyang Chen 0003, Kaisheng Ma |
AAAI | 2 |
| 2024 | Guiding a Harsh-Environments Robust Detector via RAW Data Characteristic MiningabstractConsumer-grade cameras capture the RAW physical description of a scene and then process the image signals to obtain high-quality RGB images that are faithful to human visual perception. Conventionally, dense prediction scenes require high-precision recognition of objects in RGB images. However, predicting RGB data to exhibit the expected adaptability and robustness in harsh environments can be challenging. By capitalizing on the broader color gamut and higher bit depth offered by RAW data, in this paper, we demonstrate that RAW data can significantly improve the accuracy and robustness of object detectors in harsh environments. Firstly, we propose a general Pipeline for RAW Detection (PRD), along with a preprocessing strategy tailored to RAW data. Secondly, we design the RAW Corruption Benchmark (RCB) to address the dearth of benchmarks that reflect realistic scenarios in harsh environments. Thirdly, we demonstrate the significant improvement of RAW images in object detection for low-light and corrupt scenes. Specifically, our experiments indicate that PRD (using FCOS) outperforms RGB detection by 13.9mAP on LOD-Snow without generating restored images. Finally, we introduce a new nonlinear method called Functional Regularization (FR), which can effectively mine the unique characteristics of RAW data. The code is available at https://github.com/DreamerCCC/RawMining. Hongyang Chen 0003, Hung-Shuo Tai, Kaisheng Ma |
AAAI | 3 |
| 2024 | Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication OptimizationabstractMemory is a critical design consideration in current data-intensive DNN accelerators, as it profoundly determines energy consumption, bandwidth requirements, and area costs. As DNN structures become more complex, a larger on-chip memory capacity is required to reduce data movement overhead, but at the expense of silicon costs. Some previous works have proposed memory-oriented optimizations, such as different data reuse and layer fusion schemes. However, these methods are not general and potent enough to cope with various graph structures. Zhanhong Tan, Kaisheng Ma |
ASPLOS (1) | 3 |
| 2024 | DRAFT: Direct Radiance Fields Editing with Composable Operations
Zhihan Cai, Kailu Wu, Dapeng Cao, Kaisheng Ma |
BMVC | 5 |
| 2024 | Orchestrate Latent Expertise: Advancing Online Continual Learning with Multi-Level Supervision and Reverse Self-DistillationabstractTo accommodate real-world dynamics, artificial intelligence systems need to cope with sequentially arriving content in an online manner. Beyond regular Continual Learning (CL) attempting to address catastrophic forgetting with offline training of each task, Online Continual Learning (OCL) is a more challenging yet realistic setting that performs CL in a one-pass data stream. Current OCL methods primarily rely on memory replay of old training samples. However, a notable gap from CL to OCL stems from the additional overfitting-underfitting dilemma associated with the use of rehearsal buffers: the inadequate learning of new training samples (underfitting) and the repeated learning of a few old training samples (overfitting). To this end, we introduce a novel approach, Multi-level Online Sequential Experts (MOSE), which cultivates the model as stacked sub-experts, integrating multi-level supervision and reverse self-distillation. Supervision signals across multiple stages facilitate appropriate convergence of the new task while gathering various strengths from experts by knowledge distillation mitigates the performance decline of old tasks. MOSE demonstrates remarkable efficacy in learning new samples and preserving past knowledge through multi-level experts, thereby significantly advancing OCL performance over state-of-the-art baselines (e.g., up to 7.3% on Split CIFAR-100 and 6.1% on Split Tiny-ImageNet).††Our code is available at https://github.com/AnAppleCore/MOSE Hongwei Yan, Kaisheng Ma |
CVPR | 3 |
| 2024 | ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
Zekun Qi, Runpei Dong, Shaochen Zhang, Chunrui Han, Zheng Ge, Li Yi 0001, Kaisheng Ma |
ECCV (43) | 8 |
| 2024 | Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsabstractChiplet technology enables the integration of an increasing number of transistors on a single accelerator with higher yield in the post-Moore era, addressing the immense computational demands arising from rapid AI advancements. However, it also introduces more expensive packaging costs and costly Die-to-Die (D2D) interfaces, which require more area, consume higher power, and offer lower bandwidth than onchip interconnects. Maximizing the benefits and minimizing the drawbacks of chiplet technology is crucial for developing largescale DNN chiplet accelerators, which poses challenges to both architecture and mapping. Despite its importance in the post-Moore era, methods to address these challenges remain scarce. To bridge the gap, we first propose a layer-centric encoding method to encode Layer-Pipeline (LP) spatial mapping for largescale DNN inference accelerators and depict the optimization space of it. Based on it, we analyze the unexplored optimization opportunities within this space, which play a more crucial role in chiplet scenarios. Based on the encoding method and a highly configurable and universal hardware template, we propose an architecture and mapping co-exploration framework, Gemini, to explore the design and mapping space of large-scale DNN chiplet accelerators while taking monetary cost (MC), performance, and energy efficiency into account. Compared to the state-of-the-art (SOTA) Simba architecture with SOTA Tangram LP Mapping, Gemini's co-optimized architecture and mapping achieve, on average, 1.98 × performance improvement and 1.41 × energy efficiency improvement simultaneously across various DNNs and batch sizes, with only a 14.3% increase in monetary cost. Moreover, we leverage Gemini to uncover intriguing insights into the methods for utilizing chiplet technology in architecture design and mapping DNN workloads under chiplet scenarios. The Gemini framework is open-sourced at https://github.com/SETScheduling-Project/GEMINI-HPCA2024. Jingwei Cai, Zuotong Wu, Sen Peng, Zhanhong Tan, Guiming Shi, Mingyu Gao 0001, Kaisheng Ma |
HPCA | 8 |
| 2024 | Partial Differential Equation Acceleration by Exploiting Value SimilarityabstractPartial Differential Equations (PDEs) are integral in modeling various scientific and engineering phenomena. One of the most common numerical solutions for these equations is the Finite Difference Method (FDM), which reformulates derivatives into stencil computation patterns. FDM's reliance on various discretized mesh points and high-precision numerical iterations necessitates substantial memory storage and floating-point computations. Kaisheng Ma |
ICCAD | 2 |
| 2024 | LACO: A Latency-Constraint Offline Neural Network Scheduler towards Reliable Self-Driving PerceptionabstractPerception is a crucial and compute-intensive self-driving stage where multiple DNN models periodically process image and point-cloud inputs with constant intervals as latency constraints. This scenario is similar to the multistream scenario in the MLPerf benchmark but with more than one query sequence of diverse intervals from different sensors, and each query also evokes more than one model. In this paper, we call this scenario multi-multistream (M2-stream). Because the periodic inputs arrive continuously, the neural network processor should process them in time to avoid information dropping and maintain self-driving safety. Kaisheng Ma, Zhanhong Tan, Mengdi Wu |
ICCAD | 1 |
| 2024 | DreamLLM: Synergistic Multimodal Comprehension and CreationabstractThis paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation. DreamLLM operates on two fundamental principles. The first focuses on the generative modeling of both language and image posteriors by direct sampling in the raw multimodal space. This approach circumvents the limitations and information loss inherent to external feature extractors like CLIP, and a more thorough multimodal understanding is obtained. Second, DreamLLM fosters the generation of raw, interleaved documents, modeling both text and image contents, along with unstructured layouts. This allows DreamLLM to learn all conditional, marginal, and joint multimodal distributions effectively. As a result, DreamLLM is the first MLLM capable of generating free-form interleaved content. Comprehensive experiments highlight DreamLLM's superior performance as a zero-shot multimodal generalist, reaping from the enhanced learning synergy. Project page: https://dreamllm.github.io. Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jianjian Sun, Xiangwen Kong, Xiangyu Zhang 0005, Kaisheng Ma, Li Yi 0001 |
ICLR | 13 |
| 2024 | NAOL: NeRF-Assisted Omnidirectional Localization
Muzhou Yu, Zhihan Cai, Kailu Wu, Dapeng Cao, Kaisheng Ma |
ICPR (4) | 5 |
| 2024 | MG-VLN: Benchmarking Multi-Goal and Long-Horizon Vision-Language Navigation with Language Enhanced Memory MapabstractVision-Language Navigation (VLN) with high-level language instructions is a crucial task in robotics. Existing VLN benchmarks, such as the REVERIE challenge which has single-goal instructions and limited navigation steps, do not fully encapsulate the complexity of real-world navigation that often require multi-objective and long-horizon navigation. To address this, we propose a new benchmark task: Multi-Goal and Long-Horizon Vision-Language Navigation (MG-VLN), extending the REVERIE benchmark to encompass multi-objective and long-horizon navigation scenarios with sequences of high-level instructions. This task aims to provide a simulation benchmark to guide the design of lifelong and long-horizon navigation robots. To initiate the exploration in this newly proposed task, we first investigate the role of long-term memory in improving navigation performance by leveraging environmental information gathered during previous sub-goals. Additionally, we examine the types of knowledge that most effectively enrich this long-term memory. Specifically, we integrate the visual contents with linguistic knowledge such as object categories, visual captions, and object attributes/relationships. Our findings indicate that: 1) the explicit long-term memory map significantly enhances navigation performance in multi-goal and long-horizon scenarios; 2) incorporating object attributes and relationships information is the most advantageous for aligning environmental cues with high-level instructions. Kaisheng Ma |
IROS | 2 |
| 2024 | Ring Road: A Scalable Polar-Coordinate-based 2D Network-on-Chip ArchitectureabstractNetworks-on-chip (NoCs) are scaled out to build large-scale multi-chip networks to meet the growing demand for computing. However, the traditional router-based NoC architecture has significant limitations: 1) The overhead of routers is high, especially for long-distance traffic that traverses numerous chips and hops; 2) On/off-chip traffic is mixed up, so the local NoC performance is dragged down by the cross-chip traffic; 3) The routing design of intra/inter-chip networks is also entangled. The alternative router-less solution based on isolated multi-ring (IMR) can address some of these issues; however, the complete lack of routers leads to wiring overhead and limited scalability thus not suitable for large-scale multi-chip networks. Therefore, we are motivated to design a new network architecture that can reduce router/wiring overhead, isolate on/off-chip traffic, and decouple inter/intra-chip routing design by combining the advantages of both routers and router-less rings. In this paper, we propose Ring Road, in which high-speed isolated multi-rings are integrated with compact routers, thus achieving flexible traffic delivery with less router usage and overhead. A polar-coordinate-based description is presented to describe the topology, and corresponding deadlock-free routing algorithms are discussed. As a standalone NoC, Ring Road is low-cost, low-latency, and high-performance. It can also be scaled out into multi-chip networks without redesigning the on-chip routing regardless of inter-chip topology. With low-cost isolated ring channels, no matter how heavy the cross-chip traffic is, the local NoC performance remains consistent. Circuit implementation and post-synthesis analysis show that Ring Road achieves$1.7\sim 2\times$bandwidth-per-area and$1.6\sim 1.9\times$bandwidth-per-power compared with the traditional mesh router. Cycle-accurate simulation shows that Ring Road achieves better performance and energy efficiency at various workloads and configurations. Yinxiao Feng, Kaisheng Ma |
MICRO | 3 |
| 2024 | Point-GCC: Universal Self-supervised 3D Scene Pre-training via Geometry-Color Contrast
Guofan Fan, Zekun Qi, Wenkai Shi, Kaisheng Ma |
ACM Multimedia | 4 |
| 2024 | Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single ImageabstractIn this work, we introduce Unique3D, a novel image-to-3D framework for efficiently generating high-quality 3D meshes from single-view images, featuring state-of-the-art generation fidelity and strong generalizability. Previous methods based on Score Distillation Sampling (SDS) can produce diversified 3D results by distilling 3D knowledge from large 2D diffusion models, but they usually suffer from long per-case optimization time with inconsistent issues. Recent works address the problem and generate better 3D results either by finetuning a multi-view diffusion model or training a fast feed-forward model. However, they still lack intricate textures and complex geometries due to inconsistency and limited generated resolution. To simultaneously achieve high fidelity, consistency, and efficiency in single image-to-3D, we propose a novel framework Unique3D that includes a multi-view diffusion model with a corresponding normal diffusion model to generate multi-view images with their normal maps, a multi-level upscale process to progressively improve the resolution of generated orthographic multi-views, as well as an instant and consistent mesh reconstruction algorithm called ISOMER, which fully integrates the color and geometric priors into mesh results. Extensive experiments demonstrate that our Unique3D significantly outperforms other image-to-3D baselines in terms of geometric and textural details. Kailu Wu, Fangfu Liu, Zhihan Cai, Runjie Yan, Hanyang Wang 0003, Yueqi Duan, Kaisheng Ma |
NeurIPS | 8 |
| 2024 | Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale IntegrationabstractExisting high-performance computing (HPC) interconnection architectures are based on high-radix switches, which limits the injection/local performance and introduces latency/energy/cost overhead. The new wafer-scale packaging and high-speed wireline technologies provide high-density, low-latency, and high-bandwidth connectivity, thus promising to support direct-connected high-radix interconnection architecture.In this paper, we propose a wafer-based interconnection architecture called Switch-Less-Dragonfly-on-Wafers. By utilizing distributed high-bandwidth networks-on-chip-on-wafer, costly high-radix switches of the Dragonfly topology are eliminated while increasing the injection/local throughput and maintaining the global throughput. Based on the proposed architecture, we also introduce baseline and improved deadlock-free minimal/non-minimal routing algorithms with only one additional virtual channel. Extensive evaluations show that the Switch-Less-Dragonfly-on-Wafers outperforms the traditional switch-based Dragonfly in both cost and performance. Similar approaches can be applied to other switch-based direct topologies, thus promising to power future large-scale supercomputers. Yinxiao Feng, Kaisheng Ma |
SC | 2 |
| 2024 | Evaluating Chiplet-based Large-Scale Interconnection Networks via Cycle-Accurate Packet-Parallel Simulation
Yinxiao Feng, Kaisheng Ma |
USENIX ATC | 4 |
| 2023 | Language-Assisted 3D Feature Learning for Semantic Scene UnderstandingabstractLearning descriptive 3D features is crucial for understanding 3D scenes with diverse objects and complex structures. However, it is usually unknown whether important geometric attributes and scene context obtain enough emphasis in an end-to-end trained 3D scene understanding network. To guide 3D feature learning toward important geometric attributes and scene context, we explore the help of textual scene descriptions. Given some free-form descriptions paired with 3D scenes, we extract the knowledge regarding the object relationships and object attributes. We then inject the knowledge to 3D feature learning through three classification-based auxiliary tasks. This language-assisted training can be combined with modern object detection and instance segmentation methods to promote 3D semantic scene understanding, especially in a label-deficient regime. Moreover, the 3D feature learned with language assistance is better aligned with the language features, which can benefit various 3D-language multimodal tasks. Experiments on several benchmarks of 3D-only and 3D-language tasks demonstrate the effectiveness of our language-assisted 3D feature learning. Code is available at https://github.com/Asterisci/Language-Assisted-3D. Guofan Fan, Guanghan Wang, Zhengyuan Su, Kaisheng Ma |
AAAI | 5 |
| 2023 | Region-aware Knowledge Distillation for Efficient Image-to-Image Translation
Linfeng Zhang 0001, Xin Chen 0071, Runpei Dong, Kaisheng Ma |
BMVC | 4 |
| 2023 | Structured Knowledge Distillation Towards Efficient Multi-View 3D Object Detection
Linfeng Zhang 0001, Yukang Shi, Hung-Shuo Tai, Kaisheng Ma |
BMVC | 7 |
| 2023 | PointDistiller: Structured Knowledge Distillation Towards Efficient and Compact 3D DetectionabstractThe remarkable breakthroughs in point cloud representation learning have boosted their usage in real-world applications such as self-driving cars and virtual reality. However, these applications usually have a strict requirement for not only accurate but also efficient 3D object detection. Recently, knowledge distillation has been proposed as an effective model compression technique, which transfers the knowledge from an over-parameterized teacher to a lightweight student and achieves consistent effectiveness in 2D vision. However, due to point clouds' sparsity and irregularity, directly applying previous image-based knowledge distillation methods to point cloud detectors usually leads to unsatisfactory performance. To fill the gap, this paper proposes PointDistiller; a structured knowledge distillation framework for point clouds-based 3D detection. Concretely, PointDistiller includes local distillation which extracts and distills the local geometric structure of point clouds with dynamic graph convolution and reweighted learning strategy, which highlights student learning on the crucial points or voxels to improve knowledge distillation efficiency. Extensive experiments on both voxels-based and raw points-based detectors have demonstrated the effectiveness of our method over seven previous knowledge distillation methods. For instance, our 4 x compressed PointPillars student achieves 2.8 and 3.4 mAP improvements on BEV and 3D object detection, outperforming its teacher by 0.9 and 1.8 mAP, respectively. Codes are available in https://github.com/RunpeiDong/PointDistiller. Linfeng Zhang 0001, Runpei Dong, Hung-Shuo Tai, Kaisheng Ma |
CVPR | 4 |
| 2023 | PHEP: Paillier Homomorphic Encryption Processors for Privacy-Preserving Applications in Cloud Computingabstract• Cloud computing has evolved into the key infrastructure of emerging applications, storing massive amounts of data. Yet, how to safely handle this sensitive data in a shared cloud is a major concern. Paillier homomorphic encryption is an important privacy protection approach that permits arithmetic operations on ciphertext without first decrypting it, offering a viable solution to the privacy dilemma. • The Paillier approach has a significant computational overhead compared to plaintext computation because computing in the ciphertext domain requires expensive large integer modular operations that are inefficient for CPUs. As a result, it is preferable to create domain-specific processors for Paillier. Paillier computing patterns are divided into two types, both of which are extensively employed in Paillier applications: independent vector operations and multiply-and-accumulate (MAC) operations. The former is primarily employed in applications such as private information retrieval and on the client side for privacy-preserving AI. In contrast, the latter is required for cloud-side AI inference, particularly computing convolution in neural networks. • We introduce PHEP: Paillier Homomorphic Encryption Processors for cloud-based privacy-preserving applications. PHEP is built on two Paillier acceleration chips: Paillier engine-1 and Paillier engine-2, both produced on the same wafer. Paillier engine-1 focuses on vector operations and attempts to increase computation as much as feasible. It contains 80 processing elements (PE) and can provide 480 TOPS (INT8) for a 16-chip Full-Height-Full-Length (FHFL) PCle card. Paillier engine-2 is designed for MAC operations and has 16 high-performance bit-serial sparse PEs. It only has 192 TOPS (INT8) for an 8-chip FHFL PCle board. However, it is specialized for matrix operations like convolutions. Both engine chips have the same hardware interface, allowing them to use the same PCB board, FPGA scheduler, and software framework design. The PHEP accelerator card also contains a host FPGA. The host FPGA schedules both data transfers and computation among these engine chips. To manage these engines, we use a complex software stack. The software stack includes an offline compiler and an online task scheduler for automatically balancing compute workload across multiple cards on the same server and even across multiple servers. The findings of the end-to-end evaluation reveal that PHEP can perform Paillier-based machine learning workloads 1–2 orders of magnitude faster than state-of-the-art CPUs (Intel Xeon Platinum 8260M with 192 cores), making these privacy-preserving applications practical. Guiming Shi, Xueqiang Wang, Zhanhong Tan, Dapeng Cao, Jingwei Cai, Wuke Zhang, Yifu Wu, Kaisheng Ma |
HCS | 12 |
| 2023 | A Scalable Multi-Chiplet Deep Learning Accelerator with Hub-Side 2.5D Heterogeneous IntegrationabstractWith the slowdown of Moore's law, the scenario diversity of specialized computing, and the rapid development of application algorithms, an efficient chip design requires modularization, flexibility, and scalability. In this study, we propose_ a Chiplet,-based deep, learning accelerator protoype that -contains oneHUB . Chipletand, six. extended SIDE- Chiplets integrated on an RDL layer for the 2.5D package. The SIDE and the HUB contain one and four AI cores, respectively. Given that our Chiplet-system targets diverse scenarios via scalable connected SIDE Chiplets, we need to handle three challenges: a) devise a flexible architecture design supporting diverse shapes, b) search for a workload mapping with low die-to-die communication, and c) adopt a high-bandwidth die-to-die interface to maintain efficient data transfer. This study proposes a flexible neural core (FNC) featuring dynamic bit-width computing and flexible parallelism. Next, we use a hierarchy-based mapping. scheme to decouple different parallelism levels and help analyze the communication. A 12Gbps,_D2D interface is introduced to achieve 192Gb/s bandwidth per D2D port with 1.04pJ/bit efficiency and 55um bump pitch. The proposed seven-Chiplet accelerator achieves a peak performance of 1 0/20/40 TOPS for INT16/8/4. When enabling 0~6 SIDE Chiplets, the system power ranges from 4.5W to 12W. The power efficiency of the FNC is 2.02TOPS/W while that of the overall system is 1.67TOPS/W. Zhanhong Tan, Yifu Wu, Yannian Zhang, Haobing Shi, Wuke Zhang, Kaisheng Ma |
HCS | 6 |
| 2023 | A Scalable Methodology for Designing Efficient Interconnection Network of ChipletsabstractThe Chiplet methodology can accelerate VLSI system development and provide better flexibility. However, it is not easy to build interconnection networks across multiple chiplets and maintain high-performance deadlock-free routing in systems of various hierarchical topologies. In particular, most on-chiplet networks are based on flat topologies such as 2D-mesh, which are inflexible and insufficient for large-scale multi-chiplet systems.To take full advantage of the multi-chiplet architecture and advanced packaging, we propose an interconnection method that can flexibly establish high-radix interconnection networks from typical 2D-mesh-NoC-based chiplets. A minus-first-based deadlock-free adaptive routing algorithm and a safe/unsafe flow control policy are introduced for these multi-chiplet interconnection networks. Additionally, a general approach network interleaving is used to balance the communication bandwidth within and between chiplets.We evaluate different architectures and traffic patterns on a cycle-accurate C++ simulator. Compared with traditional adaptive routing in 2D-mesh, our methodology can significantly improve network performance in various cases. The more chiplets there are, the more effective the method is. For 64 4×4-2D-mesh-based chiplets, The maximum injection rate increase is up to 2×, and the average latency reduction is up to 45%. Yinxiao Feng, Kaisheng Ma |
HPCA | 3 |
| 2023 | Tiny Updater: Towards Efficient Neural Network-Driven Software UpdatingabstractSignificant advancements have been accomplished with deep neural networks in diverse visual tasks, which have substantially elevated their deployment in edge device software. However, during the update of neural network-based software, users are required to download all the parameters of the neural network anew, which harms the user experience. Motivated by previous progress in model compression, we propose a novel training methodology named Tiny Updater to address this issue. Specifically, by adopting the variant of pruning and knowledge distillation methods, Tiny Updater can update the neural network-based software by only downloading a few parameters (10%~20%) instead of all the parameters in the neural network. Experiments on eleven datasets of three tasks, including image classification, image-to-image translation, and video recognition have demonstrated its effectiveness. Codes have been released in https://github.com/ArchipLab-LinfengZhang/TinyUpdater. Linfeng Zhang 0001, Kaisheng Ma |
ICCV | 2 |
| 2023 | Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?
Runpei Dong, Zekun Qi, Linfeng Zhang 0001, Jianjian Sun, Zheng Ge, Li Yi 0001, Kaisheng Ma |
ICLR | 8 |
| 2023 | Hebbian and Gradient-based Plasticity Enables Robust Memory and Rapid Learning in RNNs
Zhongfan Jia, Kaisheng Ma |
ICLR | 5 |
| 2023 | Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative PretrainingabstractMainstream 3D representation learning approaches are built upon contrastive or generative modeling pretext tasks, where great improvements in performance on various downstream tasks have been achieved. However, we find these two paradigms have different characteristics: (i) contrastive models are data-hungry that suffer from a representation over-fitting issue; (ii) generative models have a data filling issue that shows inferior data scaling capacity compared to contrastive models. This motivates us to learn 3D representations by sharing the merits of both paradigms, which is non-trivial due to the pattern difference between the two paradigms. In this paper, we propose contrast with reconstruct (ReCon) that unifies these two paradigms. ReCon is trained to learn from both generative modeling teachers and cross-modal contrastive teachers through ensemble distillation, where the generative student is used to guide the contrastive student. An encoder-decoder style ReCon-block is proposed that transfers knowledge through cross attention with stop-gradient, which avoids pretraining over-fitting and pattern difference issues. ReCon achieves a new state-of-the-art in 3D representation learning, e.g., 91.26% accuracy on ScanObjectNN. Codes have been released at https://github.com/qizekun/ReCon. Zekun Qi, Runpei Dong, Guofan Fan, Zheng Ge, Xiangyu Zhang 0005, Kaisheng Ma, Li Yi 0001 |
ICML | 6 |
| 2023 | Revisiting Data Augmentation in Model Compression: An Empirical and Comprehensive StudyabstractThe excellent performance of deep neural networks is usually accompanied by a large number of parameters and computations, which have limited their usage on the resource- limited edge devices. To address this issue, abundant methods such as pruning, quantization and knowledge distillation have been proposed to compress neural networks and achieved significant breakthroughs. However, most of these compression methods focus on the architecture or the training method of neural networks but ignore the influence from data augmentation. In this paper, we revisit the usage of data augmentation in model compression and give a comprehensive study on the relation between model sizes and their optimal data augmentation policy. To sum up, we mainly have the following three observations: (A) Models in different sizes prefer data augmentation with different magnitudes. Hence, in iterative pruning, data augmentation with varying magnitudes leads to better performance than data augmentation with a consistent magnitude. (B) Data augmentation with a high magnitude may significantly improve the performance of large models but harm the performance of small models. Fortunately, small models can still benefit from strong data augmentations by firstly learning them with “additional parameters” and then discard these “additional parameters” during inference. (C) The prediction of a pre-trained large model can be utilized to measure the difficulty of data augmentation. Thus it can be utilized as a criterion to design better data augmentation policies. We hope this paper may promote more research on the usage of data augmentation in model compression. Muzhou Yu, Linfeng Zhang 0001, Kaisheng Ma |
IJCNN | 3 |
| 2023 | Inter-layer Scheduling Space Definition and Exploration for Tiled AcceleratorsabstractWith the continuous expansion of the DNN accelerator scale, inter-layer scheduling, which studies the allocation of computing resources to each layer and the computing order of all layers in a DNN, plays an increasingly important role in maintaining a high utilization rate and energy efficiency of DNN inference accelerators. However, current inter-layer scheduling is mainly conducted based on some heuristic patterns. The space of inter-layer scheduling has not been clearly defined, resulting in significantly limited optimization opportunities and a lack of understanding on different inter-layer scheduling choices and their consequences. Jingwei Cai, Zuotong Wu, Sen Peng, Kaisheng Ma |
ISCA | 5 |
| 2023 | Heterogeneous Die-to-Die Interfaces: Enabling More Flexible Chiplet Interconnection SystemsabstractThe chiplet architecture is one of the emerging methodologies and is believed to be scalable and economical. However, most current multi-chiplet systems are based on one uniform die-to-die interface, which severely limits flexibility. First, any interface has specific applicable workloads/scales/scenarios; therefore, chiplets with a uniform interface cannot be freely reused in different systems. Second, since modern computing systems must deal with complex and mixed tasks, the uniform interface does not cope well with flexible workloads, especially for large-scale systems. Yinxiao Feng, Kaisheng Ma |
MICRO | 3 |
| 2023 | VPP: Efficient Conditional 3D Generation via Voxel-Point Progressive RepresentationabstractConditional 3D generation is undergoing a significant advancement, enabling the free creation of 3D content from inputs such as text or 2D images. However, previous approaches have suffered from low inference efficiency, limited generation categories, and restricted downstream applications. In this work, we revisit the impact of different 3D representations on generation quality and efficiency. We propose a progressive generation method through Voxel-Point Progressive Representation (VPP). VPP leverages structured voxel representation in the proposed Voxel Semantic Generator and the sparsity of unstructured point representation in the Point Upsampler, enabling efficient generation of multi-category objects. VPP can generate high-quality 8K point clouds within 0.2 seconds. Additionally, the masked generation Transformer allows for various 3D downstream tasks, such as generation, editing, completion, and pre-training. Extensive experiments demonstrate that VPP efficiently generates high-fidelity and diverse 3D shapes across different categories, while also exhibiting excellent representation transfer performance. Codes will be released at https://github.com/qizekun/VPP. Zekun Qi, Muzhou Yu, Runpei Dong, Kaisheng Ma |
NeurIPS | 4 |
| 2023 | Structured Knowledge Distillation for Accurate and Efficient Object DetectionabstractKnowledge distillation, which aims to transfer the knowledge learned by a cumbersome teacher model to a lightweight student model, has become one of the most popular and effective techniques in computer vision. However, many previous knowledge distillation methods are designed for image classification and fail in more challenging tasks such as object detection. In this paper, we first suggest that the failure of knowledge distillation on object detection is mainly caused by two reasons: (1) the imbalance between pixels of foreground and background and (2) lack of knowledge distillation on the relation among different pixels. Then, we propose a structured knowledge distillation scheme, includingattention-guided distillationandnon-local distillationto address the two issues, respectively. Attention-guided distillation is proposed to find the crucial pixels of foreground objects with an attention mechanism and then make the students take more effort to learn their features. Non-local distillation is proposed to enable students to learn not only the feature of an individual pixel but also the relation between different pixels captured by non-local modules. Experimental results have demonstrated the effectiveness of our method on thirteen kinds of object detection models with twelve comparison methods for both object detection and instance segmentation. For instance, Faster RCNN with our distillation achieves 43.9 mAP on MS COCO2017, which is 4.1 higher than the baseline. Additionally, we show that our method is also beneficial to the robustness and domain generalization ability of detectors. Codes and model weights have been released on GitHub† Linfeng Zhang 0001, Kaisheng Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Learn From Unpaired Data for Image Restoration: A Variational Bayes ApproachabstractCollecting paired training data is difficult in practice, but the unpaired samples broadly exist. Current approaches aim at generating synthesized training data from unpaired samples by exploring the relationship between the corrupted and clean data. This work proposes LUD-VAE, a deep generative method to learn the joint probability density function from data sampled from marginal distributions. Our approach is based on a carefully designed probabilistic graphical model in which the clean and corrupted data domains are conditionally independent. Using variational inference, we maximize the evidence lower bound (ELBO) to estimate the joint probability density function. Furthermore, we show that the ELBO is computable without paired samples under the inference invariant assumption. This property provides the mathematical rationale of our approach in the unpaired setting. Finally, we apply our method to real-world image denoising, super-resolution, and low-light image enhancement tasks and train the models using the synthetic data generated by the LUD-VAE. Experimental results validate the advantages of our method over other approaches. Dihan Zheng, Kaisheng Ma, Chenglong Bao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Reliable and Efficient Parallel Checkpointing Framework for Nonvolatile Processor With Concurrent PeripheralsabstractIntermittent systems powered by ambient energy harvesting are becoming popular for the benefits of an infinite lifetime and minimum maintenance requirements. Nonvolatile processors (NVPs) enable continual task executions under an unstable power supply with an efficient reactive checkpointing strategy. However, recovering concurrent peripherals in an intermittent system may incur significant overhead once power failures take place, and the recovery of interrupts also lacks discussion in existing works. Noticing the different optimization directions between responsive checkpointing within NVPs and proactive checkpointing required to recover concurrent peripherals, this paper proposes REMARK, an NVP architecture enabling hybrid checkpointing and efficient peripheral recovery. REMARK expands the current NVP structure with a hybrid backup/restore module, a peripheral handler and an interrupt handler, which addresses the recovery problem of both peripherals and interrupts efficiently. A REMARK chip is fabricated to verify the proposed architecture. Results show that the execution efficiency is improved by$13\times $compared with the state-of-the-art. With programming optimization, another 36.5% performance improvement can be achieved. Tongda Wu, Kaisheng Ma, Jingtong Hu, Chun Jason Xue, Jinyang Li 0002, Huazhong Yang, Yongpan Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | A Good Data Augmentation Policy is not All You Need: A Multi-Task Learning PerspectiveabstractData augmentation, which improves the diversity of datasets by applying image transformations, has become one of the most effective techniques in visual representation learning. Usually, the design of augmentation policies faces a diversity-difficulty trade-off. On the one hand, a simple augmentation leads to a low training set diversity, which can not improve model performance significantly. On the other hand, an excessively hard augmentation has an overlarge regularization effect which harms model performance. Recently, automatic augmentation methods have been proposed to address this issue by searching the optimal data augmentation policy from a predefined searching space. However, these methods still suffer from heavy searching overhead or complex optimization objectives. In this paper, instead of searching the optimal augmentation policy, we propose to break the diversity-difficulty trade-off from a multi-task learning perspective. By formulating model learning on the augmented images and the original images as the auxiliary task and the primary task in multi-task learning respectively, the hard augmentation does not directly influence the training of the primary branch and thus its negative influence can be alleviated. Hence, neural networks can learn valuable semantic information even with a totally random augmentation policy. Experimental results on ten datasets for four tasks demonstrate the superiority of our method over the other twelve methods. Codes have been released inhttps://github.com/ArchipLab-LinfengZhang/data-augmentation-multi-task. Linfeng Zhang 0001, Kaisheng Ma |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | LW-ISP: A Lightweight Model with ISP and Deep Learning
Hongyang Chen 0003, Kaisheng Ma |
BMVC | 2 |
| 2022 | Wavelet Knowledge Distillation: Towards Efficient Image-to-Image TranslationabstractRemarkable achievements have been attained with Generative Adversarial Networks (GANs) in image-to-image translation. However, due to a tremendous amount of parameters, state-of-the-art GANs usually suffer from low efficiency and bulky memory usage. To tackle this challenge, firstly, this paper investigates GANs performance from a frequency perspective. The results show that GANs, especially small GANs lack the ability to generate high-quality high frequency information. To address this problem, we propose a novel knowledge distillation method referred to as wavelet knowledge distillation. Instead of directly distilling the generated images of teachers, wavelet knowledge distillation first decomposes the images into different frequency bands with discrete wavelet transformation and then only distills the high frequency bands. As a result, the student GAN can pay more attention to its learning on high frequency bands. Experiments demonstrate that our method leads to 7.08× compression and 6.80× acceleration on CycleGAN with almost no performance drop. Additionally, we have studied the relation between discriminators and generators which shows that the compression of discriminators can promote the performance of compressed generators. Linfeng Zhang 0001, Xin Chen 0071, Xiaobing Tu, Pengfei Wan 0001, Kaisheng Ma |
CVPR | 6 |
| 2022 | Rethinking the Augmentation Module in Contrastive Learning: Learning Hierarchical Augmentation Invariance with Expanded ViewsabstractA data augmentation module is utilized in contrastive learning to transform the given data example into two views, which is considered essential and irreplaceable. However, the pre-determined composition of multiple data augmentations brings two drawbacks. First, the artificial choice of augmentation types brings specific representational invariances to the model, which have different de-grees of positive and negative effects on different down-stream tasks. Treating each type of augmentation equally during training makes the model learn non-optimal repre-sentations for various downstream tasks and limits the flex-ibility to choose augmentation types beforehand. Second, the strong data augmentations used in classic contrastive learning methods may bring too much invariance in some cases, and fine- grained information that is essential to some downstream tasks may be lost. This paper proposes a gen-eral method to alleviate these two problems by considering “where” and “what” to contrast in a general contrastive learning framework. We first propose to learn different aug-mentation invariances at different depths of the model ac-cording to the importance of each data augmentation in-stead of learning representational invariances evenly in the backbone. We then propose to expand the contrast content with augmentation embeddings to reduce the misleading ef-fects of strong data augmentations. Experiments based on several baseline methods demonstrate that we learn better representations for various benchmarks on classification, detection, and segmentation downstream tasks. Kaisheng Ma |
CVPR | 2 |
| 2022 | YOLoC: deploy large-scale neural network by ROM-based computing-in-memory using residual branch on a chipabstractComputing-in-memory (CiM) is a promising technique to achieve high energy efficiency in data-intensive matrix-vector multiplication (MVM) by relieving the memory bottleneck. Unfortunately, due to the limited SRAM capacity, existing SRAM-based CiM needs to reload the weights from DRAM in large-scale networks. This undesired fact weakens the energy efficiency significantly. This work, for the first time, proposes the concept, design, and optimization of computing-in-ROM to achieve much higher on-chip memory capacity, and thus less DRAM access and lower energy consumption. Furthermore, to support different computing scenarios with varying weights, a weight fine-tune technique, namely Residual Branch (ReBranch), is also proposed. ReBranch combines ROM-CiM and assisting SRAM-CiM to achieve high versatility. YOLoC, a ReBranch-assisted ROM-CiM framework for object detection is presented and evaluated. With the same area in 28nm CMOS, YOLoC for several datasets has shown significant energy efficiency improvement by 14.8x for YOLO (DarkNet-19) and 4.8x for ResNet-18, with <8% latency overhead and almost no mean average precision (mAP) loss (−0.5% ~ +0.2%), compared with the fully SRAM-based CiM. Guodong Yin, Zhanhong Tan, Mingyen Lee, Yongpan Liu, Huazhong Yang, Kaisheng Ma, Xueqing Li 0002 |
DAC | 8 |
| 2022 | Chiplet actuary: a quantitative cost model and multi-chiplet architecture explorationabstractMulti-chip integration is widely recognized as the extension of Moore's Law. Cost-saving is a frequently mentioned advantage, but previous works rarely present quantitative demonstrations on the cost superiority of multi-chip integration over monolithic SoC. In this paper, we build a quantitative cost model and put forward an analytical method for multi-chip systems based on three typical multi-chip integration technologies to analyze the cost benefits from yield improvement, chiplet and package reuse, and heterogeneity. We re-examine the actual cost of multi-chip systems from various perspectives and show how to reduce the total cost of the VLSI system through appropriate multi-chiplet architecture. Yinxiao Feng, Kaisheng Ma |
DAC | 2 |
| 2022 | Contrastive Deep Supervision
Linfeng Zhang 0001, Xin Chen 0071, Runpei Dong, Kaisheng Ma |
ECCV (26) | 5 |
| 2022 | Finding the Task-Optimal Low-Bit Sub-Distribution in Deep Neural NetworksabstractQuantized neural networks typically require smaller memory footprints and lower computation complexity, which is crucial for efficient deployment. However, quantization inevitably leads to a distribution divergence from the original network, which generally degrades the performance. To tackle this issue, massive efforts have been made, but most existing approaches lack statistical considerations and depend on several manual configurations. In this paper, we present an adaptive-mapping quantization method to learn an optimal latent sub-distribution that is inherent within models and smoothly approximated with a concrete Gaussian Mixture (GM). In particular, the network weights are projected in compliance with the GM-approximated sub-distribution. This sub-distribution evolves along with the weight update in a co-tuning schema guided by the direct task-objective optimization. Sufficient experiments on image classification and object detection over various modern architectures demonstrate the effectiveness, generalization property, and transferability of the proposed method. Besides, an efficient deployment flow for the mobile CPU is developed, achieving up to 7.46$\times$ inference acceleration on an octa-core ARM CPU. Our codes have been publicly released at https://github.com/RunpeiDong/DGMS. Runpei Dong, Zhanhong Tan, Mengdi Wu, Linfeng Zhang 0001, Kaisheng Ma |
ICML | 5 |
| 2022 | MemSR: Training Memory-efficient Lightweight Model for Image Super-ResolutionabstractMethods based on deep neural networks with a massive number of layers and skip-connections have made impressive improvements on single image super-resolution (SISR). The skip-connections in these complex models boost the performance at the cost of a large amount of memory. With the increase of camera resolution from 1 million pixels to 100 million pixels on mobile phones, the memory footprint of these algorithms also increases hundreds of times, which restricts the applicability of these models on memory-limited devices. A plain model consisting of a stack of 3{\texttimes}3 convolutions with ReLU, in contrast, has the highest memory efficiency but poorly performs on super-resolution. This paper aims at calculating a winning initialization from a complex teacher network for a plain student network, which can provide performance comparable to complex models. To this end, we convert the teacher model to an equivalent large plain model and derive the plain student’s initialization. We further improve the student’s performance through initialization-aware feature distillation. Extensive experiments suggest that the proposed method results in a model with a competitive trade-off between accuracy and speed at a much lower memory footprint than other state-of-the-art lightweight approaches. Kailu Wu, Chung-Kuei Lee, Kaisheng Ma |
ICML | 3 |
| 2022 | SAFIT: Segmentation-Aware Scene Flow with Improved TransformerabstractScene flow prediction is a challenging task that aims at jointly estimating the 3D structure and 3D motion of dynamic scenes. The previous methods concentrate more on point-wise estimation instead of considering the correspondence between objects as well as lacking the sensation of high-level semantic knowledge. In this paper, we propose a concise yet effective method for scene flow prediction. The key idea is to extend the view of all points for computing point cloud features into object-level, thus simultaneously modeling the relationships of the object-level and point-level via an improved transformer. In addition, we introduce a novel unsupervised loss called segmentation-aware loss, which can model semanticaware details to help predict scene flow more accurately and robustly. Since this loss can be trained without any ground truth, it can be used in both supervised training and self-supervised training. Experiments on both supervised training and self-supervised training demonstrate the effectiveness of our method. On supervised training, 3.8%, 22.58%, 10.90% and 21.82 % accuracy boosts than FLOT [23] can be observed on FT3Ds, KITTIs, FT3Do and KITTIo datasets. On self-supervised scheme, 48.23% and 48.96% accuracy boost than PointPWC-Net [40] can be observed on KITTIo and KITTIs datasets. Yukang Shi, Kaisheng Ma |
ICRA | 2 |
| 2022 | Self-Distillation: Towards Efficient and Compact Neural NetworksabstractRemarkable achievements have been obtained by deep neural networks in the last several years. However, the breakthrough in neural networks accuracy is always accompanied by explosive growth of computation and parameters, which leads to a severe limitation of model deployment. In this paper, we propose a novel knowledge distillation technique named self-distillation to address this problem. Self-distillation attaches several attention modules and shallow classifiers at different depths of neural networks and distills knowledge from the deepest classifier to the shallower classifiers. Different from the conventional knowledge distillation methods where the knowledge of the teacher model is transferred to another student model, self-distillation can be considered as knowledge transfer in the same model - from the deeper layers to the shallow layers. Moreover, the additional classifiers in self-distillation allow the neural network to work in a dynamic manner, which leads to a much higher acceleration. Experiments demonstrate that self-distillation has consistent and significant effectiveness on various neural networks and datasets. On average, 3.49 and 2.32 percent accuracy boost are observed on CIFAR100 and ImageNet. Besides, experiments show that self-distillation can be combined with other model compression methods, including knowledge distillation, pruning and lightweight model design. Linfeng Zhang 0001, Chenglong Bao, Kaisheng Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Non-Structured DNN Weight Pruning - Is It Beneficial in Any Platform?abstractLarge deep neural network (DNN) models pose the key challenge to energy efficiency due to the significantly higher energy consumption of off-chip DRAM accesses than arithmetic or SRAM operations. It motivates the intensive research on model compression with two main approaches. Weight pruning leverages the redundancy in the number of weights and can be performed in a non-structured, which has higher flexibility and pruning rate but incurs index accesses due to irregular weights, or structured manner, which preserves the full matrix structure with a lower pruning rate. Weight quantization leverages the redundancy in the number of bits in weights. Compared to pruning, quantization is much more hardware-friendly and has become a "must-do" step for FPGA and ASIC implementations. Thus, any evaluation of the effectiveness of pruning should be on top of quantization. The key open question is, with quantization, what kind of pruning (non-structured versus structured) is most beneficial? This question is fundamental because the answer will determine the design aspects that we should really focus on to avoid the diminishing return of certain optimizations. This article provides a definitive answer to the question for the first time. First, we build ADMM-NN-S by extending and enhancing ADMM-NN, a recently proposed joint weight pruning and quantization framework, with the algorithmic supports for structured pruning, dynamic ADMM regulation, and masked mapping and retraining. Second, we develop a methodology for fair and fundamental comparison of non-structured and structured pruning in terms of both storage and computation efficiency. Our results show that ADMM-NN-S consistently outperforms the prior art: 1) it achieves 348× , 36× , and 8× overall weight pruning on LeNet-5, AlexNet, and ResNet-50, respectively, with (almost) zero accuracy loss and 2) we demonstrate the first fully binarized (for all layers) DNNs can be lossless in accuracy in many cases. These results provide a strong baseline and credibility of our study. Based on the proposed comparison framework, with the same accuracy and quantization, the results show that non-structured pruning is not competitive in terms of both storage and computation efficiency. Thus, we conclude that structured pruning has a greater potential compared to non-structured pruning. We encourage the community to focus on studying the DNN inference acceleration with structured sparsity. Sheng Lin 0001, Shaokai Ye, Zhezhi He, Linfeng Zhang 0001, Geng Yuan, Sia Huat Tan, Zhengang Li 0001, Deliang Fan, Xuehai Qian, Xue Lin 0001, Kaisheng Ma, Yanzhi Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 12 |
| 2021 | Multi-Glimpse Network: A Robust and Efficient Classification Architecture based on Recurrent Downsampled Attention
Sia Huat Tan, Runpei Dong, Kaisheng Ma |
BMVC | 3 |
| 2021 | Improve Object Detection with Feature-based Knowledge Distillation: Towards Accurate and Efficient Detectors
Linfeng Zhang 0001, Kaisheng Ma |
ICLR | 2 |
| 2021 | An Unsupervised Deep Learning Approach for Real-World Image Denoising
Dihan Zheng, Sia Huat Tan, Zuoqiang Shi, Kaisheng Ma, Chenglong Bao |
ICLR | 5 |
| 2021 | Wavelet J-Net: A Frequency Perspective on Convolutional Neural NetworksabstractIt is well acknowledged in image processing domain that the information can be decomposed into different frequency parts and each part has its own merits. However, existing neural networks always ignore the distinctions and straightforwardly feed all the information into neural networks together, treating them equally. In this paper, we propose a novel neural networks framework named J-Net that decomposes images into different frequency bands and then processes them sequentially. Concretely, the images have been decomposed by wavelet transformation and then the wavelet coefficients are fed into neural networks gradually in different depth according to their decomposition levels. An attention module is utilized to facilitate the fusion of neural network features and injected information, yielding significant performance gain. Furthermore, we show how does the information with different frequency impact the accuracy of neural networks. Experiments show that 5.91%, 5.32% and 2.00% accuracy improvements on Caltech 101, Caltech256 and ImageNet, respectively. Linfeng Zhang 0001, Xiaoman Zhang, Chenglong Bao, Kaisheng Ma |
IJCNN | 4 |
| 2021 | NN-Baton: DNN Workload Orchestration and Chiplet Granularity Exploration for Multichip AcceleratorsabstractThe revolution of machine learning poses an unprecedented demand for computation resources, urging more transistors on a single monolithic chip, which is not sustainable in the Post-Moore era. The multichip integration with small functional dies, called chiplets, can reduce the manufacturing cost, improve the fabrication yield, and achieve die-level reuse for different system scales. DNN workload mapping and hardware design space exploration on such multichip systems are critical, but missing in the current stage.This work provides a hierarchical and analytical framework to describe the DNN mapping on a multichip accelerator and analyze the communication overhead. Based on this framework, we propose an automatic tool called NN-Baton with a pre-design flow and a post-design flow. The pre-design flow aims to guide the chiplet granularity exploration with given area and performance budgets for the target workload. The post-design flow focuses on the workload orchestration on different computation levels -package, chiplet, and core - in the hierarchy. Compared to Simba, NN-Baton generates mapping strategies that save 22.5%∼44% energy under the same computation and memory configurations.The architecture exploration demonstrates that area is a decisive factor for the chiplet granularity. For a 2048-MAC system under a 2 mm2chiplet area constraint, the 4-chiplet implementation with 4 cores and 16 lanes of 8-size vector-MAC is always the top-pick computation allocation across several benchmarks. In contrast, the optimal memory allocation policy in the hierarchy typically depends on the neural network models. Zhanhong Tan, Hongyu Cai, Runpei Dong, Kaisheng Ma |
ISCA | 4 |
| 2021 | AFEC: Active Forgetting of Negative Transfer in Continual LearningabstractContinual learning aims to learn a sequence of tasks from dynamic data distributions. Without accessing to the old training samples, knowledge transfer from the old tasks to each new task is difficult to determine, which might be either positive or negative. If the old knowledge interferes with the learning of a new task, i.e., the forward knowledge transfer is negative, then precisely remembering the old tasks will further aggravate the interference, thus decreasing the performance of continual learning. By contrast, biological neural networks can actively forget the old knowledge that conflicts with the learning of a new experience, through regulating the learning-triggered synaptic expansion and synaptic convergence. Inspired by the biological active forgetting, we propose to actively forget the old knowledge that limits the learning of new tasks to benefit continual learning. Under the framework of Bayesian continual learning, we develop a novel approach named Active Forgetting with synaptic Expansion-Convergence (AFEC). Our method dynamically expands parameters to learn each new task and then selectively combines them, which is formally consistent with the underlying mechanism of biological active forgetting. We extensively evaluate AFEC on a variety of continual learning benchmarks, including CIFAR-10 regression tasks, visual classification tasks and Atari reinforcement tasks, where AFEC effectively improves the learning of new tasks and achieves the state-of-the-art performance in a plug-and-play way. Mingtian Zhang, Zhongfan Jia, Qian Li 0040, Chenglong Bao, Kaisheng Ma, Jun Zhu 0001 |
NeurIPS | 6 |
| 2020 | PCONV: The Missing but Desirable Sparsity in DNN Weight Pruning for Real-Time Execution on Mobile DevicesabstractModel compression techniques on Deep Neural Network (DNN) have been widely acknowledged as an effective way to achieve acceleration on a variety of platforms, and DNN weight pruning is a straightforward and effective method. There are currently two mainstreams of pruning methods representing two extremes of pruning regularity: non-structured, fine-grained pruning can achieve high sparsity and accuracy, but is not hardware friendly; structured, coarse-grained pruning exploits hardware-efficient structures in pruning, but suffers from accuracy drop when the pruning rate is high. In this paper, we introduce PCONV, comprising a new sparsity dimension, – fine-grained pruning patterns inside the coarse-grained structures. PCONV comprises two types of sparsities, Sparse Convolution Patterns (SCP) which is generated from intra-convolution kernel pruning and connectivity sparsity generated from inter-convolution kernel pruning. Essentially, SCP enhances accuracy due to its special vision properties, and connectivity sparsity increases pruning rate while maintaining balanced workload on filter computation. To deploy PCONV, we develop a novel compiler-assisted DNN inference framework and execute PCONV models in real-time without accuracy compromise, which cannot be achieved in prior work. Our experimental results show that, PCONV outperforms three state-of-art end-to-end DNN frameworks, TensorFlow-Lite, TVM, and Alibaba Mobile Neural Network with speedup up to 39.2 ×, 11.4 ×, and 6.3 ×, respectively, with no accuracy loss. Mobile devices can achieve real-time inference on large-scale DNNs. Fu-Ming Guo, Wei Niu 0002, Xue Lin 0001, Jian Tang 0008, Kaisheng Ma, Bin Ren 0002, Yanzhi Wang 0001 |
AAAI | 6 |
| 2020 | Light-weight Calibrator: A Separable Component for Unsupervised Domain AdaptationabstractExisting domain adaptation methods aim at learning features that can be generalized among domains. These methods commonly require to update source classifier to adapt to the target domain and do not properly handle the trade-off between the source domain and the target domain. In this work, instead of training a classifier to adapt to the target domain, we use a separable component called data calibrator to help the fixed source classifier recover discrimination power in the target domain, while preserving the source domain's performance. When the difference between two domains is small, the source classifier's representation is sufficient to perform well in the target domain and outperforms GAN-based methods in digits. Otherwise, the proposed method can leverage synthetic images generated by GANs to boost performance and achieve state-of-the-art performance in digits datasets and driving scene semantic segmentation. Our method also empirically suggests the potential connection between domain adaptation and adversarial attacks. Shaokai Ye, Kailu Wu, Mu Zhou, Sia Huat Tan, Kaidi Xu, Jiebo Song, Chenglong Bao, Kaisheng Ma |
CVPR | 9 |
| 2020 | Auxiliary Training: Towards Accurate and Robust ModelsabstractTraining process is crucial for the deployment of the network in applications which have two strict requirements on both accuracy and robustness. However, most existing approaches are in a dilemma, i.e. model accuracy and robustness form an embarrassing tradeoff - the improvement of one leads to the drop of the other. The challenge remains as for we try to improve the accuracy and robustness simultaneously. In this paper, we propose a novel training method via introducing the auxiliary classifiers for training on corrupted samples, while the clean samples are normally trained with the primary classifier. In the training stage, a novel distillation method named input-aware self distillation is proposed to facilitate the primary classifier to learn the robust information from auxiliary classifiers. Along with it, a new normalization method - selective batch normalization is proposed to prevent the model from the negative influence of corrupted images. At the end of training period, a L2-norm penalty is applied to the weights of primary and auxiliary classifiers such that their weights are asymptotically identical. In the stage of inference, only the primary classifier is used and thus no extra computation and storage are needed. Extensive experiments on CIFAR10, CIFAR100 and ImageNet show that noticeable improvements on both accuracy and robustness can be observed by the proposed auxiliary training. On average, auxiliary training achieves 2.21% accuracy and 21.64% robustness (measured by corruption error) improvements over traditional training methods on CIFAR100. Codes has been released on github. Linfeng Zhang 0001, Muzhou Yu, Zuoqiang Shi, Chenglong Bao, Kaisheng Ma |
CVPR | 6 |
| 2020 | PCNN: Pattern-based Fine-Grained Regular Pruning Towards Optimizing CNN AcceleratorsabstractWeight pruning is a powerful technique to realize model compression. We propose PCNN, a fine-grained regular 1D pruning method. A novel index format called Sparsity Pattern Mask (SPM) is presented to encode the sparsity in PCNN. Leveraging SPM with limited pruning patterns and non-zero sequences with equal length, PCNN can be efficiently employed in hardware. Evaluated on VGG-16 and ResNet-18, our PCNN achieves the compression rate up to 8.4× with only 0.2% accuracy loss. We also implement a pattern-aware architecture in 55nm process, achieving up to 9.0× speedup and 28.39 TOPS/W efficiency with only 3.1% on-chip memory overhead of indices. Zhanhong Tan, Jiebo Song, Sia Huat Tan, Hongyang Chen 0003, Yuanqing Miao, Yifu Wu, Shaokai Ye, Yanzhi Wang 0001, Dehui Li, Kaisheng Ma |
DAC | 11 |
| 2020 | An Image Enhancing Pattern-Based Sparsity for Real-Time Inference on Mobile Devices
Wei Niu 0002, Tianyun Zhang, Sijia Liu 0001, Sheng Lin 0001, Hongjia Li 0003, Wujie Wen, Xiang Chen 0010, Jian Tang 0008, Kaisheng Ma, Bin Ren 0002, Yanzhi Wang 0001 |
ECCV (13) | 10 |
| 2020 | Design Insights of Non-volatile Processors and Accelerators in Energy Harvesting SystemsabstractThere is growing interest in deploying energy harvesting processors and accelerators in Internet of Things (IoT). Energy harvesting harnesses the energy scavenged from the environment to power a system. Although it has many advantages over battery-operated systems such as lightweight, compact size, and no necessity of recharging and maintenance, it may suffer frequently power-down and a fluctuating power supply even with power on. Non-volatile processor (NVP) is a promising architecture for effective computing in energy harvesting scenarios. Recently, non-volatile accelerators (NVA) have been proposed to perform computations of deep learning algorithms. In this paper, we overview the recent studies of NVP and NVA across the layers of hardware, architecture, software and their co-design. Especially, we present the design insights of how the state-of-the-art works adapt their specific designs to the intermittent and fluctuating power conditions with the energy harvesting technology. Finally, we discuss recent trends using NVP and NVA in energy harvesting scenarios. Keni Qiu, Mengying Zhao, Zhenge Jia, Jingtong Hu, Chun Jason Xue, Kaisheng Ma, Xueqing Li 0002, Yongpan Liu, Narayanan Vijaykrishnan |
ACM Great Lakes Symposium on VLSI | 6 |
| 2020 | Task-Oriented Feature DistillationabstractFeature distillation, a primary method in knowledge distillation, always leads to significant accuracy improvements. Most existing methods distill features in the teacher network through a manually designed transformation. In this paper, we propose a novel distillation method named task-oriented feature distillation (TOFD) where the transformation is convolutional layers that are trained in a data-driven manner by task loss. As a result, the task-oriented information in the features can be captured and distilled to students. Moreover, an orthogonal loss is applied to the feature resizing layer in TOFD to improve the performance of knowledge distillation. Experiments show that TOFD outperforms other distillation methods by a large margin on both image classification and 3D classification tasks. Codes have been released in Github. Linfeng Zhang 0001, Yukang Shi, Zuoqiang Shi, Kaisheng Ma, Chenglong Bao |
NeurIPS | 4 |
| 2019 | Adversarial Robustness vs. Model Compression, or Both?abstractIt is well known that deep neural networks (DNNs) are vulnerable to adversarial attacks, which are implemented by adding crafted perturbations onto benign examples. Min-max robust optimization based adversarial training can provide a notion of security against adversarial attacks. However, adversarial robustness requires a significantly larger capacity of the network than that for the natural training with only benign examples. This paper proposes a framework of concurrent adversarial training and weight pruning that enables model compression while still preserving the adversarial robustness and essentially tackles the dilemma of adversarial training. Furthermore, this work studies two hypotheses about weight pruning in the conventional setting and finds that weight pruning is essential for reducing the network model size in the adversarial setting; training a small model from scratch even with inherited initialization from the large model cannot achieve neither adversarial robustness nor high standard accuracy. Code is available at https://github.com/yeshaokai/Robustness-Aware-Pruning-ADMM. Shaokai Ye, Xue Lin 0001, Kaidi Xu, Sijia Liu 0001, Hao Cheng 0015, Jan-Henrik Lambrechts, Huan Zhang 0001, Aojun Zhou, Kaisheng Ma, Yanzhi Wang 0001 |
ICCV | 9 |
| 2019 | Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationabstractConvolutional neural networks have been widely deployed in various application scenarios. In order to extend the applications' boundaries to some accuracy-crucial domains, researchers have been investigating approaches to boost accuracy through either deeper or wider network structures, which brings with them the exponential increment of the computational and storage cost, delaying the responding time. In this paper, we propose a general training framework named self distillation, which notably enhances the performance (accuracy) of convolutional neural networks through shrinking the size of the network rather than aggrandizing it. Different from traditional knowledge distillation - a knowledge transformation methodology among networks, which forces student neural networks to approximate the softmax layer outputs of pre-trained teacher neural networks, the proposed self distillation framework distills knowledge within network itself. The networks are firstly divided into several sections. Then the knowledge in the deeper portion of the networks is squeezed into the shallow ones. Experiments further prove the generalization of the proposed self distillation framework: enhancement of accuracy at average level is 2.65%, varying from 0.61% in ResNeXt as minimum to 4.07% in VGG19 as maximum. In addition, it can also provide flexibility of depth-wise scalable inference on resource-limited edge devices. Our codes have been released on github. Linfeng Zhang 0001, Jiebo Song, Anni Gao, Chenglong Bao, Kaisheng Ma |
ICCV | 6 |
| 2019 | Gated Contiguous Memory U-Net for Single Image Dehazing
Hang Dong 0001, Fei Wang 0008, Yu Guo 0006, Kaisheng Ma |
ICONIP (2) | 5 |
| 2019 | SCAN: A Scalable Neural Networks Framework Towards Compact and Efficient ModelsabstractRemarkable achievements have been attained by deep neural networks in various applications. However, the increasing depth and width of such models also lead to explosive growth in both storage and computation, which has restricted the deployment of deep neural networks on resource-limited edge devices. To address this problem, we propose the so-called SCAN framework for networks training and inference, which is orthogonal and complementary to existing acceleration and compression methods. The proposed SCAN firstly divides neural networks into multiple sections according to their depth and constructs shallow classifiers upon the intermediate features of different sections. Moreover, attention modules and knowledge distillation are utilized to enhance the accuracy of shallow classifiers. Based on this architecture, we further propose a threshold controlled scalable inference mechanism to approach human-like sample-specific inference. Experimental results show that SCAN can be easily equipped on various neural networks without any adjustment on hyper-parameters or neural networks architectures, yielding significant performance gain on CIFAR100 and ImageNet. Codes will be released on github soon. Linfeng Zhang 0001, Zhanhong Tan, Jiebo Song, Chenglong Bao, Kaisheng Ma |
NeurIPS | 6 |
| 2018 | NEOFog: Nonvolatility-Exploiting Optimizations for Fog ComputingabstractNonvolatile processors have emerged as one of the promising solutions for energy harvesting scenarios, among which Wireless Sensor Networks (WSN) provide some of the most important applications. In a typical distributed sensing system, due to difference in location, energy harvester angles, power sources, etc. different nodes may have different amount of energy ready for use. While prior approaches have examined these challenges, they have not done so in the context of the features offered by nonvolatile computing approaches, which disrupt certain foundational assumptions. We propose a new set of nonvolatility-exploiting optimizations and embody them in the NEOFog system architecture. We discuss shifts in the tradeoffs in data and program distribution for nonvolatile processing-based WSNs, showing how non-volatile processing and non-volatile RF support alter the benefits of computation and communication-centric approaches. We also propose a new algorithm specific to nonvolatile sensing systems for load balancing both computation and communication demands. Collectively, the NV-aware optimizations in NEOFog increase the ability to perform in-fog processing by 4.2X and can increase this to 8X if virtualized nodes are 3X multiplexed. Kaisheng Ma, Xueqing Li 0002, Mahmut T. Kandemir, Jack Sampson, Narayanan Vijaykrishnan, Jinyang Li 0002, Tongda Wu, Zhibo Wang 0004, Yongpan Liu, Yuan Xie 0001 |
ASPLOS | 1 |
| 2018 | Dual Mode Ferroelectric Transistor based Non-Volatile Flip-Flops for Intermittently-Powered SystemsabstractIn this work, we propose dual mode ferroelectric transistors (D-FEFETs) that exhibit dynamic tuning of operation between volatile and non-volatile modes with the help of a control signal. We utilize the unique features of D-FEFET to design two variants of non-volatile flip-flops (NVFFs). In both designs, D-FEFETs are operated in the volatile mode for normal operations and in the non-volatile mode to backup the state of the flip-flop during a power outage. The first design comprises of a truly embedded non-volatile element (D-FEFET) which enables a fully automatic backup operation. In the second design, we introduce need-based backup, which lowers energy during normal operation at the cost of area with respect to the first design. Compared to a previously proposed FEFET based NVFF, the first design achieves 19% area reduction along with 96% lower backup energy and 9% lower restore energy, but at 14%-35% larger operation energy. The second design shows 11% lower area, 21% lower backup energy, 16% decrease in backup delay and similar operation energy but with a penalty of 17% and 19% in the restore energy and delay, respectively. System-level analysis of the proposed NVFFs in context of a state-of-the-art intermittently-powered system using real benchmarks yielded 5%-33% energy savings. Sandeep Krishna Thirumala, Arnab Raha, Hrishikesh Jayakumar, Kaisheng Ma, Narayanan Vijaykrishnan, Vijay Raghunathan, Sumeet Kumar Gupta |
ISLPED | 4 |
| 2018 | Symmetric 2-D-Memory Access to Multidimensional Data
Sumitha George, Xueqing Li 0002, Minli Julie Liao, Kaisheng Ma, Srivatsa Rangachar Srinivasa, Karthik Mohan, Ahmedullah Aziz, Jack Sampson, Sumeet Kumar Gupta, Narayanan Vijaykrishnan |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Spendthrift: Machine learning based resource and frequency scaling for ambient energy harvesting nonvolatile processorsabstractBatteryless energy harvesting systems face a twofold challenge in converting incoming energy into forward progress. Not only must such systems contend with inherently weak and fluctuating power sources, but they have very limited temporal windows for capitalizing on transitory periods of above-average power. To maximize forward progress, such systems should aggressively consume energy when it is available, rather than optimizing for peak averagecase efficiency. However, there are multiple ways that a processor can trade between consumption and performance. In this paper, we examine two approaches, frequency scaling and resource scaling, and develop a predictor-driven scheme for dynamically allocating future power budgets between the two techniques. We show that our solution can achieve forward progress equal to 2.08X of the baseline Out-of-Order (OoO) processor with the best static configuration of frequency and resources. The combined technique outperforms either technique in isolation, with frequency-only and resource-only approaches achieving 1.43X and 1.61X forward progress improvements, respectively. Kaisheng Ma, Xueqing Li 0002, Srivatsa Rangachar Srinivasa, Yongpan Liu, Jack Sampson, Yuan Xie 0001, Narayanan Vijaykrishnan |
ASP-DAC | 1 |
| 2017 | Nonvolatile processors: Why is it trending?abstractEnergy harvesting has become a promising solution to power up Internet-of-Things (IoT) devices. In this scenario, the constrained power budget and frequent absence of ambient energy cause severe reliability issues and performance degradation on conventional CMOS computing circuits. Fortunately, the advent of nonvolatile processor (NVP) opens the possibility to compute continuously using an intermittent power supply. It is considered as a key component of the next generation IoT edge devices. In this work, we provide insights to the evolution of the NVP and its application in real world scenarios. Efforts on improving the performance of NVP and future research prospects are also discussed in this paper. Fang Su, Kaisheng Ma, Xueqing Li 0002, Tongda Wu, Yongpan Liu, Narayanan Vijaykrishnan |
DATE | 2 |
| 2017 | Incidental computing on IoT nonvolatile processorsabstractBatteryless IoT devices powered through energy harvesting face a fundamental imbalance between the potential volume of collected data and the amount of energy available for processing that data locally. However, many such devices perform similar operations across each new input record, which provides opportunities for mining the potential information in buffered historical data, at potentially lower effort, while processing new data rather than abandoning old inputs due to limited computational energy. We call this approach incidental computing, and highlight synergies between this approach and approximation techniques when deployed on a non-volatile processor platform (NVP). In addition to incidental computations, the backup and restore operations in an incidental NVP provide approximation opportunities and optimizations that are unique to NVPs. Kaisheng Ma, Xueqing Li 0002, Jinyang Li 0002, Yongpan Liu, Yuan Xie 0001, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan |
MICRO | 1 |
| 2017 | Dynamic Power and Energy Management for Energy Harvesting Nonvolatile Processor SystemsabstractSelf-powered systems running on scavenged energy will be a key enabler for pervasive computing across the Internet of Things. The variability of input power in energy-harvesting systems limits the effectiveness of static optimizations aimed at maximizing the input-energy-to-computation ratio. We show that the resultant gap between available and exploitable energy is significant, and that energy storage optimizations alone do not significantly close the gap. We characterize these effects on a real, fabricated energy-harvesting system based on a nonvolatile processor. We introduce a unified energy-oriented approach to first optimize the number of backups, by more aggressively using the stored energy available when power failure occurs, and then optimize forward progress via improving the rate of input energy to computation via dynamic voltage and frequency scaling and self-learning techniques. We evaluate combining these schemes and show capture of up to 75.5% of all input energy toward processor computation, an average of 1.54 × increase over the best static “Forward Progress” baseline system. Notably, our energy-optimizing policy combinations simultaneously improve both the rate of forward progress and the rate of backup events (by up to 60.7% and 79.2% for RF power, respectively, and up to 231.2% and reduced to zero, respectively, for solar power). This contrasts with static frequency optimization approaches in which these two metrics are antagonistic. Kaisheng Ma, Xueqing Li 0002, Huichu Liu, Xiao Sheng, Karthik Swaminathan, Yongpan Liu, Yuan Xie 0001, Jack Sampson, Narayanan Vijaykrishnan |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | Nonvolatile memory design based on ferroelectric FETsabstractFerroelectric FETs (FEFETs) offer intriguing possibilities for the design of low power nonvolatile memories by virtue of their three-terminal structure coupled with the ability of the ferroelectric (FE) material to retain its polarization in the absence of an electric field. Utilizing the distinct features of FEFETs, we propose a 2-transistor (2T) FEFET-based nonvolatile memory with separate read and write paths. With proper co-design at the device, cell and array levels, the proposed design achieves non-destructive read and lower write power at iso-write speed compared to standard FERAM. In addition, the FEFET-based memory exhibits high distinguishability with six orders of magnitude difference in the read currents corresponding to the two states. Comparative analysis based on experimentally calibrated models shows significant improvement of access energy-delay. For example, at a fixed write time of 550ps, the write voltage and energy are 58.5% and 67.7% lower than FERAM, respectively. These benefits are achieved with 2.4 times the area overhead. Further exploration of the proposed FEFET memory in energy harvesting nonvolatile processors shows an average improvement of 27% in forward progress over FERAM. Sumitha George, Kaisheng Ma, Ahmedullah Aziz, Xueqing Li 0002, Asif Islam Khan, Sayeef S. Salahuddin, Meng-Fan Chang, Suman Datta, Jack Sampson, Sumeet Kumar Gupta, Narayanan Vijaykrishnan |
DAC | 2 |
| 2016 | Enabling Internet-of-Things: Opportunities brought by emerging devices, circuits, and architecturesabstractThe Internet-of-Things (IoT) has excited low-power design from device, circuits, to architectures levels. This paper talks about how recent emerging beyond-CMOS devices, such as tunnel field effect transistor (TFET), negative capacitance FET (NCFET), and phase transition devices (PTD), could extend the low-power design space to enable IoT applications with beyond-CMOS features. Xueqing Li 0002, Kaisheng Ma, Sumitha George, Jack Sampson, Narayanan Vijaykrishnan |
VLSI-SoC | 2 |
| 2015 | Ambient energy harvesting nonvolatile processors: from circuit to systemabstractEnergy harvesting is gaining more and more attentions due to its characteristics of ultra-long operation time without maintenance. However, frequent unpredictable power failures from energy harvesters bring performance and reliability challenges to traditional processors. Nonvolatile processors are promising to solve such a problem due to their advantage of zero leakage and efficient backup and restore operations. To optimize the nonvolatile processor design, this paper proposes new metrics of nonvolatile processors to consider energy harvesting factors for the first time. Furthermore, we explore the nonvolatile processor design from circuit to system level. A prototype of energy harvesting nonvolatile processor is set up and experimental results show that the proposed performance metric meets the measured results by less than 6.27% average errors. Finally, the energy consumption of nonvolatile processor is analyzed under different benchmarks. Yongpan Liu, Hehe Li, Xueqing Li 0002, Kaisheng Ma, Shuangchen Li, Meng-Fan Chang, Jack Sampson, Yuan Xie 0001, Jiwu Shu, Huazhong Yang |
DAC | 6 |
| 2015 | Architecture exploration for ambient energy harvesting nonvolatile processorsabstractEnergy harvesting has been widely investigated as a promising method of providing power for ultra-low-power applications. Such energy sources include solar energy, radio-frequency (RF) radiation, piezoelectricity, thermal gradients, etc. However, the power supplied by these sources is highly unreliable and dependent upon ambient environment factors. Hence, it is necessary to develop specialized systems that are tolerant to this power variation, and also capable of making forward progress on the computation tasks. The simulation platform in this paper is calibrated using measured results from a fabricated nonvolatile processor and used to explore the design space for a nonvolatile processor with different architectures, different input power sources, and policies for maximizing forward progress. Kaisheng Ma, Shuangchen Li, Karthik Swaminathan, Xueqing Li 0002, Yongpan Liu, Jack Sampson, Yuan Xie 0001, Narayanan Vijaykrishnan |
HPCA | 1 |
| 2015 | Dynamic Machine Learning Based Matching of Nonvolatile Processor Microarchitecture to Harvested Energy ProfileabstractEnergy harvesting systems without an energy storage device have to efficiently harness the fluctuating and weak power sources to ensure the maximum computational progress. While a simpler processor enables a higher turn-on potential with a weak source, a more powerful processor can utilize more energy that is harvested. Earlier work shows that different complexity levels of nonvolatile microarchitectures provide best fit for different power sources, and even different trails within same power source. In this work, we propose a dynamic nonvolatile microarchitecture by integrating all non-pipelined (NP), N-stage-pipeline (NSP), and Out of Order (OoO) cores together. Neural network machine learning algorithms are also integrated to dynamically adjust the microarchitecture to achieve the maximum forward progress. This integrated solution can achieve forward progress equal to 2.4× of the baseline NP architecture (1.82× of an OoO core). Kaisheng Ma, Xueqing Li 0002, Yongpan Liu, Jack Sampson, Yuan Xie 0001, Narayanan Vijaykrishnan |
ICCAD | 1 |
| 2015 | Key characterization factors of accurate power modeling for FinFET circuits
Kaisheng Ma, Xiaoxin Cui, Nan Liao, Dunshan Yu |
Sci. China Inf. Sci. | 1 |
| 2014 | Low power adiabatic logic based on FinFETs
Nan Liao, Xiaoxin Cui, Kaisheng Ma, Dunshan Yu |
Sci. China Inf. Sci. | 4 |
| 2014 | Ultra-low power dissipation of improved complementary pass-transistor adiabatic logic circuits based on FinFETs
Xiaoxin Cui, Nan Liao, Kaisheng Ma, Dunshan Yu |
Sci. China Inf. Sci. | 4 |
| 2013 | A Dynamic-Adjusting Threshold-Voltage Scheme for FinFETs low power designsabstractIn this paper, a novel device/circuit co-design scheme, namely Dynamic-Adjusting Threshold-Voltage Scheme (DATS) for independent-gate mode FinFET circuits has been proposed. The main idea of this scheme is that a pair of back-gate bias of FinFETs is adjusted dynamically to change threshold voltage according to the system operating frequency and operating mode, which could optimize circuit power, especially leakage power. The experimental and simulation result shows that the leakage power dissipation reduced greatly when circuits operate at the lower frequency, and the energy-delay product of FinFET circuits is reduced by 30% approximately. Xiaoxin Cui, Kaisheng Ma, Nan Liao, Dunshan Yu |
ISCAS | 2 |