EDBT 2026 Demo / reviewers in the wild / expert
Keyi Zhang
dblp:164/6584
· DBLP profile ↗
21ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 2 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FilmSceneDesigner: Chaining Set Design for Procedural Film Scene GenerationabstractFilm set design plays a pivotal role in cinematic storytelling and shaping the visual atmosphere. However, the traditional process depends on expert-driven manual modeling, which is labor-intensive and time-consuming. To address this issue, we introduce FilmSceneDesigner, an automated scene generation system that emulates professional film set design workflow. Given a natural language description, including scene type, historical period, and style, we design an agent-based chaining framework to generate structured parameters aligned with film set design workflow, guided by prompt strategies that ensure parameter accuracy and coherence. On the other hand, we propose a procedural generation pipeline which executes a series of dedicated functions with the structured parameters for floorplan and structure generation, material assignment, door and window placement, and object retrieval and layout, ultimately constructing a complete film scene from scratch. Moreover, to enhance cinematic realism and asset diversity, we construct SetDepot-Pro, a curated dataset of 6,862 film-specific 3D assets and 733 materials. Experimental results and human evaluations demonstrate that our system produces structurally sound scenes with strong cinematic fidelity, supporting downstream tasks such as virtual previs, construction drawing and mood board creation. Keyi Zhang, Yiye Yan, Jiting Zhou |
AAAI | 2 |
| 2026 | Non-Linear Design Epistemology - - Future Archaeology: Toward a Study of the Variability of Spatial Perception under Dynamic Civilizational ChangeabstractThis research investigates how emerging computational technologies reconfigure architectural epistemology, shifting architecture from a fixed representation of history toward a dynamic and interpretive medium. This work focuses on how AI-mediated systems enable non-linear and multi-temporal interpretations of spatial perception, fostering reflective and multi-dimensional engagement with public space for both designers and the public. With this framework, interactive systems and workshops are developed around the historical evolution, contemporary condition, and speculative futures of the prison as a spatial typology. Through AI-generated narratives, tangible interaction, and augmented reality, participants engage with shifting spatial and social meanings as coexisting temporal constructs. This framework supports multiple, coexisting interpretations of spatial and social meaning, enabling a non-linear encounter and visual expression with the past, present, and speculative futures of a given public spatial node. This study investigates (1) how technologies (computation and AI) support Non-Linear Design Epistemology, redefining architectural cognition as continuous, adaptive, and generative through comprehensive literature review; (2) how embedding meaning of spatiality evolve along the evolution of architectural typology’s spatial structure and social meaning through site research and case study; (3) how co-interactive workshops and installations through combining role-playing, tangible engagement, and AI narratives can reveal novel insights into spatial cognition and collaborative sensemaking; (4) and extends toward how technology-mediated experiences can be applied to actively shape the structure and spatial expression of architectural entities in accordance with Non-Linear Design Epistemology. Collectively, the research demonstrates how computationally empowered frameworks can expand architectural theory, mediate human perception, and inform the design of immersive, interactive, and AI-enhanced engagements with heritage and social space. Keyi Zhang |
Creativity & Cognition | 1 |
| 2026 | Onycho Lab: Designing Wearable Fingertip Prostheses for Feminist Hapticism to Reframe Artificial Nails as Bodily TechnologyabstractWe present Onycho Lab, a performative “manicure surgery” reframing artificial nails as bodily technology. Staged as a temporary clinic, participants book appointments, complete consent forms, and undergo a tactile procedure while an overhead camera projects the close-up transformation. The system includes a wearable fingertip prosthesis that senses pressure and delivers vibrotactile feedback via a soft bracelet. Across a three-day exhibition (58 procedures), fingertip micro-augmentations redirected attention from visual display toward somatic memory and negotiated agency, with implications for feminist HCI and soma design. Keyi Zhang, Patrick Hartono |
Creativity & Cognition | 1 |
| 2025 | NUPEA: Optimizing Critical Loads on Spatial Dataflow Architectures via Non-Uniform Processing-Element AccessabstractData movement is the dominant energy, performance, and scalability bottleneck in modern architectures.Systems have tackled data movement by distributing data, e.g., via non-uniform memory access (NUMA) architectures.However, to reduce data movement, these architectures must identify critical data and place it closer to compute.Clever data placement is complex and often ineffective.Spatial dataflow architectures (SDAs) present a new opportunity to tackle data movement.SDAs distribute program instructions across a spatial fabric of processing elements (PEs).On large SDAs, some PEs are necessarily closer to memory than others, giving rise to non-uniform processing-element access (NUPEA).Clever instruction placement can thus reduce data movement by, e.g., placing critical loads close to memory.This paper introduces NUPEA and contrasts it with prior datacentric approaches to scaling data movement.We find that it is often easier for the compiler to identify critical loads than the data they access, making NUPEA applicable where NUMA is not.We present simple architecture and compiler optimizations for NUPEA and implement them on the Monaco SDA architecture and effcc compiler, both industry products by Efficient Computer.On Monaco, across a range of important kernels, NUPEA yields an avg 28% speedup over a uniform-PE-access (UPEA) SDA and an avg 20% speed over a UPEA SDA with NUMA. Souradip Ghosh, Graham Gobieski, Keyi Zhang, Brandon Lucia, Nathan Beckmann, Tony Nowatzki |
ISCA | 3 |
| 2024 | Onyx: A Programmable Accelerator for Sparse Tensor Algebraabstract•Applications ranging from scientific computing to machine learning can have extremely sparse inputs Kalhan Koul, Maxwell Strange, Jackson Melchert, Alex Carsello, Yuchen Mei, Olivia Hsu, Taeyoung Kong, Huifeng Ke, Keyi Zhang, Qiaoyi Liu, Gedeon Nyengele, Akhilesh Balasingam, Jayashree Adivarahan, Ritvik Sharma, Zhouhua Xie, Christopher Torng, Joel S. Emer, Fredrik Kjolstad, Mark Horowitz, Priyanka Raina |
HCS | 10 |
| 2024 | Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing
Jiatong Shi, Yueqian Lin, Xinyi Bai, Keyi Zhang, Yuning Wu 0001, Yuxun Tang, Qin Jin, Shinji Watanabe 0001 |
INTERSPEECH | 4 |
| 2024 | Quasi-Fractal UCA-Based OAM for Highly Efficient Orthogonal TransmissionabstractThe development of orbital angular momentum (OAM)-based radio vortex transmission presents a promising opportunity for increasing the capacity of wireless communication in correlated channels due to its inherent orthogonality among different OAM modes. One of the most popular schemes for high-efficient OAM transmission is the digital baseband associated with uniform circular array (UCA) based transceiver. However, the periodicity of complex-exponential feed makes the maximum number of orthogonal signals carried by multiple OAM modes generally restricted to the array-element number of UCA antenna, which poses an open question of how to employ more OAM modes given a fixed number of array elements. Furthermore, signals modulated with high-order OAM modes are difficult to be captured by the receiver due to their serious divergence as propagating in free space, thus severely limiting the capacity of radio vortex communications. To overcome the above challenges, in this paper based on the partly element-overlapped fractal geometry layout and effectively using low-order OAM modes, we propose the quasi-fractal UCA (QF-UCA) antenna based OAM multiplexing transmission. We perform the two-dimension OAM modulation (TOM) and demodulation (TOD) schemes with the orthogonal OAM mode number exceeding the array-element number, which is beyond the traditional concept of multiple antennas based wireless communications. Simulation results show that our proposed scheme can achieve more number of orthogonal multiplexing streams than the maximum number of orthogonal multiplexing corresponding to traditional multiple antenna systems. Wenchi Cheng, Haiyue Jing, Wei Zhang 0001, Keyi Zhang, Hailin Zhang 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | AMDNet: Adaptive Fall Detection Based on Multi-scale Deformable Convolution Network
Minghua Jiang, Keyi Zhang, Yongkang Ma, Li Liu 0047, Tao Peng 0006, Xinrong Hu, Feng Yu 0017 |
CGI (3) | 2 |
| 2023 | AHA: An Agile Approach to the Design of Coarse-Grained Reconfigurable Accelerators and CompilersabstractWith the slowing of Moore’s law, computer architects have turned to domain-specific hardware specialization to continue improving the performance and efficiency of computing systems. However, specialization typically entails significant modifications to the software stack to properly leverage the updated hardware. The lack of a structured approach for updating the compiler and the accelerator in tandem has impeded many attempts to systematize this procedure. We propose a new approach to enable flexible and evolvable domain-specific hardware specialization based on coarse-grained reconfigurable arrays (CGRAs). Our agile methodology employs a combination of new programming languages and formal methods to automatically generate the accelerator hardware and its compiler from a single source of truth. This enables the creation of design-space exploration frameworks that automatically generate accelerator architectures that approach the efficiencies of hand-designed accelerators, with a significantly lower design effort for both hardware and compiler generation. Our current system accelerates dense linear algebra applications but is modular and can be extended to support other domains. Our methodology has the potential to significantly improve the productivity of hardware-software engineering teams and enable quicker customization and deployment of complex accelerator-rich computing systems. Kalhan Koul, Jackson Melchert, Kavya Sreedhar, Leonard Truong, Gedeon Nyengele, Keyi Zhang, Qiaoyi Liu, Jeff Setter, Yuchen Mei, Maxwell Strange, Ross Daly, Caleb Donovick, Alex Carsello, Taeyoung Kong, Kathleen Feng, Dillon Huff, Ankita Nayak, Rajsekhar Setaluri, James Thomas 0003, Nikhil Bhagdikar, David Durst, Zachary A. Myers, Nestan Tsiskaridze, Stephen Richardson, Rick Bahr, Kayvon Fatahalian, Pat Hanrahan, Clark W. Barrett, Mark Horowitz, Christopher Torng, Fredrik Kjolstad, Priyanka Raina |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2023 | Improving Energy Efficiency of CGRAs with Low-Overhead Fine-Grained Power DomainsabstractTo effectively minimize static power for a wide range of applications, power domains for coarse-grained reconfigurable array (CGRA) architectures need to be more fine-grained than those found in a typical application-specific integrated circuit. However, the special isolation logic needed to ensure electrical protection between off and on domains makes fine-grained power domains area- and timing-inefficient. We propose a novel design of the CGRA routing fabric that reduces the area overhead of power domain boundary protection from around 9% to less than 1% without incurring any extra timing delay from the isolation cells. Conventional Unified Power Format based flow for power domain boundary protection does not support this design choice. Therefore, we create our own compiler-like passes that iteratively introduce the needed design changes, and formally verify the transformations using methods based on satisfiability modulo theories. These passes also let us optimize how we handle test and debug signals through the off tiles in the CGRA. Using our framework, we add power domains to a CGRA that we designed and taped out. The CGRA has 32 × 16 processing element and memory tiles and 4-MB secondary memory. We address the implementation challenges encountered due to the introduction of fine-grained power domains, including the addressing of the CGRA tiles, the power grid design, well substrate connections, and distribution of global signals. Our CGRA achieves up to 83% reduction in leakage power and 26% reduction in total power versus an identical CGRA without multiple power domains, for a range of image processing and machine learning applications. Ankita Nayak, Keyi Zhang, Rajsekhar Setaluri, Alex Carsello, Makai Mann, Christopher Torng, Stephen Richardson, Rick Bahr, Pat Hanrahan, Mark Horowitz, Priyanka Raina |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | Bringing source-level debugging frameworks to hardware generatorsabstractHigh-level hardware generators have significantly increased the productivity of design engineers. They use software engineering constructs to reduce the repetition required to express complex designs and enable more composability. However, these benefits are undermined by a lack of debugging infrastructure, requiring hardware designers to debug generated, usually incomprehensible, RTL code. This paper describes a framework that connects modern software source-level debugging frameworks to RTL created from hardware generators. Our working prototype offers an Integrated Development Environment (IDE) experience for generators such as RocketChip (Chisel), allowing designers to set breakpoints in complex source code, relate RTL simulation state back to source-level variables, and do forward and backward debugging, with almost no simulation overhead (less than 5%). Keyi Zhang, Zain Asgar, Mark Horowitz |
DAC | 1 |
| 2022 | Amber: Coarse-Grained Reconfigurable Array-Based SoC for Dense Linear Algebra AccelerationabstractDedicated hardware accelerators popular for imaging, vision, and machine learning (ML) applications Kathleen Feng, Alex Carsello, Taeyoung Kong, Kalhan Koul, Qiaoyi Liu, Jackson Melchert, Gedeon Nyengele, Maxwell Strange, Keyi Zhang, Ankita Nayak, Jeff Setter, James Thomas 0003, Kavya Sreedhar, Nikhil Bhagdikar, Zachary A. Myers, Brandon D'Agostino, Pranil Joshi, Stephen Richardson, Rick Bahr, Christopher Torng, Mark Horowitz, Priyanka Raina |
HCS | 9 |
| 2020 | fault: A Python Embedded Domain-Specific Language for Metaprogramming Portable Hardware Verification ComponentsabstractWhile hardware generators have drastically improved design productivity, they have introduced new challenges for the task of verification. To effectively cover the functionality of a sophisticated generator, verification engineers require tools that provide the flexibility of metaprogramming. However, flexibility alone is not enough; components must also be portable in order to encourage the proliferation of verification libraries as well as enable new methodologies. This paper introduces fault , a Python embedded hardware verification language that aims to empower design teams to realize the full potential of generators. Leonard Truong, Steven Herbst, Rajsekhar Setaluri, Makai Mann, Ross Daly, Keyi Zhang, Caleb Donovick, Daniel Stanley, Mark Horowitz, Clark W. Barrett, Pat Hanrahan |
CAV (1) | 6 |
| 2020 | Creating an Agile Hardware Design FlowabstractAlthough an agile approach is standard for software design, how to properly adapt this method to hardware is still an open question. This work addresses this question while building a system on chip (SoC) with specialized accelerators. Rather than using a traditional waterfall design flow, which starts by studying the application to be accelerated, we begin by constructing a complete flow from an application expressed in a high-level domain-specific language (DSL), in our case Halide, to a generic coarse-grained reconfigurable array (CGRA). As our under-standing of the application grows, the CGRA design evolves, and we have developed a suite of tools that tune application code, the compiler, and the CGRA to increase the efficiency of the resulting implementation. To meet our continued need to update parts of the system while maintaining the end-to-end flow, we have created DSL-based hardware generators that not only provide the Verilog needed for the implementation of the CGRA, but also create the collateral that the compiler/mapper/place and route system needs to configure its operation. This work provides a systematic approach for desiging and evolving high-performance and energy-efficient hardware-software systems for any application domain. Rick Bahr, Clark W. Barrett, Nikhil Bhagdikar, Alex Carsello, Ross Daly, Caleb Donovick, David Durst, Kayvon Fatahalian, Kathleen Feng, Pat Hanrahan, Teguh Hofstee, Mark Horowitz, Dillon Huff, Fredrik Kjolstad, Taeyoung Kong, Qiaoyi Liu, Makai Mann, Jackson Melchert, Ankita Nayak, Aina Niemetz, Gedeon Nyengele, Priyanka Raina, Stephen Richardson, Rajsekhar Setaluri, Jeff Setter, Kavya Sreedhar, Maxwell Strange, James Thomas 0003, Christopher Torng, Leonard Truong, Nestan Tsiskaridze, Keyi Zhang |
DAC | 32 |
| 2020 | A Framework for Adding Low-Overhead, Fine-Grained Power Domains to CGRAsabstractTo effectively minimize static power for a wide range of applications, power domains for a coarse-grained reconfigurable array (CGRA) need to be finer-grained than a typical ASIC. However, the special isolation logic needed to ensure electrical protection between off and on domains makes fine-grained power domains area- and timing-inefficient. We propose a novel design of the CGRA routing fabric that intrinsically provides boundary protection. This technique reduces the area overhead of boundary protection between power domains for the CGRA from around 9% to less than 1% and removes the delay from the isolation cells. However, with this design choice, we cannot leverage the conventional UPF-based flow to introduce power domain boundary protection. We create compiler-like passes that iteratively introduce the needed design transformations, and formally verify the passes with satisfiability modulo theories (SMT) methods. These passes also allow us to optimize how we handle test and debug signals through the off tiles. We use our framework to insert power domains into an SoC with an ARM Cortex M3 processor and a CGRA with 32 × 16 processing element (PE) and memory tiles and 4MB secondary memory. Depending on the size of the applications mapped, our CGRA achieves up to an 83% reduction in leakage power and 26% reduction in total power versus a CGRA without multiple power domains, for a range of image processing and machine learning applications. Ankita Nayak, Keyi Zhang, Rajsekhar Setaluri, Alex Carsello, Makai Mann, Stephen Richardson, Rick Bahr, Pat Hanrahan, Mark Horowitz, Priyanka Raina |
DATE | 2 |
| 2020 | Learning in situ: a randomized experiment in video streaming
Francis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi, Keyi Zhang, Philip Alexander Levis, Keith Winstein |
NSDI | 6 |
| 2019 | Sequence Scrambling for Non-Hollow-OAM Based Wireless CommunicationsabstractOrbital angular momentum (OAM), which is inherently possessed by the electromagnetic (EM) beams, bridges a new way for multiple access using multiple orthogonal OAM-modes in wireless communications. Uniform circular array (UCA), as a convenient antenna structure for supporting the transmission and reception of OAM based signals, has been paid much attention in recent years. However, the OAM beams generated by the UCAs are centrally hollow and divergent, limiting the efficient reception as well as the long-distance transmission for OAM based signals. To solve this problem, in this paper we propose the Zadoff-Chu (ZC) sequence based scrambling scheme to generate the non-hollow OAM beams while maintaining the orthogonality among different OAM-modes. Then, based on the properties of applying discrete Fourier transform (DFT) for circulant matrices and zero-free sequences, we develop the scrambling-tolerant detection scheme to efficiently obtain the transmit symbols. Performance evaluations show the non-hollow irregular OAM beams and validate that our developed ZC sequence based scrambling scheme works well for long-distance transmission in OAM based wireless communications. Keyi Zhang, Wenchi Cheng, Runyu Lyu, Wei Zhang 0001, Hailin Zhang 0001, Fan Qin 0002 |
ICC | 1 |
| 2019 | Online Model Distillation for Efficient Video InferenceabstractHigh-quality computer vision models typically address the problem of understanding the general distribution of real-world images. However, most cameras observe only a very small fraction of this distribution. This offers the possibility of achieving more efficient inference by specializing compact, low-cost models to the specific distribution of frames observed by a single camera. In this paper, we employ the technique of model distillation (supervising a low-cost student model using the output of a high-cost teacher) to specialize accurate, low-cost semantic segmentation models to a target video stream. Rather than learn a specialized student model on offline data from the video stream, we train the student in an online fashion on the live video, intermittently running the teacher to provide a target for learning. Online model distillation yields semantic segmentation models that closely approximate their Mask R-CNN teacher with 7 to 17× lower inference runtime cost (11 to 26× in FLOPs), even when the target video's distribution is non-stationary. Our method requires no offline pretraining on the target video stream, achieves higher accuracy and lower cost than solutions based on flow or video object segmentation, and can exhibit better temporal stability than the original teacher. We also provide a new video dataset for evaluating the efficiency of inference over long running video streams. Ravi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan, Kayvon Fatahalian |
ICCV | 3 |
| 2018 | Playing SNES Games With NeuroEvolution of Augmenting TopologiesabstractTeaching a computer to play video games has generally been seen as a reasonable benchmark for developing new AI techniques. In recent years, extensive research has been completed to develop reinforcement learning (RL) algorithms to play various Atari 2600 games, resulting in new applications of algorithms such as Deep Q-Learning or Policy Gradient that outperform humans. However, games from Super Nintendo Entertainment System (SNES) are far more complicated than Atari 2600 games as many of these state-of-the-art algorithms still struggle to perform on this platform. In this paper, we present a new platform to research algorithms on SNES games and investigate NeuroEvolution of Augmenting Topologies (NEAT) as a possible approach to develop algorithms that outperform humans in SNES games. Son Pham, Keyi Zhang, Tung Phan, Jasper Ding, Christopher L. Dancy |
AAAI | 2 |
| 2017 | Crowdsourcing low-power wide-area IoT networksabstractThe Internet of Things (IoT) promises to allow everyday objects to connect to the Internet and seamlessly interact with users and other machines. For this vital Internet connection, most current IoT devices use a personal gateway device such as a smartphone or a home WiFi access point. The necessity of configuring and maintaining these gateways presents an additional burden for both users and developers of IoT applications. Our vision for IoT connectivity is to eliminate the need for the personal gateway by developing crowdsourced low-power wide area networks (csLPWAN). Recent technologies such as RPMA, LoRa, and R-FDMA enable links to reach 15km using ISM-band transceivers, making csLPWANs an attractive option. In this paper, we investigate the practicality of csLPWANs and develop the first csLPWAN planning tool, PlanIt, which combines topography-aware RF signal analysis with demographic data to predict LPWAN coverage in specific geographic areas. Using PlanIt, we find that most cities achieve 99% coverage by deploying a single LPWAN base station within the city. To provide better service on the csLPWAN, we propose and evaluate D-QN, a near-optimal MAC protocol to enable efficient bandwidth sharing in highly utilized networks. In the future, csLPWANs could accommodate a heterogeneous set of IoT applications, simplifying the IoT application development cycle, reducing total system cost, improving application reliability, and enhancing the user experience. Keyi Zhang, Alan Marchiori |
PerCom | 1 |
| 2016 | New Filter Design for Static Neural Networks with Mixed Time-Varying Delays
Guoquan Liu, Shumin Zhou, Xianxi Luo, Keyi Zhang |
ICIC (2) | 4 |