Feng Min

dblp:88/3148 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AsymVPU: A Scalable and Area-Efficient Vector Architecture via Intra-Lane Asymmetry and Hierarchical Co-Design
abstract
Modern AI inference interleaves compute-dense matrix kernels with memory-sensitive element-wise and reduction operators. Symmetric RISC-V vector processors execute this mix with a uniform lane design, replicating heavy arithmetic units even when many instructions need only lightweight arithmetic. This paper presents AsymVPU, a hardware-software co-designed vector architecture that introduces fine-grained intra-lane asymmetry: one main processing element retains full FP64/FMA capability, while three auxiliary processing elements provide dense support for lightweight vector operations. AsymVPU couples this lane organization with a two-level hierarchical reduction network, a coordinated global load/store path, and compiler-inserted density hints that preserve the RVV programming abstraction. Implemented in SystemVerilog and evaluated in a same-flow 12 nm comparison against an Ara-derived symmetric baseline, AsymVPU achieves up to 2.4 × higher compute density for quantized workloads and 60% lower reduction latency, while retaining competitive performance on FMA-dominated kernels.
Junzhe Jing, Feng Min, Ying Wang 0001, Yinhe Han 0001
ACM Great Lakes Symposium on VLSI3
2025 Dadu-Corki: Algorithm-Architecture Co-Design for Embodied AI-powered Robotic Manipulation
abstract
Embodied AI robots have the potential to fundamentally improve the way human beings live and manufacture.Continued progress in the burgeoning field of using large language models to control robots depends critically on an efficient computing substrate, and this trend is strongly evident in manipulation tasks.In particular, today's computing systems for embodied AI robots for manipulation tasks are designed purely based on the interest of algorithm developers, where robot actions are divided into a discrete frame basis.Such an execution pipeline creates high latency and energy consumption.This paper proposes Corki, an algorithm-architecture co-design framework for real-time embodied AI-powered robotic manipulation applications.We aim to decouple LLM inference, robotic control, and data communication in the embodied AI robots' compute pipeline.Instead of predicting action for one single frame, * equal contribution.
Yiyang Huang 0002, Yuhui Hao, Bo Yu 0014, Yuxin Yang 0002, Feng Min, Yinhe Han 0001, Lin Ma 0002, Shaoshan Liu, Qiang Liu 0011, Yiming Gan
ISCA6
2025 Generalizable person re-identification method using bi-stream interactive learning with feature reconstruction
Feng Min, Yuhui Liu, Yixin Mao
Pattern Recognit.1
2025 A Data-Centric Software-Hardware Co-Designed Architecture for Large-Scale Graph Processing
abstract
Graph processing plays an important role in many practical applications. However, the inherent characteristics of graph processing, including random memory access and the low computation-to-communication ratio, make it difficult to efficiently execute on traditional computing architectures, such as CPUs and GPUs. Near-memory computing has the characteristics of low latency and high bandwidth. It is widely regarded as a promising direction for designing graph processing accelerators. However, the storage space of a single device cannot meet the demand of large-scale graph processing. Using multiple devices will bring lots of inter-device data transmission, which may counteract the benefits of near-memory computing. To fundamentally reduce the data transmission overhead, we propose a data-centric graph processing framework for systems with multiple near-memory computing devices. The framework uses a data-centric programming model as the software hardware interface. For software, we propose an optimized data flow and a heuristic multi-step weighted maximum matching algorithm to achieve efficient inter-device communication and ensure load balancing. For hardware, we design a data reuse driven task controller and a data type-aware on-chip memory, which can effectively improve the utilization of the on-chip memory. Compared with the two most recent near-memory graph accelerators, our framework significantly reduces energy consumption and inter-device communication.
Zerun Li, Xiaoming Chen 0003, Yuxin Yang 0002, Feng Min, Xiaoyu Zhang 0009, Yinhe Han 0001
IEEE Trans. Computers4
2024 MemSort: In-Memory Sorting Architecture
abstract
Sorting is one of the most fundamental operations in computer programming and used in countless algorithms. The performance of traditional von Neumann computers running sorting is limited by the bandwidth between memories and processors. Computing-in-memory (CiM) is a promising technology which has the potential to solve the “memory wall” bottleneck. CiM is suitable for data-intensive applications, and it is ideal for accelerating large-scale data sorting. In this paper, we propose a novel in-memory sorting accelerator, named MemSort, based on a proposed in-memory comparison array design based on emerging non-volatile devices. MemSort supports three sort operations including counting sort, merging sort, and the combination of counting sort and merging sort. We build a performance model for the combination sort which enables flexible allocation of resources under given constraints to meet the requirements of various applications for sorting. The evaluation results show that MemSort shows significant performance improvement and energy efficiency at both the system level and application level when processing large-scale data sorting. Compared with the CPU implementation, MemSort achieves energy savings of 19.69-72.75x and speedups of 24.48-38.58 x with the same power constraint. MemSort's throughput is at least 4.86 x higher than that of the recent FPGA-based sorting accelerator FANS. MemSort exhibits more than 11 x throughput and 4.03 x area efficiency, compared with the recent CiM - based sorting accelerator, RIME.
Rui Liu 0045, Xiaoyu Zhang 0009, Xinyu Wang 0040, Feng Min, Zhejian Luo, Xiaoming Chen 0003, Yinhe Han 0001, Minghua Tang
ICCD4
2024 Action recognition based on adaptive region perception
Tongwei Lu, Feng Min, Yanduo Zhang
Neural Comput. Appl.3
2023 Kora: A Cloud-Native Event Streaming Platform for Kafka
abstract
Event streaming is an increasingly critical infrastructure service used in many industries and there is growing demand for cloud-native solutions. Confluent Cloud provides a massive scale event streaming platform built on top of Apache Kafka with tens of thousands of clusters running in 70+ regions across AWS, Google Cloud, and Azure. This paper introduces Kora , the cloud-native platform for Apache Kafka at the core of Confluent Cloud. We describe Kora's design that enables it to meet its cloud-native goals, such as reliability, elasticity, and cost efficiency. We discuss Kora's abstractions which allow users to think in terms of their workload requirements and not the underlying infrastructure, and we discuss how Kora is designed to provide consistent, predictable performance across cloud environments with diverse capabilities.
Anna Povzner, Prince Mahajan, Jason Gustafson, Jun Rao, Ismael Juma, Feng Min, Shriram Sridharan, Nikhil Bhatia, Gopi K. Attaluri, Adithya Chandra, Stanislav Kozlovski, Rajini Sivaram, Lucas Bradstreet, Bob Barrett, Dhruvil Shah, David Jacot, David Arthur, Manveer Chawla, Ron Dagostino, Colin Mccabe, Manikumar Reddy Obili, Kowshik Prakasam, Jose Garcia Sancio, Alok Nikhil
Proc. VLDB Endow.6
2021 An effective weighted vector median filter for impulse noise reduction based on minimizing the degree of aggregation
abstract
Abstract Impulse noise is regarded as an outlier in the local window of an image. To detect noise, many proposed methods are based on aggregated distance, including spatially weighted aggregated distance, n nearest neighbour distance, local density, and angle‐weighted quaternion aggregated distance. However, these methods ignore the weight of each pixel or have limited adaptability. This study introduces the concept of degree of aggregation and proposes a weighting method to obtain the weight vector of the pixels by minimizing the degree of aggregation. The weight vector obtained gives larger components on the signal pixels than on the noisy pixels. Then it is fused with the aggregated distance to form a weighted aggregated distance that can reasonably characterise the noise and signal. The weighted aggregated distance, along with an adaptive segmentation method, can effectively detect the noise. To further enhance the effect of noise detection and removal, an adaptive selection strategy is incorporated to reduce the noise density in the local window. At last, noisy pixels detected are replaced with the weighted channel combination optimization values. The experimental results exhibit the validity of the proposed method by showing better performance in terms of both objective criteria and visual effects.
Tongwei Lu, Feng Min, Tao Lu 0001
IET Image Process.3
2021 Dadu-Eye: A 5.3 TOPS/W, 30 fps/1080p High Accuracy Stereo Vision Accelerator
abstract
Stereo vision is widely deployed on robots and drones to enable depth estimation at a low cost. The combination of lightweight deep neural network (DNN) and cost volumes algorithm is proved to possess the advantages of both high depth estimation accuracy and speed. However, currently there is no accelerator architecture compatible with both efficient DNN inference and cost generation algorithms such as stereo matching. This work proposes a stereo vision accelerator called Dadu-eye, dedicated to real-time processing of high-resolution image streams. The proposed architecture adopts a pipelined hardware design with the techniques of operation approximation and scheduling-level optimization. First, a cost estimation block is designed to generate cost volumes from both luminance and color information. Second, a super pipelined multiplication and accumulation array with a row scan-based fused-layer convolution scheduling is proposed to perform the encoding and decoding neural network efficiently. Finally, an optical flow block is designed and cooperates with the array to approximately predict half of the frames’ depth to achieve real-time (30fps) processing on 1080p view. Based on the SMIC 40 nm CMOS process, this stereo vision accelerator achieves 5.3 TOPS/W power efficiency and significantly reduces 81% off-chip memory access.
Feng Min, Ying Wang 0001, Xingqi Zou, Yinhe Han 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2016 An Image-Based Approach to Automatic Crop Organ Extraction via Low-Rank Matrix Recovery
abstract
Automatic extraction of crop organ from images is a crucial step for quantitatively acquiring crop growth information in precision agriculture. There has been some attempt on this task, but the performance is not satisfactory. In this paper, we proposed an image-based method based on low-rank matrix recovery to extract organ accurately. In our method, a crop image is considered to be compose of two factors: background and organ. In a certain feature space, the image is represented as a low-rank matrix plus sparse noises. The organ is then extracted by identifying the sparse noises when using low-rank matrix recovery algorithm. In order to ensure the rank of background is low, a linear transform for the feature space is introduced and needs to be learned from historical data. Dynamic threshold segmentation followed by vegetation removing techniques are ultimately adopted in the final step. The experimental results on the benchmark farmland dataset show that our method achieve competitive performance, compared with the other well-established methods, yielding the highest performance of 93.9% with the lowest standard deviation of 2.86%, which means our method is more robust and not sensitive to the complex environmental elements and different cultivars.
Zhenghong Yu, Haichang Yin, Haijie Feng, Minfang Chen, Huabing Zhou, Tongwei Lu, Feng Min
ISPDC7
2010 Automatic Face Replacement in Video Based on 2D Morphable Model
abstract
This paper presents an automatic face replacement approach in video based on 2D morphable model. Our approach includes three main modules: face alignment, face morph, and face fusion. Given a source image and target video, the Active Shape Models (ASM) is adopted to source image and target frames for face alignment. Then the source face shape is warped to match the target face shape by a 2D morphable model. The color and lighting of source face are adjusted to keep consistent with those of target face, and seamlessly blended in the target face. Our approach is fully automatic without user interference, and generates natural and realistic results.
Feng Min, Nong Sang, Zhefu Wang
ICPR1
2008 An empirical study of facial components classification by integrating dimensionality reduction and clustering
abstract
In this paper we present an empirical study of facial components classification by integrating dimensionality reduction and unsupervised clustering. The proposed framework contains two iterative steps: 1) Fixing cluster labels, the facial samples are projected onto lower dimensional subspace through dimensionality reduction method; 2) Fixing the subspace, the clustering algorithm is performed to generate cluster labels. Through iterative steps, clusters are discovered in the lower dimensional subspaces to avoid the curse of dimensionality, while the subspaces are adaptively re-adjusted for global optimality. In order to achieve an effective and robust system, we compare the PCA with the LLE for dimensionality reduction, and compare K-means with LDA-guided K-means(LDA-Km) for unsupervised clustering. The quantitative experimental results prove the LLE and LDA-Km are superior for facial data on public Lotus Hill Institute(LHI) dataset. We also apply the presented study to improve the portrait sketching results.
Feng Min, Nong Sang
ICPR1
2007 A Multi-Resolution Dynamic Model for Face Aging Simulation
abstract
In this paper we present a dynamic model for simulating face aging process. We adopt a high resolution grammatical face model[1] and augment it with age and hair features. This model represents all face images by a multi-layer And-Or graph and integrates three most prominent aspects related to aging changes: global appearance changes in hair style and shape, deformations and aging effects of facial components, and wrinkles appearance at various facial zones. Then face aging is modeled as a dynamic Markov process on this graph representation which is learned from a large dataset. Given an input image, we firstly compute the graph representation, and then sample the graph structures over various age groups according to the learned dynamic model. Finally we generate new face images with the sampled graphs. Our approach has three novel aspects: (1) the aging model is learned from a dataset of 50,000 adult faces at different ages; (2) we explicitly model the uncertainty in face aging and can sample multiple plausible aged faces for an input image; and (3) we conduct a simple human experiment to validate the simulated aging process.
Jin-Li Suo, Feng Min, Song-Chun Zhu, Shiguang Shan, Xilin Chen 0001
CVPR2