EDBT 2026 Demo / reviewers in the wild / expert
Fengyu Sun
dblp:52/10863
· DBLP profile ↗
15ranked-venue papers
1as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and GenerationabstractRecent research explores the potential of Diffusion Models (DMs) for consistent object editing, which aims to modify object position, size, and composition, etc., while preserving the consistency of objects and background without changing their texture and attributes. Current inference-time methods often rely on DDIM inversion, which inherently compromises efficiency and the achievable consistency of edited images. Recent methods also utilize energy guidance which iteratively updates the predicted noise and can drive the latents away from the original image, resulting in distortions. In this paper, we propose PixelMan, an inversion-free and training-free method for achieving consistent object editing via Pixel Manipulation and generation, where we directly create a duplicate copy of the source object at target location in the pixel space, and introduce an efficient sampling approach to iteratively harmonize the manipulated object into the target location and inpaint its original location, while ensuring image consistency by anchoring the edited image to be generated to the pixel-manipulated image as well as by introducing various consistency-preserving optimization techniques during inference. Experimental evaluations based on benchmark datasets as well as extensive visual comparisons show that in as few as 16 inference steps, PixelMan outperforms a range of state-of-the-art training-based and training-free methods (usually requiring 50 steps) on multiple consistent object editing tasks. Liyao Jiang, Negar Hassanpour, Mohammad Salameh, Mohammadreza Samadi, Jiao He, Fengyu Sun, Di Niu 0002 |
AAAI | 6 |
| 2025 | FunEditor: Achieving Complex Image Edits via Function Aggregation with Diffusion ModelsabstractDiffusion models have demonstrated outstanding performance in generative tasks, making them ideal candidates for image editing. Recent studies highlight their ability to apply desired edits effectively by following textual instructions, yet with two key challenges remaining. First, these models struggle to apply multiple edits simultaneously, resulting in computational inefficiencies due to their reliance on sequential processing. Second, relying on textual prompts to determine the editing region can lead to unintended alterations to the image. We introduce FunEditor, an efficient diffusion model designed to learn atomic editing functions and perform complex edits by aggregating simpler functions. This approach enables complex editing tasks, such as object movement, by aggregating multiple functions and applying them simultaneously to specific areas. Our experiments demonstrate that FunEditor significantly outperforms recent inference-time optimization methods and fine-tuned models, either quantitatively across various metrics or through visual comparisons or both, on complex tasks like object movement and object pasting. In the meantime, with only 4 steps of inference, FunEditor achieves 5--24 times inference speedups over existing popular methods. Mohammadreza Samadi, Fred X. Han, Mohammad Salameh, Fengyu Sun, Chunhua Zhou, Di Niu 0002 |
AAAI | 5 |
| 2025 | LPerceptual Quality Assessment of AI Generated Content Videos: a Dataset and BenchmarkabstractIn recent years, artificial intelligence (AI) driven video generation has garnered significant attention due to advancements in large language model techniques. Thus, there is a great demand to explore the effectiveness of video quality assessment (VQA) models in evaluating the perceptual quality of AI-generated content (AIGC) videos and in optimizing video generation techniques. Therefore, in this paper, we try to systemically investigate the AIGC-VQA problem from both subjective and objective quality assessment perspectives. For the subjective perspective, we construct a Large-scale Generated Video Quality assessment (LGVQ) dataset, consisting of 2,808 AIGC videos generated by 6 video generation models using 468 carefully selected text prompts. We evaluate the perceptual quality of AIGC videos from three dimensions: spatial quality, temporal quality, and text-to-video alignment, which hold the utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset, which fully demonstrates the performance of current mainstream VQA methods in evaluating AIGV quality. We hope that this work can contribute to the advancement of AIGC video generation technology as well as the evaluation techniques for AIGC videos. The LGVQ dataset will release publicly. Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ISCAS | 10 |
| 2025 | Human-Activity AGV Quality Assessment: A Benchmark Dataset and an Objective Evaluation MetricabstractAI-driven video generation techniques have made significant progress in recent years. However, AI-generated videos (AGVs) involving human activities often exhibit substantial visual and semantic distortions, hindering the practical application of video generation technologies in real-world scenarios. To address this challenge, we conduct a pioneering study on human activity AGV quality assessment, focusing on visual quality evaluation and the identification of semantic distortions. First, we construct the AI-Generated Human activity Video Quality Assessment (Human-AGVQA) dataset, consisting of 6,000 AGVs derived from 15 popular text-to-video (T2V) models using 400 text prompts that describe diverse human activities. We conduct a subjective study to evaluate the human appearance quality, action continuity quality, and overall video quality of AGVs, and identify semantic issues of human body parts. Based on Human-AGVQA, we benchmark the performance of T2V models and analyze their strengths and weaknesses in generating different categories of human activities. Second, we develop an objective evaluation metric, named AI-Generated Human activity Video Quality metric (GHVQ), to automatically analyze the quality of human activity AGVs. GHVQ systematically extracts human-focused quality features, AI-generated content-aware quality features, and temporal continuity features, making it a comprehensive and explainable quality metric for human activity AGVs. The extensive experimental results show that GHVQ outperforms existing quality metrics on the Human-AGVQA dataset by a large margin, demonstrating its efficacy in assessing the quality of human activity AGVs. The Human-AGVQA dataset and GHVQ metric will be released at https://github.com/zczhang-sjtu/GHVQ.git. Wei Sun 0029, Xinyue Li 0001, Qihang Ge, Jun Jia, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 9 |
| 2025 | Physics-informed neural networks for the prediction of robot dynamics considering motor and external force couplingsabstractIn recent years, physics-informed neural networks (PINNs) have shown remarkable potential in modeling conservative systems of rigid-body dynamics. However, when applied to practical interaction tasks of manipulators (e.g., part assembly and medical operations), existing PINN frameworks lack effective external force modeling mechanisms, resulting in significantly degraded prediction accuracy in dynamic interaction scenarios. Additionally, because industrial robots (including UR5 and UR10e robots) are generally not equipped with joint torque sensors, obtaining precise dynamics training data remains challenging. To address these issues, this study proposes two enhanced PINNs that integrate motor dynamics and external force modeling. First, two data-driven Jacobian matrix estimation methods are introduced to incorporate external forces: one learns the mapping between end-effector velocity and joint velocity to approximate the Jacobian matrix, while the other first learns the system’s kinematic behavior and then derives the Jacobian matrix through analytical differentiation of the forward kinematics model. Second, current-to-torque mapping is embedded as physical prior knowledge to establish direct correlations between system motion states and motor currents. Experimental results on two different manipulators demonstrate that both models achieve high-precision torque estimation in complex external force scenarios without requiring joint torque sensors. Compared with state-of-the-art methods, the proposed models improve overall modeling accuracy by 31.12% and 37.07% on average across various complex scenarios, while reducing joint trajectory tracking errors by 40.31% and 51.79%, respectively. Fengyu Sun, Peilin Xiong |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2025 | LMM-VQA: Advancing Video Quality Assessment With Large Multimodal ModelsabstractThe explosive growth of videos on streaming media platforms has underscored the urgent need for effective video quality assessment (VQA) algorithms to monitor and perceptually optimize the quality of streaming videos. However, VQA remains an extremely challenging task due to the diverse video content and the complex spatial and temporal distortions, thus necessitating more advanced methods to address these issues. Nowadays, large multimodal models (LMMs), such as GPT-4V, have exhibited strong capabilities for various visual understanding tasks, motivating us to leverage the powerful multimodal representation ability of LMMs to solve the VQA task. Therefore, we propose anLargeMulti-Modal basedVideoQuality Assessment (LMM-VQA) model, which introduces a novel spatiotemporal visual modeling strategy for quality-aware feature extraction. Specifically, we reformulate the quality regression problem into a question and answering (Q&A) task and construct Q&A prompts for VQA instruction tuning. Then, we design a spatiotemporal vision encoder to extract spatial and temporal features to represent the quality characteristics of videos, which are subsequently mapped into the language space by the spatiotemporal projector for modality alignment. Finally, the aligned visual tokens and the quality-inquired text tokens are aggregated as inputs for the large language model (LLM) to generate the quality score as well as the quality level. Extensive experiments demonstrate thatLMM-VQAachieves state-of-the-art performance across five VQA benchmarks, exhibiting an average improvement of 5% in generalization ability over existing methods. Furthermore, due to the advanced design of the spatiotemporal encoder and projector, LMM-VQA also performs exceptionally well on general video understanding tasks, further validating its effectiveness. Our code will be released at https://github.com/Sueqk/LMM-VQA. Qihang Ge, Wei Sun 0029, Yu Zhang 0133, Zhongpeng Ji, Fengyu Sun, Shangling Jui, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified ModelabstractIn recent years, AI-driven video generation has gained significant attention due to great advancements in visual and language generative techniques. Consequently, there is a growing need for accurate Video Quality Assessment (VQA) metrics to evaluate the perceptual quality of AI-generated content (AIGC) videos and optimize video generation models. However, assessing the quality of AIGC videos remains a significant challenge because these videos often exhibit highly complex distortions, such as unnatural actions and irrational objects. To address this challenge, we systematically investigate the AIGC-VQA problem in this article, considering both subjective and objective quality assessment perspectives. For the subjective perspective, we construct the L arge-scale G enerated V ideo Q uality Assessment (LGVQ) dataset, consisting of \(2,\!808\) AIGC videos generated by six video generation models using 468 carefully curated text prompts. Unlike previous subjective VQA experiments, we evaluate the perceptual quality of AIGC videos from three critical dimensions: spatial quality, temporal quality, and text-video alignment, which hold utmost importance for current video generation techniques. For the objective perspective, we establish a benchmark for evaluating existing quality assessment metrics on the LGVQ dataset. Our findings show that current metrics perform poorly on this dataset, highlighting a gap in effective evaluation tools. To bridge this gap, we propose the U nify G enerated V ideo Q uality Assessment (UGVQ) model, designed to accurately evaluate the multi-dimensional quality of AIGC videos. The UGVQ model integrates the visual and motion features of videos with the textual features of their corresponding prompts, forming a unified quality-aware feature representation tailored to AIGC videos. Experimental results demonstrate that UGVQ achieves state-of-the-art performance on the LGVQ dataset across all three quality dimensions, validating its effectiveness as an accurate quality metric for AIGC videos. We hope that our benchmark can promote the development of AIGC-VQA studies. Both the LGVQ dataset and the UGVQ model are publicly available on https://github.com/zczhang-sjtu/UGVQ.git . Wei Sun 0029, Xinyue Li 0001, Jun Jia, Xiongkuo Min, Chunyi Li 0001, Zijian Chen 0001, Puyi Wang, Fengyu Sun, Shangling Jui, Guangtao Zhai |
ACM Trans. Multim. Comput. Commun. Appl. | 10 |
| 2024 | Building Optimal Neural Architectures Using Interpretable KnowledgeabstractNeural Architecture Search is a costly practice. The fact that a search space can span a vast number of design choices with each architecture evaluation taking nontrivial overhead makes it hard for an algorithm to sufficiently explore candidate networks. In this paper, we propose Auto-Build, a scheme which learns to align the latent embeddings of operations and architecture modules with the ground-truth performance of the architectures they appear in. By doing so, AutoBuild is capable of assigning interpretable importance scores to architecture modules, such as individual operation features and larger macro operation sequences such that high-performance neural networks can be constructed without any need for search. Through experiments performed on state-of-the-art image classification, segmentation, and Stable Diffusion models, we show that by mining a relatively small set of evaluated architectures, AutoBuild can learn to build high-quality architectures directly or help to reduce search space to focus on relevant areas, finding better architectures that outperform both the original labeled ones and ones found by search baselines. Code available at https://github.com/Ascend-Research/AutoBuild Keith G. Mills, Fred X. Han, Mohammad Salameh, Shengyao Lu, Chunhua Zhou, Jiao He, Fengyu Sun, Di Niu 0002 |
CVPR | 7 |
| 2024 | CASCO: Cascaded Co-Optimization for Holistic Neural Network AccelerationabstractAutomatic design space exploration of neural accelerators has become essential to maintain high productivity in deep learning acceleration. Previous accelerator co-design approaches mainly focus on exploring hardware and software design spaces assuming individual mapping of neural network layers is deployed on hardware independently while neglecting the vast yet crucial joint effect of layer fusion. To address this shortcoming, we propose CASCO, a cascaded and holistic co-optimization aware of all co-design aspects, including the accelerator's hardware architecture, on-chip operator mapping, and off-chip layer fusion optimization. We then propose an efficient joint-optimization algorithm framework for exploring the joint-design space efficiently by adaptive search resource allocation. Empirical results show that CASCO outperforms the existing co-design framework HASCO by a large margin. Especially, when training on the same networks, CASCO produces generalizable HW design on unseen neural network applications with EDP reduction from 1.4× to 3.2×. Bahador Rashidi, Kiarash Aghakasiri, Fred X. Han, Laiyuan Gong, Fengyu Sun |
DATE | 8 |
| 2024 | A Multiscale Objective Function for Camera Color CorrectionabstractColor correction (CC) plays a pivotal role in camera imaging. Existing approaches usually conduct CC tuning by minimizing ∆E (e.g. ∆E2000), a standard metric proposed by CIE for representing color differences in LAB space. However, we observe that not all the colors with identical ∆E error to the target color have with same perceptual preference. Consequently, optimizing CC by minimizing ∆E solely does not always produce satisfactory color-rendition accuracy. To deal with the problem, in this paper, we propose a new score function, namely Ψ, for a more accurate discrimination of different color-rendition mappings. This is achieved by a multi-scale objective incorporating not only ∆E, but also ∆H and ∆C, which respectively indicate color differences from hue and chroma perspectives. We describe the details of Ψ and show how to adjust its parameters for different preferences. We verify the usefulness of Ψ in experiments by embedding it in various CC tuning algorithms. The empirical results show that Ψ consistently leads to better color-rendition accuracy not only in training but also in validation sets. Finally, we deploy our new objective for tuning a real-world commercial digital camera and show that it delivers improved performance. Bahador Rashidi, Kiarash Aghakasiri, Yue Zhang 0025, Fengyu Sun |
ICASSP | 7 |
| 2023 | RUPQ: Improving low-bit quantization by equalizing relative updates of quantization parameters
Valentin Buchnev, Jiao He, Fengyu Sun, Ivan Koryakovskiy |
BMVC | 3 |
| 2023 | RepQ: Generalizing Quantization-Aware Training for Re-Parametrized Architectures
Anastasiia Prutianova, Alexey Zaytsev 0002, Chung-Kuei Lee, Fengyu Sun, Ivan Koryakovskiy |
BMVC | 4 |
| 2023 | UNICO: Unified Hardware Software Co-Optimization for Robust Neural Network AccelerationabstractSpecialized hardware has become an indispensable component to deep neural network (DNN) acceleration. To keep up with the rapid evolution of neural networks, holistic and automated solutions for jointly optimizing both hardware (HW) architectures and software (SW) mapping have been studied. These studies face two major challenges. First, the combined HW-SW design space is vast, which hinders the finding of optimal or near-optimal designs. This issue is exacerbated for industrial cases when cycle accurate models are used for design evaluation in the joint optimization. Second, HW design is prone to overfitting to the input DNNs used in the HW-SW co-optimization. To address these issues, in this paper, we propose UNICO, an efficient Unified Co-Optimization framework with a novel Robustness metric for better HW generalization. Guided by a high-fidelity surrogate model, UNICO employs multi-objective Bayesian optimization to effectively explore the HW design space, and conducts adaptive, parallel and scalable software mapping search based on successive halving. To reduce HW overfitting, we propose a HW robustness metric by relating a HW configuration’s quality to its sensitivity in software mapping search, and quantitatively incorporate this metric to search for more robust HW design(s). We implement UNICO in open source accelerator platform, and compare it with the state-of-the-art solution HASCO. Experiments show that UNICO significantly outperforms HASCO; it finds design(s) with similar quality to HASCO up to 4 × faster, and eventually converges to better and more robust designs. Finally, we deploy UNICO for optimizing an industrial accelerator, and show that it generates enhanced HW design(s) for key real-world DNNs. Bahador Rashidi, Chao Gao 0012, Chunhua Zhou, Di Niu 0002, Fengyu Sun |
MICRO | 7 |
| 2023 | AutoGO: Automated Computation Graph Optimization for Neural Network EvolutionabstractOptimizing Deep Neural Networks (DNNs) to obtain high-quality models for efficient real-world deployment has posed multi-faceted challenges to machine learning engineers. Existing methods either search for neural architectures in heuristic design spaces or apply low-level adjustments to computation primitives to improve inference efficiency on hardware. We present Automated Graph Optimization (AutoGO), a framework to evolve neural networks in a low-level Computation Graph (CG) of primitive operations to improve both its performance and hardware friendliness. Through a tokenization scheme, AutoGO performs variable-sized segment mutations, making both primitive changes and larger-grained changes to CGs. We introduce our segmentation and mutation algorithms, efficient frequent segment mining technique, as well as a pretrained context-aware predictor to estimate the impact of segment replacements. Extensive experimental results show that AutoGO can automatically evolve several typical large convolutional networks to achieve significant task performance improvement and FLOPs reduction on a range of CV tasks, ranging from Classification, Semantic Segmentation, Human Pose Estimation, to Super Resolution, yet without introducing any newer primitive operations. We also demonstrate the lightweight deployment results of AutoGO-optimized super-resolution and denoising U-Nets on a cycle simulator for a Neural Processing Unit (NPU), achieving PSNR improvement and latency/power reduction simultaneously. Code available at https://github.com/Ascend-Research/AutoGO. Mohammad Salameh, Keith G. Mills, Negar Hassanpour, Fred X. Han, Wei Lu 0023, Shangling Jui, Chunhua Zhou, Fengyu Sun, Di Niu 0002 |
NeurIPS | 9 |
| 2020 | A Robust Audio-Visual Speech Enhancement ModelabstractMost existing audio-visual speech enhancement (AVSE) methods work well in conditions with strong noise, however when applied to conditions with a medium SNR, serious performance degradations are often observed. These degradations can be partly attributed to the feature-fusion(early fusion etc.) architecture that tightly couples the audio information that is very strong and the visual information that is relatively weak. In this paper, we present a safe AVSE approach that can make the visual stream contribute to audio speech enhancment(ASE) safely in conditions of various SNRs by late fusion.The key novelty is two-fold: Firstly, we define power binary masks (PBMs) as a rough representation of speech signals. This rough representation admits the weakness of the visual information and so can be easily predicted from the visual stream. Secondly, we design a posterior augmentation architecture that integrate the visual-derived PBMs to the audio-derived masks via a gating network. By this architecture, the entire performance is lower-bounded by the audio-based component. Our experiments on the Grid dataset demonstrated that this new approach consistently outperforms the audio-based system in all noise conditions, confirming that it is a safe way to incorporate visual knowledge in speech enhancement. Wupeng Wang, Dong Wang 0013, Xiao Chen 0012, Fengyu Sun |
ICASSP | 5 |