EDBT 2026 Demo / reviewers in the wild / expert
Qi Zuo
dblp:65/880
· DBLP profile ↗
26ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Systems, architecture and hardware · 9Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-faceted contrastive learning with inter-frame difference for traffic video question answering
Kan Guo, Qi Zuo, Yongli Hu, Lanping Qian, Daxin Tian, Jiapu Wang, Guixian Qu, Tingzheng Jia, Junbin Gao |
Knowl. Based Syst. | 2 |
| 2026 | MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control FlowabstractGenerating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. To further boost the performance, we advance the motion diffusion to motion-aligned temporal latent diffusion by developing a novel motion VAE. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available. Zhe Li 0038, Yisheng He, Weichao Shen, Qi Zuo, Lingteng Qiu, Shenhao Zhu, Zilong Dong, Laurence T. Yang, Chang Xu 0002, Weihao Yuan 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian ReconstructionabstractGenerating animatable human avatars from a single image is essential for various digital human modeling applications. Existing 3D reconstruction methods often struggle to capture fine details in animatable models, while generative approaches for controllable animation, though avoiding explicit 3D modeling, suffer from viewpoint inconsistencies in extreme poses and computational inefficiencies. In this paper, we address these challenges by leveraging the power of generative models to produce detailed multi-view canonical pose images, which help resolve ambiguities in animatable human reconstruction. We then propose a robust method for 3D reconstruction of inconsistent images, enabling real-time rendering during inference. Specifically, we adapt a transformer-based video generation model to generate multi-view canonical pose images and normal maps, pretraining on a large-scale video dataset to improve generalization. To handle view inconsistencies, we recast the reconstruction problem as a 4D task and introduce an efficient 3D modeling approach using 4D Gaussian Splatting. Experiments demonstrate that our method achieves photorealistic, real-time animation of 3D human avatars from in-the-wild images, showcasing its effectiveness and generalization capability. Our code will be available on https://github.com/aigc3d/AniGS. Lingteng Qiu, Shenhao Zhu, Qi Zuo, Xiaodong Gu 0004, Zhe Li 0038, Weihao Yuan 0001, Liefeng Bo, Guanying Chen, Zilong Dong |
CVPR | 3 |
| 2025 | LHM: Large Animatable Human Reconstruction Model for Single Image to 3D in Seconds
Lingteng Qiu, Xiaodong Gu 0004, Peihao Li 0003, Qi Zuo, Weichao Shen, Kejie Qiu, Weihao Yuan 0001, Guanying Chen, Zilong Dong, Liefeng Bo |
ICCV | 4 |
| 2025 | Spatial-Temporal Traffic Prediction Based on Multi-Scale Time Difference
Yongli Hu, Qi Zuo, Kan Guo, Zhongfan Sun, Tingzheng Jia |
ICIC (11) | 2 |
| 2025 | HyPlaneHead: Rethinking Tri-plane-like Representations in Full-Head Image SynthesisabstractTri-plane-like representations have been widely adopted in 3D-aware GANs for head image synthesis and other 3D object/scene modeling tasks due to their efficiency. However, querying features via Cartesian coordinate projection often leads to feature entanglement, which results in mirroring artifacts. A recent work, SphereHead, attempted to address this issue by introducing spherical tri-planes based on a spherical coordinate system. While it successfully mitigates feature entanglement, SphereHead suffers from uneven mapping between the square feature maps and the spherical planes, leading to inefficient feature map utilization during rendering and difficulties in generating fine image details.Moreover, both tri-plane and spherical tri-plane representations share a subtle yet persistent issue: feature penetration across convolutional channels can cause interference between planes, particularly when one plane dominates the others (see Fig. 1). These challenges collectively prevent tri-plane-based methods from reaching their full potential. In this paper, we systematically analyze these problems for the first time and propose innovative solutions to address them. Specifically, we introduce a novel hybrid-plane (hy-plane for short) representation that combines the strengths of both planar and spherical planes while avoiding their respective drawbacks. We further enhance the spherical plane by replacing the conventional theta-phi warping with a novel near-equal-area warping strategy, which maximizes the effective utilization of the square feature map. In addition, our generator synthesizes a single-channel unified feature map instead of multiple feature maps in separate channels, thereby effectively eliminating feature penetration. With a series of technical improvements, our hy-plane representation enables our method, HyPlaneHead, to achieve state-of-the-art performance in full-head image synthesis. Heyuan Li, Kenkun Liu, Lingteng Qiu, Qi Zuo, Keru Zheng, Zilong Dong, Xiaoguang Han 0001 |
NeurIPS | 4 |
| 2024 | GPLD3D: Latent Diffusion of 3D Shape Generative Models by Enforcing Geometric and Physical PriorsabstractState-of-the-art man-made shape generative models usually adopt established generative models under a suitable implicit shape representation. A common theme is to perform distribution alignment, which does not explicitly model important shape priors. As a result, many synthetic shapes are not connected. Other synthetic shapes present problems of physical stability and geometric feasibility. This paper introduces a novel latent diffusion shape-generative model regularized by a quality checker that outputs a score of a latent code. The scoring function employs a learned function that provides a geometric feasibility score and a deterministic procedure to quantify a physical stability score. The key to our approach is a new diffusion procedure that combines the discrete empirical data distribution and a continuous distribution induced by the quality checker. We introduce a principled approach to determine the trade-off parameters for learning the denoising network at different noise levels. Experimental results show that our approach outperforms state-of-the-art shape generations quantitatively and qualitatively on ShapeNet-v2. Qi Zuo, Xiaodong Gu 0004, Weihao Yuan 0001, Zilong Dong, Liefeng Bo, Qixing Huang |
CVPR | 2 |
| 2024 | RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3DabstractLifting 2D diffusion for 3D generation is a challenging problem due to the lack of geometric prior and the complex entanglement of materials and lighting in natural images. Existing methods have shown promise by first creating the geometry through score-distillation sampling (SDS) applied to rendered surface normals, followed by appearance modeling. However, relying on a 2D RGB diffusion model to optimize surface normals is suboptimal due to the distribution discrepancy between natural images and normals maps, leading to instability in optimization. In this paper, recognizing that the normal and depth information effectively describe scene geometry and be auto-matically estimated from images, we propose to learn a generalizable Normal-Depth diffusion model for 3D generation. We achieve this by training on the large-scale LAION dataset together with the generalizable image-to-depth and normal prior models. In an attempt to alleviate the mixed illumination effects in the generated materials, we introduce an albedo diffusion model to impose data-driven constraints on the albedo component. Our experiments show that when integrated into existing text-to-3D pipelines, our models significantly enhance the detail richness, achieving state-of-the-art results. Our project page is at https://aigc3d.github.io/richdreamer/. Lingteng Qiu, Guanying Chen, Xiaodong Gu 0004, Qi Zuo, Mutian Xu, Yushuang Wu, Weihao Yuan 0001, Zilong Dong, Liefeng Bo, Xiaoguang Han 0001 |
CVPR | 4 |
| 2024 | An Optimization Framework to Enforce Multi-view Consistency for Texturing 3D Meshes
Xiaodong Gu 0004, Qi Zuo, Weihao Yuan 0001, Liefeng Bo, Zilong Dong, Qixing Huang |
ECCV (36) | 5 |
| 2024 | High-Fidelity 3D Textured Shapes Generation by Sparse Encoding and Adversarial Decoding
Qi Zuo, Xiaodong Gu 0004, Weihao Yuan 0001, Lingteng Qiu, Liefeng Bo, Zilong Dong |
ECCV (10) | 1 |
| 2024 | StableNormal: Reducing Diffusion Variance for Stable and Sharp NormalabstractThis work addresses the challenge of high-quality surface normal estimation from monocular colored inputs (i.e., images and videos), a field which has recently been revolutionized by repurposing diffusion priors. However, previous attempts still struggle with stochastic inference, conflicting with the deterministic nature of the Image2Normal task, and costly ensembling step, which slows down the estimation process. Our method, StableNormal, mitigates the stochasticity of the diffusion process by reducing inference variance, thus producing "Stable-and-Sharp" normal estimates without any additional ensembling process. StableNormal works robustly under challenging imaging conditions, such as extreme lighting, blurring, and low quality. It is also robust against transparent and reflective surfaces, as well as cluttered scenes with numerous objects. Specifically, StableNormal employs a coarse-to-fine strategy, which starts with a one-step normal estimator (YOSO) to derive an initial normal guess, that is relatively coarse but reliable, then followed by a semantic-guided refinement process (SG-DRN) that refines the normals to recover geometric details. The effectiveness of StableNormal is demonstrated through competitive performance in standard datasets such as DIODE-indoor, iBims, ScannetV2 and NYUv2, and also in various downstream tasks, such as surface reconstruction and normal enhancement. These results evidence that StableNormal retains both the "stability" and "sharpness" for accurate normal estimation. StableNormal represents a baby attempt to repurpose diffusion priors for deterministic estimation. To democratize this, code and models have been publicly available in hf.co/Stable-X. Chongjie Ye, Lingteng Qiu, Xiaodong Gu 0004, Qi Zuo, Yushuang Wu, Zilong Dong, Liefeng Bo, Yuliang Xiu, Xiaoguang Han 0001 |
ACM Trans. Graph. | 4 |
| 2023 | DG3D: Generating High Quality 3D Textured Shapes by Learning to Discriminate Multi-Modal Diffusion-RenderingsabstractMany virtual reality applications require massive 3D content, which impels the need for low-cost and efficient modeling tools in terms of quality and quantity. In this paper, we present a Diffusion-augmented Generative model to generate high-fidelity 3D textured meshes that can be directly used in modern graphics engines. Challenges in directly generating textured mesh arise from the instability and texture incompleteness of a hybrid framework which contains conversion between 2D features and 3D space. To alleviate these difficulties, DG3D incorporates a diffusion-based augmentation module into the min-max game between the 3D tetrahedral mesh generator and 2D renderings discriminators, which stabilizes network optimization and prevents mode collapse in vanilla GANs. We also suggest using multi-modal renderings in discrimination to further increase the aesthetics and completeness of generated textures. Extensive experiments on the public benchmark and real scans show that our proposed DG3D outperforms existing state-of-the-art methods by a large margin, i.e., 5% ∼ 40% in FID-3D score and 5%∼10% in geometry-related metrics. Code is available at https://github.com/seakforzq/DG3D. Qi Zuo, Jianfang Li 0001, Liefeng Bo |
ICCV | 1 |
| 2020 | A Survey on Access Control in the Age of Internet of ThingsabstractWith the development of Internet-of-Things (IoT) technology, various types of information, such as social resources and physical resources, are deeply integrated for different comprehensive applications. Social networking, car networking, medical services, video surveillance, and other forms of the IoT information service model gradually change people's daily lives. Facing the vast amounts of IoT information data, the IoT search technology is used to quickly find accurate information to meet the real-time search needs of users. However, IoT search requires using a large amount of user private information, such as personal health information, location information, and social relations information, to provide personalized services. Employing private information from users will encounter security problems if an effective access control mechanism is missing during the IoT search process. An access control mechanism can effectively monitor the access activities of resources and ensure that authorized users access information resources under legitimate conditions. This survey examines the growing literature on access control for an IoT search. Problems and challenges of access control mechanisms are analyzed to facilitate the adoption of access control solutions in real-life settings. This article aims to provide theoretical, methodological, and technical guidance for IoT search access control mechanisms in large-scale dynamic heterogeneous environments. Based on a literature review, we also analyzed the future development direction of access control in the age of IoT. Jing Qiu 0002, Zhihong Tian 0001, Chunlai Du, Qi Zuo, Shen Su, Binxing Fang |
IEEE Internet Things J. | 4 |
| 2013 | A work-stealing scheduling framework supporting fault toleranceabstractFault tolerance and load balancing are critical points for executing long-running parallel applications on multicore clusters. This paper addresses both fault tolerance and load balancing on multicore clusters by presenting a novel work-stealing task scheduling framework which supports hardware fault tolerance. In this framework, both transient and permanent faults are detected and recovered at task granularity. We incorporate task-based fault detection and recovery mechanisms into a hierarchical work-stealing scheme to establish the framework. This framework provides low-overhead fault-tolerance and optimal load balancing by fully exploiting task parallelism. Yizhuo Wang 0001, Weixing Ji, Feng Shi 0009, Qi Zuo |
DATE | 4 |
| 2012 | Knowledge-Based Adaptive Self-Scheduling
Yizhuo Wang 0001, Weixing Ji, Feng Shi 0009, Qi Zuo, Ning Deng 0002 |
NPC | 4 |
| 2012 | A Hierarchical Work-Stealing Framework for Multi-core ClustersabstractWork-stealing has been widely used in task-based parallel programming for dynamic load balancing. The overhead of work-stealing on distributed memory systems is much higher than that on shared memory systems. To minimize the overhead of work-stealing on a multi-core cluster, we propose a hierarchical work-stealing framework, in which work-stealing is performed inside a node before across the node boundary. Two key techniques used in our framework to reduce the inter-node steals are: a) adaptive initial partitioning for different task parallel patterns; b) centralized control for inter-node work-stealing, which improves the efficiency of victim selection and termination detection. We compare our technique to the classical work-stealing scheme and a state-of-the-art work-stealing scheme [1] for multi-core clusters. Our technique outperforms them by 19% and 8% respectively. Yizhuo Wang 0001, Weixing Ji, Qi Zuo, Feng Shi 0009 |
PDCAT | 3 |
| 2011 | A Semi-automatic Scratchpad Memory Management Framework for CMP
Ning Deng 0002, Weixing Ji, Qi Zuo |
APPT | 4 |
| 2011 | Floorplanning exploration and performance evaluation of a new Network-on-ChipabstractThe Network-on-Chip (NoC) paradigm has emerged as a revolutionary methodology in current System-on-Chips (SoCs) for integrating a large number of processing elements in a single die. It has the advantage of enhanced performance, scalability and modularity, compared with previous bus-based communication architectures. Recently, A new Triplet-based Hierarchical Interconnection Network (THIN) has been proposed. In this paper, we explore the three-dimensional (3D) floor-planning of THIN and present two different floorplanning and routing methods using both the Manhattan routing and the Y-architecture routing architectures. A cycle-accurate simulator is developed based on Noxim NoC simulator and ORION 2.0 energy model. The latency, power consumption and area requirement of both THIN and Mesh are evaluated. The experimental results indicate that the proposed design provides 24.95% reduction in average power consumption and 16.84% improvement in area requirement. Licheng Xue, Weixing Ji, Qi Zuo |
DATE | 3 |
| 2011 | Dynamic and adaptive SPM management for a multi-task environment
Weixing Ji, Ning Deng 0002, Feng Shi 0009, Qi Zuo |
J. Syst. Archit. | 4 |
| 2009 | Group-caching for NoC based multicore cache coherent systemsabstractMost CMPs use on-chip networks to connect cores and tend to integrate more simple cores on a single die. Low-radix networks, such as 2D-MESH, are widely used in tiled CMPs since they can be mapped to on-chip networks efficiently. However, low-radix networks introduce high network latency caused by long diameter. In this paper, we propose the use of group-caching design in NoC based multicore cache coherent systems. In our design, on-chip L2 banks are organized to form multiple groups. Each cache group behaves like a shared L2 cache for the cores inside cache group while the cache coherence between cache groups is maintained by coherence messages. Besides, group-caching also adopts the new cache replacement policy to improve the inefficient use of the aggregate L2 cache capacity. Compared to banked and shared L2 design, as most L2 accesses are served by local cache group, the hop count is significantly reduced. Experiment results based on full-system simulation show that for 2D-MESH, group-caching can increase the performance by 2%∼8% compared to banked and shared L2 design, with network energy consumption reduced by 11%∼13%. Experiment results also show that the communication overhead inside cache group plays an important role in the performance of groupcaching. Feng Shi 0009, Qi Zuo, Weixing Ji, Ning Deng 0002, Licheng Xue, Yu-an Tan 0001 |
DATE | 3 |
| 2009 | N-port memory mapping for LUT-based FPGAsabstractAs current FPGAs grow in logic capacity, they are widely used to implement entire systems. In some specific applications, such as our embedded multi-core processor TriBA[1],user memory models are not limited to single-port or dual-port. Thus, we need a cost-effective way to realize N-port memory on FPGA since most commercial products do not provide N-port physical arrays. In this paper, we propose a hierarchical N-port memory architecture for LUT-based FPGAs. The principle of this architecture is to create a two-level memory hierarchy formed by different resources. We map the memory resources inside LUTs as 1-port memory banks, and interleave these banks to create N-port L1 memory. We also interleave physical dual-port arrays to build N-port L2 memory. We also provide the data transfer between L1 and L2 memories and assume that such data transfer is managed by software control just like the strategy used by SPM. Compared to L1 memory, L2 memory has the advantage in cost and also has several disadvantages, such as longer access time and higher conflict probability. If most accesses are served by its L1 memory portion, hierarchical memory architecture will achieve both goals in cost and access time. We implement this architecture on Xilinx Virtex-II chips to measure its cost and also use the memory trace collected from multi-core simulator to measure its average access time. The product of cost and average access time shows that, hierarchical memory architecture is a cost-effective way to realize N-port memory on FPGA. Feng Shi 0009, Qi Zuo, Weixing Ji, Mengxiao Liu |
FPGA | 3 |
| 2009 | Performance prediction based on hierarchy parallel features captured in multi-processing systemabstractAs the computing ability of high performance computers are improved by increasing the number of computing elements, how to utilize the available computing resources becomes an important issue. Different strategies to solve an problem based on a multi-processing system can bring about distinct performance. In this paper, we propose a method to predict the performance of parallel applications. The method describes the parallel features of the multi-processing systems in a hierarchy way, and evaluates solutions based on the description. In this way, programmers can find the better solution of an application before real programming. Feng Shi 0009, Ning Deng 0002, Qi Zuo |
HPDC | 4 |
| 2008 | Unsupervised learning of categories from sets of partially matching image features for power line inspection robotabstractObject recognition and categorization are considered as fundamental steps in the vision based navigation for inspection robot as it must plan its behaviors based on various kinds of obstacles detected from the complex background. However, current approaches typically require some amount of supervision, which is viewed as a expensive burden and restricted to relatively small number of applications in practice. For this purpose, we present an computationally efficient approach that does not need supervision and is capable of learning object categories automatically from unlabeled images which are represented by an set of local features, and all sets are clustered according to their partial-match feature correspondences, which is done by a enhanced Spatial Pyramid Match algorithm (E-SPK). Then a graph-theoretic clustering method is applied to seek the primary grouping among the images. The consistent subsets within the groups are identified by inferring category templates. Given the input, the output of the approach is a partition of the images into a set of learned categories. We demonstrate this approach on a field experiment for a powerline inspection robot. Si-Yao Fu, Qi Zuo, Zeng-Guang Hou, Zi-ze Liang, Min Tan 0001, Xiaoling Fu |
IJCNN | 2 |
| 2007 | Performance Evaluation of a Self-Maintained Memory ModuleabstractHardware approach emerges as one of the candidate in improving the performance of dynamic memory management. This paper presents measurements of a self-maintained memory module subjected to several different workloads. This memory module supporting explicit dynamic memory management takes advantage of the high speed of a pure hardware implementation. Object allocation and deletion are strictly bounded in time. The whole heap space is divided into two semi-spaces, and a concurrent bidirectional memory compaction algorithm is exploited, so that memory compaction can be done while mutator process is running on the processor concurrently. Reported measurements demonstrate that hardware-assisted memory management is a viable alternative to traditional explicit memory management techniques. Experimental results show that more than 60% of memory traffic is saved by the proposed memory compaction scheme compared to software-only approach. Both processor delay and program execution time are greatly reduced. Weixing Ji, Feng Shi 0009, Qi Zuo |
RTSS | 4 |
| 2006 | Neural Network Based Modeling for Oil Well Pressure Data Compensation System
Jian-long Tang, En Li 0001, Zeng-Guang Hou, Qi Zuo, Zi-ze Liang, Min Tan 0001 |
ICIC (2) | 4 |
| 2006 | Structure-Constrained Obstacles Recognition for Power Transmission Line Inspection RobotabstractInspection robot must plan its behavior to detect the obstacles from the complex background according to their types when it is crawling along the power transmission line in order to negotiate reliably. However, in most instances, detecting the obstacles from the complex background is a hard task. For this purpose, a novel and fast visual obstacle recognition algorithm is designed based on the structure of the 220 KV power transmission line. Basic principle and architecture of the algorithm are given. By this approach, three typical obstacles on the power transmission line such as insulator strings, counterweights and suspension clamps can be recognized with high accuracy. Experiments in the real power transmission line show its effectiveness. This method can contribute to the process of the mobile robot negotiating obstacles Si-Yao Fu, Yun-Chu Zhang, Zi-ze Liang, Zeng-Guang Hou, Min Tan 0001, Wenbo Ye, Lian Bo, Qi Zuo |
IROS | 9 |