EDBT 2026 Demo / reviewers in the wild / expert
Yian Zhao
dblp:51/8416
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-7943-3164ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BiHiTo: Biomolecular Hierarchy-inspired TokenizationabstractThree-dimensional atomic arrangements of biomolecules are key to demystifying biological functions. The rapid expansion of accessible structural data, driven by advances in AI for science, highlights the critical challenge of efficiently modeling large-scale biomolecular structures, which are high-dimensional systems shaped by biological assembly principles. To address this, we introduce BiHiTo, a multi-level Biomolecular Hierarchy-inspired Tokenizer that intrinsically mimics natural biological assembly hierarchies. Specifically, we design a multi-codebook quantizer that mirrors the natural hierarchy of biomolecular structure, enabling simultaneous capture of representations spanning atomic motifs to global conformational variations. This hierarchical alignment markedly improves the biological interpretability and reconstruction fidelity of biomolecular structure.Extensive experiments demonstrate that BiHiTo delivers state-of-the-art performance and robust generalization across molecular dynamics trajectories and macromolecular complexes, facilitating advances in structure generation and dynamic conformation exploration. In the reconstruction of the CASP14 and OOD test set FastFolding protein multi-conformation data, our method achieves a 17% and 51% reduction in RMSD compared to Bio2Token, respectively. Ruochong Zheng, Yutian Liu 0004, Yian Zhao, Zhiwei Nie, Xuehan Hou, Youdong Mao, Jie Chen 0001 |
AAAI | 3 |
| 2026 | Unified Granularity Controller for Interactive SegmentationabstractInteractive Segmentation (IS) segments specific objects or parts by deducing human intent from sparse input prompts. However, the sparse-to-dense mapping is ambiguous, making it challenging for users to obtain segmentations at the desired granularity and causing them to engage in trial-and-error cycles. Although existing multi-granularity IS models (e.g., SAM) alleviate the ambiguity of single-granularity methods by predicting multiple masks simultaneously, this approach has limited scalability and produces redundant results. To address this issue, we introduce a creative granularity-controllable IS paradigm that resolves ambiguity by enabling users to precisely control the segmentation granularity. Specifically, we propose a Unified Granularity Controller (UniGraCo) that supports multi-type optional granularity control signals to pursue unified control over diverse segmentation requirements, effectively overcoming the limitation of single-type control in adapting to different needs, thus boosting the system efficiency and practicality. To mitigate the excessive cost of annotating the multi-granularity masks and the corresponding granularity control signals for training UniGraCo, we construct an automated data engine capable of generating high-quality and granularity-abundant mask-granularity data pairs at low cost. To enable UniGraCo to learn unified granularity controllability in an efficient and stable manner, we further design a granularity-controllable learning strategy. This strategy leverages the generated data pairs to incrementally equip the pre-trained IS model with granularity controllability while preserving its segmentation capability. Extensive experiments on intricate scenarios at both instance and part level demonstrate that our UniGraCo has significant advantages over previous methods, highlighting its potential as a practical interactive tool. Yian Zhao, Kehan Li 0002, Pengchong Qiao, Chang Liu 0047, Rongrong Ji, Jie Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | iSegMan: Interactive Segment-and-Manipulate 3D GaussiansabstractThe efficient rendering and explicit nature of 3DGS promote the advancement of 3D scene manipulation. However, existing methods typically encounter challenges in controlling the manipulation region and are unable to furnish the user with interactive feedback, which inevitably leads to unexpected results. Intuitively, incorporating interactive 3D segmentation tools can compensate for this deficiency. Nevertheless, existing segmentation frameworks impose a preprocessing step of scene-specific parameter training, which limits the efficiency and flexibility of scene manipulation. To deliver a 3D region control module that is well-suited for scene manipulation with reliable efficiency, we propose interactive Segment-and-Manipulate 3D Gaussians (iSegMan), an interactive segmentation and manipulation framework that only requires simple 2D user interactions in any view. To propagate user interactions to other views, we propose Epipolar-guided Interaction Propagation (EIP), which innovatively exploits epipolar constraint for efficient and robust interaction matching. To avoid scene-specific training to maintain efficiency, we further propose the novel Visibility-based Gaussian Voting (VGV), which obtains 2D segmentations from SAM and models the region extraction as a voting game between 2D Pixels and 3D Gaussians based on Gaussian visibility. Taking advantage of the efficient and precise region control of EIP and VGV, we put forth a Manipulation Toolbox to implement various functions on selected regions, enhancing the controllability, flexibility and practicality of scene manipulation. Extensive results on 3D scene manipulation and segmentation tasks fully demonstrate the significant advantages of iSegMan. Project page is available at https://zhao-yian.github.io/iSegMan. Yian Zhao, Wanshi Xu, Ruochong Zheng, Pengchong Qiao |
CVPR | 1 |
| 2025 | Temporal-Aware Query Routing for Real-Time Video Instance Segmentation
Zesen Cheng, Kehan Li 0002, Yian Zhao, Jie Chen 0001 |
ICCV | 3 |
| 2025 | Tune-Your-Style: Intensity-Tunable 3D Style Transfer with Gaussian Splatting
Yian Zhao, Rushi Ye, Ruochong Zheng, Zesen Cheng, Jiashu Yang, Pengchong Qiao, Jie Chen 0001 |
ICCV | 1 |
| 2025 | E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event CamerasabstractNovel view synthesis and 4D reconstruction techniques predominantly rely on RGB cameras, thereby inheriting inherent limitations such as the dependence on adequate lighting, susceptibility to motion blur, and a limited dynamic range. Event cameras, offering advantages of low power, high temporal resolution and high dynamic range, have brought a new perspective to addressing the scene reconstruction challenges in high-speed motion and low-light scenes. To this end, we propose E-4DGS, the first event-driven dynamic Gaussian Splatting approach, for novel view synthesis from multi-view event streams with fast-moving cameras. Specifically, we introduce an event-based initialization scheme to ensure stable training and propose event-adaptive slicing splatting for time-aware reconstruction. Additionally, we employ intensity importance pruning to eliminate floating artifacts and enhance 3D consistency, while incorporating an adaptive contrast threshold for more precise optimization. We design a synthetic multi-view camera setup with six moving event cameras surrounding the object in a 360-degree configuration and provide a benchmark multi-view event stream dataset that captures challenging motion scenarios. Our approach outperforms both event-only and event-RGB fusion baselines and paves the way for the exploration of multi-view event-based reconstruction as a novel approach for rapid scene capture. Chaoran Feng 0001, Zhenyu Tang 0004, Wangbo Yu, Yatian Pang, Yian Zhao, Jianbin Zhao, Li Yuan 0007, Yonghong Tian 0001 |
ACM Multimedia | 5 |
| 2025 | Referent-Aligned Training for Unsupervised Interactive Segmentation
Luna Wang, Yian Zhao |
PRCV (1) | 2 |
| 2024 | GraCo: Granularity-Controllable Interactive SegmentationabstractInteractive Segmentation (IS) segments specific objects or parts in the image according to user input. Current IS pipelines fall into two categories: single-granularity out-put and multi-granularity output. The latter aims to allevi-ate the spatial ambiguity present in the former. However, the multi-granularity output pipeline suffers from limited interaction flexibility and produces redundant results. In this work, we introduce Granularity-Controllable Interactive Segmentation (GraCo), a novel approach that allows precise control of prediction granularity by introducing ad-ditional parameters to input. This enhances the customization of the interactive system and eliminates redundancy while resolving ambiguity. Nevertheless, the exorbitant cost of annotating multi-granularity masks and the lack of avail-able datasets with granularity annotations make it difficult for models to acquire the necessary guidance to control out-put granularity. To address this problem, we design an any-granularity mask generator that exploits the semantic property of the pre-trained IS model to automatically gen-erate abundant mask-granularity pairs without requiring additional manual annotation. Based on these pairs, we propose a granularity-controllable learning strategy that efficiently imparts the granularity controllability to the IS model. Extensive experiments on intricate scenarios at ob-ject and part levels demonstrate that our GraCo has signifi-cant advantages over previous methods. This highlights the potential of GraCo to be a flexible annotation tool, capable of adapting to diverse segmentation scenarios. The project page: https://zhao-yian.github.io/GraCo. Yian Zhao, Kehan Li 0002, Zesen Cheng, Pengchong Qiao, Xiawu Zheng, Rongrong Ji, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
CVPR | 1 |
| 2024 | DETRs Beat YOLOs on Real-time Object DetectionabstractThe YOLO series has become the most popular frame-work for real-time object detection due to its reasonable trade-off between speed and accuracy. However, we observe that the speed and accuracy of YOLOs are negatively affected by the NMS. Recently, end-to-end Transformer-based detectors (DETRs) have provided an alternative to eliminating NMS. Nevertheless, the high computational cost limits their practicality and hinders them from fully exploiting the advantage of excluding NMS. In this paper, we propose the Real-Time DEtection TRansformer (RT-DETR), the first real-time end-to-end object detector to our best knowledge that addresses the above dilemma. We build RT-DETR in two steps, drawing on the advanced DETR: first we focus on maintaining accuracy while improving speed, followed by maintaining speed while improving accuracy. Specifically, we design an efficient hybrid encoder to expeditiously process multi-scale features by decoupling intra-scale interaction and cross-scale fusion to improve speed. Then, we propose the uncertainty-minimal query selection to provide high-quality initial queries to the decoder, thereby improving accuracy. In addition, RT-DETR supports flexible speed tuning by adjusting the number of decoder layers to adapt to various scenarios without retraining. Our RT-DETR-R50 /R101 achieves 53.1% 154.3% AP on COCO and 108 /74 FPS on T4 GPU, outperforming previously advanced YOLOs in both speed and accuracy. Furthermore, RT-DETR-R50 outperforms DINO-R50 by 2.2% AP in accuracy and about 21 times in FPS. After pre-training with Objects365, RT-DETR-R50 / R101 achieves 55.3% 156.2% AP. The project page: https://zhao-yian.github.io/RTDEtr. Yian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei, Guanzhong Wang, Qingqing Dang |
CVPR | 1 |
| 2024 | Federated Unlearning With Momentum DegradationabstractData privacy is becoming increasingly important as data becomes more valuable, as evidenced by the enactment of right-to-be-forgotten laws and regulations. However, in a federated learning (FL) system, simply deleting data from the database when a user requests data revocation is not sufficient, as the training data is already implicitly contained in the parameter distribution of the models trained with it. Furthermore, the global model in the FL system is vulnerable to data poisoning attacks by malicious nodes. Exploring a reliable data poisoning reversal method can effectively counter such attacks. In this article, we analyze the necessity of decoupling the processes of unlearning and training and propose a training-agnostic and efficient method that can effectively perform two types of unlearning tasks: 1) client revocation and 2) category removal. Specifically, we decompose the unlearning process into two steps: 1) knowledge erasure and 2) memory guidance. We first propose a novel knowledge erasure strategy called momentum degradation (MoDe) which realizes the erasure of implicit knowledge in the model and ensures that the model can move smoothly to the early state of the retrained model. To mitigate the performance degradation caused by the first step, the memory guidance strategy implements guided fine-tuning of the model on different data points, which can effectively restore the discriminability of the model on the remaining data points. Extensive experiments demonstrate that our method outperforms the existing task-specific algorithms and matches the performance of retraining, accelerating the execution time by 5–20 times compared to retraining on different data sets. Yian Zhao, Pengfei Wang 0013, Heng Qi, Jianguo Huang, Zongzheng Wei, Qiang Zhang 0008 |
IEEE Internet Things J. | 1 |
| 2023 | ACSeg: Adaptive Conceptualization for Unsupervised Semantic SegmentationabstractRecently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship among pixel-level representations typically contains rich class-aware information that semantically identical pixel embeddings in the representation space gather together to form sophisticated concepts. However, leveraging the learned models to ascertain semantically consistent pixel groups or regions in the image is non-trivial since over/ under-clustering overwhelms the conceptualization procedure under various semantic distributions of different images. In this work, we investigate the pixel-level semantic aggregation in self-supervised ViT pre-trained models as image Segmentation and propose the Adaptive Conceptualization approach for USS, termed ACSeg. Concretely, we explicitly encode concepts into learnable prototypes and design the Adaptive Concept Generator (ACG), which adaptively maps these prototypes to informative concepts for each image. Meanwhile, considering the scene complexity of different images, we propose the modularity loss to optimize ACG independent of the concept number based on estimating the intensity of pixel pairs belonging to the same concept. Finally, we turn the USS task into classifying the discovered concepts in an unsupervised manner. Extensive experiments with state-of-the-art results demonstrate the effectiveness of the proposed ACSeg. Kehan Li 0002, Zhennan Wang 0001, Zesen Cheng, Runyi Yu 0002, Yian Zhao, Guoli Song, Chang Liu 0030, Li Yuan 0007, Jie Chen 0001 |
CVPR | 5 |
| 2023 | Multi-granularity Interaction Simulation for Unsupervised Interactive SegmentationabstractInteractive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and medical image analysis. Typically, massive and expansive pixel-level annotations are spent to train deep models by object-oriented interactions with manually labeled object masks. In this work, we reveal that informative interactions can be made by simulation with semantic-consistent yet diverse region exploration in an unsupervised paradigm. Concretely, we introduce a Multi-granularity Interaction Simulation (MIS) approach to open up a promising direction for unsupervised interactive segmentation. Drawing on the high-quality dense features produced by recent self-supervised models, we propose to gradually merge patches or regions with similar features to form more extensive regions and thus, every merged region serves as a semantic-meaningful multi-granularity proposal. By randomly sampling these proposals and simulating possible interactions based on them, we provide meaningful interaction at multiple granularities to teach the model to understand interactions. Our MIS significantly outperforms non-deep learning unsupervised methods and is even comparable with some previous deep-supervised methods without any annotation. Kehan Li 0002, Yian Zhao, Zhennan Wang 0001, Zesen Cheng, Peng Jin 0001, Xiangyang Ji, Li Yuan 0007, Chang Liu 0030, Jie Chen 0001 |
ICCV | 2 |
| 2022 | A2E2: Aerial-assisted energy-efficient edge sensing in intelligent public transportation systems
Pengfei Wang 0013, Zhaohong Yan, Guangjie Han, Yian Zhao, Chi Lin 0001, Ning Wang 0002, Qiang Zhang 0008 |
J. Syst. Archit. | 5 |
| 2022 | Blockchain-Enhanced Federated Learning Market With Social Internet of ThingsabstractThe machine learning performance usually could be improved by training with massive data. However, requesters can only select a subset of devices with limited training data to execute federated learning (FL) tasks as a result of their limited budgets in today’s IoT scenario. To resolve this pressing issue, we devise a blockchain-enhanced FL market (BFL) to$(i)$make data in computationally bounded devices available for training with social Internet of things,$(ii)$maximize the amount of training data with given budgets for an FL task, and$(iii)$decentralize the FL market with blockchain. To achieve these goals, we firstly propose a trust-enhanced collaborative learning strategy (TCL) and a quality-oriented task allocation algorithm (QTA), where TCL enables training data sharing among trusted devices with social Internet of things, and QTA allocates suitable devices to execute FL tasks while maximizing the training quality with fixed budgets. Then, we devise an encrypted model training scheme (EMT) based on a simple but countervailable differential privacy methodology to prevent attacks from malicious devices. In addition, we also propose a contribution-driven delegated proof of stake (DPoS) consensus mechanism to guarantee the fairness of reward distribution in the block generation process. Finally, extensive evaluations are conducted to verify the proposed BFL could improve the total utility of requesters and average accuracy of FL models significantly. Pengfei Wang 0013, Yian Zhao, Mohammad S. Obaidat, Zongzheng Wei, Heng Qi, Chi Lin 0001, Yunming Xiao, Qiang Zhang 0008 |
IEEE J. Sel. Areas Commun. | 2 |