VLDB 2026 Research / reviewers in the wild / expert
Yiyang Zhou
dblp:175/1589
· DBLP profile ↗
31ranked-venue papers
9as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 8 first-author · 22 since 2021Systems, architecture and hardware · 7 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LSE-Codec: An Arbitrary Frame-Rate Compliant Lossless Compression Model for Neuromorphic Spike CameraabstractSpike cameras represent a novel class of neuromorphic imaging devices that capture visual scenes as binary spike streams through temporal integration of light intensity. While offering exceptional dynamic range and energy efficiency, they generate massive binary data volumes that require efficient compression. Existing codecs fail to exploit the unique structure of spike data and lack support for arbitrary temporal resolution. We propose LSE-Codec, a neural lossless compression framework specifically designed for spike cameras that operates independently of frame-rate, as shown in Fig. 1. Our approach introduces Local Spike Embedding (LSE) to reorganize sparse binary spike patterns into compact 8 -bit symbols while preserving spatial structure, followed by a hierarchical autoencoder with autoregressive entropy modeling to predict spike distributions. Evaluated on different datasets under aligned test conditions, our method achieves state-of-the-art performance, reducing bit rates by over 11% compared to JPEG-XL (best-effort mode). This work bridges a critical gap in spike vision systems with frame-rate agnostic lossless coding, enabling efficient storage and transmission of spike data without sacrificing fidelity. Fanke Dong, Yiyang Zhou, Chuanmin Jia |
DCC | 2 |
| 2025 | GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?abstractYiyang Zhou, Linjie Li, Shi Qiu, Zhengyuan Yang, Yuyang Zhao, Siwei Han, Yangfan He, Kangqi Li, Haonian Ji, Zihao Zhao, Haibo Tong, Lijuan Wang, Huaxiu Yao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yiyang Zhou, Shi Qiu 0016, Zhengyuan Yang, Siwei Han, Yangfan He, Kangqi Li, Haonian Ji, Haibo Tong, Huaxiu Yao |
EMNLP | 1 |
| 2025 | Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language AlignmentabstractThe recent advancements in large language models (LLMs) and pre-trained vision models have accelerated the development of vision-language large models (VLLMs), enhancing the interaction between visual and linguistic modalities. Despite their notable success across various domains, VLLMs face challenges in modality alignment, which can lead to issues like hallucinations and unsafe content generation. Current alignment techniques often rely on coarse feedback and external datasets, limiting scalability and performance. In this paper, we propose FiSAO (Fine-Grained Self-Alignment Optimization), a novel self-alignment method that utilizes the model’s own visual encoder as a fine-grained verifier to improve vision-language alignment without the need for additional data. By leveraging token-level feedback from the vision encoder, FiSAO significantly improves vision-language alignment, even surpassing traditional preference tuning methods that require additional data. Through both theoretical analysis and experimental validation, we demonstrate that FiSAO effectively addresses the misalignment problem in VLLMs, marking the first instance of token-level rewards being applied to such models. Our code is avaliable at \url{https://anonymous.4open.science/r/FISAO-57F0/}. Chenhang Cui, An Zhang 0003, Yiyang Zhou, Zhaorun Chen, Gelei Deng, Huaxiu Yao, Tat-Seng Chua |
ICLR | 3 |
| 2025 | MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsabstractInterleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancements, the evaluation of this capability remains insufficient. Existing benchmarks suffer from limitations in data scale, scope, and evaluation depth, while current evaluation metrics are often costly or biased, lacking in reliability for practical applications. To address these challenges, we introduce MMIE, a large-scale knowledge-intensive benchmark for evaluating interleaved multimodal comprehension and generation in Large Vision-Language Models (LVLMs). MMIE comprises 20K meticulously curated multimodal queries, spanning 3 categories, 12 fields, and 102 subfields, including mathematics, coding, physics, literature, health, and arts. It supports both interleaved inputs and outputs, offering a mix of multiple-choice and open-ended question formats to evaluate diverse competencies. Moreover, we propose a reliable automated evaluation metric, leveraging a scoring model fine-tuned with human-annotated data and systematic evaluation criteria, aimed at reducing bias and improving evaluation accuracy. Extensive experiments demonstrate the effectiveness of our benchmark and metrics in providing a comprehensive evaluation of interleaved LVLMs. Specifically, we evaluate eight LVLMs, revealing that even the best models show significant room for improvement, with most achieving only moderate results. We believe MMIE will drive further advancements in the development of interleaved LVLMs. Peng Xia 0005, Siwei Han, Shi Qiu 0016, Yiyang Zhou, Zhaoyang Wang 0004, Zhaorun Chen, Chenhang Cui, Mingyu Ding, Huaxiu Yao |
ICLR | 4 |
| 2025 | Anyprefer: An Agentic Framework for Preference Data SynthesisabstractHigh-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consuming and costly. Recent methods often adopt a self-rewarding approach, where the target model generates and annotates its own preference data, but this can lead to inaccuracies since the reward model shares weights with the target model, thereby amplifying inherent biases. To address these issues, we propose Anyprefer, a framework designed to synthesize high-quality preference data for aligning the target model. Anyprefer frames the data synthesis process as a cooperative two-player Markov Game, where the target model and the judge model collaborate together. Here, a series of external tools are introduced to assist the judge model in accurately rewarding the target model’s responses, mitigating biases in the rewarding process. In addition, a feedback mechanism is introduced to optimize prompts for both models, enhancing collaboration and improving data quality.
The synthesized data is compiled into a new preference dataset, Anyprefer-V1, consisting of 58K high-quality preference pairs.
Extensive experiments show that Anyprefer significantly improves model alignment performance across four main applications, covering 21 datasets, achieving average improvements of 18.55% in five natural language generation datasets, 3.66% in nine vision-language understanding datasets, 30.05% in three medical image analysis datasets, and 16.00% in four visuo-motor control tasks. Yiyang Zhou, Zhaoyang Wang 0004, Tianle Wang 0009, Shangyu Xing, Peng Xia 0005, Bo Li 0026, Zijian Zhang 0010, Zhaorun Chen, Xuchao Zhang, Chetan Bansal, Mohit Bansal, Huaxiu Yao |
ICLR | 1 |
| 2025 | MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?abstractWhile text-to-image models like GPT-4o-Image and FLUX are rapidly proliferating, they often encounter challenges such as hallucination, bias, and the production of unsafe, low-quality output. To effectively address these issues, it is crucial to align these models with desired behaviors based on feedback from a multimodal judge. Despite their significance, current multimodal judges frequently undergo inadequate evaluation of their capabilities and limitations, potentially leading to misalignment and unsafe fine-tuning outcomes. To address this issue, we introduce MJ-Bench, a novel benchmark which incorporates a comprehensive preference dataset to evaluate multimodal judges in providing feedback for image generation models across six key perspectives: alignment, safety, image quality, bias, composition, and visualization. Specifically, we evaluate a large variety of multimodal judges including smaller-sized CLIP-based scoring models, open-source VLMs, and close-source VLMs on each decomposed subcategory of our preference dataset. Experiments reveal that close-source VLMs generally provide better feedback, with GPT-4o outperforming other judges in average. Compared with open-source VLMs, smaller-sized scoring models can provide better feedback regarding text-image alignment and image quality, while VLMs provide more accurate feedback regarding safety and generation bias due to their stronger reasoning capabilities. Further studies in feedback scale reveal that VLM judges can generally provide more accurate and stable feedback in natural language than numerical scales. Notably, human evaluations on end-to-end and fine-tuned models using separate feedback from these multimodal judges provide similar conclusions, further confirming the effectiveness of MJ-Bench. Zhaorun Chen, Zichen Wen, Yichao Du, Yiyang Zhou, Chenhang Cui, Siwei Han, Jen Weng, Chaoqi Wang, Zhengwei Tong, Leria Huang, Canyu Chen, Haoqin Tu, Qinghao Ye, Zhihong Zhu 0001, Zhuokai Zhao, Rafael Rafailov, Chelsea Finn, Huaxiu Yao |
NeurIPS | 4 |
| 2025 | MJ-Video: Benchmarking and Rewarding Video Generation with Fine-Grained Video PreferenceabstractRecent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing models still struggle with key challenges such as instruction misalignment, content hallucination, safety concerns, and generation bias. To address these limitations, we introduce MJ-BENCH-VIDEO, a large-scale video preference benchmark designed to evaluate video generation across five critical aspects: Alignment, Safety, Fineness, Coherence & Consistency, and Bias & Fairness. This benchmark further incorporates 28 fine-grained criteria to provide a comprehensive evaluation of video preference. Building upon this dataset, we propose MJ-VIDEO, a Mixture-of-Experts (MoE)-based video reward model designed to deliver fine-grained reward. MJ-VIDEO can dynamically select relevant experts to accurately judge the preference based on the input text-video pair. This architecture enables more precise and adaptable preference judgments. Through extensive benchmarking on MJ-BENCH-VIDEO, we analyze the limitations of existing video reward models and demonstrate the superior performance of MJ-VIDEO in video preference assessment, achieving 17.58% and 15.87% improvements in overall and fine-grained preference judgments, respectively. Additionally, MJ-VIDEO is able to improve the alignment performance in video generation via preference fine-tuning. Haibo Tong, Zhaoyang Wang 0004, Zhaorun Chen, Haonian Ji, Shi Qiu 0016, Siwei Han, Kexin Geng, Zhongkai Xue, Yiyang Zhou, Peng Xia 0005, Mingyu Ding, Rafael Rafailov, Chelsea Finn, Huaxiu Yao |
NeurIPS | 9 |
| 2025 | ReAgent-V: A Reward-Driven Multi-Agent Framework for Video UnderstandingabstractVideo understanding is fundamental to tasks such as action recognition, video reasoning, and robotic control. Early video understanding methods based on large vision-language models (LVLMs) typically adopt a single-pass reasoning paradigm without dynamic feedback, limiting the model’s capacity to self-correct and adapt in complex scenarios. Recent efforts have attempted to address this limitation by incorporating reward models and reinforcement learning to enhance reasoning, or by employing tool-agent frameworks. However, these approaches face several challenges, including high annotation costs, reward signals that fail to capture real-time reasoning states, and low inference efficiency. To overcome these issues, we propose ReAgent-V, a novel agentic video understanding framework that integrates efficient frame selection with real-time reward generation during inference. These reward signals not only guide iterative answer refinement through a multi-perspective reflection mechanism—adjusting predictions from conservative, neutral, and aggressive viewpoints—but also enable automatic filtering of high-quality data for supervised fine-tuning (SFT), direct preference optimization (DPO), and group relative policy optimization (GRPO). ReAgent-V is lightweight, modular, and extensible, supporting flexible tool integration tailored to diverse tasks. Extensive experiments on 12 datasets across three core applications—video understanding, video reasoning enhancement, and vision-language-action model alignment—demonstrate significant gains in generalization and reasoning, with improvements of up to 6.9%, 2.1%, and 9.8%, respectively, highlighting the effectiveness and versatility of the proposed framework. Yiyang Zhou, Yangfan He, Yaofeng Su, Siwei Han, Joel Jang, Gedas Bertasius, Mohit Bansal, Huaxiu Yao |
NeurIPS | 1 |
| 2025 | Improving UI responsiveness in Android by restructured renderingabstractMobile operating systems, such as Android, are increasingly used across diverse applications, where ensuring high responsiveness to user interactions is critical, particularly in mission-critical and real-time scenarios. Mobile operating systems typically process user interaction events and UI rendering on the same thread, commonly referred to as the main thread of a mobile application. As a result, user interaction handling can face significant delays when blocked by overloaded UI rendering tasks, compromising responsiveness. Existing mobile operating systems lack effective mechanisms to mitigate this issue. This paper addresses the problem by restructuring the UI rendering workflow to improve responsiveness in the presence of heavy rendering workloads. Specifically, two techniques are proposed that are tailored to whether the event handling results require screen display. Experimental results demonstrate improvements in both average-case and worst-case response times of event handling, enhancing the UI responsiveness. Although the implementation focuses on Android, the proposed approaches are adaptable to other mobile operating systems with similar rendering architectures, such as iOS and HarmonyOS. Mingsong Lv, Tao Hu 0018, Menglong Cui, Tao Yang 0024, Yiyang Zhou, Qingxu Deng, Nan Guan |
J. Syst. Archit. | 5 |
| 2025 | Neighbor-Based Completion for Addressing Incomplete Multiview ClusteringabstractDriven by the complementarity and consistency inherent in multiview data, multiview clustering (MVC) has garnered widespread attention in various domains. Real-world data often encounters the issue of missing information, leading to a surge of interest in the domain of incomplete MVC (IMVC). Despite existing approaches having made significant progress in addressing IMVC, two significant challenges persist: 1) many alignment-based methodologies tend to overlook the topological relationships among instances and 2) the view representations based on completion lack reconstructive properties, casting doubt on their alignment with the actual view representations. In response, we present a novel approach termed neighbor-based completion for addressing IMVC (NBIMVC), which capitalizes on the topological information among instances and the consistent information across views. Specifically, our method uses autoencoders to learn feature representations for each view and leverages nearest-neighbor relationships between unique and complete instances to complete missing features in missing views. Subsequently, we enforce hard negative alignment constraints on complete paired instances in the feature space. Finally, we ensure the consistency of views in the semantic space by employing cluster information and a shared clustering network, which facilitates the final multiview categories output and effectively resolves the IMVC problem. Extensive experimental evaluations validate the efficacy of our proposed method, showcasing comparable or superior performance to existing approaches. Wenbiao Yan, Jihua Zhu, Yiyang Zhou, Jinqian Chen, Haozhe Cheng, Kun Yue, Qinghai Zheng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | How Many Are in This Image A Safety Evaluation Benchmark for Vision LLMs
Haoqin Tu, Chenhang Cui, Yiyang Zhou, Bingchen Zhao, Junlin Han, Wangchunshu Zhou, Huaxiu Yao, Cihang Xie |
ECCV (51) | 4 |
| 2024 | Analyzing and Mitigating Object Hallucination in Large Vision-Language ModelsabstractLarge vision-language models (LVLMs) have shown remarkable abilities in understanding visual information with human languages. However, LVLMs still suffer from object hallucination, which is the problem of generating descriptions that include objects that do not actually exist in the images. This can negatively impact many vision-language tasks, such as visual summarization and reasoning. To address this issue, we propose a simple yet powerful algorithm, LVLM Hallucination Revisor (LURE), to post-hoc rectify object hallucination in LVLMs by reconstructing less hallucinatory descriptions. LURE is grounded in a rigorous statistical analysis of the key factors underlying object hallucination, including co-occurrence (the frequent appearance of certain objects alongside others in images), uncertainty (objects with higher uncertainty during LVLM decoding), and object position (hallucination often appears in the later part of the generated text). LURE can also be seamlessly integrated with any LVLMs. We evaluate LURE on six open-source LVLMs and found it outperforms the previous best approach in both general object hallucination evaluation metrics, GPT, and human evaluations. Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, Huaxiu Yao |
ICLR | 1 |
| 2024 | VHELM: A Holistic Evaluation of Vision Language ModelsabstractCurrent benchmarks for assessing vision-language models (VLMs) often focus on their perception or problem-solving capabilities and neglect other critical aspects such as fairness, multilinguality, or toxicity. Furthermore, they differ in their evaluation procedures and the scope of the evaluation, making it difficult to compare models. To address these issues, we extend the HELM framework to VLMs to present the Holistic Evaluation of Vision Language Models (VHELM). VHELM aggregates various datasets to cover one or more of the 9 aspects: visual perception, knowledge, reasoning, bias, fairness, multilinguality, robustness, toxicity, and safety. In doing so, we produce a comprehensive, multi-dimensional view of the capabilities of the VLMs across these important factors. In addition, we standardize the standard inference parameters, methods of prompting, and evaluation metrics to enable fair comparisons across models. Our framework is designed to be lightweight and automatic so that evaluation runs are cheap and fast. Our initial run evaluates 22 VLMs on 21 existing datasets to provide a holistic snapshot of the models. We uncover new key findings, such as the fact that efficiency-focused models (e.g., Claude 3 Haiku or Gemini 1.5 Flash) perform significantly worse than their full models (e.g., Claude 3 Opus or Gemini 1.5 Pro) on the bias benchmark but not when evaluated on the other aspects. For transparency, we release the raw model generations and complete results on our website at https://crfm.stanford.edu/helm/vhelm/v2.0.1. VHELM is intended to be a living benchmark, and we hope to continue adding new datasets and models over time. Haoqin Tu, Chi Heem Wong, Yiyang Zhou, Yifan Mai 0001, Josselin Somerville Roberts, Michihiro Yasunaga, Huaxiu Yao, Cihang Xie, Percy Liang |
NeurIPS | 5 |
| 2024 | CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language ModelsabstractArtificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized healthcare. However, the trustworthiness of Med-LVLMs remains unverified, posing significant risks for future model deployment. In this paper, we introduce CARES and aim to comprehensively evaluate the Trustworthiness of Med-LVLMs across the medical domain. We assess the trustworthiness of Med-LVLMs across five dimensions, including trustfulness, fairness, safety, privacy, and robustness. CARES comprises about 41K question-answer pairs in both closed and open-ended formats, covering 16 medical image modalities and 27 anatomical regions. Our analysis reveals that the models consistently exhibit concerns regarding trustworthiness, often displaying factual inaccuracies and failing to maintain fairness across different demographic groups. Furthermore, they are vulnerable to attacks and demonstrate a lack of privacy awareness. We publicly release our benchmark and code in https://github.com/richard-peng-xia/CARES. Peng Xia 0005, Juanxi Tian, Yangrui Gong, Ruibo Hou, Zhenbang Wu, Zhiyuan Fan, Yiyang Zhou, Kangyu Zhu, Zhaoyang Wang 0004, Xiao Wang 0044, Xuchao Zhang, Chetan Bansal, Marc Niethammer, Junzhou Huang, Hongtu Zhu, Yun Li 0010, Jimeng Sun 0001, ZongYuan Ge, Gang Li 0001, James Zou 0001, Huaxiu Yao |
NeurIPS | 9 |
| 2024 | Calibrated Self-Rewarding Vision Language ModelsabstractLarge Vision-Language Models (LVLMs) have made substantial progress by integrating pre-trained large language models (LLMs) and vision models through instruction tuning. Despite these advancements, LVLMs often exhibit the hallucination phenomenon, where generated text responses appear linguistically plausible but contradict the input image, indicating a misalignment between image and text pairs. This misalignment arises because the model tends to prioritize textual information over visual input, even when both the language model and visual representations are of high quality. Existing methods leverage additional models or human annotations to curate preference data and enhance modality alignment through preference optimization. These approaches are resource-intensive and may not effectively reflect the target LVLM's preferences, making the curated preferences easily distinguishable. Our work addresses these challenges by proposing the Calibrated Self-Rewarding (CSR) approach, which enables the model to self-improve by iteratively generating candidate responses, evaluating the reward for each response, and curating preference data for fine-tuning. In the reward modeling, we employ a step-wise strategy and incorporate visual constraints into the self-rewarding process to place greater emphasis on visual input. Empirical results demonstrate that CSR significantly enhances performance and reduces hallucinations across twelve benchmarks and tasks, achieving substantial improvements over existing methods by 7.62\%. Our empirical results are further supported by rigorous theoretical analysis, under mild assumptions, verifying the effectiveness of introducing visual constraints into the self-rewarding paradigm. Additionally, CSR shows compatibility with different vision-language models and the ability to incrementally improve performance through iterative fine-tuning. Yiyang Zhou, Zhiyuan Fan, Dongjie Cheng, Sihan Yang 0001, Zhaorun Chen, Chenhang Cui, Linjun Zhang, Huaxiu Yao |
NeurIPS | 1 |
| 2024 | MCoCo: Multi-level Consistency Collaborative multi-view clustering
Yiyang Zhou, Qinghai Zheng, Wenbiao Yan, Jihua Zhu |
Expert Syst. Appl. | 1 |
| 2024 | Multi-view Semantic Consistency based Information Bottleneck for Clustering
Wenbiao Yan, Yiyang Zhou, Qinghai Zheng, Jihua Zhu |
Knowl. Based Syst. | 2 |
| 2024 | S2TNet: Spectral-Spatial Triplet Network for Few-Shot Hyperspectral Image ClassificationabstractDeep learning (DL) has shown great potential for hyperspectral image (HSI) classification. However, DL models easily get trapped into overfitting due to limited training samples. To overcome this issue, a novel spectral–spatial triplet network (S2TNet) is proposed for few-shot HSI classification. First, a lightweight spectral–spatial network (SSN) composed of 1-D and 2-D convolution is introduced to extract spectral–spatial features. Second, a hard sample selection strategy is proposed by integrating classification and contrast training to deal with unbalanced positive and negative samples in traditional triplet networks. Third, an enhanced triplet loss function is proposed by considering the relationship between positive and negative sample pairs to ensure the distance between homogeneous samples is smaller than that of heterogeneous samples, which effectively improves the discrimination ability of the model. Experiments conducted on two widely used hyperspectral datasets demonstrate that S2TNet significantly outperforms other related methods, with 0.81%–16.83% and 1.40%–13.83% improvements (under 20 labeled samples per class for training) in terms of overall accuracy (OA) in Indian Pine (IP) data and University of Pavia (PU), respectively. Guijie Yue, Yiyang Zhou, Zhaohui Xue |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Ghostbuster: A Software Approach for Reducing Ghosting Effect on Electrophoretic DisplaysabstractElectrophoretic displays (EPDs), also known as e-paper, offer a paper-like visual experience by reflecting ambient light, making them distinct from traditional LCD or LED displays. They are favored for their eye comfort, energy efficiency, and material flexibility, which make them appealing for a wide range of embedded devices, including eReaders, smartphones, tablets, and wearables. However, EPDs face a significant challenge: the necessity for a fast refresh rate (to maintain an acceptable display performance) introduces a pronounced ghosting effect. This effect results in noticeable color discrepancies between the displayed and source images, harming the user experience and hindering EPDs’ broader application in devices requiring dynamic content display. This article proposes a software-based solution to address the ghosting issue in EPDs. Our approach involves developing analytical models to predict the occurrence of ghosting effects and adjusting the source images to counteract the anticipated color deviations, which can reduce the perceivable ghosts on the display. Experimental evaluation conducted on real-world EPDs validates the effectiveness of our proposed approach in reducing the ghosting effect. Tao Hu 0018, Menglong Cui, Mingsong Lv, Tao Yang 0024, Yiyang Zhou, Qingxu Deng, Chun Jason Xue, Nan Guan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Center Feature Fusion: Selective Multi-Sensor Fusion of Center-based ObjectsabstractLeveraging multi-modal fusion, especially between camera and LiDAR, has become essential for building accurate and robust 3D object detection systems for autonomous vehicles. Until recently, point decorating approaches, in which point clouds are augmented with camera features, have been the dominant approach in the field. However, these approaches fail to utilize the higher resolution images from cameras. Recent works projecting camera features to the bird's-eye-view (BEV) space for fusion have also been proposed, however they require projecting millions of pixels, most of which only contain background information. In this work, we propose a novel approach Center Feature Fusion (CFF), in which we leverage center-based detection networks in both the camera and LiDAR streams to identify relevant object locations. We then use the center-based detection to identify the locations of pixel features relevant to object locations, a small fraction of the total number in the image. These are then projected and fused in the BEV frame. On the nuScenes dataset, we outperform the LiDAR-only baseline by 4.9% mAP while fusing up to 100x fewer features than other fusion methods. Philip L. Jacobson, Yiyang Zhou, Masayoshi Tomizuka, Ming C. Wu |
ICRA | 2 |
| 2023 | Contrastive Label EnhancementabstractLabel distribution learning (LDL) is a new machine learning paradigm for solving label ambiguity. Since it is difficult to directly obtain label distributions, many studies are focusing on how to recover label distributions from logical labels, dubbed label enhancement (LE). Existing LE methods estimate label distributions by simply building a mapping relationship between features and label distributions under the supervision of logical labels. They typically overlook the fact that both features and logical labels are descriptions of the instance from different views. Therefore, we propose a novel method called Contrastive Label Enhancement (ConLE) which integrates features and logical labels into the unified projection space to generate high-level features by contrastive learning strategy. In this approach, features and logical labels belonging to the same sample are pulled closer, while those of different samples are projected farther away from each other in the projection space. Subsequently, we leverage the obtained high-level features to gain label distributions through a well-designed training strategy that considers the consistency of label attributes. Extensive experiments on LDL benchmark datasets demonstrate the effectiveness and superiority of our method. Yiyang Zhou, Jihua Zhu, Xinyuan Liu 0001, Wenbiao Yan |
IJCAI | 2 |
| 2023 | Semantically consistent multi-view representation learning
Yiyang Zhou, Qinghai Zheng, Shunshun Bai, Jihua Zhu |
Knowl. Based Syst. | 1 |
| 2022 | What Matters for 3D Scene Flow Network
Guangming Wang 0001, Yunzhe Hu, Zhe Liu 0022, Yiyang Zhou, Masayoshi Tomizuka, Hesheng Wang 0001 |
ECCV (33) | 4 |
| 2022 | DetMatch: Two Teachers are Better than One for Joint 2D and 3D Semi-Supervised Object Detection
Jinhyung Park, Chenfeng Xu, Yiyang Zhou, Masayoshi Tomizuka |
ECCV (10) | 3 |
| 2022 | S3Net: Spectral-Spatial Siamese Network for Few-Shot Hyperspectral Image ClassificationabstractDeep learning (DL) has shown great potentials for hyperspectral image (HSI) classification due to its powerful ability of nonlinear modeling and end-to-end optimization. However, DL models are easily get trapped into overfitting due to limited training labels since the labeling process is time-consuming and laborious in real classification scenario. To overcome this issue, we propose a novel spectral-spatial siamese network (S3Net) for few-shot HSI classification. Firstly, a lightweight spectral-spatial network (SSN) composed of 1-D and 2-D convolution is proposed to extract spectral-spatial features. Secondly, S3Net is constructed by two SSNs in dual branches, which can augment training set by feeding sample pairs into each branch, and thus enhancing the model separability. To provide more features for the model, differentiated patches are fed into each branch, where negative samples are random selected to avoid redundancy. Finally, a weighted contrastive loss is designed to promote the model to fit in the right direction by focusing on sample pairs that are hardly to be identified. Moreover, another adaptive cross entropy loss is conceived to learn the fusion ratio of the two branches. Experiments based on three commonly used HSI data sets demonstrate that S3Net outperforms traditional and state-of-the-art DL-based HSI classification methods under few-shot training scenario. In addition, the weighted contrastive loss and the adaptive cross entropy loss jointly improve the discrimination power of the model. Zhaohui Xue, Yiyang Zhou, Peijun Du |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Labels are Not Perfect: Inferring Spatial Uncertainty in Object DetectionabstractThe availability of many real-world driving datasets is a key reason behind the recent progress of object detection algorithms in autonomous driving. However, there exist ambiguity or even failures in object labels due to error-prone annotation process or sensor observation noise. Current public object detection datasets only provide deterministic object labels without considering their inherent uncertainty, as does the common training process or evaluation metrics for object detectors. As a result, an in-depth evaluation among different object detection methods remains challenging, and the training process of object detectors is sub-optimal, especially in probabilistic object detection. In this work, we infer the uncertainty in bounding box labels from LiDAR point clouds based on a generative model, and define a new representation of the probabilistic bounding box through a spatial uncertainty distribution. Comprehensive experiments show that the proposed model reflects complex environmental noises in LiDAR perception and the label quality. Furthermore, we propose Jaccard IoU (JIoU) as a new evaluation metric that extends IoU by incorporating label uncertainty. We conduct an in-depth comparison among several LiDAR-based object detectors using the JIoU metric. Finally, we incorporate the proposed label uncertainty in a loss function to train a probabilistic object detector and to improve its detection accuracy. We verify our proposed methods on two public datasets (KITTI, Waymo), as well as on simulation data. Code is released athttps://github.com/ZiningWang/Inferring-Spatial-Uncertainty-in-Object-Detection. Di Feng, Yiyang Zhou, Lars Rosenbaum, Fabian Timm, Klaus Dietmayer, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | A Simple and Efficient Multi-task Network for 3D Object Detection and Road UnderstandingabstractDetecting dynamic objects and predicting static road information such as drivable areas and ground heights are crucial for safe autonomous driving. Previous works studied each perception task separately, and lacked a collective quantitative analysis. In this work, we show that it is possible to perform all perception tasks via a simple and efficient multi-task network. Our proposed network, LidarMTL, takes raw LiDAR point cloud as inputs, and predicts six perception outputs for 3D object detection and road understanding. The network is based on an encoder-decoder architecture with 3D sparse convolution and deconvolution operations. Extensive experiments verify the proposed method with competitive accuracies compared to state-of-the-art object detectors and other task-specific networks. LidarMTL is also leveraged for online localization. Code and pre-trained model have been made available at https://github.com/frankfengdi/LidarMTL. Di Feng, Yiyang Zhou, Chenfeng Xu, Masayoshi Tomizuka |
IROS | 2 |
| 2021 | Automatic Construction of Lane-level HD Maps for Urban ScenesabstractHigh definition (HD) maps have demonstrated their essential roles in enabling full autonomy, especially in complex urban scenarios. As a crucial layer of the HD map, lane-level maps are particularly useful: they contain geometrical and topological information for both lanes and intersections. However, large scale construction of HD maps is limited by tedious human labeling and high maintenance costs, especially for urban scenarios with complicated road structures and irregular markings. This paper proposes an approach based on semantic-particle filter to tackle the automatic lane-level mapping problem in urban scenes. The map skeleton is firstly structured as a directed cyclic graph from online mapping database OpenStreetMap. Our proposed method then performs semantic segmentation on 2D front-view images from ego vehicles and explores the lane semantics on a birds-eye-view domain with true topographical projection. Exploiting OpenStreetMap, we further infer lane topology and reference trajectory at intersections with the aforementioned lane semantics. The proposed algorithm has been tested in densely urbanized areas, and the results demonstrate accurate and robust reconstruction of the lane-level HD map. Yiyang Zhou, Yuichi Takeda, Masayoshi Tomizuka |
IROS | 1 |
| 2020 | UrbanLoco: A Full Sensor Suite Dataset for Mapping and Localization in Urban ScenesabstractMapping and localization is a critical module of autonomous driving, and significant achievements have been reached in this field. Beyond Global Navigation Satellite System (GNSS), research in point cloud registration, visual feature matching, and inertia navigation has greatly enhanced the accuracy and robustness of mapping and localization in different scenarios. However, highly urbanized scenes are still challenging: LIDAR- and camera-based methods perform poorly with numerous dynamic objects; the GNSS-based solutions experience signal loss and multi-path problems; the inertia measurement units (IMU) suffer from drifting. Unfortunately, current public datasets either do not adequately address this urban challenge or do not provide enough sensor information related to map-ping and localization. Here we present UrbanLoco: a mapping/localization dataset collected in highly-urbanized environments with a full sensor-suite. The dataset includes 13 trajectories collected in San Francisco and Hong Kong, covering a total length of over 40 kilometers. Our dataset includes a wide variety of urban terrains: urban canyons, bridges, tunnels, sharp turns, etc. More importantly, our dataset includes information from LIDAR, cameras, IMU, and GNSS receivers. Now the dataset is publicly available through the link in the footnote1. Weisong Wen, Yiyang Zhou, Guohao Zhang, Saman Fahandezh-Saadi, Xiwei Bai, Masayoshi Tomizuka, Li-Ta Hsu |
ICRA | 2 |
| 2020 | Inferring Spatial Uncertainty in Object DetectionabstractThe availability of real-world datasets is the prerequisite for developing object detection methods for autonomous driving. While ambiguity exists in object labels due to error-prone annotation process or sensor observation noises, current object detection datasets only provide deterministic annotations without considering their uncertainty. This precludes an in-depth evaluation among different object detection methods, especially for those that explicitly model predictive probability. In this work, we propose a generative model to estimate bounding box label uncertainties from LiDAR point clouds, and define a new representation of the probabilistic bounding box through spatial distribution. Comprehensive experiments show that the proposed model represents uncertainties commonly seen in driving scenarios. Based on the spatial distribution, we further propose an extension of IoU, called the Jaccard IoU (JIoU), as a new evaluation metric that incorporates label uncertainty. Experiments on the KITTI and the Waymo Open Datasets show that JIoU is superior to IoU when evaluating probabilistic object detectors. Di Feng, Yiyang Zhou, Lars Rosenbaum, Fabian Timm, Klaus Dietmayer, Masayoshi Tomizuka |
IROS | 3 |
| 2017 | Visual Robotic Object Grasping Through Combining RGB-D Data and 3D Meshes
Yiyang Zhou, Wenhai Wang, Wenjie Guan, Yirui Wu, Heng Lai, Tong Lu 0002, Min Cai |
MMM (1) | 1 |