EDBT 2026 Demo / reviewers in the wild / expert
Tianjun Zhang
dblp:193/9376
· DBLP profile ↗
45ranked-venue papers
15as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 9 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 12 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Agentix: An Efficient Serving Engine for LLM Agents as General Programs
Michael Luo, Xiaoxiang Shi, Colin Cai, Tianjun Zhang, Justin Wong, Yanping Huang, Joseph Gonzalez 0001, Ion Stoica |
NSDI | 4 |
| 2026 | HiGS-Calib: A Hierarchical 3D Gaussian Splatting-Based Targetless Local-Consistent LiDAR-Camera Calibration Methodabstract3D Gaussian Splatting (3DGS) has emerged as a powerful scene representation, offering geometrically dense and photometrically accurate modeling capabilities that present a promising new paradigm for accurate targetless sensor calibration. Current 3DGS-based LiDAR-camera calibration methods usually highly rely on the joint optimization of the Gaussian model and extrinsics, and thus suffer two critical limitations. On the one hand, the global optimization is usually sensitive to the accumulated localization error of LiDAR. On the other hand, the inaccurate extrinsics may cause oscillations during the joint optimization. Specifically, there is a fundamental dilemma: accurate extrinsic calibration requires accurate scene models, while constructing accurate models itself depends on accurate extrinsics. This dilemma frequently triggers oscillatory optimization trajectories, significantly increasing vulnerability to premature convergence at suboptimal states. To address these challenges, we propose HiGS-Calib, a novel 3DGS calibration pipeline featuring the integration of our proposed Local-Consistent Photometric-Geometric (LCPG) error model and the hierarchical architecture. The LCPG error leverages the spatial consistency within local windows to quantify pose misalignment using only geometric attributes of the 3DGS model, bypassing color-pose reliance constraints. Besides, diverging from joint optimization paradigms, HiGS-Calib implements coarse-to-fine iterative optimization, decoupling scene modeling from extrinsic refinement and thereby achieving stable and accurate calibration. Extensive evaluation demonstrates the significantly improved calibration accuracy and stability of our HiGS-Calib over other state-of-the-art methods. To make our results reproducible, the source code has been released at https://github.com/IRMVLab/HiGS-Calib. Tianjun Zhang, Lin Zhang 0014, Hesheng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | SuperCorrect: Advancing Small LLM Reasoning with Thought Template Distillation and Self-CorrectionabstractLarge language models (LLMs) like GPT-4, DeepSeek-R1, and ReasonFlux have shown significant improvements in various reasoning tasks. However, smaller LLMs still struggle with complex mathematical reasoning because they fail to effectively identify and correct reasoning errors. Recent reflection-based methods aim to address these issues by enabling self-reflection and self-correction, but they still face challenges in independently detecting errors in their reasoning steps. To overcome these limitations, we propose SuperCorrect, a novel two-stage framework that uses a large teacher model to supervise and correct both the reasoning and reflection processes of a smaller student model. In the first stage, we extract hierarchical high-level and detailed thought templates from the teacher model to guide the student model in eliciting more fine-grained reasoning thoughts. In the second stage, we introduce cross-model collaborative direct preference optimization (DPO) to enhance the self-correction abilities of the student model by following the teacher's correction traces during training. This cross-model DPO approach teaches the student model to effectively locate and resolve erroneous thoughts with error-driven insights from the teacher model, breaking the bottleneck of its thoughts and acquiring new skills and knowledge to tackle challenging problems. Extensive experiments consistently demonstrate our superiority over previous methods. Notably, our SuperCorrect-7B model significantly surpasses powerful DeepSeekMath-7B by 7.8\%/5.3\% and Qwen2.5-Math-7B by 15.1\%/6.3\% on MATH/GSM8K benchmarks, achieving new SOTA performance among all 7B models. Code is available at: https://github.com/YangLing0818/SuperCorrect-llm Ling Yang 0006, Zhaochen Yu, Tianjun Zhang, Minkai Xu, Joseph Gonzalez 0001, Bin Cui 0001, Shuicheng Yan |
ICLR | 3 |
| 2025 | LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeabstractLarge Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from academia and industry. However, as new and improved LLMs are developed, existing evaluation benchmarks (e.g., HumanEvla, MBPP) are no longer sufficient for assessing their capabilities suffering from data contamination, overfitting, saturation, and focus on merely code generation. In this work, we propose LiveCodeBench, a comprehensive and contamination-free evaluation of LLMs for code, which collects new problems over time from contests across three competition platforms, Leetcode, Atcoder, and Codeforces. Notably, our benchmark also focuses on a broader range of code-related capabilities, such as self-repair, code execution, and test output prediction, beyond just code generation. Currently, LiveCodeBench hosts over six hundred coding problems that were published between May 2023 and Aug 2024. We evaluate over 50 LLMs on LiveCodeBench (LCB for brevity) presenting the largest evaluation study of code LLMs on competition problems. Based on the study, we present novel empirical findings on contamination, overfitting, and holistic evaluations. We demonstrate that time-segmented evaluations serve as a robust approach to evade contamination; they are successful at detecting contamination across a wide range of open and closed models including GPT-4O, Claude, Deepseek, and Codestral. Next, we highlight overfitting and saturation of traditional coding benchmarks like HumanEvla and demonstrate LCB allows more reliable evaluations. Finally, our holistic evaluation scenarios allow for measuring the different capabilities of programming agents in isolation. Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida I. Wang, Armando Solar-Lezama, Koushik Sen, Ion Stoica |
ICLR | 6 |
| 2025 | Copilot Arena: A Platform for Code LLM Evaluation in the WildabstractEvaluating in-the-wild coding capabilities of large language models (LLMs) is a challenging endeavor with no existing solution. We introduce Copilot Arena, a platform to collect user preferences through native integration into a developer's working environment. Copilot Arena comprises a novel interface for comparing pairs of model outputs, a sampling strategy to reduce experienced latency, and a prompting scheme to enable code completion functionality. Copilot Arena has served over 4.5 million suggestions from 10 models and collected over 11k pairwise judgements. Our results highlight the importance of model evaluations in integrated settings. We find that model rankings from Copilot Arena differ from those of existing evaluations, which we attribute to the unique distribution of data and tasks contained in Copilot Arena. We also identify novel insights into human preferences on code such as an observed consistency in user preference across programming languages yet significant variation in preference due to task category. We open-source Copilot Arena and release data to enable human-centric evaluations and improve understanding of coding assistants. Wayne Chi, Valerie Chen, Anastasios Angelopoulos, Wei-Lin Chiang, Aditya Mittal, Naman Jain, Tianjun Zhang, Ion Stoica, Chris Donahue, Ameet Talwalkar |
ICML | 7 |
| 2025 | Digital twin based intelligent control system on gas extraction from boreholes and experimental researchabstractIntelligent gas extraction in mines is a critical enabler for the realization of smart mine development. As an essential component of this process, the intelligent control of borehole gas extraction relies on real-time monitoring data from numerous extraction parameters, integrated with next-generation information technologies, to achieve objectives such as intelligent deployment of negative pressure, control of inefficient boreholes, and evaluate extraction effectiveness. This paper constructs an intelligent control system for borehole gas extraction based on a digital twin “Four-Dimensional” framework, enabling bidirectional mapping between physical entities and virtual digital twins through the integration of physical entities (PE) as foundational carriers, virtual entities (VE) as three-dimensional models, digital twin data (DD) as the control core, and services (SS) and connectivity (CN) as the methodological framework. Key technologies include developing a control model for the gas flow process from coal seam to borehole, processing multi-source extraction data using data fusion methods, constructing virtual twins with SolidWorks, and designing control schemes based on Model Predictive Control (MPC) algorithm. In order to verify the rationality and feasibility of the system, a pilot study on the digital twin for intelligent control of borehole gas extraction was carried out. The results show that the concentration of gas extraction after control rises significantly, and the gas extraction concentration of borehole 1# reaches the optimal range when the valve is opened to 75 %, while that of borehole 2# reaches the maximum concentration of the adjustable range when the valve is opened to 70 %, which meets the control target. Suinan He, Hongyu Pan, Tianjun Zhang, Xinshuang Cao, Yilun Xue |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | I-DACS: Always Maintaining Consistency Between Poses and the Field for Radiance Field Construction Without Pose PriorabstractThe radiance field, emerging as a novel 3D scene representation, has found widespread application across diverse fields. Standard radiance field construction approaches rely on the ground-truth poses of key-frames, while building the field without pose prior remains a formidable challenge. Recent advancements have made strides in mitigating this challenge, albeit to a limited extent, by jointly optimizing poses and the radiance field. However, in these schemes, the consistency between the radiance field and poses is achieved completely by training. Once the poses of key-frames undergo changes, long-term training is required to readjust the field to fit them. To address such a limitation, we propose a new solution for radiance field construction without pose prior, namely I-DACS (Incremental radiance field construction with Direction-Aware Color Sampling). Diverging from most of the existing global optimization solutions, we choose to incrementally solve the poses and construct a radiance field within a sliding-window framework. The poses are unequivocally retrieved from the radiance field, devoid of any constraints and accompanying noise from other observation models, so as to achieve the consistency of poses to the field. Besides, in the radiance field, the color information is much higher-frequency and more time-consuming to learn compared with the density. To accelerate training, we isolate the color information to a distinct color field, and construct the color field based on an innovative direction-aware color sampling strategy, by which the color field can be derived directly from images without training. The color field obtained in this way is always consistent with the poses, and intricate details of training images can be retained to the utmost extent. Extensive experimental results evidently showcase both the remarkable training speed and the outstanding performance in rendering quality and localization accuracy achieved by I-DACS. To make our results reproducible, the source code has been released athttps://cslinzhang.github.io/I-DACS-MainPage/. Tianjun Zhang, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Towards a Robust Visual-Inertial-Surround-View SLAM System for Autonomous Indoor ParkingabstractAn autonomous parking system is a low-speed unmanned driving system applied in indoor parking environments. Real-time and high-precision vehicle localization and map construction of the environment are two core functional modules of the system. Camera and IMU (Inertial Measurement Unit) sensors provide complementary data to create a Visual-Inertial Simultaneous Localization and Mapping (VI-SLAM) system. However, existing SLAM systems face challenges in complex parking environments. Moreover, limitations inherent in VI-SLAM systems further compromise their perception accuracy, affecting both localization and optimization. This article addresses the shortcomings of current VI-SLAM systems by proposing the RVIS SLAM system. This robust semantic SLAM system integrates data from three sensors: a front-view camera, an IMU, and a surround-view system. To ensure localization accuracy, the system utilizes metric information from common semantic objects on the ground. These objects include parking-slots, speed bumps, and parking-slot numbers captured in surround-view images to build scale-aware constraints. These constraints refine the initial scale of the SLAM system, which is often compromised under low IMU excitation conditions. Additionally, in optimization, SLAM systems ideally assume that the front-end produces optimization graphs without data association outliers. However, in real-world indoor parking environments, sensor noise and vehicle vibrations make this assumption unrealistic. To mitigate the adverse effects of outliers in SLAM systems, this article proposes a robust surround-view semantic data association strategy. This strategy quantifies the uncertainty of surround-view semantic landmarks for the first time, ensuring reliable localization and mapping in challenging environments. Extensive experiments in typical indoor parking environments validate the effectiveness and efficiency of the proposed RVIS SLAM system. Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Shengjie Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | LLoCO: Learning Long Contexts OfflineabstractSijun Tan, Xiuyu Li, Shishir G Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph E. Gonzalez, Raluca Ada Popa. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sijun Tan, Xiuyu Li, Shishir G. Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph Gonzalez 0001, Raluca A. Popa |
EMNLP | 5 |
| 2024 | AgentBench: Evaluating LLMs as AgentsabstractThe potential of Large Language Model (LLM) as agents has been widely acknowledged recently.
Thus, there is an urgent need to quantitatively evaluate LLMs as agents on challenging tasks in interactive environments.
We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities.
Our extensive test over 29 API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B.
We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents.
Improving instruction following and training on high quality multi-round alignment data could improve agent performance.
And different from existing assumptions, training on code present ambivalent impacts on different agent tasks.
Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench. Xiao Liu 0036, Hao Yu 0030, Hanchen Zhang, Yifan Xu 0014, Xuanyu Lei, Hanyu Lai, Yu Gu 0016, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng 0001, Aohan Zeng, Zhengxiao Du, Sheng Shen 0001, Tianjun Zhang, Yu Su 0001, Huan Sun 0001, Minlie Huang, Yuxiao Dong, Jie Tang 0001 |
ICLR | 17 |
| 2024 | LLM-Assisted Code Cleaning For Training Accurate Code GeneratorsabstractNatural language to code generation is an important application area of LLMs and has received wide attention from the community.
The majority of relevant studies have exclusively concentrated on increasing the quantity and functional correctness of training sets while disregarding other stylistic elements of programs. More recently, data quality has garnered a lot of interest and multiple works have showcased its importance for improving performance. In this work, we investigate data quality for code and find that making the code more structured and readable leads to improved code generation performance of the system. We build a novel data-cleaning pipeline that uses these principles to transform existing programs by 1.) renaming variables, 2.) modularizing and decomposing complex code into smaller helper sub-functions, and 3.) inserting natural-language based planning annotations. We evaluate our approach on two challenging algorithmic code generation benchmarks and find that fine-tuning CodeLLaMa-7B on our transformed programs improves the performance by up to \textbf{30\%} compared to fine-tuning on the original dataset. Additionally, we demonstrate improved performance from using a smaller amount of higher-quality data, finding that a model fine-tuned on the entire original dataset is outperformed by a model trained on one-eighth of our cleaned dataset. Even in comparison to closed-source models, our models outperform the much larger AlphaCode models. Naman Jain, Tianjun Zhang, Wei-Lin Chiang, Joseph Gonzalez 0001, Koushik Sen, Ion Stoica |
ICLR | 2 |
| 2024 | R2E: Turning any Github Repository into a Programming Agent EnvironmentabstractWhile Large Language Models’ (LLMs) coding capabilities have advanced rapidly, corresponding evaluation benchmarks on real-world programming setups are yet to catch up. Building a scalable and interactive testbed for evaluating general-purpose AI coding agents for real-world code has been challenging, particularly due to a lack of high-quality test suites available. In this paper, we present Repository to Environment (R2E), a framework that can turn any GitHub repository into a test environment to evaluate the performance of code-generating systems, both static and interactive. R2E is powered by a synergistic combination of program analysis and LLMs to construct equivalence test harnesses for any GitHub function. We instantiate our framework to build the first large-scale benchmark, R2E-Eval1, for building realistic environments for AI coding assistants. Our results demonstrate that even when SOTA models cannot generate correct solutions with advanced prompting techniques, they can effectively use environment feedback highlighting the need to move from static functional coding to interactive programming paradigm. We hope that our framework (and the instantiated benchmark) can motivate research directions by providing web-scale open-ended coding environments. R2E code is available at https://r2e.dev/ Naman Jain, Manish Shetty, Tianjun Zhang, King Han, Koushik Sen, Ion Stoica |
ICML | 3 |
| 2024 | In-Context Principle Learning from MistakesabstractIn-context learning (ICL, also known as few-shot prompting) has been the standard method of adapting LLMs to downstream tasks, by learning from a few input-output examples. Nonetheless, all ICL-based approaches only learn from correct input-output pairs. In this paper, we revisit this paradigm, by learning more from the few given input-output examples. We introduce Learning Principles (LEAP): First, we intentionally induce the model to make mistakes on these few examples; then we reflect on these mistakes, and learn explicit task-specific “principles” from them, which help solve similar problems and avoid common mistakes; finally, we prompt the model to answer unseen test questions using the original few-shot examples and these learned general principles. We evaluate LEAP on a wide range of benchmarks, including multi-hop question answering (Hotpot QA), textual QA (DROP), Big-Bench Hard reasoning, and math problems (GSM8K and MATH); in all these benchmarks, LEAP improves the strongest available LLMs such as GPT-3.5-turbo, GPT-4, GPT-4-turbo and Claude-2.1. For example, LEAP improves over the standard few-shot prompting using GPT-4 by 7.5% in DROP, and by 3.3% in HotpotQA. Importantly, LEAP does not require any more input or examples than the standard few-shot prompting settings. Tianjun Zhang, Aman Madaan, Luyu Gao, Steven Zheng, Swaroop Mishra, Yiming Yang 0002, Niket Tandon, Uri Alon 0002 |
ICML | 1 |
| 2024 | Gorilla: Large Language Model Connected with Massive APIsabstractLarge Language Models (LLMs) have seen an impressive wave of advances, with
models now excelling in a variety of tasks, such as mathematical reasoning and
program synthesis. However, their potential to effectively use tools via API calls
remains unfulfilled. This is a challenging task even for today’s state-of-the-art
LLMs such as GPT-4 largely due to their unawareness of what APIs are available
and how to use them in a frequently updated tool set. We develop Gorilla, a
finetuned LLaMA model that surpasses the performance of GPT-4 on writing API
calls. Trained with the novel Retriever Aware Training (RAT), when combined
with a document retriever, Gorilla demonstrates a strong capability to adapt to
test-time document changes, allowing flexible user updates or version changes.
It also substantially mitigates the issue of hallucination, commonly encountered
when prompting LLMs directly. To evaluate the model’s ability, we introduce
APIBench, a comprehensive dataset consisting of HuggingFace, TorchHub, and
TensorHub APIs. The successful integration of the retrieval system with Gorilla
demonstrates the potential for LLMs to use tools more accurately, keep up with
frequently updated documentation, and consequently increase the reliability and
applicability of their outputs. Gorilla’s code, model, data, and demo are available
at: https://gorilla.cs.berkeley.edu Shishir G. Patil, Tianjun Zhang, Xin Wang 0066, Joseph Gonzalez 0001 |
NeurIPS | 2 |
| 2024 | Recursive Introspection: Teaching Language Model Agents How to Self-ImproveabstractA central piece in enabling intelligent agentic behavior in foundation models is to make them capable of introspecting upon their behavior, reasoning, and correcting their mistakes as more computation or interaction is available. Even the strongest proprietary large language models (LLMs) do not quite exhibit the ability of continually improving their responses sequentially. In this paper, we develop $\textbf{RISE:}$ $\textbf{R}$ecursive $\textbf{I}$ntro$\textbf{S}$p$\textbf{E}$ction, an approach for fine-tuning LLMs to introduce this capability, despite prior work hypothesizing that this capability may not be possible to attain. Our approach prescribes an iterative fine-tuning procedure, which attempts to teach the model how to alter its response after having executed previously unsuccessful attempts to solve a hard test-time problem, with optionally additional environment feedback. RISE poses fine-tuning for a single-turn prompt as solving a multi-turn Markov decision process (MDP), where the initial state is the prompt. Inspired by principles in online imitation and offline reinforcement learning, we propose strategies for multi-turn data collection and training so as to imbue an LLM with the capability to recursively detect and correct its previous mistakes in subsequent iterations. Our experiments show that RISE enables Llama2, Llama3, and Mistral models to improve themselves with more turns on reasoning tasks, outperforming several single-turn strategies given an equal amount of inference-time computation. We also find that RISE scales well, often attaining larger benefits with more capable models, without disrupting one-turn abilities as a result of expressing more complex distributions. Yuxiao Qu, Tianjun Zhang, Naman Garg, Aviral Kumar |
NeurIPS | 2 |
| 2024 | Buffer of Thoughts: Thought-Augmented Reasoning with Large Language ModelsabstractWe introduce Buffer of Thoughts (BoT), a novel and versatile thought-augmented reasoning approach for enhancing accuracy, efficiency and robustness of large language models (LLMs). Specifically, we propose meta-buffer to store a series of informative high-level thoughts, namely thought-template, distilled from the problem-solving processes across various tasks. Then for each problem, we retrieve a relevant thought-template and adaptively instantiate it with specific reasoning structures to conduct efficient reasoning. To guarantee the scalability and stability, we further propose buffer-manager to dynamically update the meta-buffer, thus enhancing the capacity of meta-buffer as more tasks are solved. We conduct extensive experiments on 10 challenging reasoning-intensive tasks, and achieve significant performance improvements over previous SOTA methods: 11\% on Game of 24, 20\% on Geometric Shapes and 51\% on Checkmate-in-One. Further analysis demonstrate the superior generalization ability and model robustness of our BoT, while requiring only 12\% of the cost of multi-query prompting methods (e.g., tree/graph of thoughts) on average. Code is available at: https://github.com/YangLing0818/buffer-of-thought-llm Ling Yang 0006, Zhaochen Yu, Tianjun Zhang, Shiyi Cao, Minkai Xu, Wentao Zhang 0001, Joseph Gonzalez 0001, Bin Cui 0001 |
NeurIPS | 3 |
| 2024 | Multitask Vision-Language Prompt TuningabstractPrompt Tuning, conditioning on task-specific learned prompt vectors, has emerged as a data-efficient and parameter-efficient method for adapting large pretrained vision-language models to multiple downstream tasks. However, existing approaches usually consider learning prompt vectors for each task independently from scratch, thereby failing to exploit the rich shareable knowledge across different vision-language tasks. In this paper, we propose multitask vision-language prompt tuning (MVLPT), which incorporates cross-task knowledge into prompt tuning for vision-language models. Specifically, (i) we demonstrate the effectiveness of learning a single transferable prompt from multiple source tasks to initialize the prompt for each target task; (ii) we show many target tasks can benefit each other from sharing prompt vectors and thus can be jointly learned via multitask prompt tuning. We benchmark the proposed MVLPT using three representative prompt tuning methods, namely text prompt tuning, visual prompt tuning, and the unified vision-language prompt tuning. Results in 20 vision tasks demonstrate that the proposed approach outperforms all single-task baseline prompt tuning methods, setting the new state-of-the-art on the few-shot Elevater benchmarks and cross-task generalization benchmarks. To understand where the cross-task knowledge is most effective, we also conduct a large-scale study on task transferability with 20 vision tasks in 400 combinations for each prompt tuning method. It shows that the most performant MVLPT for each prompt tuning method prefers different task combinations and many tasks can benefit each other, depending on their visual similarity and label similarity. Sheng Shen 0001, Shijia Yang, Tianjun Zhang, Bohan Zhai, Joseph Gonzalez 0001, Kurt Keutzer, Trevor Darrell |
WACV | 3 |
| 2024 | An Underwater Organism Image Dataset and a Lightweight Module Designed for Object Detection NetworksabstractLong-term monitoring and recognition of underwater organism objects are of great significance in marine ecology, fisheries science and many other disciplines. Traditional techniques in this field, including manual fishing-based ones and sonar-based ones, are usually flawed. Specifically, the method based on manual fishing is time-consuming and unsuitable for scientific researches, while the sonar-based one, has the defects of low acoustic image accuracy and large echo errors. In recent years, the rapid development of deep learning and its excellent performance in computer vision tasks make vision-based solutions feasible. However, the researches in this area are still relatively insufficient in mainly two aspects. First, to our knowledge, there is still a lack of large-scale datasets of underwater organism images with accurate annotations. Second, in consideration of the limitation on hardware resources of underwater devices, an underwater organism detection algorithm that is both accurate and lightweight enough to be able to infer in real time is still lacking. As an attempt to fill in the aforementioned research gaps to some extent, we established the Multiple Kinds of Underwater Organisms (MKUO) dataset with accurate bounding box annotations of taxonomic information, which consists of 10,043 annotated images, covering eighty-four underwater organism categories. Based on our benchmark dataset, we evaluated a series of existing object detection algorithms to obtain their accuracy and complexity indicators as the baseline for future reference. In addition, we also propose a novel lightweight module, namely Sparse Ghost Module, designed especially for object detection networks. By substituting the standard convolution with our proposed one, the network complexity can be significantly reduced and the inference speed can be greatly improved without obvious detection accuracy loss. To make our results reproducible, the dataset and the source code are available online at https://cslinzhang.github.io/MKUO-and-Sparse-Ghost-Module/ . Jiafeng Huang, Tianjun Zhang, Shengjie Zhao 0001, Lin Zhang 0014, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Efficient Planning in a Compact Latent Action Space
Zhengyao Jiang, Tianjun Zhang, Michael Janner, Tim Rocktäschel, Edward Grefenstette, Yuandong Tian |
ICLR | 2 |
| 2023 | Latent Variable Representation for Reinforcement Learning
Tongzheng Ren, Chenjun Xiao, Tianjun Zhang, Na Li 0002, Zhaoran Wang 0001, Sujay Sanghavi, Dale Schuurmans, Bo Dai 0001 |
ICLR | 3 |
| 2023 | Spectral Decomposition Representation for Reinforcement Learning
Tongzheng Ren, Tianjun Zhang, Lisa Lee, Joseph Gonzalez 0001, Dale Schuurmans, Bo Dai 0001 |
ICLR | 2 |
| 2023 | TEMPERA: Test-Time Prompt Editing via Reinforcement Learning
Tianjun Zhang, Xuezhi Wang 0002, Denny Zhou, Dale Schuurmans, Joseph Gonzalez 0001 |
ICLR | 1 |
| 2023 | Adaptively Hashing 3DLUTs for Lightweight Real-time Image EnhancementabstractImage enhancement is an essential and longstanding task in computer vision, in which the 3D lookup table (3DLUT) is widely used due to its powerful mapping capability and high time efficiency. However, the 3DLUT generally requires a large parameter amount since it caches the mapping results for all colors of the entire discrete color space with a 3D array. As a result, standard 3DLUT-based enhancement methods suffer from heavy memory footprints, limiting their practical applications. Based on the analyses of the inherent low grid utilization rate of 3DLUT, we propose HashLUT, an efficient hash form of the standard 3DLUT, and further build a lightweight real-time image enhancement network that adaptively learns HashLUTs and handles hash collisions end-to-end. Experiments on two benchmarks demonstrate that our model achieves comparable enhancement performance to the state-of-the-art methods with significantly fewer parameters. Source codes are available at https://github.com/Xian-Bei. Lin Zhang 0014, Tianjun Zhang, Dongqing Wang |
ICME | 3 |
| 2023 | The Wisdom of Hindsight Makes Language Models Better Instruction FollowersabstractReinforcement learning has seen wide success in finetuning large language models to better align with instructions via human feedback. The so-called algorithm, Reinforcement Learning with Human Feedback (RLHF) demonstrates impressive performance on the GPT series models. However, the underlying reinforcement learning algorithm is complex and requires additional training for reward and value networks. In this paper, we consider an alternative approach: converting feedback to instruction by relabeling the original one and training the model for better alignment in a supervised manner. Such an algorithm doesn't require any additional parameters except for the original language model and maximally reuses the pretraining pipeline. To achieve this, we formulate instruction alignment problem for language models as a goal-reaching problem in decision making. We propose Hindsight Instruction Relabeling (HIR), a novel algorithm for aligning language models with instructions. The resulting two-stage algorithm shed light to a family of reward-free approaches that utilize the hindsightly relabeled instructions based on feedback. We evaluate the performance of HIR extensively on 12 challenging BigBench reasoning tasks and show that HIR outperforms the baseline algorithms and is comparable to or even surpasses supervised fine-tuning. The implementation of HIR is available at https://github.com/tianjunz/HIR. Tianjun Zhang, Fangchen Liu, Justin Wong, Pieter Abbeel, Joseph Gonzalez 0001 |
ICML | 1 |
| 2023 | Energy-based Predictive Representations for Partially Observed Reinforcement LearningabstractIn real-world applications, handling partial observability is a common requirement for reinforcement learning algorithms, which is not captured by a Markov decision process (MDP). Although partially observable Markov decision processes (POMDPs) have been specifically designed to address this requirement, they present significant computational and statistical challenges in learning and planning. In this work, we introduce the Energy-based Predictive Representation (EPR) to provide a unified approach for designing practical reinforcement learning algorithms in both the MDP and POMDP settings. This framework enables coherent handling of learning, exploration, and planning tasks. The proposed framework leverages a powerful neural energy-based model to extract an adequate representation, allowing for efficient approximation of Q-functions. This representation facilitates the efficient computation of confidence, enabling the implementation of optimism or pessimism in planning when faced with uncertainty. Consequently, it effectively manages the trade-off between exploration and exploitation. Experimental investigations demonstrate that the proposed algorithm achieves state-of-the-art performance in both MDP and POMDP settings. Tianjun Zhang, Tongzheng Ren, Chenjun Xiao, Joseph Gonzalez 0001, Dale Schuurmans, Bo Dai 0001 |
UAI | 1 |
| 2022 | C-Planning: An Automatic Curriculum for Learning Goal-Reaching Tasks
Tianjun Zhang, Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine, Joseph Gonzalez 0001 |
ICLR | 1 |
| 2022 | Multi-objective Optimization by Learning Space Partition
Linnan Wang, Kevin Yang, Tianjun Zhang, Tian Guo 0001, Yuandong Tian |
ICLR | 4 |
| 2022 | Towards Underwater Image Restoration: A Physical-Accurate Pipeline and a Large Scale Full-Reference BenchmarkabstractUnderwater images always present low-quality features such as low contrast, blurred edges and color distortion, which brings great challenges to high-level underwater vision tasks. In this paper, a novel underwater image restoration method, namely MonoUIR (Monocular Underwater Image Restoration), is proposed, which is based on a more physical-accurate imaging model compared to existing schemes. And with the monocular depth estimation, MonoUIR has no dependence on extra ranging equipment or specific shooting operations. Experimental results demonstrate that MonoUIR overwhelmingly outperforms other physical model-based competitors. In addition, the Real-world Undersea Color Board (RUCB) dataset is established, providing the illconditioned underwater images collected in the East China Sea and the corresponding high-quality references. To our knowledge, this is the first full-reference underwater benchmark dataset collected entirely in a real-world marine environment, which will further support the fullreference evaluation of underwater image restoration approaches. The source code and the dataset are available at https://TongJiayan.github.io/MonoUIR-Homepage. Jiayan Tong, Tianjun Zhang, Lin Zhang 0014 |
ICME | 2 |
| 2022 | Making Linear MDPs Practical via Contrastive Representation LearningabstractIt is common to address the curse of dimensionality in Markov decision processes (MDPs) by exploiting low-rank representations. This motivates much of the recent theoretical study on linear MDPs. However, most approaches require a given representation under unrealistic assumptions about the normalization of the decomposition or introduce unresolved computational challenges in practice. Instead, we consider an alternative definition of linear MDPs that automatically ensures normalization while allowing efficient representation learning via contrastive estimation. The framework also admits confidence-adjusted index algorithms, enabling an efficient and principled approach to incorporating optimism or pessimism in the face of uncertainty. To the best of our knowledge, this provides the first practical representation learning method for linear MDPs that achieves both strong theoretical guarantees and empirical performance. Theoretically, we prove that the proposed algorithm is sample efficient in both the online and offline settings. Empirically, we demonstrate superior performance over existing state-of-the-art model-based and model-free algorithms on several benchmarks. Tianjun Zhang, Tongzheng Ren, Sherry Yang 0001, Joseph Gonzalez 0001, Dale Schuurmans, Bo Dai 0001 |
ICML | 1 |
| 2022 | CLUT-Net: Learning Adaptively Compressed Representations of 3DLUTs for Lightweight Image EnhancementabstractLearning-based image enhancement has made great progress recently, among which the 3-Dimensional LookUp Table (3DLUT) based methods achieve a good balance between enhancement performance and time-efficiency. Generally, the more basis 3DLUTs are used in such methods, the more application scenarios could be covered, and thus the stronger enhancement capability could be achieved. However, more 3DLUTs would also lead to the rapid growth of the parameter amount, since a single 3DLUT has as many as D3 parameters where D is the table length. A large parameter amount not only hinders the practical application of the 3DLUT-based schemes but also gives rise to the training difficulty and does harm to the effectiveness of the basis 3DLUTs, leading to even worse performances with more utilized 3DLUTs. Through in-depth analysis of the inherent compressibility of 3DLUT, we propose an effective Compressed representation of 3-dimensional LookUp Table (CLUT) which maintains the powerful mapping capability of 3DLUT but with a significantly reduced parameter amount. Based on CLUT, we further construct a lightweight image enhancement network, namely CLUT-Net, in which image-adaptive and compression-adaptive CLUTs are learned in an end-to-end manner. Extensive experimental results on three benchmark datasets demonstrate that our proposed CLUT-Net outperforms the existing state-of-the-art image enhancement methods with orders of magnitude smaller parameter amounts. The source codes are available at https://github.com/Xian-Bei/CLUT-Net. Tianjun Zhang, Lin Zhang 0014 |
ACM Multimedia | 3 |
| 2022 | Contrastive Learning as Goal-Conditioned Reinforcement LearningabstractIn reinforcement learning (RL), it is easier to solve a task if given a good representation. While deep RL should automatically acquire such good representations, prior work often finds that learning representations in an end-to-end fashion is unstable and instead equip RL algorithms with additional representation learning parts (e.g., auxiliary losses, data augmentation). How can we design RL algorithms that directly acquire good representations? In this paper, instead of adding representation learning parts to an existing RL algorithm, we show (contrastive) representation learning methods are already RL algorithms in their own right. To do this, we build upon prior work and apply contrastive representation learning to action-labeled trajectories, in such a way that the (inner product of) learned representations exactly corresponds to a goal-conditioned value function. We use this idea to reinterpret a prior RL method as performing contrastive learning, and then use the idea to propose a much simpler method that achieves similar performance. Across a range of goal-conditioned RL tasks, we demonstrate that contrastive RL methods achieve higher success rates than prior non-contrastive methods. We also show that contrastive RL outperforms prior methods on image-based tasks, without using data augmentation or auxiliary objectives Benjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan Salakhutdinov |
NeurIPS | 2 |
| 2022 | A free lunch from the noise: Provable and practical exploration for representation learningabstractRepresentation learning lies at the heart of the empirical success of deep learning for dealing with the curse of dimensionality. However, the power of representation learning has not been fully exploited yet in reinforcement learning (RL), due to i), the trade-off between expressiveness and tractability; and ii), the coupling between exploration and representation learning. In this paper, we first reveal the fact that under some noise assumption in the stochastic control model, we can obtain the linear spectral feature of its corresponding Markov transition operator in closed-form for free. Based on this observation, we propose Spectral Dynamics Embedding (SPEDE), which breaks the tradeoff and completes optimistic exploration for representation learning by exploiting the structure of the noise. We provide rigorous theoretical analysis of SPEDE, and demonstrate the practical superior performance over the existing state-of-the-art empirical algorithms on several benchmarks. Tongzheng Ren, Tianjun Zhang, Csaba Szepesvári, Bo Dai 0001 |
UAI | 2 |
| 2022 | MOFISSLAM: A Multi-Object Semantic SLAM System With Front-View, Inertial, and Surround-View Sensors for Indoor ParkingabstractThe semantic SLAM (Simultaneous Localization And Mapping) system is a crucial module for autonomous indoor parking. Visual cameras (monocular/binocular) and IMU (Inertial Measurement Unit) constitute the basic configuration to build such a system. The performance of existing SLAM systems typically deteriorates in the presence of dynamically movable objects or objects with little texture. By contrast, semantic objects on the ground embody the most salient and stable features in the indoor parking environment. Due to their inabilities to perceive such features on the ground, existing SLAM systems are prone to tracking inconsistency during navigation. In this paper, we present MOFISSLAM, a novel tightly-coupled${M}$ulti-${O}$bject semantic SLAM system integrating${F}$ront-view,${I}$nertial, and${S}$urround-view sensors for autonomous indoor parking. The proposed system moves beyond existing semantic SLAM systems by complementing the sensor configuration with a surround-view system capturing images from a top-down viewpoint. In MOFISSLAM, apart from low-level visual features and inertial motion data, typical semantic objects (parking-slots, parking-slot IDs and speed bumps) detected in surround-views are also incorporated in optimization, forming robust surround-view constraints. Specifically, each surround-view feature imposes a surround-view constraint that can be split into a contact term and a registration term. The former pre-defines the position of each individual surround-view feature subject to whether it has semantic contact with other surround-view features. Three contact modes, defined ascomplementary,adjacentandcoincident, are identified to guarantee a unified form of all contact terms. The latter further constrains by registering each surround-view observation and its position in the world coordinate system. In parallel, to objectively evaluate SLAM studies for autonomous indoor parking, a large-scale dataset with groundtruth trajectories is collected, which is the first of its kind. Its groundtruth trajectories, commonly unavailable, are obtained by tracking artificial features scattered in the indoor parking environment, whose 3D coordinates are measured with an ETS (Electronic Total Station). The collected dataset has been made publicly available athttps://shaoxuan92.github.io/MOFIS. Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | CVIDS: A Collaborative Localization and Dense Mapping Framework for Multi-Agent Based Visual-Inertial SLAMabstractNowadays, visual SLAM (Simultaneous Localization And Mapping) has become a hot research topic due to its low costs and wide application scopes. Traditional visual SLAM frameworks are usually designed for single-agent systems, completing both the localization and the mapping with sensors equipped on a single robot or a mobile device. However, the mobility and work capacity of the single agent are usually limited. In reality, robots or mobile devices sometimes may be deployed in the form of clusters, such as drone formations, wearable motion capture systems, and so on. As far as we know, existing SLAM systems designed for multi-agents are still sporadic, and most of them have non-negligible limitations in functions. Specifically, on one hand, most of the existing multi-agent SLAM systems can only extract some key features and build sparse maps. On the other hand, schemes that can reconstruct the environment densely cannot get rid of the dependence on depth sensors, such as RGBD cameras or LiDARs. Systems that can yield high-density maps just with monocular camera suites are temporarily lacking. As an attempt to fill in the research gap to some extent, we design a novel collaborative SLAM system, namely CVIDS (Collaborative Visual-Inertial Dense SLAM), which follows a centralized and loosely coupled framework and can be integrated with any existing Visual-Inertial Odometry (VIO) to accomplish the co-localization and the dense reconstruction. Integrating our proposed robust loop closure detection module and two-stage pose-graph optimization pipeline, the co-localization module of CVIDS can estimate the poses of different agents in a unified coordinate system efficiently from the packed images and local poses sent by the client-ends of different agents. Besides, our motion-based dense mapping module can effectively recover the 3D structures of selected keyframes and then fuse their depth information to the global map for reconstruction. The superior performance of CVIDS is corroborated by both quantitative and qualitative experimental results. To make our results reproducible, the source code has been released at https://cslinzhang.github.io/CVIDS. Tianjun Zhang, Lin Zhang 0014, Yang Chen 0037, Yicong Zhou |
IEEE Trans. Image Process. | 1 |
| 2022 | Online Correction of Camera Poses for the Surround-view System: A Sparse Direct ApproachabstractThe surround-view module is an indispensable component of a modern advanced driving assistance system. By calibrating the intrinsics and extrinsics of the surround-view cameras accurately, a top-down surround-view can be generated from raw fisheye images. However, poses of these cameras sometimes may change. At present, how to correct poses of cameras in a surround-view system online without re-calibration is still an open issue. To settle this problem, we introduce the sparse direct framework and propose a novel optimization scheme of a cascade structure. This scheme is actually composed of two levels of optimization and two corresponding photometric error based models are proposed. The model for the first-level optimization is called the ground model, as its photometric errors are measured on the ground plane. For the second level of the optimization, it’s based on the so-called ground-camera model, in which photometric errors are computed on the imaging planes. With these models, the pose correction task is formulated as a nonlinear least-squares problem to minimize photometric errors in overlapping regions of adjacent bird’s-eye-view images. With a cascade structure of these two levels of optimization, an appropriate balance between the speed and the accuracy can be achieved. Experiments show that our method can effectively eliminate the misalignment caused by cameras’ moderate pose changes in the surround-view system. Source code and test cases are available online at https://cslinzhang.github.io/CamPoseCorrection/ . Tianjun Zhang, Hao Deng 0002, Lin Zhang 0014, Shengjie Zhao 0001, Xiao Liu 0030, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Contrastive Code Representation LearningabstractRecent work learns contextual representations of source code by reconstructing tokens from their context.For downstream semantic understanding tasks like code clone detection, these representations should ideally capture program functionality.However, we show that the popular reconstruction-based RoBERTa model is sensitive to source code edits, even when the edits preserve semantics.We propose Con-traCode: a contrastive pre-training task that learns code functionality, not form.Con-traCode pre-trains a neural network to identify functionally similar variants of a program among many non-equivalent distractors.We scalably generate these variants using an automated source-to-source compiler as a form of data augmentation.Contrastive pretraining outperforms RoBERTa on an adversarial code clone detection benchmark by 39% AUROC.Surprisingly, improved adversarial robustness translates to better accuracy over natural code; ContraCode improves summarization and TypeScript type inference accuracy by 2 to 13 percentage points over competitive baselines.All source is available at https://github.com/parasj/contracode. Paras Jain 0001, Ajay Jain, Tianjun Zhang, Pieter Abbeel, Joseph Gonzalez 0001, Ion Stoica |
EMNLP (1) | 3 |
| 2021 | ROECS: A Robust Semi-direct Pipeline Towards Online Extrinsics Correction of the Surround-view SystemabstractGenerally, a surround-view system (SVS), which is an indispensable component of advanced driving assistant systems (ADAS), consists of four to six wide-angle fisheye cameras. As long as both intrinsics and extrinsics of all cameras have been calibrated, a top-down surround-view with the real scale can be synthesized at runtime from fisheye images captured by these cameras. However, when the vehicle is driving on the road, relative poses between cameras in the SVS may change from the initial calibrated states due to bumps or collisions. In case that extrinsics' representations are not adjusted accordingly, on the surround-view, obvious geometric misalignment will appear. Currently, the researches on correcting the extrinsics of the SVS in an online manner are quite sporadic, and a mature and robust pipeline is still lacking. As an attempt to fill this research gap to some extent, in this work, we present a novel extrinsics correction pipeline designed specially for the SVS, namely ROECS (Robust Online Extrinsics Correction of the Surround-view system). Specifically, a "refined bi-camera error" model is firstly designed. Then, by minimizing the overall "bi-camera error" within a sparse and semi-direct framework, the SVS's extrinsics can be iteratively optimized and become accurate eventually. Besides, an innovative three-step pixel selection strategy is also proposed. The superior robustness and the generalization capability of ROECS are validated by both quantitative and qualitative experimental results. To make the results reproducible, the collected data and the source code have been released at https://cslinzhang.github.io/ROECS/. Tianjun Zhang, Brian Nlong Zhao, Ying Shen 0005, Xuan Shao, Lin Zhang 0014, Yicong Zhou |
ACM Multimedia | 1 |
| 2021 | Learning Space Partitions for Path PlanningabstractPath planning, the problem of efficiently discovering high-reward trajectories, often requires optimizing a high-dimensional and multimodal reward function. Popular approaches like CEM and CMA-ES greedily focus on promising regions of the search space and may get trapped in local maxima. DOO and VOOT balance exploration and exploitation, but use space partitioning strategies independent of the reward function to be optimized. Recently, LaMCTS empirically learns to partition the search space in a reward-sensitive manner for black-box optimization. In this paper, we develop a novel formal regret analysis for when and why such an adaptive region partitioning scheme works. We also propose a new path planning method LaP3 which improves the function value estimation within each sub-region, and uses a latent representation of the search space. Empirically, LaP3 outperforms existing path planning methods in 2D navigation tasks, especially in the presence of difficult-to-escape local optima, and shows benefits when plugged into the planning components of model-based RL such as PETS. These gains transfer to highly multimodal real-world tasks, where we outperform strong baselines in compiler phase ordering by up to 39% on average across 9 tasks, and in molecular design by up to 0.4 on properties on a 0-1 scale. Code is available at https://github.com/yangkevin2/neurips2021-lap3. Kevin Yang, Tianjun Zhang, Chris Cummins, Brandon Cui, Benoit Steiner, Linnan Wang, Joseph Gonzalez 0001, Daniel Klein 0001, Yuandong Tian |
NeurIPS | 2 |
| 2021 | MADE: Exploration via Maximizing Deviation from Explored RegionsabstractIn online reinforcement learning (RL), efficient exploration remains particularly challenging in high-dimensional environments with sparse rewards. In low-dimensional environments, where tabular parameterization is possible, count-based upper confidence bound (UCB) exploration methods achieve minimax near-optimal rates. However, it remains unclear how to efficiently implement UCB in realistic RL tasks that involve non-linear function approximation. To address this, we propose a new exploration approach via maximizing the deviation of the occupancy of the next policy from the explored regions. We add this term as an adaptive regularizer to the standard RL objective to balance exploration vs. exploitation. We pair the new objective with a provably convergent algorithm, giving rise to a new intrinsic reward that adjusts existing bonuses. The proposed intrinsic reward is easy to implement and combine with other existing RL algorithms to conduct exploration. As a proof of concept, we evaluate the new intrinsic reward on tabular examples across a variety of model-based and model-free algorithms, showing improvements over count-only exploration strategies. When tested on navigation and locomotion tasks from MiniGrid and DeepMind Control Suite benchmarks, our approach significantly improves sample efficiency over state-of-the-art methods. Tianjun Zhang, Paria Rashidinejad, Jiantao Jiao, Yuandong Tian, Joseph Gonzalez 0001, Stuart Russell 0001 |
NeurIPS | 1 |
| 2021 | NovelD: A Simple yet Effective Exploration CriterionabstractEfficient exploration under sparse rewards remains a key challenge in deep reinforcement learning. Previous exploration methods (e.g., RND) have achieved strong results in multiple hard tasks. However, if there are multiple novel areas to explore, these methods often focus quickly on one without sufficiently trying others (like a depth-wise first search manner). In some scenarios (e.g., four corridor environment in Sec 4.2), we observe they explore in one corridor for long and fail to cover all the states. On the other hand, in theoretical RL, with optimistic initialization and the inverse square root of visitation count as a bonus, it won't suffer from this and explores different novel regions alternatively (like a breadth-first search manner). In this paper, inspired by this, we propose a simple but effective criterion called NovelD by weighting every novel area approximately equally. Our algorithm is very simple but yet shows comparable performance or even outperforms multiple SOTA exploration methods in many hard exploration tasks. Specifically, NovelD solves all the static procedurally-generated tasks in Mini-Grid with just 120M environment steps, without any curriculum learning. In comparison, the previous SOTA only solves 50% of them. NovelD also achieves SOTA on multiple tasks in NetHack, a rogue-like game that contains more challenging procedurally-generated environments. In multiple Atari games (e.g., MonteZuma's Revenge, Venture, Gravitar), NovelD outperforms RND. We analyze NovelD thoroughly in MiniGrid and found that empirically it helps the agent explore the environment more uniformly with a focus on exploring beyond the boundary. Tianjun Zhang, Huazhe Xu, Xiaolong Wang 0004, Yi Wu 0013, Kurt Keutzer, Joseph Gonzalez 0001, Yuandong Tian |
NeurIPS | 1 |
| 2020 | Oecs: Towards Online Extrinsics Correction For The Surround-View SystemabstractA typical surround-view system consists of four fisheye cameras. By performing an offline calibration that determines both the intrinsics and extrinsics of the system, surround-view images can be synthesized at runtime. However, poses of calibrated cameras sometimes may change. In such a case, if cameras' extrinsics are not updated accordingly, observable geometric misalignment will appear in surround-views. Most existing solutions to this problem resort to re-calibration, which is quite cumbersome. Thus, how to correct cameras' extrinsics in an online manner without using re-calibration is still an open issue. In this paper, we attempt to propose a novel solution to this problem and the proposed solution is referred to as “Online Extrinsics Correction for the Surround-view system OECS for short. We first design a Bi-Camera error model, measuring the photometric discrepancy between two corresponding pixels on images captured by two adjacent cameras. Then, by minimizing the system's overall BiCamera error, cameras' extrinsics can be optimized and the optimization is conducted within a sparse direct framework. The efficacy and efficiency of OECS are validated by experiments. Data and source code used in this work are publicly available at https://z619850002.github.io/OECage/. Tianjun Zhang, Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou |
ICME | 1 |
| 2020 | A Tightly-coupled Semantic SLAM System with Visual, Inertial and Surround-view Sensors for Autonomous Indoor ParkingabstractThe semantic SLAM (simultaneous localization and mapping) system is an indispensable module for autonomous indoor parking. Monocular and binocular visual cameras constitute the basic configuration to build such a system. Features used in existing SLAM systems are often dynamically movable, blurred and repetitively textured. By contrast, semantic features on the ground are more stable and consistent in the indoor parking environment. Due to their inabilities to perceive salient features on the ground, existing SLAM systems are prone to tracking loss during navigation. Therefore, a surround-view camera system capturing images from a top-down viewpoint is necessarily called for. To this end, this paper proposes a novel tightly-coupled semantic SLAM system by integrating Visual, Inertial, and Surround-view sensors, VIS SLAM for short, for autonomous indoor parking. In VIS SLAM, apart from low-level visual features and IMU (inertial measurement unit) motion data, parking-slots in surround-view images are also detected and geometrically associated, forming semantic constraints. Specifically, each parking-slot can impose a surround-view constraint that can be split into an adjacency term and a registration term. The former pre-defines the position of each individual parking-slot subject to whether it has an adjacent neighbor. The latter further constrains by registering between each observed parking-slot and its position in the world coordinate system. To validate the effectiveness and efficiency of VIS SLAM, a large-scale dataset composed of synchronous multi-sensor data collected from typical indoor parking sites is established, which is the first of its kind. The collected dataset has been made publicly available at https://cslinzhang.github.io/VISSLAM/. Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Hongyu Li 0001, Yicong Zhou |
ACM Multimedia | 3 |
| 2019 | Synetgy: Algorithm-hardware Co-design for ConvNet Accelerators on Embedded FPGAsabstractUsing FPGAs to accelerate ConvNets has attracted significant attention in recent years. However, FPGA accelerator design has not leveraged the latest progress of ConvNets. As a result, the key application characteristics such as frames-per-second (FPS) are ignored in favor of simply counting GOPs, and results on accuracy, which is critical to application success, are often not even reported. In this work, we adopt an algorithm-hardware co-design approach to develop a ConvNet accelerator called Synetgy and a novel ConvNet model called DiracDeltaNet. Both the accelerator and ConvNet are tailored to FPGA requirements. DiracDeltaNet, as the name suggests, is a ConvNet with only $1\times 1$ convolutions while spatial convolutions are replaced by more efficient shift operations. DiracDeltaNet achieves competitive accuracy on ImageNet (89.0% top-5), but with 48× fewer parameters and 65× fewer OPs than VGG16. We further quantize DiracDeltaNet's weights to 1-bit and activations to 4-bits, with less than 1% accuracy loss. These quantizations exploit well the nature of FPGA hardware. In short, DiracDeltaNet's small model size, low computational OP count, ultra-low precision and simplified operators allow us to co-design a highly customized computing unit for an FPGA. We implement the computing units for DiracDeltaNet on an Ultra96 SoC system through high-level synthesis. Our accelerator's final top-5 accuracy of 88.2% on ImageNet, is higher than all the previously reported embedded FPGA accelerators. In addition, the accelerator reaches an inference speed of 96.5 FPS on the ImageNet classification task, surpassing prior works with similar accuracy by at least 16.9×. Qijing Huang 0001, Bichen Wu, Tianjun Zhang, Liang Ma 0003, Giulio Gambardella, Michaela Blott, Luciano Lavagno, Kees A. Vissers, John Wawrzynek, Kurt Keutzer |
FPGA | 4 |
| 2019 | ANODEV2: A Coupled Neural ODE FrameworkabstractIt has been observed that residual networks can be viewed as the explicit Euler discretization of an Ordinary Differential Equation (ODE). This observation motivated the introduction of so-called Neural ODEs, in which other discretization schemes and/or adaptive time stepping techniques can be used to improve the performance of residual networks. Here, we propose \OURS, which extends this approach by introducing a framework that allows ODE-based evolution for both the weights and the activations, in a coupled formulation. Such an approach provides more modeling flexibility, and it can help with generalization performance. We present the formulation of \OURS, derive optimality conditions, and implement the coupled framework in PyTorch. We present empirical results using several different configurations of \OURS, testing them on the CIFAR-10 dataset. We report results showing that our coupled ODE-based framework is indeed trainable, and that it achieves higher accuracy, compared to the baseline ResNet network and the recently-proposed Neural ODE approach. Tianjun Zhang, Zhewei Yao, Amir Gholami, Joseph Gonzalez 0001, Kurt Keutzer, Michael W. Mahoney, George Biros |
NeurIPS | 1 |
| 2018 | GenAx: A Genome Sequencing AcceleratorabstractGenomics can transform health-care through precision medicine. Plummeting sequencing costs would soon make genome testing affordable to the masses. Compute efficiency, however, has to improve by orders of magnitude to sequence and analyze the raw genome data. Sequencing software used today can take several hundreds to thousands of CPU hours to align reads to a reference sequence. This paper presents GenAx, an accelerator for read alignment, a time-consuming step in genome sequencing. It consists of a seeding and seed-extension accelerator. The latter is based on an innovative automata design that was designed from the ground-up to enable hardware acceleration. Unlike conventional Levenshtein automata, it is string independent and scales quadratically with edit distance, instead of string length. It supports critical features commonly used in sequencing such as affine gap scoring and traceback. GenAx provides a throughput of 4,058K reads/s for Illumina 101 bp reads. GenAx achieves 31.7× speedup over the standard BWA-MEM sequence aligner running on a 56-thread dualsocket 14-core Xeon E5 server processor, while reducing power consumption by 12× and area by 5.6×. Daichi Fujiki, Arun Subramaniyan 0001, Tianjun Zhang, Reetuparna Das, David T. Blaauw, Satish Narayanasamy |
ISCA | 3 |