VLDB 2026 Research / reviewers in the wild / expert
Xishan Zhang
dblp:133/6391
· DBLP profile ↗
37ranked-venue papers
5as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 10 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Diffusion Planning with Temporal DiffusionabstractDiffusion planning is a promising method for learning high-performance policies from offline data. To avoid the impact of discrepancies between planning and reality on performance, previous works generate new plans at each time step. However, this incurs significant computational overhead and leads to lower decision frequencies, and frequent plan switching may also affect performance. In contrast, humans might create detailed short-term plans and more general, sometimes vague, long-term plans, and adjust them over time. Inspired by this, we propose the Temporal Diffusion Planner (TDP) which improves decision efficiency by distributing the denoising steps across the time dimension. TDP begins by generating an initial plan that becomes progressively more vague over time. At each subsequent time step, rather than generating an entirely new plan, TDP updates the previous one with a small number of denoising steps. This reduces the average number of denoising steps, improving decision efficiency. Additionally, we introduce an automated replanning mechanism to prevent significant deviations between the plan and reality. Experiments on D4RL show that, compared to previous works that generate new plans every time step, TDP significantly improves the decision-making frequency by 11-24.8 times while achieving higher or comparable performance. Jiaming Guo, Rui Zhang 0040, Zerun Li, Yunkai Gao 0001, Shaohui Peng, Siming Lan, Xing Hu 0001, Zidong Du, Xishan Zhang, Ling Li 0001 |
AAAI | 9 |
| 2026 | StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs Through Knowledge-Reasoning FusionabstractAutoformalization aims to translate natural-language mathematical statements into a formal language. While LLMs have accelerated progress in this area, existing methods still suffer from low accuracy. We identify two key abilities for effective autoformalization: comprehensive mastery of formal-language domain knowledge, and reasoning capability of natural language problem understanding and informal-formal alignment. Without the former, a model cannot identify the correct formal objects; without the latter, it struggles to interpret real-world contexts and map them precisely into formal expressions. To address these gaps, we introduce ThinkingF, a data synthesis and training pipeline that improves both abilities. First, we construct two datasets: one by distilling and selecting large-scale examples rich in formal knowledge, and another by generating informal-to-formal reasoning trajectories guided by expert-designed templates. We then apply SFT and RLVR with these datasets to further fuse and refine the two abilities. The resulting 7B and 32B models exhibit both comprehensive formal knowledge and strong informal-to-formal reasoning. Notably, StepFun-Formalizer-32B achieves SOTA BEq@1 scores of 40.5% on FormalMATH-Lite and 26.7% on ProverBench, surpassing all prior general-purpose and specialized models. Ruosi Wan, Shijie Shang, Chenrui Cao, Rui Zhang 0040, Xishan Zhang, Zidong Du, Jie Yang 0002, Xing Hu 0001 |
AAAI | 9 |
| 2026 | PE Loss: Perception-Enhanced Distortion-Oriented Loss for Image RestorationabstractImage restoration is the inverse problem of recovering high-quality images from knowledge of degraded images; it includes image super-resolution, image denoising, image deblurring, etc. The objective of image restoration methods is to minimize the error defined by the loss function between the network output image and the corresponding ground truth. Distortion-oriented loss functions are fundamental and commonly used in image restoration. However, these functions treat all pixels as equally important without differentiating between sharp and blurred edge areas, which does not match human visual perception. As a result, these methods can produce accurate but blurred images. To address this issue and achieve both accurate and perceptually satisfactory results, we propose a novel perception-enhanced distortion-oriented loss (PE loss) for image restoration, inspired by the Mach band effect. This effect demonstrates that sharp edges are perceived as having better quality than blurred edges by the human visual system. Our approach includes designing a blur factor map that detects blurred pixels and penalizes them by amplifying their error. The PE loss is a simple yet effective plug-and-play method, and we apply it to state-of-the-art networks. Extensive quantitative and qualitative experiments show that our method can restore images with sharp edges and high perceptual quality. Xishan Zhang, Shaoli Liu |
Comput. Vis. Media | 4 |
| 2026 | Cambricon-QM: A Hybrid Architecture for Microscaling Format Training
Yongwei Zhao 0001, Chang Liu 0021, Zidong Du, Xing Hu 0001, Yimin Zhuang, Yifan Hao 0001, Xinkai Song, Wei Li 0008, Xishan Zhang, Ling Li 0001, Zhiwei Xu 0002, Tianshi Chen 0002, Qi Guo 0001 |
IEEE Trans. Computers | 12 |
| 2026 | CodeV: Empowering LLMs With HDL Generation Through Multilevel SummarizationabstractThe design flow of processors, particularly in hardware description languages (HDL) like Verilog and Chisel, is complex and costly. While recent advances in large language models (LLMs) have significantly improved coding tasks in software languages such as Python, their application in HDL generation remains limited due to the scarcity of high-quality HDL data. Traditional methods of adapting LLMs for hardware design rely on synthetic HDL datasets, which often suffer from low quality because even advanced LLMs like GPT perform poorly in the HDL domain. Moreover, these methods focus solely on chat tasks and the Verilog language, limiting their application scenarios. In this paper, we observe that: (1) HDL code collected from the real world is of higher quality than code generated by LLMs. (2) LLMs like GPT-3.5 excel in summarizing HDL code rather than generating it. (3) An explicit language tag can help LLMs better adapt to the target language when there is insufficient data. Based on these observations, we propose an efficient LLM fine-tuning pipeline for HDL generation that integrates a multi-level summarization data synthesis process with a novel Chat-FIM-Tag supervised fine-tuning method. The pipeline enhances the generation of HDL code from natural language descriptions and enables the handling of various tasks such as chat and infilling incomplete code. Utilizing this pipeline, we introduce CodeV, a series of HDL generation LLMs. Among them, CodeV-All not only possesses a more diverse range of language abilities (Verilog and Chisel) and a broader scope of tasks (Chat and FIM), but also achieves performance on VerilogEval that is comparable to that of CodeV-Verilog fine-tuned on Verilog only, making them the first series of open-source LLMs designed for multi-scenario HDL generation. Code, models, and dataset: https://github.com/IPRC-DIP/CodeV. Yang Zhao 0013, Chongxiao Li, Pengwei Jin, Muxin Song, Yinan Xu 0001, Ziyuan Nan, Mingju Gao, Tianyun Ma, Yansong Pan, Rui Zhang 0040, Xishan Zhang, Zidong Du, Qi Guo 0001, Xing Hu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 14 |
| 2025 | InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-InstructabstractRecent advancements in open-source code large language models (LLMs) have been driven by fine-tuning on the data generated from powerful closed-source LLMs, which are expensive to obtain. This paper explores whether it is possible to use a fine-tuned open-source model to generate additional data to augment its instruction-tuning dataset. We make two observations: (1) A code snippet can serve as the response to different instructions. (2) Instruction-tuned code LLMs perform better at translating code into instructions than the reverse. Based on these observations, we propose Inverse-Instruct, a data augmentation technique that uses a fine-tuned LLM to generate additional instructions of code responses from its own training dataset. The additional instruction-response pairs are added to the original dataset, and a stronger code LLM can be obtained by fine-tuning on the augmented dataset. We empirically validate Inverse-Instruct on a range of open-source code models (e.g. CodeLlama-Python and DeepSeek-Coder) and benchmarks (e.g., HumanEval(+), MBPP(+), DS-1000 and MultiPL-E), showing it consistently improves the base models. Yewen Pu, Lingzhe Gao, Ziyuan Nan, Kaizhao Yuan, Rui Zhang 0040, Xishan Zhang, Zidong Du, Qi Guo 0001, Dawei Yin 0001, Xing Hu 0001, Yunji Chen |
AAAI | 11 |
| 2025 | QiMeng-CodeV-R1: Reasoning-Enhanced Verilog GenerationabstractLarge language models (LLMs) trained via reinforcement learning with verifiable reward (RLVR) have achieved breakthroughs on tasks with explicit, automatable verification, such as software programming and mathematical problems. Extending RLVR to electronic design automation (EDA), especially automatically generating hardware description languages (HDLs) like Verilog from natural-language (NL) specifications, however, poses three key challenges: the lack of automated and accurate verification environments, the scarcity of high-quality NL-code pairs, and the prohibitive computation cost of RLVR. To this end, we introduce CodeV-R1, an RLVR framework for training Verilog generation LLMs. First, we develop a rule-based testbench generator that performs robust equivalence checking against golden references. Second, we propose a round-trip data synthesis method that pairs open-source Verilog snippets with LLM-generated NL descriptions, verifies code–NL–code consistency via the generated testbench, and filters out inequivalent examples to yield a high-quality dataset. Third, we employ a two-stage "distill-then-RL" training pipeline: distillation for the cold start of reasoning abilities, followed by adaptive DAPO, our novel RLVR algorithm that can reduce training cost by adaptively adjusting sampling rate. The resulting model, CodeV-R1-7B, achieves 68.6 \% and 72.9 \% pass@1 on VerilogEval v2 and RTLLM v1.1, respectively, surpassing prior state-of-the-art by 12$\sim$20 \%, while even exceeding the performance of 671B DeepSeek-R1 on RTLLM. We have released our model, training code, and dataset to facilitate research in EDA and LLM communities. Yaoyu Zhu, Han-Qi Lyu, Chongxiao Li, Jianan Mu, Yang Zhao 0013, Pengwei Jin, Shuyao Cheng, Shengwen Liang, Xishan Zhang, Rui Zhang 0040, Zidong Du, Qi Guo 0001, Xing Hu 0001, Yunji Chen |
NeurIPS | 14 |
| 2024 | Emergent Communication for Numerical Concepts GeneralizationabstractResearch on emergent communication has recently gained significant traction as a promising avenue for the linguistic community to unravel human language's origins and explore artificial intelligence's generalization capabilities. Current research has predominantly concentrated on recognizing qualitative patterns of object attributes(e.g., shape and color) and paid little attention to the quantitative relationship among object quantities which is known as the part of numerical concepts. The ability to generalize numerical concepts, i.e., counting and calculations with unseen quantities, is essential, as it mirrors humans' foundational abstract reasoning abilities. In this work, we introduce the NumGame, leveraging the referential game framework, forcing agents to communicate and generalize the numerical concepts effectively. Inspired by the human learning process of numbers, we present a two-stage training approach that sequentially fosters a rudimentary numerical sense followed by the ability of arithmetic calculation, ultimately aiding agents in generating semantically stable and unambiguous language for numerical concepts. The experimental results indicate the impressive generalization capabilities to unseen quantities and regularity of the language emergence from communication. Enshuai Zhou, Yifan Hao 0001, Rui Zhang 0040, Zidong Du, Xishan Zhang, Xinkai Song, Chao Wang 0003, Xuehai Zhou, Jiaming Guo, Qi Yi, Shaohui Peng, Ruizhi Chen, Qi Guo 0001, Yunji Chen |
AAAI | 6 |
| 2024 | Ex3: Automatic Novel Writing by Extracting, Excelsior and ExpandingabstractHuang Lei, Jiaming Guo, Guanhua He, Xishan Zhang, Rui Zhang, Shaohui Peng, Shaoli Liu, Tianshi Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Huang Lei, Jiaming Guo, Guanhua He, Xishan Zhang, Rui Zhang 0040, Shaohui Peng, Shaoli Liu, Tianshi Chen 0002 |
ACL (1) | 4 |
| 2024 | RecurrentBEV: A Long-Term Temporal Fusion Framework for Multi-view 3D Detection
Xishan Zhang, Rui Zhang 0040, Guanhua He, Shaoli Liu |
ECCV (72) | 2 |
| 2024 | AutoMiner: Reinforcement Learning-Based Mining Attack Simulator
Lide Xue, Ziyang Han, Bingren Chen, Xishan Zhang, Xuehai Zhou |
ICA3PP (1) | 5 |
| 2024 | Automated CPU Design by Learning from Input-Output Examples
Shuyao Cheng, Pengwei Jin, Qi Guo 0001, Zidong Du, Rui Zhang 0040, Xing Hu 0001, Yongwei Zhao 0001, Yifan Hao 0001, Xiangtao Guan, Husheng Han, Zhengyue Zhao, Xishan Zhang, Yuejie Chu, Weilong Mao, Tianshi Chen 0002, Yunji Chen |
IJCAI | 13 |
| 2024 | ECFO: An Efficient Edge Classification-Based Fusion Optimizer for Deep Learning CompilersabstractOperation fusion is a critical technique in optimizing deep learning compilers as it enhances computational efficiency by integrating multiple operations into a single computational graph. However, finding an effective fusion strategy is challenging, requiring the definition of an optimization search space and identification of the best strategy within this space. Existing methods, such as heuristic searches and learning-based searches, have significant limitations. Heuristic searches are complex, labor-intensive, and often lack generalizability across different network architectures. On the other hand, learning-based methods demand extensive training and pro-longed search time. To address these challenges, we introduce the Edge Classification-Based Fusion Optimizer (ECFO), a novel approach that reconceptualizes operation fusion as an edge classification problem. By leveraging Graph Neural Networks (GNNs) for efficient graph feature encoding, ECFO streamline the optimization process and significantly reduces computational overhead. Comprehensive evaluations across diverse neural networks demonstrate that ECFO decrease search time by up to 23x and improves inference performance by 3.2%, representing a substantial advancement over existing strategies. Wei Li 0008, Kangcheng Liu, Lide Xue, Zidong Du, Xishan Zhang, Xuehai Zhou |
SMC | 7 |
| 2024 | Real-Time Robust Video Object Detection System Against Physical-World Adversarial AttacksabstractDNN-based video object detection (VOD) powers autonomous driving and video surveillance industries with rising importance and promising opportunities. However, adversarial patch attack yields huge concern in live vision tasks because of its practicality, feasibility, and powerful attack effectiveness. This work proposes Themis, a software/hardware system to defend against adversarial patches for real-time robust VOD. We observe that adversarial patches exhibit extremely localized superficial feature importance in a small region with nonrobust predictions, and thus propose the adversarial region detection algorithm for adversarial effect elimination. Themis also proposes a systematic design to efficiently support the algorithm by eliminating redundant computations and memory traffics. Experimental results show that the proposed methodology can effectively recover the system from the adversarial attack with negligible hardware overhead. Husheng Han, Xing Hu 0001, Yifan Hao 0001, Kaidi Xu, Pucheng Dang, Ying Wang 0001, Yongwei Zhao 0001, Zidong Du, Qi Guo 0001, Yanzhi Wang 0001, Xishan Zhang, Tianshi Chen 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 11 |
| 2023 | Online Prototype Alignment for Few-shot Policy TransferabstractDomain adaptation in RL mainly deals with the changes of observation when transferring the policy to a new environment. Many traditional approaches of domain adaptation in RL manage to learn a mapping function between the source and target domain in explicit or implicit ways. However, they typically require access to abundant data from the target domain. Besides, they often rely on visual clues to learn the mapping function and may fail when the source domain looks quite different from the target domain. To address these problems, in this paper, we propose a novel framework Online Prototype Alignment (OPA) to learn the mapping function based on the functional similarity of elements and is able to achieve few-shot policy transfer within only several episodes. The key insight of OPA is to introduce an exploration mechanism that can interact with the unseen elements of the target domain in an efficient and purposeful manner, and then connect them with the seen elements in the source domain according to their functionalities (instead of visual clues). Experimental results show that when the target domain looks visually different from the source domain, OPA can achieve better transfer performance even with much fewer samples from the target domain, outperforming prior methods. Qi Yi, Rui Zhang 0040, Shaohui Peng, Jiaming Guo, Yunkai Gao 0001, Kaizhao Yuan, Ruizhi Chen, Siming Lan, Xing Hu 0001, Zidong Du, Xishan Zhang, Qi Guo 0001, Yunji Chen |
ICML | 11 |
| 2023 | Non-autoregressive Machine Translation with Probabilistic Context-free GrammarabstractNon-autoregressive Transformer(NAT) significantly accelerates the inference of neural machine translation. However, conventional NAT models suffer from limited expression power and performance degradation compared to autoregressive (AT) models due to the assumption of conditional independence among target tokens. To address these limitations, we propose a novel approach called PCFG-NAT, which leverages a specially designed Probabilistic Context-Free Grammar (PCFG) to enhance the ability of NAT models to capture complex dependencies among output tokens. Experimental results on major machine translation benchmarks demonstrate that PCFG-NAT further narrows the gap in translation quality between NAT and AT models. Moreover, PCFG-NAT facilitates a deeper understanding of the generated sentences, addressing the lack of satisfactory explainability in neural machine translation. Code is publicly available at https://github.com/ictnlp/PCFG-NAT. Shangtong Gui, Chenze Shao, Zhengrui Ma, Xishan Zhang, Yunji Chen, Yang Feng 0004 |
NeurIPS | 4 |
| 2023 | Efficient Symbolic Policy Learning with Differentiable Symbolic ExpressionabstractDeep reinforcement learning (DRL) has led to a wide range of advances in sequential decision-making tasks. However, the complexity of neural network policies makes it difficult to understand and deploy with limited computational resources. Currently, employing compact symbolic expressions as symbolic policies is a promising strategy to obtain simple and interpretable policies. Previous symbolic policy methods usually involve complex training processes and pre-trained neural network policies, which are inefficient and limit the application of symbolic policies. In this paper, we propose an efficient gradient-based learning method named Efficient Symbolic Policy Learning (ESPL) that learns the symbolic policy from scratch in an end-to-end way. We introduce a symbolic network as the search space and employ a path selector to find the compact symbolic policy. By doing so we represent the policy with a differentiable symbolic expression and train it in an off-policy manner which further improves the efficiency. In addition, in contrast with previous symbolic policies which only work in single-task RL because of complexity, we expand ESPL on meta-RL to generate symbolic policies for unseen tasks. Experimentally, we show that our approach generates symbolic policies with higher performance and greatly improves data efficiency for single-task RL. In meta-RL, we demonstrate that compared with neural network policies the proposed symbolic policy achieves higher performance and efficiency and shows the potential to be interpretable. Jiaming Guo, Rui Zhang 0040, Shaohui Peng, Qi Yi, Xing Hu 0001, Ruizhi Chen, Zidong Du, Xishan Zhang, Ling Li 0001, Qi Guo 0001, Yunji Chen |
NeurIPS | 8 |
| 2023 | Emergent Communication for Rules ReasoningabstractResearch on emergent communication between deep-learning-based agents has received extensive attention due to its inspiration for linguistics and artificial intelligence.
However, previous attempts have hovered around emerging communication under perception-oriented environmental settings,
that forces agents to describe low-level perceptual features intra image or symbol contexts.
In this work, inspired by the classic human reasoning test (namely Raven's Progressive Matrix), we propose the Reasoning Game, a cognition-oriented environment that encourages agents to reason and communicate high-level rules, rather than perceived low-level contexts.
Moreover, we propose 1) an unbiased dataset (namely rule-RAVEN) as a benchmark to avoid overfitting, 2) and a two-stage curriculum agent training method as a baseline for more stable convergence in the Reasoning Game,
where contexts and semantics are bilaterally drifting.
Experimental results show that, in the Reasoning Game, a semantically stable and compositional language emerges to solve reasoning problems.
The emerged language helps agents apply the extracted rules to the generalization of unseen context attributes, and to the transfer between different context attributes or even tasks. Yifan Hao 0001, Rui Zhang 0040, Enshuai Zhou, Zidong Du, Xishan Zhang, Xinkai Song, Yuanbo Wen 0001, Yongwei Zhao 0001, Xuehai Zhou, Jiaming Guo, Qi Yi, Shaohui Peng, Ruizhi Chen, Qi Guo 0001, Yunji Chen |
NeurIPS | 6 |
| 2023 | Contrastive Modules with Temporal Attention for Multi-Task Reinforcement LearningabstractIn the field of multi-task reinforcement learning, the modular principle, which involves specializing functionalities into different modules and combining them appropriately, has been widely adopted as a promising approach to prevent the negative transfer problem that performance degradation due to conflicts between tasks. However, most of the existing multi-task RL methods only combine shared modules at the task level, ignoring that there may be conflicts within the task. In addition, these methods do not take into account that without constraints, some modules may learn similar functions, resulting in restricting the model's expressiveness and generalization capability of modular methods.
In this paper, we propose the Contrastive Modules with Temporal Attention(CMTA) method to address these limitations. CMTA constrains the modules to be different from each other by contrastive learning and combining shared modules at a finer granularity than the task level with temporal attention, alleviating the negative transfer within the task and improving the generalization ability and the performance for multi-task RL.
We conducted the experiment on Meta-World, a multi-task RL benchmark containing various robotics manipulation tasks. Experimental results show that CMTA outperforms learning each task individually for the first time and achieves substantial performance improvements over the baselines. Siming Lan, Rui Zhang 0040, Qi Yi, Jiaming Guo, Shaohui Peng, Yunkai Gao 0001, Ruizhi Chen, Zidong Du, Xing Hu 0001, Xishan Zhang, Ling Li 0001, Yunji Chen |
NeurIPS | 11 |
| 2022 | Neural Program Synthesis with Query
Rui Zhang 0040, Xing Hu 0001, Xishan Zhang, Pengwei Jin, Zidong Du, Qi Guo 0001, Yunji Chen |
ICLR | 4 |
| 2022 | Causality-driven Hierarchical Structure Discovery for Reinforcement LearningabstractHierarchical reinforcement learning (HRL) has been proven to be effective for tasks with sparse rewards, for it can improve the agent's exploration efficiency by discovering high-quality hierarchical structures (e.g., subgoals or options). However, automatically discovering high-quality hierarchical structures is still a great challenge. Previous HRL methods can only find the hierarchical structures in simple environments, as they are mainly achieved through the randomness of agent's policies during exploration. In complicated environments, such a randomness-driven exploration paradigm can hardly discover high-quality hierarchical structures because of the low exploration efficiency. In this paper, we propose CDHRL, a causality-driven hierarchical reinforcement learning framework, to build high-quality hierarchical structures efficiently in complicated environments. The key insight is that the causalities among environment variables are naturally fit for modeling reachable subgoals and their dependencies; thus, the causality is suitable to be the guidance in building high-quality hierarchical structures. Roughly, we build the hierarchy of subgoals based on causality autonomously, and utilize the subgoal-based policies to unfold further causality efficiently. Therefore, CDHRL leverages a causality-driven discovery instead of a randomness-driven exploration for high-quality hierarchical structure construction. The results in two complex environments, 2D-Minecraft and Eden, show that CDHRL can discover high-quality hierarchical structures and significantly enhance exploration efficiency. Shaohui Peng, Xing Hu 0001, Rui Zhang 0040, Ke Tang 0001, Jiaming Guo, Qi Yi, Ruizhi Chen, Xishan Zhang, Zidong Du, Ling Li 0001, Qi Guo 0001, Yunji Chen |
NeurIPS | 8 |
| 2022 | Object-Category Aware Reinforcement LearningabstractObject-oriented reinforcement learning (OORL) is a promising way to improve the sample efficiency and generalization ability over standard RL. Recent works that try to solve OORL tasks without additional feature engineering mainly focus on learning the object representations and then solving tasks via reasoning based on these object representations. However, none of these works tries to explicitly model the inherent similarity between different object instances of the same category. Objects of the same category should share similar functionalities; therefore, the category is the most critical property of an object. Following this insight, we propose a novel framework named Object-Category Aware Reinforcement Learning (OCARL), which utilizes the category information of objects to facilitate both perception and reasoning. OCARL consists of three parts: (1) Category-Aware Unsupervised Object Discovery (UOD), which discovers the objects as well as their corresponding categories; (2) Object-Category Aware Perception, which encodes the category information and is also robust to the incompleteness of (1) at the same time; (3) Object-Centric Modular Reasoning, which adopts multiple independent and object-category-specific networks when reasoning based on objects. Our experiments show that OCARL can improve both the sample efficiency and generalization in the OORL domain. Qi Yi, Rui Zhang 0040, Shaohui Peng, Jiaming Guo, Xing Hu 0001, Zidong Du, Xishan Zhang, Qi Guo 0001, Yunji Chen |
NeurIPS | 7 |
| 2022 | RTSfM: Real-Time Structure From Motion for Mosaicing and DSM Mapping of Sequential Aerial Images With Low OverlapabstractInspired by simultaneous localization and mapping (SLAM) style workflow, this article presented an online sequential structure from motion (SfM) solution for high-frequency video and large baseline high-resolution aerial images with high efficiency and novel precision. First, as traditional SLAM systems are not good in processing low overlap images, based on our novel hierarchical feature matching paradigm with multihomography and BoW, we proposed a robust tracking method where the relative pose and its scale are estimated separately followed by a joint optimization by considering both perspective-n-point (PnP) and epipolar constraints. Second, to further optimize the camera poses for the sparse map and dense pointcloud reconstruction, we provided a graph-based optimization with reprojection and GPS constraints, which make the camera trajectory and map georeferenced. We also incrementally generated the dense point cloud in real time from keyframes after local mapping optimization. Finally, we use a publicly available aerial image dataset with sequences of different environments, to evaluate the effectiveness of the proposed method, meanwhile, the robust performance of our solution is demonstrated with applications of high-quality aerial images mosaic and digital surface model (DSM) reconstruction in real time. Compared with the state-of-the-art SLAM and traditional SfM methods, the presented system can output large-scale high-quality ortho-mosaic and DSM in real time with the low computational cost. Lin Chen 0042, Xishan Zhang, Shibiao Xu, Shuhui Bu, Hongkai Jiang, Pengcheng Han, Ke Li 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Rethinking the Importance of Quantization Bias, Toward Full Low-Bit TrainingabstractQuantization is a promising technique to reduce the computation and storage costs of DNNs. Low-bit ( ≤ 8 bits) precision training remains an open problem due to the difficulty of gradient quantization. In this paper, we find two long-standing misunderstandings of the bias of gradient quantization noise. First, the large bias of gradient quantization noise, instead of the variance, is the key factor of training accuracy loss. Second, the widely used stochastic rounding cannot solve the training crash problem caused by the gradient quantization bias in practice. Moreover, we find that the asymmetric distribution of gradients causes a large bias of gradient quantization noise. Based on our findings, we propose a novel adaptive piecewise quantization method to effectively limit the bias of gradient quantization noise. Accordingly, we propose a new data format, Piecewise Fixed Point (PWF), to present data after quantization. We apply our method to different applications including image classification, machine translation, optical character recognition, and text classification. We achieve approximately 1.9 ∼ 3.5× speedup compared with full precision training with an accuracy loss of less than 0.5%. To the best of our knowledge, this is the first work to quantize gradients of all layers to 8 bits in both large-scale CNN and RNN training with negligible accuracy loss. Chang Liu 0021, Xishan Zhang, Rui Zhang 0040, Ling Li 0001, Shiyi Zhou, Zidong Du, Shaoli Liu, Tianshi Chen 0002 |
IEEE Trans. Image Process. | 2 |
| 2021 | Domain-Specific Suppression for Adaptive Object DetectionabstractDomain adaptation methods face performance degradation in object detection, as the complexity of tasks require more about the transferability of the model. We propose a new perspective on how CNN models gain the transferability, viewing the weights of a model as a series of motion patterns. The directions of weights, and the gradients, can be divided into domain-specific and domain-invariant parts, and the goal of domain adaptation is to concentrate on the domain-invariant direction while eliminating the disturbance from domain-specific one. Current UDA object detection methods view the two directions as a whole while optimizing, which will cause domain-invariant direction mismatch even if the output features are perfectly aligned. In this paper, we propose the domain-specific suppression, an exemplary and generalizable constraint to the original convolution gradients in backpropagation to detach the two parts of directions and suppress the domain-specific one. We further validate our theoretical analysis and methods on several domain adaptive object detection tasks, including weather, camera configuration, and synthetic to real-world adaptation. Our experiment results show significant advance over the state-of-the-art methods in the UDA object detection field, performing a promotion of 10.2 ∼ 12.2% mAP on all these domain adaptation scenarios. Rui Zhang 0040, Yangyang Xia, Xishan Zhang, Shaoli Liu |
CVPR | 6 |
| 2021 | Hindsight Value Function for Variance Reduction in Stochastic Dynamic EnvironmentabstractPolicy gradient methods are appealing in deep reinforcement learning but suffer from high variance of gradient estimate. To reduce the variance, the state value function is applied commonly. However, the effect of the state value function becomes limited in stochastic dynamic environments, where the unexpected state dynamics and rewards will increase the variance. In this paper, we propose to replace the state value function with a novel hindsight value function, which leverages the information from the future to reduce the variance of the gradient estimate for stochastic dynamic environments. Particularly, to obtain an ideally unbiased gradient estimate, we propose an information-theoretic approach, which optimizes the embeddings of the future to be independent of previous actions. In our experiments, we apply the proposed hindsight value function in stochastic dynamic environments, including discrete-action environments and continuous-action environments. Compared with the standard state value function, the proposed hindsight value function consistently reduces the variance, stabilizes the training, and improves the eventual policy. Jiaming Guo, Rui Zhang 0040, Xishan Zhang, Shaohui Peng, Qi Yi, Zidong Du, Xing Hu 0001, Qi Guo 0001, Yunji Chen |
IJCAI | 3 |
| 2021 | Cambricon-Q: A Hybrid Architecture for Efficient TrainingabstractDeep neural network (DNN) training is notoriously time-consuming, and quantization is promising to improve the training efficiency with reduced bandwidth/storage requirements and computation costs. However, state-of-the-art quantized algorithms with negligible training accuracy loss, which require on-the-fly statistic-based quantization over a great amount of data (e.g., neurons and weights) and high-precision weight update, cannot be effectively deployed on existing DNN accelerators. To address this problem, we propose the first customized architecture for efficient quantized training with negligible accuracy loss, which is named as Cambricon-Q. Cambricon-Q features a hybrid architecture consisting of an ASIC acceleration core and a near-data-processing (NDP) engine. The acceleration core mainly targets at improving the efficiency of statistic-based quantization with specialized computing units for both statistical analysis (e.g., determining maximum) and data reformating, while the NDP engine avoids transferring the high-precision weights from the off-chip memory to the acceleration core. Experimental results show that on the evaluated benchmarks, Cambricon-Q improves the energy efficiency of DNN training by 6.41× and 1.62×, performance by 4.20× and 1.70× compared to GPU and TPU, respectively, with only ⩽ 0.4% accuracy degradation compared with full precision training. Yongwei Zhao 0001, Chang Liu 0021, Zidong Du, Qi Guo 0001, Xing Hu 0001, Yimin Zhuang, Xinkai Song, Wei Li 0008, Xishan Zhang, Ling Li 0001, Zhiwei Xu 0002, Tianshi Chen 0002 |
ISCA | 10 |
| 2021 | Distilling Object Detectors with Feature RichnessabstractIn recent years, large-scale deep models have achieved great success, but the huge computational complexity and massive storage requirements make it a great challenge to deploy them in resource-limited devices. As a model compression and acceleration method, knowledge distillation effectively improves the performance of small models by transferring the dark knowledge from the teacher detector. However, most of the existing distillation-based detection methods mainly imitating features near bounding boxes, which suffer from two limitations. First, they ignore the beneficial features outside the bounding boxes. Second, these methods imitate some features which are mistakenly regarded as the background by the teacher detector. To address the above issues, we propose a novel Feature-Richness Score (FRS) method to choose important features that improve generalized detectability during distilling. The proposed method effectively retrieves the important features outside the bounding boxes and removes the detrimental features within the bounding boxes. Extensive experiments show that our methods achieve excellent performance on both anchor-based and anchor-free detectors. For example, RetinaNet with ResNet-50 achieves 39.7% in mAP on the COCO2017 dataset, which even surpasses the ResNet-101 based teacher detector 38.9% by 0.8%. Our implementation is available at https://github.com/duzhixing/FRS. Zhixing Du, Rui Zhang 0040, Xishan Zhang, Shaoli Liu, Tianshi Chen 0002, Yunji Chen |
NeurIPS | 4 |
| 2021 | A Decomposable Winograd Method for N-D Convolution Acceleration in Video Analysis
Rui Zhang 0040, Xishan Zhang, Xianzhuo Wang, Pengwei Jin, Shaoli Liu, Ling Li 0001, Yunji Chen |
Int. J. Comput. Vis. | 3 |
| 2020 | DWM: A Decomposable Winograd Method for Convolution AccelerationabstractWinograd's minimal filtering algorithm has been widely used in Convolutional Neural Networks (CNNs) to reduce the number of multiplications for faster processing. However, it is only effective on convolutions with kernel size as 3x3 and stride as 1, because it suffers from significantly increased FLOPs and numerical accuracy problem for kernel size larger than 3x3 and fails on convolution with stride larger than 1. In this paper, we propose a novel Decomposable Winograd Method (DWM), which breaks through the limitation of original Winograd's minimal filtering algorithm to a wide and general convolutions. DWM decomposes kernels with large size or large stride to several small kernels with stride as 1 for further applying Winograd method, so that DWM can reduce the number of multiplications while keeping the numerical accuracy. It enables the fast exploring of larger kernel size and larger stride value in CNNs for high performance and accuracy and even the potential for new CNNs. Comparing against the original Winograd, the proposed DWM is able to support all kinds of convolutions with a speedup of ∼2, without affecting the numerical accuracy. Xishan Zhang, Rui Zhang 0040, Tian Zhi, Deyuan He, Jiaming Guo, Chang Liu 0021, Qi Guo 0001, Zidong Du, Shaoli Liu, Tianshi Chen 0002, Yunji Chen |
AAAI | 2 |
| 2020 | Fixed-Point Back-Propagation TrainingabstractRecent emerged quantization technique (i.e., using low bit-width fixed-point data instead of high bit-width floating-point data) has been applied to inference of deep neural networks for fast and efficient execution. However, directly applying quantization in training can cause significant accuracy loss, thus remaining an open challenge. In this paper, we propose a novel training approach, which applies a layer-wise precision-adaptive quantization in deep neural networks. The new training approach leverages our key insight that the degradation of training accuracy is attributed to the dramatic change of data distribution. Therefore, by keeping the data distribution stable through a layer-wise precision-adaptive quantization, we are able to directly train deep neural networks using low bit-width fixed-point data and achieve guaranteed accuracy, without changing hyper parameters. Experimental results on a wide variety of network architectures (e.g., convolution and recurrent networks) and applications (e.g., image classification, object detection, segmentation and machine translation) show that the proposed approach can train these neural networks with negligible accuracy losses (-1.40%-1.3%, 0.02% on average), and speed up training by 252% on a state-of-the-art Intel CPU. Xishan Zhang, Shaoli Liu, Rui Zhang 0040, Chang Liu 0021, Shiyi Zhou, Jiaming Guo, Qi Guo 0001, Zidong Du, Tian Zhi, Yunji Chen |
CVPR | 1 |
| 2020 | Point in: Counting Trees with Weakly Supervised Segmentation NetworkabstractFor tree counting tasks, since traditional image processing methods require expensive feature engineering and are not end-to-end frameworks, this will cause additional noise and cannot be optimized overall, so this method has not been widely used in recent trends of tree counting application. Recently, many deep learning based approaches are designed for this task because of the powerful feature extracting ability. The representative way is bounding box based supervised method, but time-consuming annotations are indispensable for them. Moreover, these methods are difficult to overcome the occlusion or overlap. To solve this problem, we propose a weakly tree counting network (WTCNet) based on deep segmentation network with only point supervision. It can simultaneously complete tree counting with localization and output mask of each tree at the same time. We first adopt a novel feature extractor network (FENet) to get features of input images, and then an effective strategy is introduced to deal with different mask predictions. In the end, we propose a basic localization guidance accompany with rectification guidance to train the network. We create two different datasets and select an existing challenging plant dataset to evaluate our method on three different tasks. Experimental results show the good performance improvement of our method compared with other existing methods. Further study shows that our method has great potential to reduce human labor and provide effective ground-truth masks and the results show the superiority of our method over the advanced methods. Pinmo Tong, Xishan Zhang, Pengcheng Han, Shuhui Bu |
ICPR | 2 |
| 2017 | Task-Driven Dynamic Fusion: Reducing Ambiguity in Video DescriptionabstractIntegrating complementary features from multiple channels is expected to solve the description ambiguity problem in video captioning, whereas inappropriate fusion strategies often harm rather than help the performance. Existing static fusion methods in video captioning such as concatenation and summation cannot attend to appropriate feature channels, thus fail to adaptively support the recognition of various kinds of visual entities such as actions and objects. This paper contributes to: 1)The first in-depth study of the weakness inherent in data-driven static fusion methods for video captioning. 2) The establishment of a task-driven dynamic fusion (TDDF) method. It can adaptively choose different fusion patterns according to model status. 3) The improvement of video captioning. Extensive experiments conducted on two well-known benchmarks demonstrate that our dynamic fusion method outperforms the state-of-the-art results on MSVD with METEOR scores 0.333, and achieves superior METEOR scores 0.278 on MSR-VTT-10K. Compared to single features, the relative improvement derived from our fusion method are 10.0% and 5.7% respectively on two datasets. Xishan Zhang, Ke Gao 0012, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001, Qi Tian 0001 |
CVPR | 1 |
| 2017 | Video Description with Spatial-Temporal AttentionabstractTemporal attention has been widely used in video description to adaptively focus on important frames. However, most existing methods based on temporal attention suffer from the problems of recognition error and detail missing, because only coarse frame-level global features are employed. Inspired by recent successful work in image description using spatial attention, we propose a spatial-temporal attention (STAT) method to address such problems. In particular, first, we take advantage of object-level local features to address the problem of detail missing. Second, the STAT method further selects relevant local features by spatial attention and then attend to important frames by temporal attention to recognize related semantics. The proposed two-stage attention mechanism can recognize the salient objects more precisely with high recall and automatically focus on the most relevant spatial-temporal segments given the sentence context. Extensive experiments on two well-known benchmarks suggest that STAT method outperforms the state-of-the-art methods on MSVD with BLEU4 score 0.511, and achieves superior BLEU4 score 0.374 on MSR-VTT-10K. Compared to the method without local features, the relative improvements derived from our STAT method are 10.1% and 0.8% respectively on two benchmarks. Compared to the method using only temporal attention, the relative improvements derived from our STAT method are 18.3% and 9.0% respectively on two benchmarks. Yunbin Tu, Xishan Zhang, Bingtao Liu, Chenggang Yan 0001 |
ACM Multimedia | 2 |
| 2017 | Trip Outfits Advisor: Location-Oriented Clothing RecommendationabstractWhen packing for a journey, have you ever asked “what clothes should I take with me?” Wearing appropriate and aesthetically pleasing clothing when traveling is a concern for many of us. Our data observation of photos from several popular travel websites reveals that people's choice of clothing items and their color combinations have strong correlations with the weather, the season, and the main type of attraction at the destination. This leads to an interesting and novel problem: can the correlation between clothing and locations be automatically learned from social photos and leveraged for location-oriented clothing recommendations? In this paper, we systematically study this problem and propose a hybrid multilabel convolutional neural network combined with the support vector machine (mCNN-SVM) approach to capture the intrinsic and complex correlations between clothing attributes and location attributes. Specifically, we adapt the CNN architecture to multilabel learning and fine-tune it using each fine-grained clothing item. Then, the recognized items are fed to the SVM to learn the correlations. Experiments on three fashion datasets and a benchmark journey outfit dataset show that our proposed approach outperforms several baselines by over 10.52-16.38% in terms of the mAP for clothing item recognition and outperforms several alternative methods by over 9.59-29.41% in terms of the mAP when ranking clothing by appropriateness for travel destinations. Finally, an interesting case study demonstrates the effectiveness of our method by answering what items to wear, how to match them, and how to dress in an aesthetically pleasing manner for a journey. Xishan Zhang, Jia Jia 0001, Ke Gao 0012, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Deep Fusion of Multiple Semantic Cues for Complex Event RecognitionabstractWe present a deep learning strategy to fuse multiple semantic cues for complex event recognition. In particular, we tackle the recognition task by answering how to jointly analyze human actions (who is doing what), objects (what), and scenes (where). First, each type of semantic features (e.g., human action trajectories) is fed into a corresponding multi-layer feature abstraction pathway, followed by a fusion layer connecting all the different pathways. Second, the correlations of how the semantic cues interacting with each other are learned in an unsupervised cross-modality autoencoder fashion. Finally, by fine-tuning a large-margin objective deployed on this deep architecture, we are able to answer the question on how the semantic cues of who, what, and where compose a complex event. As compared with the traditional feature fusion methods (e.g., various early or late strategies), our method jointly learns the essential higher level features that are most effective for fusion and recognition. We perform extensive experiments on two real-world complex event video benchmarks, MED'11 and CCV, and demonstrate that our method outperforms the best published results by 21% and 11%, respectively, on an event recognition task. Xishan Zhang, Hanwang Zhang, Yongdong Zhang 0001, Yang Yang 0002, Meng Wang 0001, Huan-Bo Luan, Jintao Li 0001, Tat-Seng Chua |
IEEE Trans. Image Process. | 1 |
| 2015 | Enhancing Video Event Recognition Using Automatically Constructed Semantic-Visual Knowledge BaseabstractThe task of recognizing events from video has attracted a lot of attention in recent years. However, due to the complex nature of user-defined events, the use of purely audio- visual content analysis without domain knowledge has been found to be grossly inadequate. In this paper, we propose to construct a semantic-visual knowledge base to encode the rich event-centric concepts and their relationships from the well- established lexical databases, including FrameNet, as well as the concept-specific visual knowledge from ImageNet. Based on this semantic-visual knowledge bases, we design an effective system for video event recognition. Specifically, in order to narrow the semantic gap between the high-level complex events and low-level visual representations, we utilize the event-centric semantic concepts encoded in the knowledge base as the intermediate-level event representation, which offers both human-perceivable and machine-interpretable semantic clues for event recognition. In addition, in order to leverage the abundant ImageNet images, we propose a robust transfer learning model to learn the noise- resistant concept classifiers for videos. Extensive experiments on various real-world video datasets demonstrate the superiority of our proposed system as compared to the state-of-the-art approaches. Xishan Zhang, Yang Yang 0002, Yongdong Zhang 0001, Huan-Bo Luan, Jintao Li 0001, Hanwang Zhang, Tat-Seng Chua |
IEEE Trans. Multim. | 1 |