VLDB 2026 Research / reviewers in the wild / expert
Zhiyuan Liang
dblp:191/0904
· DBLP profile ↗
19ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Segmentation and scene understanding · 20% Efficient and distributed learning · 18% Language models and text generation · 16% | |
| Network and information security
1 paper |
Digital forensics and information hiding · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Database system architecture and tuning · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 30 heaviest of 38, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
semantic segmentation |
1.8 | 3 | 2023 | Multi-Granularity Context Network for Efficient Video Semantic Segmentation · IEEE Trans. Image Process. 2023 Person Foreground Segmentation by Learning Multi-Domain Networks · IEEE Trans. Image Process. 2022 Tree Energy Loss: Towards Sparsely Annotated Semantic Segmentation · CVPR 2022 |
Machine learning › Efficient and distributed learning
model compression |
1.5 | 2 | 2025 | Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths · NeurIPS 2025 Multi-Granularity Context Network for Efficient Video Semantic Segmentation · IEEE Trans. Image Process. 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
theory of mind |
1.0 | 1 | 2026 | GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs · ACL (1) 2026 |
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
diffusion transformer |
0.9 | 1 | 2025 | REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training · NeurIPS 2025 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.9 | 1 | 2025 | CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights · NeurIPS 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation › preference optimization
self-rewarding language model |
0.9 | 1 | 2025 | CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025 |
Machine learning › Deep learning architectures and training
state space model |
0.9 | 1 | 2025 | Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › efficient training
training acceleration |
0.9 | 1 | 2025 | REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › state space model
vision mamba |
0.9 | 1 | 2025 | Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
zero-shot domain adaptation |
0.9 | 1 | 2025 | Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights · NeurIPS 2025 |
Database system architecture and tuning
hybrid transactional and analytical processing |
0.8 | 1 | 2024 | OSSInsight: Scalable GitHub Analysis · Proc. VLDB Endow. 2024 |
Empirical software engineering
mining software repositories |
0.8 | 1 | 2024 | OSSInsight: Scalable GitHub Analysis · Proc. VLDB Endow. 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.7 | 1 | 2023 | Multi-Granularity Context Network for Efficient Video Semantic Segmentation · IEEE Trans. Image Process. 2023 |
Computer vision › Segmentation and scene understanding › video segmentation
video semantic segmentation |
0.7 | 1 | 2023 | Multi-Granularity Context Network for Efficient Video Semantic Segmentation · IEEE Trans. Image Process. 2023 |
Computer vision › Segmentation and scene understanding › object segmentation
human segmentation |
0.6 | 1 | 2022 | Person Foreground Segmentation by Learning Multi-Domain Networks · IEEE Trans. Image Process. 2022 |
Machine learning › Transfer learning and domain adaptation › cross-domain learning
multi-domain learning |
0.6 | 1 | 2022 | Person Foreground Segmentation by Learning Multi-Domain Networks · IEEE Trans. Image Process. 2022 |
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation |
0.6 | 1 | 2022 | Tree Energy Loss: Towards Sparsely Annotated Semantic Segmentation · CVPR 2022 |
Computer vision › Video understanding and tracking › video analytics › behavior analysis
multi-person video analysis |
0.5 | 1 | 2021 | Face Forensics in the Wild · CVPR 2021 |
Digital forensics and information hiding
digital forensics |
0.5 | 1 | 2021 | Face Forensics in the Wild · CVPR 2021 |
Digital forensics and information hiding › forgery detection
face forgery detection |
0.5 | 1 | 2021 | Face Forensics in the Wild · CVPR 2021 |
Computer vision › Segmentation and scene understanding
human parsing |
0.4 | 1 | 2020 | Self-Learning With Rectification Strategy for Human Parsing · CVPR 2020 |
Computer vision › Video understanding and tracking
object tracking |
0.4 | 1 | 2020 | Local Semantic Siamese Networks for Fast Tracking · IEEE Trans. Image Process. 2020 |
Computer vision › Segmentation and scene understanding › pseudo-label learning
pseudo-label refinement |
0.4 | 1 | 2020 | Self-Learning With Rectification Strategy for Human Parsing · CVPR 2020 |
Machine learning › Representation and self-supervised learning
self-learning |
0.4 | 1 | 2020 | Self-Learning With Rectification Strategy for Human Parsing · CVPR 2020 |
Computer vision › Video understanding and tracking › object tracking › deep tracking
siamese tracking |
0.4 | 1 | 2020 | Local Semantic Siamese Networks for Fast Tracking · IEEE Trans. Image Process. 2020 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.3 | 1 | 2025 | Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
knowledge distillation · 1.5SQL query generation · 1.5LLM-based data analysis · 1.5pseudo-labeling · 1.0benchmark construction · 1.0text encoder · 0.9multi-scale scanning · 0.9hyper-convolutional decoder · 0.9direct preference optimization · 0.9consistency regularization · 0.9LoRA · 0.9LLM-as-a-judge · 0.9multiple instance learning · 0.5domain-adversarial quality assessment · 0.5density-based clustering · 0.2DBSCAN clustering · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMsabstractWeidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Can Zhang, Xinyan Wan, Zhiyuan Liang, Pengfei Zhou, Yang You, Wangbo Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Xinyan Wan, Zhiyuan Liang, Yang You 0001, Wangbo Zhao |
ACL (1) | 7 |
| 2026 | EC-RAG: Towards Efficient Edge-Cloud Retrieval-Augmented Generation Systems
Liang Wang 0001, Ranjun Jia, Kai Lu 0002, Jiguang Wan, Hao Huo, Yulong Zhai, Zhiyuan Liang |
ICDE | 8 |
| 2026 | ByteDance: Let bytes perform brilliantly in multi-view encrypted traffic classification
Yuwei Xu 0001, Zhiyuan Liang, Xiaotian Fang, Kehui Song, Qiao Xiang, Guang Cheng 0001 |
Comput. Networks | 2 |
| 2025 | CREAM: Consistency Regularized Self-Rewarding Language ModelsabstractRecent self-rewarding large language models (LLM) have successfully applied LLM-as-a-Judge to iteratively improve the alignment performance without the need of human annotations for preference data. These methods commonly utilize the same LLM to act as both the policy model (which generates responses) and the reward model (which scores and ranks those responses). The ranked responses are then used as preference pairs to train the LLM via direct alignment technologies (e.g. DPO). However, it is noteworthy that throughout this process, there is no guarantee of accuracy in the rewarding and ranking, which is critical for ensuring accurate rewards and high-quality preference data. Empirical results from relatively small LLMs (e.g., 7B parameters) also indicate that improvements from self-rewarding may diminish after several iterations in certain situations, which we hypothesize is due to accumulated bias in the reward system. This bias can lead to unreliable preference data for training the LLM. To address this issue, we first formulate and analyze the generalized iterative preference fine-tuning framework for self-rewarding language model. We then introduce the regularization to this generalized framework to mitigate the overconfident preference labeling in the self-rewarding process. Based on this theoretical insight, we propose a Consistency Regularized sElf-rewarding lAnguage Model (CREAM) that leverages the consistency of rewards across different iterations to regularize the self-rewarding training, helping the model to learn from more reliable preference data. With this explicit regularization, our empirical results demonstrate the superiority of CREAM in improving both reward consistency and alignment performance. The code is publicly available at https://github.com/Raibows/CREAM. Zhaoyang Wang 0004, Weilei He, Zhiyuan Liang, Xuchao Zhang, Chetan Bansal, Huaxiu Yao |
ICLR | 3 |
| 2025 | Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsabstractModern Parameter-Efficient Fine-Tuning (PEFT) methods such as low-rank adaptation (LoRA) reduce the cost of customizing large language models (LLMs), yet still require a separate optimization run for every downstream dataset. We introduce \textbf{Drag-and-Drop LLMs (\textit{DnD})}, a prompt-conditioned parameter generator that eliminates per-task training by mapping a handful of unlabeled task prompts directly to LoRA weight updates. A lightweight text encoder distills each prompt batch into condition embeddings, which are then transformed by a cascaded hyper-convolutional decoder into the full set of LoRA matrices. Once trained in a diverse collection of
prompt-checkpoint pairs, DnD produces task-specific parameters in seconds, yielding i) up to
\textbf{12,000$\times$} lower overhead than full fine-tuning, ii) average gains up to \textbf{30\%} in performance over the strongest training LoRAs on unseen common-sense reasoning, math, coding, and multimodal benchmarks, and iii) robust cross-domain generalization improving \textbf{40\%} performance without access to the target data or labels. Our results demonstrate that prompt-conditioned parameter generation is a viable alternative to gradient-based adaptation for rapidly specializing LLMs.
We open source \href{https://jerryliang24.github.io/DnD}{our project} in support of future research. Zhiyuan Liang, Dongwen Tang, Yuhao Zhou 0004, Xuanlei Zhao, Mingjia Shi, Wangbo Zhao, Peihao Wang, Konstantin Schürholt, Damian Borth, Michael M. Bronstein, Yang You 0001, Zhangyang Wang, Kai Wang 0036 |
NeurIPS | 1 |
| 2025 | Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning PathsabstractMamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's chain-like scanning mechanism, which we hypothesize not only induces cascading losses in token connectivity but also limits the diversity of spatial receptive fields.
In this paper, we propose Asymmetric Multi-scale Vision Mamba (AMVim), a novel architecture designed to enhance pruning robustness. AMVim employs a dual-path structure, integrating a window-aware scanning mechanism into one path while retaining sequential scanning in the other. This asymmetry design promotes token connection diversity and enables multi-scale information flow, reinforcing spatial awareness.
Empirical results demonstrate that AMVim achieves state-of-the-art pruning robustness. During token reduction, AMVim-T achieves a substantial 34\% improvement in training-free accuracy with identical model sizes and FLOPs. Meanwhile, AMVim-S exhibits only a 1.5\% accuracy drop, performing comparably to ViT. Notably, AMVim also delivers superior performance during pruning-free settings, further validating its architectural advantages. Jindi Lv, Yuhao Zhou 0004, Mingjia Shi, Zhiyuan Liang, Xiaojiang Peng, Wangbo Zhao, Jiancheng Lv 0001, Kai Wang 0036 |
NeurIPS | 4 |
| 2025 | REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion TrainingabstractDiffusion Transformers (DiTs) deliver state-of-the-art image quality, yet their training remains notoriously slow. A recent remedy---representation alignment (REPA) that matches DiT hidden features to those of a non-generative teacher (e.g., DINO)---dramatically accelerates the early epochs but plateaus or even degrades performance later. We trace this failure to the capacity mismatch: once the generative student begins modeling the joint data distribution, the teacher's lower-dimensional embeddings and attention patterns become a straitjacket rather than a guide. We then introduce HASTE (Holistic Alignment with Stage-wise Termination for Efficient training), a two-phase schedule that keeps the help and drops the hindrance. Phase I applies a holistic alignment loss that simultaneously distills attention maps (relational priors) and feature projections (semantic anchors) from the teacher into mid-level layers of the DiT, yielding rapid convergence. Phase II then performs one-shot termination that deactivates the alignment loss, once a simple trigger such as a fixed iteration is hit, freeing the DiT to focus on denoising and exploit its generative capacity. HASTE speeds up training of diverse DiTs without architecture changes. On ImageNet 256×256, it reaches the vanilla SiT-XL/2 baseline FID in 50 epochs and matches REPA’s best FID in 500 epochs, amounting to a 28× reduction in optimization steps. HASTE also improves text-to-image DiTs on MS-COCO, proving to be a simple yet principled recipe for efficient diffusion training across various tasks. Wangbo Zhao, Yuhao Zhou 0004, Zhiyuan Liang, Mingjia Shi, Xuanlei Zhao, Kaipeng Zhang, Zhangyang Wang, Kai Wang 0036, Yang You 0001 |
NeurIPS | 5 |
| 2024 | FullView: Using Bidirectional Group Sequences to Achieve Accurate Encrypted Traffic Classification
Yuwei Xu 0001, Zhiyuan Liang, Zhengxin Xu, Kehui Song, Qiao Xiang, Guang Cheng 0001 |
SecureComm (2) | 2 |
| 2024 | OSSInsight: Scalable GitHub AnalysisabstractGitHub is a platform hosting code, enabling collaboration, and supporting version control for a global community of over 100 million developers. The need for free tools is crucial for researching open-source software. Based on our research, we found out that existing tools lack real-time GitHub data processing or have limited functionalities. This demonstration presents OSSInsight, an open source tool for researching and analyzing GitHub repositories. We first present the architecture of the tool including its access to nearly 7 billion archived & real time data and how it is powered by TiDB. The demonstration shows how OSSInsight provides analysis of GitHub data along three dimensions: developers, repositories and organizations. All these analysis are based on generated SQL queries submitted to TiDB database. TiDB possesses HTAP capabilities, utilizing its row store for simple SQL queries while relying on its column store for more complex queries. Users can view and edit these SQL queries and also view their execution plan. Finally, OSSInsight provides an innovative tool based on OpenAI, that conducts data analysis using input in English text, yielding visual representations in the form of charts and graphs. Ahmad Ghazal, Zhiyuan Liang, Sunny Bains, Hanumath Maduri |
Proc. VLDB Endow. | 2 |
| 2023 | Blind Super-Resolution of Single Remotely Sensed Hyperspectral ImageabstractHyperspectral image (HSI) super-resolution has recently advanced with significant progress by utilizing the powerful representation capabilities of deep neural networks. These approaches, however, inevitably rely on a sizable amount of training data which can be difficult to acquire for remotely sensed HSIs. In many cases, these methods are designed and tailored for only one or a few specific super-resolution scenarios, making them inflexible for handling images with different unknown degradations. In this paper, we introduce a two-step framework for blind remotely sensed HSI super-resolution, where the degradation is unknown. Specifically, in the first step, we propose to leverage the abundant remotely sensed color images to address the data insufficiency for remotely sensed HSI super-resolution. It is achieved by exploring the spatial knowledge from remotely sensed color images with a super-resolution network for a predefined degradation, which is then transferred to HSIs via band-by-band super-resolution. Direct use of the results from the transferred super-resolution network is suboptimal as it neglects the spectral correlations of different bands and the gap between predefined degradation and the real one. To make further refinements, we present an unsupervised scheme that simultaneously refines the super-resolved HSI and the unknown degradation by a non-negative matrix factorization network and a learnable degradation prior. To validate the effectiveness of our method, we conducted extensive experiments on a variety of remotely sensed HSI datasets. The results demonstrate that our method could generalize on various unknown degradations with superior performance against the state-of-the-art methods. Zhiyuan Liang, Shuai Wang 0049, Tao Zhang 0042, Ying Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Multi-Granularity Context Network for Efficient Video Semantic SegmentationabstractCurrent video semantic segmentation tasks involve two main challenges: how to take full advantage of multi-frame context information, and how to improve computational efficiency. To tackle the two challenges simultaneously, we present a novel Multi-Granularity Context Network (MGCNet) by aggregating context information at multiple granularities in a more effective and efficient way. Our method first converts image features into semantic prototypes, and then conducts a non-local operation to aggregate the per-frame and short-term contexts jointly. An additional long-term context module is introduced to capture the video-level semantic information during training. By aggregating both local and global semantic information, a strong feature representation is obtained. The proposed pixel-to-prototype non-local operation requires less computational cost than traditional non-local ones, and is video-friendly since it reuses the semantic prototypes of previous frames. Moreover, we propose an uncertainty-aware and structural knowledge distillation strategy to boost the performance of our method. Experiments on Cityscapes and CamVid datasets with multiple backbones demonstrate that the proposed MGCNet outperforms other state-of-the-art methods with high speed and low latency. Zhiyuan Liang, Xiangdong Dai, Xiaogang Jin 0001, Jianbing Shen |
IEEE Trans. Image Process. | 1 |
| 2022 | Tree Energy Loss: Towards Sparsely Annotated Semantic SegmentationabstractSparsely annotated semantic segmentation (SASS) aims to train a segmentation network with coarse-grained (i.e., point-, scribble-, and block-wise) supervisions, where only a small proportion of pixels are labeled in each image. In this paper, we propose a novel tree energy loss for SASS by providing semantic guidance for unlabeled pixels. The tree energy loss represents images as minimum spanning trees to model both low-level and high-level pair-wise affini-ties. By sequentially applying these affinities to the net-work prediction, soft pseudo labels for unlabeled pixels are generated in a coarse-to-fine manner, achieving dynamic online self-training. The tree energy loss is effective and easy to be incorporated into existing frameworks by com-bining it with a traditional segmentation loss. Compared with previous SASS methods, our method requires no multi-stage training strategies, alternating optimization proce-dures, additional supervised data, or time-consuming post-processing while outperforming them in all SASS settings. Code is available at https://github.com/megvii-research/TreeEnergyLoss. Zhiyuan Liang, Tiancai Wang, Xiangyu Zhang 0005, Jian Sun 0001, Jianbing Shen |
CVPR | 1 |
| 2022 | Person Foreground Segmentation by Learning Multi-Domain NetworksabstractSeparating the dominant person from the complex background is significant to the human-related research and photo-editing based applications. Existing segmentation algorithms are either too general to separate the person region accurately, or not capable of achieving real-time speed. In this paper, we introduce the multi-domain learning framework into a novel baseline model to construct the Multi-domain TriSeNet Networks for the real-time single person image segmentation. We first divide training data into different subdomains based on the characteristics of single person images, then apply a multi-branch Feature Fusion Module (FFM) to decouple the networks into the domain-independent and the domain-specific layers. To further enhance the accuracy, a self-supervised learning strategy is proposed to dig out domain relations during training. It helps transfer domain-specific knowledge by improving predictive consistency among different FFM branches. Moreover, we create a large-scale single person image segmentation dataset named MSSP20k, which consists of 22,100 pixel-level annotated images in the real world. The MSSP20k dataset is more complex and challenging than existing public ones in terms of scalability and variety. Experiments show that our Multi-domain TriSeNet outperforms state-of-the-art approaches on both public and the newly built datasets with real-time speed. Zhiyuan Liang, Kan Guo, Xiaogang Jin 0001, Jianbing Shen |
IEEE Trans. Image Process. | 1 |
| 2021 | Face Forensics in the WildabstractOn existing public benchmarks, face forgery detection techniques have achieved great success. However, when used in multi-person videos, which often contain many people active in the scene with only a small subset having been manipulated, their performance remains far from being satisfactory. To take face forgery detection to a new level, we construct a novel large-scale dataset, called FFIW10K, which comprises 10,000 high-quality forgery videos, with an average of three human faces in each frame. The manipulation procedure is fully automatic, controlled by a domain-adversarial quality assessment network, making our dataset highly scalable with low human cost. In addition, we propose a novel algorithm to tackle the task of multi-person face forgery detection. Supervised by only video-level label, the algorithm explores multiple instance learning and learns to automatically attend to tampered faces. Our algorithm outperforms representative approaches for both forgery classification and localization on FFIW10K, and also shows high generalization ability on existing benchmarks. We hope that our dataset and study will help the community to explore this new field in more depth. Tianfei Zhou, Wenguan Wang, Zhiyuan Liang, Jianbing Shen |
CVPR | 3 |
| 2020 | Self-Learning With Rectification Strategy for Human ParsingabstractIn this paper, we solve the sample shortage problem in the human parsing task. We begin with the self-learning strategy, which generates pseudo-labels for unlabeled data to retrain the model. However, directly using noisy pseudo-labels will cause error amplification and accumulation. Considering the topology structure of human body, we propose a trainable graph reasoning method that establishes internal structural connections between graph nodes to correct two typical errors in the pseudo-labels, i.e., the global structural error and the local consistency error. For the global error, we first transform category-wise features into a high-level graph model with coarse-grained structural information, and then decouple the high-level graph to reconstruct the category features. The reconstructed features have a stronger ability to represent the topology structure of the human body. Enlarging the receptive field of features can effectively reducing the local error. We first project feature pixels into a local graph model to capture pixel-wise relations in a hierarchical graph manner, then reverse the relation information back to the pixels. With the global structural and local consistency modules, these errors are rectified and confident pseudo-labels are generated for retraining. Extensive experiments on the LIP and the ATR datasets demonstrate the effectiveness of our global and local rectification modules. Our method outperforms other state-of-the-art methods in supervised human parsing tasks. Zhiyuan Liang, Sanyuan Zhao, Jiahao Gong, Jianbing Shen |
CVPR | 2 |
| 2020 | Local Semantic Siamese Networks for Fast TrackingabstractLearning a powerful feature representation is critical for constructing a robust Siamese tracker. However, most existing Siamese trackers learn the global appearance features of the entire object, which usually suffers from drift problems caused by partial occlusion or non-rigid appearance deformation. In this paper, we propose a new Local Semantic Siamese (LSSiam) network to extract more robust features for solving these drift problems, since the local semantic features contain more fine-grained and partial information. We learn the semantic features during offline training by adding a classification branch into the classical Siamese framework. To further enhance the representation of features, we design a generally focal logistic loss to mine the hard negative samples. During the online tracking, we remove the classification branch and propose an efficient template updating strategy to avoid aggressive computing load. Thus, the proposed tracker can run at a high-speed of 100 Frame-per-Second (FPS) far beyond real-time requirement. Extensive experiments on popular benchmarks demonstrate the proposed LSSiam tracker achieves the state-of-the-art performance with a high-speed. Our source code is available at. Zhiyuan Liang, Jianbing Shen |
IEEE Trans. Image Process. | 1 |
| 2019 | Multiobject Tracking by Submodular OptimizationabstractIn this paper, we propose a new multiobject visual tracking algorithm by submodular optimization. The proposed algorithm is composed of two main stages. At the first stage, a new selecting strategy of tracklets is proposed to cope with occlusion problem. We generate low-level tracklets using overlap criteria and min-cost flow, respectively, and then integrate them into a candidate tracklets set. In the second stage, we formulate the multiobject tracking problem as the submodular maximization problem subject to related constraints. The submodular function selects the correct tracklets from the candidate set of tracklets to form the object trajectory. Then, we design a connecting process which connects the corresponding trajectories to overcome the occlusion problem. Experimental results demonstrate the effectiveness of our tracking algorithm. Our source code is available at https://github.com/shenjianbing/submodulartrack. Jianbing Shen, Zhiyuan Liang, Jianhong Liu, Hanqiu Sun, Ling Shao 0001, Dacheng Tao |
IEEE Trans. Cybern. | 2 |
| 2018 | Robust Stereoscopic Crosstalk PredictionabstractWe propose a new metric to predict perceived crosstalk using the original images rather than both the original and ghosted images. The proposed metrics are based on color information. First, we extract a disparity map, a color difference map, and a color contrast map from original image pairs. Then, we use those maps to construct two new metrics (Vdispc and Vdlogc). Metric Vdispc considers the effect of the disparity map and the color difference map, while Vdlogc addresses the influence of the color contrast map. The prediction performance is evaluated using various types of stereoscopic crosstalk images. By incorporating Vdispc and Vdlogc, the new metric Vpdlc is proposed to achieve a higher correlation with the perceived subject crosstalk scores. Experimental results show that the new metrics achieve better performance than previous methods, which indicate that color information is one key factor for crosstalk visible prediction. Furthermore, we construct a new data set to evaluate our new metrics. Jianbing Shen, Yan Zhang 0094, Zhiyuan Liang, Chang Liu 0071, Hanqiu Sun, Xiaopeng Hao, Jianhong Liu, Jian Yang 0009, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Real-Time Superpixel Segmentation by DBSCAN Clustering AlgorithmabstractIn this paper, we propose a real-time image superpixel segmentation method with 50 frames/s by using the density-based spatial clustering of applications with noise (DBSCAN) algorithm. In order to decrease the computational costs of superpixel algorithms, we adopt a fast two-step framework. In the first clustering stage, the DBSCAN algorithm with color-similarity and geometric restrictions is used to rapidly cluster the pixels, and then, small clusters are merged into superpixels by their neighborhood through a distance measurement defined by color and spatial features in the second merging stage. A robust and simple distance function is defined for obtaining better superpixels in these two steps. The experimental results demonstrate that our real-time superpixel algorithm (50 frames/s) by the DBSCAN clustering outperforms the state-of-the-art superpixel segmentation methods in terms of both accuracy and efficiency. Jianbing Shen, Xiaopeng Hao, Zhiyuan Liang, Yu Liu 0074, Wenguan Wang, Ling Shao 0001 |
IEEE Trans. Image Process. | 3 |