Zhiyuan Liang

dblp:191/0904 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Segmentation and scene understanding · 20% Efficient and distributed learning · 18% Language models and text generation · 16%
Network and information security
1 paper
Digital forensics and information hiding · 100%
Databases, data mining, and information retrieval
1 paper
Database system architecture and tuning · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
1.832023
Multi-Granularity Context Network for Efficient Video Semantic Segmentation · IEEE Trans. Image Process. 2023
Person Foreground Segmentation by Learning Multi-Domain Networks · IEEE Trans. Image Process. 2022
Tree Energy Loss: Towards Sparsely Annotated Semantic Segmentation · CVPR 2022
Machine learning › Efficient and distributed learning
model compression
1.522025
Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths · NeurIPS 2025
Multi-Granularity Context Network for Efficient Video Semantic Segmentation · IEEE Trans. Image Process. 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
theory of mind
1.012026
GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs · ACL (1) 2026
Natural language and speech › Language models and text generation
alignment
0.912025
CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025
Machine learning › Generative modeling
diffusion model
0.912025
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.912025
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training · NeurIPS 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.912025
CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights · NeurIPS 2025
Natural language and speech › Language models and text generation
preference optimization
0.912025
CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025
Natural language and speech › Language models and text generation › preference optimization
self-rewarding language model
0.912025
CREAM: Consistency Regularized Self-Rewarding Language Models · ICLR 2025
Machine learning › Deep learning architectures and training
state space model
0.912025
Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths · NeurIPS 2025
Machine learning › Efficient and distributed learning › efficient training
training acceleration
0.912025
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training · NeurIPS 2025
Machine learning › Deep learning architectures and training › state space model
vision mamba
0.912025
Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
zero-shot domain adaptation
0.912025
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights · NeurIPS 2025
Database system architecture and tuning
hybrid transactional and analytical processing
0.812024
OSSInsight: Scalable GitHub Analysis · Proc. VLDB Endow. 2024
Empirical software engineering
mining software repositories
0.812024
OSSInsight: Scalable GitHub Analysis · Proc. VLDB Endow. 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.712023
Multi-Granularity Context Network for Efficient Video Semantic Segmentation · IEEE Trans. Image Process. 2023
Computer vision › Segmentation and scene understanding › video segmentation
video semantic segmentation
0.712023
Multi-Granularity Context Network for Efficient Video Semantic Segmentation · IEEE Trans. Image Process. 2023
Computer vision › Segmentation and scene understanding › object segmentation
human segmentation
0.612022
Person Foreground Segmentation by Learning Multi-Domain Networks · IEEE Trans. Image Process. 2022
Machine learning › Transfer learning and domain adaptation › cross-domain learning
multi-domain learning
0.612022
Person Foreground Segmentation by Learning Multi-Domain Networks · IEEE Trans. Image Process. 2022
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation
0.612022
Tree Energy Loss: Towards Sparsely Annotated Semantic Segmentation · CVPR 2022
Computer vision › Video understanding and tracking › video analytics › behavior analysis
multi-person video analysis
0.512021
Face Forensics in the Wild · CVPR 2021
Digital forensics and information hiding
digital forensics
0.512021
Face Forensics in the Wild · CVPR 2021
Digital forensics and information hiding › forgery detection
face forgery detection
0.512021
Face Forensics in the Wild · CVPR 2021
Computer vision › Segmentation and scene understanding
human parsing
0.412020
Self-Learning With Rectification Strategy for Human Parsing · CVPR 2020
Computer vision › Video understanding and tracking
object tracking
0.412020
Local Semantic Siamese Networks for Fast Tracking · IEEE Trans. Image Process. 2020
Computer vision › Segmentation and scene understanding › pseudo-label learning
pseudo-label refinement
0.412020
Self-Learning With Rectification Strategy for Human Parsing · CVPR 2020
Machine learning › Representation and self-supervised learning
self-learning
0.412020
Self-Learning With Rectification Strategy for Human Parsing · CVPR 2020
Computer vision › Video understanding and tracking › object tracking › deep tracking
siamese tracking
0.412020
Local Semantic Siamese Networks for Fast Tracking · IEEE Trans. Image Process. 2020
Machine learning › Deep learning architectures and training
attention mechanism
0.312025
Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 1.5SQL query generation · 1.5LLM-based data analysis · 1.5pseudo-labeling · 1.0benchmark construction · 1.0text encoder · 0.9multi-scale scanning · 0.9hyper-convolutional decoder · 0.9direct preference optimization · 0.9consistency regularization · 0.9LoRA · 0.9LLM-as-a-judge · 0.9multiple instance learning · 0.5domain-adversarial quality assessment · 0.5density-based clustering · 0.2DBSCAN clustering · 0.2
YearPublicationVenuePosition
2026 GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs
abstract
Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Can Zhang, Xinyan Wan, Zhiyuan Liang, Pengfei Zhou, Yang You, Wangbo Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Weidong Tang, Jierui Li, Yueling Hou, Zihan Mei, Xinyan Wan, Zhiyuan Liang, Yang You 0001, Wangbo Zhao
ACL (1)7
2026 EC-RAG: Towards Efficient Edge-Cloud Retrieval-Augmented Generation Systems
Liang Wang 0001, Ranjun Jia, Kai Lu 0002, Jiguang Wan, Hao Huo, Yulong Zhai, Zhiyuan Liang
ICDE8
2026 ByteDance: Let bytes perform brilliantly in multi-view encrypted traffic classification
Yuwei Xu 0001, Zhiyuan Liang, Xiaotian Fang, Kehui Song, Qiao Xiang, Guang Cheng 0001
Comput. Networks2
2025 CREAM: Consistency Regularized Self-Rewarding Language Models
abstract
Recent self-rewarding large language models (LLM) have successfully applied LLM-as-a-Judge to iteratively improve the alignment performance without the need of human annotations for preference data. These methods commonly utilize the same LLM to act as both the policy model (which generates responses) and the reward model (which scores and ranks those responses). The ranked responses are then used as preference pairs to train the LLM via direct alignment technologies (e.g. DPO). However, it is noteworthy that throughout this process, there is no guarantee of accuracy in the rewarding and ranking, which is critical for ensuring accurate rewards and high-quality preference data. Empirical results from relatively small LLMs (e.g., 7B parameters) also indicate that improvements from self-rewarding may diminish after several iterations in certain situations, which we hypothesize is due to accumulated bias in the reward system. This bias can lead to unreliable preference data for training the LLM. To address this issue, we first formulate and analyze the generalized iterative preference fine-tuning framework for self-rewarding language model. We then introduce the regularization to this generalized framework to mitigate the overconfident preference labeling in the self-rewarding process. Based on this theoretical insight, we propose a Consistency Regularized sElf-rewarding lAnguage Model (CREAM) that leverages the consistency of rewards across different iterations to regularize the self-rewarding training, helping the model to learn from more reliable preference data. With this explicit regularization, our empirical results demonstrate the superiority of CREAM in improving both reward consistency and alignment performance. The code is publicly available at https://github.com/Raibows/CREAM.
Zhaoyang Wang 0004, Weilei He, Zhiyuan Liang, Xuchao Zhang, Chetan Bansal, Huaxiu Yao
ICLR3
2025 Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
abstract
Modern Parameter-Efficient Fine-Tuning (PEFT) methods such as low-rank adaptation (LoRA) reduce the cost of customizing large language models (LLMs), yet still require a separate optimization run for every downstream dataset. We introduce \textbf{Drag-and-Drop LLMs (\textit{DnD})}, a prompt-conditioned parameter generator that eliminates per-task training by mapping a handful of unlabeled task prompts directly to LoRA weight updates. A lightweight text encoder distills each prompt batch into condition embeddings, which are then transformed by a cascaded hyper-convolutional decoder into the full set of LoRA matrices. Once trained in a diverse collection of prompt-checkpoint pairs, DnD produces task-specific parameters in seconds, yielding i) up to \textbf{12,000$\times$} lower overhead than full fine-tuning, ii) average gains up to \textbf{30\%} in performance over the strongest training LoRAs on unseen common-sense reasoning, math, coding, and multimodal benchmarks, and iii) robust cross-domain generalization improving \textbf{40\%} performance without access to the target data or labels. Our results demonstrate that prompt-conditioned parameter generation is a viable alternative to gradient-based adaptation for rapidly specializing LLMs. We open source \href{https://jerryliang24.github.io/DnD}{our project} in support of future research.
Zhiyuan Liang, Dongwen Tang, Yuhao Zhou 0004, Xuanlei Zhao, Mingjia Shi, Wangbo Zhao, Peihao Wang, Konstantin Schürholt, Damian Borth, Michael M. Bronstein, Yang You 0001, Zhangyang Wang, Kai Wang 0036
NeurIPS1
2025 Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths
abstract
Mamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's chain-like scanning mechanism, which we hypothesize not only induces cascading losses in token connectivity but also limits the diversity of spatial receptive fields. In this paper, we propose Asymmetric Multi-scale Vision Mamba (AMVim), a novel architecture designed to enhance pruning robustness. AMVim employs a dual-path structure, integrating a window-aware scanning mechanism into one path while retaining sequential scanning in the other. This asymmetry design promotes token connection diversity and enables multi-scale information flow, reinforcing spatial awareness. Empirical results demonstrate that AMVim achieves state-of-the-art pruning robustness. During token reduction, AMVim-T achieves a substantial 34\% improvement in training-free accuracy with identical model sizes and FLOPs. Meanwhile, AMVim-S exhibits only a 1.5\% accuracy drop, performing comparably to ViT. Notably, AMVim also delivers superior performance during pruning-free settings, further validating its architectural advantages.
Jindi Lv, Yuhao Zhou 0004, Mingjia Shi, Zhiyuan Liang, Xiaojiang Peng, Wangbo Zhao, Jiancheng Lv 0001, Kai Wang 0036
NeurIPS4
2025 REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
abstract
Diffusion Transformers (DiTs) deliver state-of-the-art image quality, yet their training remains notoriously slow. A recent remedy---representation alignment (REPA) that matches DiT hidden features to those of a non-generative teacher (e.g., DINO)---dramatically accelerates the early epochs but plateaus or even degrades performance later. We trace this failure to the capacity mismatch: once the generative student begins modeling the joint data distribution, the teacher's lower-dimensional embeddings and attention patterns become a straitjacket rather than a guide. We then introduce HASTE (Holistic Alignment with Stage-wise Termination for Efficient training), a two-phase schedule that keeps the help and drops the hindrance. Phase I applies a holistic alignment loss that simultaneously distills attention maps (relational priors) and feature projections (semantic anchors) from the teacher into mid-level layers of the DiT, yielding rapid convergence. Phase II then performs one-shot termination that deactivates the alignment loss, once a simple trigger such as a fixed iteration is hit, freeing the DiT to focus on denoising and exploit its generative capacity. HASTE speeds up training of diverse DiTs without architecture changes. On ImageNet 256×256, it reaches the vanilla SiT-XL/2 baseline FID in 50 epochs and matches REPA’s best FID in 500 epochs, amounting to a 28× reduction in optimization steps. HASTE also improves text-to-image DiTs on MS-COCO, proving to be a simple yet principled recipe for efficient diffusion training across various tasks.
Wangbo Zhao, Yuhao Zhou 0004, Zhiyuan Liang, Mingjia Shi, Xuanlei Zhao, Kaipeng Zhang, Zhangyang Wang, Kai Wang 0036, Yang You 0001
NeurIPS5
2024 FullView: Using Bidirectional Group Sequences to Achieve Accurate Encrypted Traffic Classification
Yuwei Xu 0001, Zhiyuan Liang, Zhengxin Xu, Kehui Song, Qiao Xiang, Guang Cheng 0001
SecureComm (2)2
2024 OSSInsight: Scalable GitHub Analysis
abstract
GitHub is a platform hosting code, enabling collaboration, and supporting version control for a global community of over 100 million developers. The need for free tools is crucial for researching open-source software. Based on our research, we found out that existing tools lack real-time GitHub data processing or have limited functionalities. This demonstration presents OSSInsight, an open source tool for researching and analyzing GitHub repositories. We first present the architecture of the tool including its access to nearly 7 billion archived & real time data and how it is powered by TiDB. The demonstration shows how OSSInsight provides analysis of GitHub data along three dimensions: developers, repositories and organizations. All these analysis are based on generated SQL queries submitted to TiDB database. TiDB possesses HTAP capabilities, utilizing its row store for simple SQL queries while relying on its column store for more complex queries. Users can view and edit these SQL queries and also view their execution plan. Finally, OSSInsight provides an innovative tool based on OpenAI, that conducts data analysis using input in English text, yielding visual representations in the form of charts and graphs.
Ahmad Ghazal, Zhiyuan Liang, Sunny Bains, Hanumath Maduri
Proc. VLDB Endow.2
2023 Blind Super-Resolution of Single Remotely Sensed Hyperspectral Image
abstract
Hyperspectral image (HSI) super-resolution has recently advanced with significant progress by utilizing the powerful representation capabilities of deep neural networks. These approaches, however, inevitably rely on a sizable amount of training data which can be difficult to acquire for remotely sensed HSIs. In many cases, these methods are designed and tailored for only one or a few specific super-resolution scenarios, making them inflexible for handling images with different unknown degradations. In this paper, we introduce a two-step framework for blind remotely sensed HSI super-resolution, where the degradation is unknown. Specifically, in the first step, we propose to leverage the abundant remotely sensed color images to address the data insufficiency for remotely sensed HSI super-resolution. It is achieved by exploring the spatial knowledge from remotely sensed color images with a super-resolution network for a predefined degradation, which is then transferred to HSIs via band-by-band super-resolution. Direct use of the results from the transferred super-resolution network is suboptimal as it neglects the spectral correlations of different bands and the gap between predefined degradation and the real one. To make further refinements, we present an unsupervised scheme that simultaneously refines the super-resolved HSI and the unknown degradation by a non-negative matrix factorization network and a learnable degradation prior. To validate the effectiveness of our method, we conducted extensive experiments on a variety of remotely sensed HSI datasets. The results demonstrate that our method could generalize on various unknown degradations with superior performance against the state-of-the-art methods.
Zhiyuan Liang, Shuai Wang 0049, Tao Zhang 0042, Ying Fu 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Multi-Granularity Context Network for Efficient Video Semantic Segmentation
abstract
Current video semantic segmentation tasks involve two main challenges: how to take full advantage of multi-frame context information, and how to improve computational efficiency. To tackle the two challenges simultaneously, we present a novel Multi-Granularity Context Network (MGCNet) by aggregating context information at multiple granularities in a more effective and efficient way. Our method first converts image features into semantic prototypes, and then conducts a non-local operation to aggregate the per-frame and short-term contexts jointly. An additional long-term context module is introduced to capture the video-level semantic information during training. By aggregating both local and global semantic information, a strong feature representation is obtained. The proposed pixel-to-prototype non-local operation requires less computational cost than traditional non-local ones, and is video-friendly since it reuses the semantic prototypes of previous frames. Moreover, we propose an uncertainty-aware and structural knowledge distillation strategy to boost the performance of our method. Experiments on Cityscapes and CamVid datasets with multiple backbones demonstrate that the proposed MGCNet outperforms other state-of-the-art methods with high speed and low latency.
Zhiyuan Liang, Xiangdong Dai, Xiaogang Jin 0001, Jianbing Shen
IEEE Trans. Image Process.1
2022 Tree Energy Loss: Towards Sparsely Annotated Semantic Segmentation
abstract
Sparsely annotated semantic segmentation (SASS) aims to train a segmentation network with coarse-grained (i.e., point-, scribble-, and block-wise) supervisions, where only a small proportion of pixels are labeled in each image. In this paper, we propose a novel tree energy loss for SASS by providing semantic guidance for unlabeled pixels. The tree energy loss represents images as minimum spanning trees to model both low-level and high-level pair-wise affini-ties. By sequentially applying these affinities to the net-work prediction, soft pseudo labels for unlabeled pixels are generated in a coarse-to-fine manner, achieving dynamic online self-training. The tree energy loss is effective and easy to be incorporated into existing frameworks by com-bining it with a traditional segmentation loss. Compared with previous SASS methods, our method requires no multi-stage training strategies, alternating optimization proce-dures, additional supervised data, or time-consuming post-processing while outperforming them in all SASS settings. Code is available at https://github.com/megvii-research/TreeEnergyLoss.
Zhiyuan Liang, Tiancai Wang, Xiangyu Zhang 0005, Jian Sun 0001, Jianbing Shen
CVPR1
2022 Person Foreground Segmentation by Learning Multi-Domain Networks
abstract
Separating the dominant person from the complex background is significant to the human-related research and photo-editing based applications. Existing segmentation algorithms are either too general to separate the person region accurately, or not capable of achieving real-time speed. In this paper, we introduce the multi-domain learning framework into a novel baseline model to construct the Multi-domain TriSeNet Networks for the real-time single person image segmentation. We first divide training data into different subdomains based on the characteristics of single person images, then apply a multi-branch Feature Fusion Module (FFM) to decouple the networks into the domain-independent and the domain-specific layers. To further enhance the accuracy, a self-supervised learning strategy is proposed to dig out domain relations during training. It helps transfer domain-specific knowledge by improving predictive consistency among different FFM branches. Moreover, we create a large-scale single person image segmentation dataset named MSSP20k, which consists of 22,100 pixel-level annotated images in the real world. The MSSP20k dataset is more complex and challenging than existing public ones in terms of scalability and variety. Experiments show that our Multi-domain TriSeNet outperforms state-of-the-art approaches on both public and the newly built datasets with real-time speed.
Zhiyuan Liang, Kan Guo, Xiaogang Jin 0001, Jianbing Shen
IEEE Trans. Image Process.1
2021 Face Forensics in the Wild
abstract
On existing public benchmarks, face forgery detection techniques have achieved great success. However, when used in multi-person videos, which often contain many people active in the scene with only a small subset having been manipulated, their performance remains far from being satisfactory. To take face forgery detection to a new level, we construct a novel large-scale dataset, called FFIW10K, which comprises 10,000 high-quality forgery videos, with an average of three human faces in each frame. The manipulation procedure is fully automatic, controlled by a domain-adversarial quality assessment network, making our dataset highly scalable with low human cost. In addition, we propose a novel algorithm to tackle the task of multi-person face forgery detection. Supervised by only video-level label, the algorithm explores multiple instance learning and learns to automatically attend to tampered faces. Our algorithm outperforms representative approaches for both forgery classification and localization on FFIW10K, and also shows high generalization ability on existing benchmarks. We hope that our dataset and study will help the community to explore this new field in more depth.
Tianfei Zhou, Wenguan Wang, Zhiyuan Liang, Jianbing Shen
CVPR3
2020 Self-Learning With Rectification Strategy for Human Parsing
abstract
In this paper, we solve the sample shortage problem in the human parsing task. We begin with the self-learning strategy, which generates pseudo-labels for unlabeled data to retrain the model. However, directly using noisy pseudo-labels will cause error amplification and accumulation. Considering the topology structure of human body, we propose a trainable graph reasoning method that establishes internal structural connections between graph nodes to correct two typical errors in the pseudo-labels, i.e., the global structural error and the local consistency error. For the global error, we first transform category-wise features into a high-level graph model with coarse-grained structural information, and then decouple the high-level graph to reconstruct the category features. The reconstructed features have a stronger ability to represent the topology structure of the human body. Enlarging the receptive field of features can effectively reducing the local error. We first project feature pixels into a local graph model to capture pixel-wise relations in a hierarchical graph manner, then reverse the relation information back to the pixels. With the global structural and local consistency modules, these errors are rectified and confident pseudo-labels are generated for retraining. Extensive experiments on the LIP and the ATR datasets demonstrate the effectiveness of our global and local rectification modules. Our method outperforms other state-of-the-art methods in supervised human parsing tasks.
Zhiyuan Liang, Sanyuan Zhao, Jiahao Gong, Jianbing Shen
CVPR2
2020 Local Semantic Siamese Networks for Fast Tracking
abstract
Learning a powerful feature representation is critical for constructing a robust Siamese tracker. However, most existing Siamese trackers learn the global appearance features of the entire object, which usually suffers from drift problems caused by partial occlusion or non-rigid appearance deformation. In this paper, we propose a new Local Semantic Siamese (LSSiam) network to extract more robust features for solving these drift problems, since the local semantic features contain more fine-grained and partial information. We learn the semantic features during offline training by adding a classification branch into the classical Siamese framework. To further enhance the representation of features, we design a generally focal logistic loss to mine the hard negative samples. During the online tracking, we remove the classification branch and propose an efficient template updating strategy to avoid aggressive computing load. Thus, the proposed tracker can run at a high-speed of 100 Frame-per-Second (FPS) far beyond real-time requirement. Extensive experiments on popular benchmarks demonstrate the proposed LSSiam tracker achieves the state-of-the-art performance with a high-speed. Our source code is available at.
Zhiyuan Liang, Jianbing Shen
IEEE Trans. Image Process.1
2019 Multiobject Tracking by Submodular Optimization
abstract
In this paper, we propose a new multiobject visual tracking algorithm by submodular optimization. The proposed algorithm is composed of two main stages. At the first stage, a new selecting strategy of tracklets is proposed to cope with occlusion problem. We generate low-level tracklets using overlap criteria and min-cost flow, respectively, and then integrate them into a candidate tracklets set. In the second stage, we formulate the multiobject tracking problem as the submodular maximization problem subject to related constraints. The submodular function selects the correct tracklets from the candidate set of tracklets to form the object trajectory. Then, we design a connecting process which connects the corresponding trajectories to overcome the occlusion problem. Experimental results demonstrate the effectiveness of our tracking algorithm. Our source code is available at https://github.com/shenjianbing/submodulartrack.
Jianbing Shen, Zhiyuan Liang, Jianhong Liu, Hanqiu Sun, Ling Shao 0001, Dacheng Tao
IEEE Trans. Cybern.2
2018 Robust Stereoscopic Crosstalk Prediction
abstract
We propose a new metric to predict perceived crosstalk using the original images rather than both the original and ghosted images. The proposed metrics are based on color information. First, we extract a disparity map, a color difference map, and a color contrast map from original image pairs. Then, we use those maps to construct two new metrics (Vdispc and Vdlogc). Metric Vdispc considers the effect of the disparity map and the color difference map, while Vdlogc addresses the influence of the color contrast map. The prediction performance is evaluated using various types of stereoscopic crosstalk images. By incorporating Vdispc and Vdlogc, the new metric Vpdlc is proposed to achieve a higher correlation with the perceived subject crosstalk scores. Experimental results show that the new metrics achieve better performance than previous methods, which indicate that color information is one key factor for crosstalk visible prediction. Furthermore, we construct a new data set to evaluate our new metrics.
Jianbing Shen, Yan Zhang 0094, Zhiyuan Liang, Chang Liu 0071, Hanqiu Sun, Xiaopeng Hao, Jianhong Liu, Jian Yang 0009, Ling Shao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2016 Real-Time Superpixel Segmentation by DBSCAN Clustering Algorithm
abstract
In this paper, we propose a real-time image superpixel segmentation method with 50 frames/s by using the density-based spatial clustering of applications with noise (DBSCAN) algorithm. In order to decrease the computational costs of superpixel algorithms, we adopt a fast two-step framework. In the first clustering stage, the DBSCAN algorithm with color-similarity and geometric restrictions is used to rapidly cluster the pixels, and then, small clusters are merged into superpixels by their neighborhood through a distance measurement defined by color and spatial features in the second merging stage. A robust and simple distance function is defined for obtaining better superpixels in these two steps. The experimental results demonstrate that our real-time superpixel algorithm (50 frames/s) by the DBSCAN clustering outperforms the state-of-the-art superpixel segmentation methods in terms of both accuracy and efficiency.
Jianbing Shen, Xiaopeng Hao, Zhiyuan Liang, Yu Liu 0074, Wenguan Wang, Ling Shao 0001
IEEE Trans. Image Process.3