Zejian Yuan

dblp:26/455 · DBLP profile ↗
← Back
85ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0001-5548-3634ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 60 · 16 since 2021Artificial intelligence and machine learning · 54 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Point-line transformer with dynamic polyline anchors for lane detection
Kangyi Hou, Yuxing He, Zejian Yuan
Pattern Recognit. Lett.4
2025 TableLoRA: Low-rank Adaptation on Table Structure Understanding for Large Language Models
abstract
Tabular data are crucial in many fields and their understanding by large language models (LLMs) under high parameter efficiency paradigm is important. However, directly applying parameter-efficient fine-tuning (PEFT) techniques to tabular tasks presents significant challenges, particularly in terms of better table serialization and the representation of two-dimensional structured information within a one-dimensional sequence. To address this, we propose TableLoRA, a module designed to improve LLMs’ understanding of table structure during PEFT. It incorporates special tokens for serializing tables with special token encoder and uses 2D LoRA to encode low-rank information on cell positions. Experiments on four tabular-related datasets demonstrate that TableLoRA consistently outperforms vanilla LoRA and surpasses various table encoding methods tested in control experiments. These findings reveal that TableLoRA, as a table-specific LoRA, enhances the ability of LLMs to process tabular data effectively, especially in low-parameter settings, demonstrating its potential as a robust solution for handling table-related tasks.
Mengyu Zhou, Yeye He, Haoyu Dong 0001, Shi Han, Zejian Yuan, Dongmei Zhang 0001
ACL (1)7
2025 Multi-scale Feature Field with Anti-brightness-sensitivity Postprocessing for Few-shot Neural Panoptic Segmentation
abstract
Neural scene segmentation, which achieves 3D scene reconstruction and segmentation via 2D posed inputs, has garnered substantial attention. However, previous methods have relied on dense-view manual labels, while some improved few-shot or zero-shot models based on semantic-aware features or generated labels still encounter two challenges: (i) Lack of comprehensive perceptual ability, which limits their performance especially for complex scene understanding. (ii) Misleading by brightness-sensitive priors, which affects the segmentation particularly for specular highlight regions. Therefore, in this work, we propose Multi-scale Distilled & Anti-brightness-sensitivity Postprocessed Neural Feature Field (MDAP-NFF), a method for few-shot neural panoptic segmentation. To enhance perceptual ability, multi-scale feature field is designed for distilling multi-scale semantic-aware, e.g., DINO, features. For the robustness improvement against brightness-sensitivity in distilled feature field, we propose the anti-brightness-sensitivity postprocessing module, incorporating our designed 3Donv-VM technique. By training the feature field via 3D distillation and further postprocessing the features under supervision of sparse-view labels, our method achieves reliable 3D-consistent panoptic segmentation. Experimental results show that our model can perform high-quality panoptic segmentation under sparse-view supervision. Code and more results will be available at https://David-Dou.github.io/MDAP-NFF
Bin Dou, Yongjia Ma, Zejian Yuan
ICMR4
2025 Hierarchical Queries for 3D Lane Detection Based on Multi-Frame Point Clouds
abstract
3D lane detection based on multi-frame point clouds is a critical task for autonomous driving. The challenge lies in efficiently performing temporal fusion using multiple data frames with incomplete yet complementary contexts. Existing methods either directly concatenate consecutive frames, avoiding intrinsic limitations of the raw data, or fuse entire feature maps, without distinguishing lane-related features from backgrounds. These solutions exhibit room for improvement in both precision and efficiency. In this paper, we propose an end-to-end lane detection network with hierarchical queries, which decodes lane features at different levels in a top-down manner for high-precision localization. This framework can be deployed on multi-frame inputs, as it efficiently achieves lane-related sequence fusion with reduced computational costs and improved inference speed. Specifically, we design semi-parametric lane geometry representations to model lanes as parametric curves and discrete points. Accordingly, hierarchical queries are proposed to focus on two-level lane geometries, including curve queries and point queries. Curve queries capture global structures of lanes projected onto the bird’s-eye-view (BEV) flat ground, while point queries aggregate multi-frame sequences obtained through curve-guided sampling, acquiring comprehensive and reliable point-level features. In the training stage, our proposed curve matching and point localization loss optimizes the detected lane geometries at both levels. Experiments conducted on the self-collected MultiBEV dataset validate that our method outperforms previously published single-frame and multi-frame methods. Codes are released at https://github.com/lrx02/HQNet
Ruixin Liu, Zejian Yuan
IEEE Trans. Intell. Transp. Syst.2
2024 Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear Queries
abstract
Tabular data analysis is crucial in various fields, and large language models show promise in this area. However, current research mostly focuses on rudimentary tasks like Text2SQL and TableQA, neglecting advanced analysis like forecasting and chart generation. To address this gap, we developed the Text2Analysis benchmark, incorporating advanced analysis tasks that go beyond the SQL-compatible operations and require more in-depth analysis. We also develop five innovative and effective annotation methods, harnessing the capabilities of large language models to enhance data quality and quantity. Additionally, we include unclear queries that resemble real-world user questions to test how well models can understand and tackle such challenges. Finally, we collect 2249 query-result pairs with 347 tables. We evaluate five state-of-the-art models using three different metrics and the results show that our benchmark presents introduces considerable challenge in the field of tabular data analysis, paving the way for more advanced research opportunities.
Mengyu Zhou, Xinrun Xu, Xiaojun Ma 0001, Rui Ding 0001, Lun Du, Yan Gao 0002, Ran Jia, Xu Chen 0022, Shi Han, Zejian Yuan, Dongmei Zhang 0001
AAAI11
2024 Compact HD Map Construction via Douglas-Peucker Point Transformer
abstract
High-definition (HD) map construction requires a comprehensive understanding of traffic environments, encompassing centimeter-level localization and rich semantic information. Previous works face challenges in redundant point representation or high-complexity curve modeling. In this paper, we present a flexible yet effective map element detector that synthesizes hierarchical information with a compact Douglas-Peucker (DP) point representation in a transformer architecture for robust and reliable predictions. Specifically, our proposed representation approximates class-agnostic map elements with DP points, which are sparsely located in crucial positions of structures and can get rid of redundancy and complexity. Besides, we design a position constraint with uncertainty to avoid potential ambiguities. Moreover, pairwise-point shape matching constraints are proposed to balance local structural information of different scales. Experiments on the public nuScenes dataset demonstrate that our method overwhelms current SOTAs. Extensive ablation studies validate each component of our methods. Codes will be released at https://github.com/sweety121/DPFormer.
Ruixin Liu, Zejian Yuan
AAAI2
2024 Multi-Granularity Sparse Relationship Matrix Prediction Network for End-to-End Scene Graph Generation
Lei Wang 0219, Zejian Yuan, Badong Chen
ECCV (82)2
2024 CoCoST: Automatic Complex Code Generation with Online Searching and Correctness Testing
abstract
Large Language Models have revolutionized code generation ability by converting natural language descriptions into executable code.However, generating complex code within realworld scenarios remains challenging due to intricate structures, subtle bugs, understanding of advanced data types, and lack of supplementary contents.To address these challenges, we introduce the CoCoST framework, which enhances complex code generation by online searching for more information with planned queries and correctness testing for code refinement.Moreover, CoCoST serializes the complex inputs and outputs to improve comprehension and generates test cases to ensure the adaptability for real-world applications.CoCoST is validated through rigorous experiments on the DS-1000 and ClassEval datasets.Experimental results show that CoCoST substantially improves the quality of complex code generation, highlighting its potential to enhance the practicality of LLMs in generating complex code.
Jiaru Zou, Mengyu Zhou, Shi Han, Zejian Yuan, Dongmei Zhang 0001
EMNLP6
2024 RD-NERF: Neural Robust Distilled Feature Fields for Sparse-View Scene Segmentation
abstract
We propose Neural Robust Distilled Feature Fields (RD-NeRF) for achieving robust 3D semantic feature distillation and 3D consistent scene segmentation with sparse-view labels. Specifically, we introduce a two-stage pipeline. In the distillation stage, we employ the pre-trained image feature extractor, DINO-ViT, as the teacher network. RD-NeRF distills semantic knowledge into 3D space and utilizes the Vector-Matrix (VM) tensor decomposition method to represent semantic field with volumetric rendering. For the training process, we utilize the distance-wise and angle-wise distillation loss. This enables the student network to capture high-level semantics, enhance scene reconstruction and segmentation performance, and improve robustness and effectiveness in distillation. In the segmentation stage, hash features and distilled semantic features are inputs for the segmentation MLP, which is supervised by the sparse-view labels. The experimental results demonstrate that our model performs well in 3D-consistent scene segmentation under sparse-view supervision.
Yongjia Ma, Bin Dou, Zejian Yuan
ICASSP4
2024 Learning Segmented 3D Gaussians via Efficient Feature Unprojection for Zero-Shot Neural Scene Segmentation
Bin Dou, Yongjia Ma, Zejian Yuan, Nanning Zheng 0001
ICONIP (8)5
2024 Unbiased scene graph generation via head-tail cooperative network with self-supervised learning
Lei Wang 0219, Zejian Yuan, Badong Chen
Image Vis. Comput.2
2024 Csf: global-local shading orders for intrinsic image decomposition
Handan Zhang, Yuanliu Liu, Zejian Yuan
Mach. Vis. Appl.4
2024 Hierarchical Competition Learning for Pairwise Wheel Grounding Points Estimation
abstract
Multi-task detection often utilizes anchors to aggregate results. However, achieving a balance between the requirements of all tasks is a significant challenge, as anchors with the optimal predictions for different tasks often vary. In this paper, we introduce PGPENet, a hierarchical competition learning network for high-precision pairwise wheel grounding point estimation in autonomous driving. To address the issue of inconsistent optimal anchors between tasks, we propose a novel hierarchical competition anchor assignment method that dynamically determines the number of positive anchors during training. During testing, PGPENet employs a hierarchical post-processing approach to generate the best predictions. The proposed network follows a top-down structure, including object level and point level. The object-level network performs vehicle detection and coarse point detection, while the point-level network utilizes local features to refine the point positions. Sufficient experiments on a fisheye image dataset confirm the superiority of our model, particularly in scenarios where there are discrepancies between the best viewpoints of the vehicles and the wheels.
Zejian Yuan
IEEE Trans. Intell. Transp. Syst.2
2023 Flexible 3D Lane Detection by Hierarchical Shape Matching
abstract
As one of the basic while vital technologies for HD map construction, 3D lane detection is still an open problem due to varying visual conditions, complex typologies, and strict demands for precision. In this paper, an end-to-end flexible and hierarchical lane detector is proposed to precisely predict 3D lane lines from point clouds. Specifically, we design a hierarchical network predicting flexible representations of lane shapes at different levels, simultaneously collecting global instance semantics and avoiding local errors. In the global scope, we propose to regress parametric curves w.r.t adaptive axes that help to make more robust predictions towards complex scenes, while in the local vision the structure of lane segment is detected in each of the dynamic anchor cells sampled along the global predicted curves. Moreover, corresponding global and local shape matching losses and anchor cell generation strategies are designed. Experiments on two datasets show that we overwhelm current top methods under high precision standards, and full ablation studies also verify each part of our method. Our codes will be released at https://github.com/Doo-do/FHLD.
Zhihao Guan, Ruixin Liu, Zejian Yuan, Ao Liu 0010, Erlong Li, Chao Zheng 0004, Shuqi Mei
AAAI3
2023 Learning to Generate an Unbiased Scene Graph by Using Attribute-Guided Predicate Features
abstract
Scene Graph Generation (SGG) aims to capture the semantic information in an image and build a structured representation, which facilitates downstream tasks. The current challenge in SGG is to tackle the biased predictions caused by the long-tailed distribution of predicates. Since multiple predicates in SGG are coupled in an image, existing data re-balancing methods cannot completely balance the head and tail predicates. In this work, a decoupled learning framework is proposed for unbiased scene graph generation by using attribute-guided predicate features to construct a balanced training set. Specifically, the predicate recognition is decoupled into Predicate Feature Representation Learning (PFRL) and predicate classifier training with a class-balanced predicate feature set, which is constructed by our proposed Attribute-guided Predicate Feature Generation (A-PFG) model. In the A-PFG model, we first define the class labels of and corresponding visual feature as attributes to describe a predicate. Then the predicate feature and the attribute embedding are mapped into a shared hidden space by a dual Variational Auto-encoder (VAE), and finally the synthetic predicate features are forced to learn the contextual information in the attributes via cross reconstruction and distribution alignment. To demonstrate the effectiveness of our proposed method, our decoupled learning framework and A-PFG model are applied to various SGG models. The empirical results show that our method is substantially improved on all benchmarks and achieves new state-of-the-art performance for unbiased scene graph generation. Our code is available at https://github.com/wanglei0618/A-PFG.
Lei Wang 0219, Zejian Yuan, Badong Chen
AAAI2
2023 PBFormer: Capturing Complex Scene Text Shape with Polynomial Band Transformer
abstract
We present PBFormer, an efficient yet powerful scene text detector that unifies the transformer with a novel text shape representation Polynomial Band (PB). The representation has four polynomial curves to fit a text's top, bottom, left, and right sides, which can capture a text with a complex shape by varying polynomial coefficients. PB has appealing features compared with conventional representations: 1) It can model different curvatures with a fixed number of parameters, while polygon-points-based methods need to utilize a different number of points. 2) It can distinguish adjacent or overlapping texts as they have apparent different curve coefficients, while segmentation-based or points-based methods suffer from adhesive spatial positions. PBFormer combines the PB with the transformer, which can directly generate smooth text contours sampled from predicted curves without interpolation. A parameter-free cross-scale pixel attention (CPA) module is employed to highlight the feature map of a suitable scale while suppressing the other feature maps. The simple operation can help detect small-scale texts and is compatible with the one-stage DETR framework, where no postprocessing exists for NMS. Furthermore, PBFormer is trained with a shape-contained loss, which not only enforces the piecewise alignment between the ground truth and the predicted curves but also makes curves' position and shapes consistent with each other. Without bells and whistles about text pre-training, our method is superior to the previous state-of-the-art text detectors on the arbitrary-shaped text datasets. Codes will be public.
Ruijin Liu, Ning Lu 0003, Dapeng Chen, Cheng Li 0040, Zejian Yuan, Wei Peng 0011
ACM Multimedia5
2023 Learning to Detect 3D Lanes by Shape Matching and Embedding
abstract
3D lane detection based on LiDAR point clouds is a challenging task that requires precise locations, accurate topologies, and distinguishable instances. In this paper, we propose a dual-level shape attention network (DSANet) with two branches for high-precision 3D lane predictions. Specifically, one branch predicts the refined lane segment shapes and the shape embeddings that encode the approximate lane instance shapes, the other branch detects the coarse-grained structures of the lane instances. In the training stage, two-level shape matching loss functions are introduced to jointly optimize the shape parameters of the twobranch outputs, which are simple yet effective for precision enhancement. Furthermore, a shape-guided segments aggregator is proposed to help local lane segments aggregate into complete lane instances, according to the differences of instance shapes predicted at different levels. Experiments conducted on our BEV-3DLanes dataset demonstrate that our method outperforms previous methods.
Ruixin Liu, Zhihao Guan, Zejian Yuan, Ao Liu 0010, Tang Kun, Erlong Li, Chao Zheng 0004, Shuqi Mei
WACV3
2022 Learning to Predict 3D Lane Shape and Camera Pose from a Single Image via Geometry Constraints
abstract
Detecting 3D lanes from the camera is a rising problem for autonomous vehicles. In this task, the correct camera pose is the key to generating accurate lanes, which can transform an image from perspective-view to the top-view. With this transformation, we can get rid of the perspective effects so that 3D lanes would look similar and can accurately be fitted by low-order polynomials. However, mainstream 3D lane detectors rely on perfect camera poses provided by other sensors, which is expensive and encounters multi-sensor calibration issues. To overcome this problem, we propose to predict 3D lanes by estimating camera pose from a single image with a two-stage framework. The first stage aims at the camera pose task from perspective-view images. To improve pose estimation, we introduce an auxiliary 3D lane task and geometry constraints to benefit from multi-task learning, which enhances consistencies between 3D and 2D, as well as compatibility in the above two tasks. The second stage targets the 3D lane task. It uses previously estimated pose to generate top-view images containing distance-invariant lane appearances for predicting accurate 3D lanes. Experiments demonstrate that, without ground truth camera pose, our method outperforms the state-of-the-art perfect-camera-pose-based methods and has the fewest parameters and computations. Codes are available at https://github.com/liuruijin17/CLGo.
Ruijin Liu, Dapeng Chen, Zhiliang Xiong, Zejian Yuan
AAAI5
2022 HAPNet: a head-aware pedestrian detection network associated with the affinity field
Jiali Ding, Zejian Yuan
Sci. China Inf. Sci.4
2022 Multikernel Correntropy for Robust Learning
abstract
As a novel similarity measure that is defined as the expectation of a kernel function between two random variables, correntropy has been successfully applied in robust machine learning and signal processing to combat large outliers. The kernel function in correntropy is usually a zero-mean Gaussian kernel. In a recent work, the concept of mixture correntropy (MC) was proposed to improve the learning performance, where the kernel function is a mixture Gaussian kernel, namely, a linear combination of several zero-mean Gaussian kernels with different widths. In both correntropy and MC, the center of the kernel function is, however, always located at zero. In the present work, to further improve the learning performance, we propose the concept of multikernel correntropy (MKC), in which each component of the mixture Gaussian kernel can be centered at a different location. The properties of the MKC are investigated and an efficient approach is proposed to determine the free parameters in MKC. Experimental results show that the learning algorithms under the maximum MKC criterion (MMKCC) can outperform those under the original maximum correntropy criterion (MCC) and the maximum MC criterion (MMCC).
Badong Chen, Yuqing Xie 0002, Zejian Yuan, Pengju Ren, Harry Qin
IEEE Trans. Cybern.4
2022 Unsupervised Occlusion-Aware Stereo Matching With Directed Disparity Smoothing
abstract
When handling occlusion in unsupervised stereo matching, existing methods tend to neglect the supportive role of occlusion and to perform inappropriate disparity smoothing around the occlusion. To address these problems, we propose an occlusion-aware stereo network that contains a specific module to first estimate occlusion as an additional depth cue. In the occlusion inference module, a pixel is classified with a three-category label based on whether an area is occluded by an object on the left, occluded by an object on the right, or unoccluded. After the occluders are detected, we introduce a directed disparity smoothing loss that allows valid disparity estimates to be propagated to fill the occluded region, while ambiguous matches in the occluded region do not affect other regions. Disparity and occlusion are trained alternately in an unsupervised manner with detached backpropagation to enable the directed smoothness. Experiments show that our method achieves 3-pixel threshold error rates of 6.51% and 5.69% on the KITTI 2015 and KITTI 2012 validation sets, state-of-the-art results among unsupervised learning networks at the time of submission.
Ang Li 0025, Zejian Yuan, Yonggen Ling, Wanchao Chi, Shenghao Zhang 0001
IEEE Trans. Intell. Transp. Syst.2
2022 Learning TBox With a Cascaded Anchor-Free Network for Vehicle Detection
abstract
Vehicle detection, the process of identifying vehicles as axis-aligned bounding boxes in still images, is widely used to estimate the range, time-to-collision, and motion of autonomous vehicles (AVs). Bounding boxes, while convenient, are too coarse to adapt well to vehicle shape and pose variations. In this work, we presentTBox(Trapezoid & Box), a novel fine-grained representation useful for both localization and recognition that extends the bounding box by restricting the spatial extent of a vehicle to a set of keypoints and indicating semantically significant local areas using subclasses. In contrast to the previous monolithic models, we propose a cascaded anchor-free architecture to estimate the bounding box and TBox. One subnetwork uses a stacked hourglass network to detect each vehicle as a pair of corners without using anchors. Specifically, it learns corner affinity fields, enabling it to perform robust corner grouping. The other subnetwork estimates a TBox as a set of keypoints. This subnetwork utilizes the bounding box results to avoid ambiguous keypoint associations and reuses existing features to reduce the number of parameters. We also propose a multitask learning strategy for training the cascaded model that implicitly integrates the global context with local details, introducing improvements for both tasks. During testing, a refinement algorithm explicitly uses robust local keypoints to correct possible global box errors, ensuring tight geometric representations for nearby critical vehicles. The experiments show that our method outperforms existing anchor-free detectors for vehicle detection and achieves better performance on the TBox task while using a small model.
Ruijin Liu, Zejian Yuan
IEEE Trans. Intell. Transp. Syst.2
2021 Order-independent Matching with Shape Similarity for Parking Slot Detection
Ziyi Yin 0001, Ruijin Liu, Zejian Yuan, Zhiliang Xiong
BMVC3
2021 Surround-view Free Space Boundary Detection with Polar Representation
Zidong Cao, Zhiliang Xiong, Zejian Yuan
BMVC4
2021 Self-Supervised Depth Completion via Adaptive Sampling and Relative Consistency
abstract
Depth sensing is crucial for many computer vision applications. Commodity-level RGB-D cameras are often unable to sense depth in distant, reflective and transparent regions, resulting in large missing areas. As the acquisition of depth annotations in missing areas is tedious, we propose a self-supervised method for the task of completing depth values of missing areas. Specifically, we sample the incomplete raw depth map via an adaptive sampling strategy to generate a more incomplete depth map as the input and use the raw depth map as the training label. To enable the network to propagate long-range depth information to fill large invalid areas, we further propose a relative consistency loss during training. Experiments validate the effectiveness of our self-supervised method, which outperforms previous unsupervised methods and even can compete with some supervised methods. Our code is available at https://github.com/caozidong/DepthCompletion.
Zidong Cao, Ang Li 0025, Zejian Yuan
ICIP3
2021 Defect Inspection using Gravitation Loss and Soft Labels
abstract
Defect inspection is a widely studied computer vision task with a vast range of applications. However, due to the great variety of defects and high difficulty of collecting ample abnormal images of rare occasions, such a task full of confusing noisy samples faces tremendous challenges of generalization. In this paper, we propose a practical framework for defect inspection to better discover and utilize the connections among these samples. Specifically, Gravitation Loss is proposed to enhance the discriminative power of embedding vectors learned by neural networks. Besides, Soft Loss is designed by introducing soft labels which provide more supervision for noisy samples to improve generalization. With joint supervision of Gravitation Loss, Soft Loss, and a regular classification loss, experimental results show that our method outperforms other state-of-the-art approaches on a newly-built resin plug-hole dataset.
Zhihao Guan, Zidong Guo, Jie Lyu 0001, Zejian Yuan
ICIP4
2021 Multimodal Transformer Networks for Pedestrian Trajectory Prediction
abstract
We consider the problem of forecasting the future locations of pedestrians in an ego-centric view of a moving vehicle. Current CNNs or RNNs are flawed in capturing the high dynamics of motion between pedestrians and the ego-vehicle, and suffer from the massive parameter usages due to the inefficiency of learning long-term temporal dependencies. To address these issues, we propose an efficient multimodal transformer network that aggregates the trajectory and ego-vehicle speed variations at a coarse granularity and interacts with the optical flow in a fine-grained level to fill the vacancy of highly dynamic motion. Specifically, a coarse-grained fusion stage fuses the information between trajectory and ego-vehicle speed modalities to capture the general temporal consistency. Meanwhile, a fine-grained fusion stage merges the optical flow in the center area and pedestrian area, which compensates the highly dynamic motion of ego-vehicle and target pedestrian. Besides, the whole network is only attention-based that can efficiently model long-term sequences for better capturing the temporal variations. Our multimodal transformer is validated on the PIE and JAAD datasets and achieves state-of-the-art performance with the most light-weight model size. The codes are available at https://github.com/ericyinyzy/MTN_trajectory.
Ziyi Yin 0001, Ruijin Liu, Zhiliang Xiong, Zejian Yuan
IJCAI4
2021 End-to-end Lane Shape Prediction with Transformers
abstract
Lane detection, the process of identifying lane markings as approximated curves, is widely used for lane departure warning and adaptive cruise control in autonomous vehicles. The popular pipeline that solves it in two steps- feature extraction plus post-processing, while useful, is too inefficient and flawed in learning the global context and lanes' long and thin structures. To tackle these issues, we propose an end-to-end method that directly outputs parameters of a lane shape model, using a network built with a transformer to learn richer structures and context. The lane shape model is formulated based on road structures and camera pose, providing physical interpretation for parameters of network output. The transformer models non-local interactions with a self-attention mechanism to capture slender structures and global context. The proposed method is validated on the TuSimple benchmark and shows state-of-the-art accuracy with the most lightweight model size and fastest speed. Additionally, our method shows excellent adaptability to a challenging self-collected lane detection dataset, showing its powerful deployment potential in real applications. Codes are available at https://github.com/liuruijin17/LSTR.
Ruijin Liu, Zejian Yuan, Zhiliang Xiong
WACV2
2020 Domain Adaptation Gaze Estimation by Embedding with Prediction Consistency
Zidong Guo, Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001
ACCV (5)2
2020 Learning End-to-End Action Interaction by Paired-Embedding Data Augmentation
Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001
ACCV (6)2
2020 Visual Saliency Oriented Vehicle Scale Estimation
abstract
Vehicle scale estimation with a single camera is a typical application for intelligent transportation and it faces the challenges from visual computing while intensity-based method and descriptor-based method should be balanced. This paper proposed a vehicle scale estimation method based on salient object detection to resolve this problem. The regularized intensity matching method is proposed in Lie Algebra to achieve robust and accurate scale estimation, and descriptor matching and intensity matching are combined to minimize the proposed loss function. The visual attention mechanism is designed to select image patches with texture and remove the occluded image patches. Then the weights are assigned to pixels from the selected image patches which alleviates the influence of noise-corrupted pixels. The experiments show that the proposed method significantly outperforms state-of-the-art methods with regard to the robustness and accuracy of vehicle scale estimation.
Jiali Ding, Qixin Chen, Zejian Yuan
ICPR4
2020 FastCompletion: A Cascade Network with Multiscale Group-Fused Inputs for Real-Time Depth Completion
abstract
Completing sparse data captured with commercial depth sensors is a vital and fundamental procedure for many computer vision applications. For execution in real-world scenarios, a good trade-off between accuracy and speed is increasingly in demand for depth completion methods. Most previous methods achieve satisfactory accuracy on standard benchmarks. However, they extensively rely on heavy models to handle diverse structures and require additional run time on multimodal data. In this paper, we present an efficient method of depth completion. We propose a grouped fusion strategy for efficiently extracting depth and guidance features in parallel and fusing them naturally in the feature spaces to achieve high performance. Instead of a monolithic architecture, we employ cascaded hourglass networks, each of which is specialized for certain structures and has a lightweight architecture. Given the sparsity of the depth maps, we downsample the inputs to multiple scales to further accelerate the computation. Our model runs at over 39 FPS on an embedded GPU with high-resolution inputs. Evaluations on the KITTI benchmark demonstrate that the proposed model is an ideal approach for real-world applications.
Ang Li 0025, Zejian Yuan, Yonggen Ling, Wanchao Chi, Shenghao Zhang 0001
ICPR2
2020 Attention-Oriented Action Recognition for Real- Time Human-Robot Interaction
abstract
Despite the notable progress made in action recognition tasks, not much work has been done in action recognition specifically for human-robot interaction. In this paper, we deeply explore the characteristics of the action recognition task in interaction scenes and propose an attention-oriented multi-level network framework to meet the need for real-time interaction. Specifically, a Pre-Attention network is employed to roughly focus on the interactor in the scene at low resolution firstly and then perform fine-grained pose estimation at high resolution. The other compact CNN receives the extracted skeleton sequence as input for action recognition, utilizing attention-like mechanisms to capture local spatial-temporal patterns and global semantic information effectively. To evaluate our approach, we construct a new action dataset specially for the recognition task in interaction scenes. Experimental results on our dataset and high efficiency (112 fps at 640 × 480 RGBD) on the mobile computing platform (Nvidia Jetson AGX Xavier) demonstrate excellent applicability of our method on action recognition in real-time human-robot interaction.
Ziyi Yin 0001, Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001
ICPR3
2020 A Multi-Scale Guided Cascade Hourglass Network for Depth Completion
abstract
Depth completion, a task to estimate the dense depth map from sparse measurement under the guidance from the high-resolution image, is essential to many computer vision applications. Most previous methods building on fully convolutional networks can not handle diverse patterns in the depth map efficiently and effectively. We propose a multi-scale guided cascade hourglass network to tackle this problem. Structures at different levels are captured by specialized hourglasses in the cascade network with sparse inputs in various sizes. An encoder extracts multi-scale features from color image to provide deep guidance for all the hourglasses. A multi-scale training strategy further activates the effect of cascade stages. With the role of each sub-module divided explicitly, we can implement components with simple architectures. Extensive experiments show that our lightweight model achieves competitive results compared with state-of-the-art in KITTI depth completion benchmark, with low complexity in run-time.
Ang Li 0025, Zejian Yuan, Yonggen Ling, Wanchao Chi, Shenghao Zhang 0001
WACV2
2020 Crowded Human Detection via an Anchor-pair Network
abstract
This paper presents an anchor-pair network for crowded human detection, which can overcome and solve the difficulties caused by occlusion in crowded scenes. Specifically, we use a function-aware network structure to extract more distinctive and discriminative features for head and full-body respectively, and then a CNN module is also exploited to fuse the features by learning the correlations between head and full-body to reduce crowd errors. Meanwhile, a novel paired form for anchors, denoted as anchor-pair, is proposed to estimate the head regions and full-body regions simultaneously. Furthermore, a new ingenious Joint-NMS is introduced to perform on the detected head and full-body box pairs, which produces significant performance improvement in heavily occluded scenarios at tiny computational cost. Our anchor-pair network achieves a state-of-the-art result on the CrowdHuman dataset which reduces the MR−2to 55.43%, achieving 11.59% relative improvement over our dataset baseline.
Jinguo Zhu, Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001
WACV2
2020 Accurate Pedestrian Detection by Human Pose Regression
abstract
Pedestrian detection with high detection and localization accuracy is increasingly important for many practical applications. Due to the flexible structure of the human body, it is hard to train a template-based pedestrian detector that achieves a high detection rate and a good localization accuracy simultaneously. In this paper, we utilize human pose estimation to improve the detection and localization accuracy of pedestrian detection. We design two kinds of pose-indexed features that can considerably improve the discriminability of the detector. In addition to employing a two-stage pipeline to carry out these two tasks, we unify pose estimation and pedestrian detection into a cascaded decision forest in which they can cooperate sufficiently. To prevent irregular positive examples, such as truncated ones, from distracting the pedestrian detection and the pose regression, we clean the positive training data by realigning the bounding boxes and rejecting the wrong positive samples. Experimental results on the Caltech test dataset demonstrate the effectiveness of our proposed method. Our detector achieves 11.1% MR-2, outperforming all existing detectors without using the convolutional neural network (CNN). Moreover, our method can be assembled with other detectors based on CNNs to improve detection and localization performance. By collaborating with the recent CNN-based method, our detector achieves 5.5% MR-2 on the Caltech test dataset, outperforming the state-of-the-art methods.
Zejian Yuan, Badong Chen
IEEE Trans. Image Process.2
2020 Training Cascade Compact CNN With Region-IoU for Accurate Pedestrian Detection
abstract
Recently, pedestrian detection has made significant advances benefiting from the region-based convolutional neural networks (R-CNN). However, training R-CNN with a holistic intersection over union (IoU) always brings many flawed positive samples. This paper introduces a strict matching metric, which is beneficial to selecting well-aligned positive samples. Specifically, this matching metric is defined on a set of region-IoUs instead of a holistic IoU, which considers the alignments of different part regions in a whole bounding box simultaneously. A positive sample matches a ground truth only if all its region-IoUs are bigger than a threshold. Secondly, an improved negative example selection strategy using both the classification and localization information is proposed to mine hard negative examples, which can further suppress the false positive detections near the pedestrians. Based on the proposed sample selection strategy, a cascade compact convolutional neural network (CC-CNN) is proposed for accurate pedestrian detection. Each stage of the CC-CNN is constructed with a compact network that only consists of a small number of parameters, thus making the detector suitable to be implemented on onboard embedded systems. Experimental results on two widely used pedestrian datasets demonstrate that the proposed training strategy and the CC-CNN based detector can effectively improve the detection rate and the localization accuracy using fewer parameters.
Zejian Yuan, Badong Chen
IEEE Trans. Intell. Transp. Syst.2
2019 Beyond Bounding Box: Fine-Grained Vehicle Detection via Single Stage Detector with Hierarchical output
abstract
Vehicle detection is a crucial module of the camera-based forward collision alert, which is usually used to calculate range and time-to-collision. Conventional methods based on bounding box are too coarse to handle challenging situations such as vehicle pose variations. In this paper, we propose a novel vehicle detector with a fine-grained output representation. The detector exploits a hierarchical tree-like output representation by introducing two subclasses and virtual control points, which not only discriminates each face of a vehicle but also locates their boundaries accurately. Our detector adopts popular single stage multi-scale CNN framework, which is equipped with the hierarchical output, and is beyond the bounding box methods. Experiments on our large-scale self-collected dataset show that our method achieves satisfactory performance.
Ruijin Liu, Fangying Luo, Zejian Yuan
ICIP3
2019 Learning to Plan Semantic Free-Space Boundary
abstract
Recently, free-space detection has attracted widespread attention. Most existing methods treat free-space detection as a semantic segmentation task. In this paper, we propose a novel approach to directly infer the boundary of the semantic free-space from a single image. Firstly, we design a multistage CNN to produce 2D belief maps with high resolution for boundary segments of different semantic classes, such as road boundary, vertical obstacles on road and so on. The proposed CNN architecture can implicitly learn boundary structure and long-range spatial context. Then, based on the 2D belief maps we address the semantic free-space detection as a dynamic programming problem to ensure the spatial smoothness of the predicted boundary. The experimental results on our dataset show that our method has a convincing performance on various quantitative metrics.
Ziyi Yin 0001, Zejian Yuan
ICIP3
2019 Maximum Correntropy Criterion-Based Sparse Subspace Learning for Unsupervised Feature Selection
abstract
High-dimensional data contain not only redundancy but also noises produced by the sensors. These noises are usually non-Gaussian distributed. The metrics based on Euclidean distance are not suitable for these situations in general. In order to select the useful features and combat the adverse effects of the noises simultaneously, a robust sparse subspace learning method in unsupervised scenario is proposed in this paper based on the maximum correntropy criterion that shows strong robustness against outliers. Furthermore, an iterative strategy based on half quadratic and an accelerated block coordinate update is proposed. The convergence analysis of the proposed method is also carried out to ensure the convergence to a reliable solution. Extensive experiments are conducted on real-world data sets to show that the new method can filter out the outliers and outperform several state-of-the-art unsupervised feature selection methods.
Nan Zhou 0010, Yangyang Xu 0005, Hong Cheng 0002, Zejian Yuan, Badong Chen
IEEE Trans. Circuits Syst. Video Technol.4
2018 Occlusion Aware Stereo Matching via Cooperative Unsupervised Learning
Ang Li 0025, Zejian Yuan
ACCV (6)2
2018 Automatic Graphics Program Generation Using Attention-Based Hierarchical Decoder
Zhihao Zhu 0001, Zhan Xue, Zejian Yuan
ACCV (6)3
2018 SymmNet: A Symmetric Convolutional Neural Network for Occlusion Detection
Zejian Yuan
BMVC2
2018 Joint Holistic and Partial CNN for Pedestrian Detection
Zejian Yuan
BMVC2
2018 Think and Tell: Preview Network for Image Captioning
Zhihao Zhu 0001, Zhan Xue, Zejian Yuan
BMVC3
2018 Improving Deep Visual Representation for Person Re-identification by Global and Local Image-language Association
Dapeng Chen, Hongsheng Li 0001, Xihui Liu, Yantao Shen 0002, Zejian Yuan, Xiaogang Wang 0001
ECCV (16)6
2018 Topic-Guided Attention for Image Captioning
abstract
Attention mechanisms have attracted considerable interest in image captioning because of its powerful performance. Existing attention-based models use feedback information from the caption generator as guidance to determine which of the image features should be attended to. A common defect of these attention generation methods is that they lack a higher-level guiding information from the image itself, which sets a limit on selecting the most informative image features. Therefore, in this paper, we propose a novel attention mechanism, called topic-guided attention, which integrates image topics in the attention model as a guiding information to help select the most important image features. Moreover, we extract image features and image topics with separate networks, which can be fine-tuned jointly in an end-to-end manner during training. The experimental results on the benchmark Microsoft COCO dataset show that our method yields state-of-art performance on various quantitative metrics.
Zhihao Zhu 0001, Zhan Xue, Zejian Yuan
ICIP3
2018 Learning Fixation Point Strategy for Object Detection and Classification
abstract
We propose a novel recurrent attentional structure to localize and recognize objects jointly. The network can learn to extract a sequence of local observations with detailed appearance and rough context, instead of sliding windows or convolutions on the entire image. Meanwhile, those observations are fused to complete detection and classification tasks. On training, we present a hybrid loss function to learn the parameters of the multi-task network end-to-end. Particularly, the combination of stochastic and object-awareness strategy, named SA, can select more abundant context and ensure the last fixation close to the object. In addition, we build a real-world dataset to verify the capacity of our method in detecting the object of interest including those small ones. Our method can predict a precise bounding box on an image, and achieve high speed on large images. Experimental results indicate that the proposed method can mine effective context by several local observations. Moreover, the precision and speed are easily improved by changing the number of recurrent steps. Source code is available at https://github.com/jielyu/RADCN.
Jie Lyu 0001, Zejian Yuan, Dapeng Chen
ICPR2
2018 Automatic salient object sequence rebuilding for video segment analysis
Haibin Duan, Zejian Yuan, Nanning Zheng 0001
Sci. China Inf. Sci.4
2017 Fast Pedestrian Detection via Random Projection Features with Shape Prior
abstract
Accurate pedestrian detection with high speed is always of great interests especially for practical application. Detectors usually follow the feature selection paradigm, and need to first construct rich and diverse features. In particular, current state-of-the-arts generate more channels of feature by convolving the basic feature channels with filter banks, which significantly improves accuracy. In this paper, we propose to apply random projection over the basic feature channels, implicitly selecting feature from a much larger feature space. Our method is more efficient than the ones employing filter banks by avoiding the convolution operation. We further impose shape prior to guide the random projection, making the generated feature be more robust to occlusion, pose variation and scale change. Experimental results on Caltech pedestrian dataset demonstrate the accuracy and efficiency of our method. Compared with thestate-of-arts, our method can achieve 5-10× speedup with comparable accuracy.
Zejian Yuan, Dapeng Chen, Jie Lyu 0001
WACV2
2017 Exemplar-Guided Similarity Learning on Polynomial Kernel Feature Map for Person Re-identification
Dapeng Chen, Zejian Yuan, Jingdong Wang 0001, Badong Chen, Gang Hua 0001, Nanning Zheng 0001
Int. J. Comput. Vis.2
2017 Salient Object Detection: A Discriminative Regional Feature Integration Approach
Jingdong Wang 0001, Huaizu Jiang, Zejian Yuan, Ming-Ming Cheng, Xiaowei Hu 0003, Nanning Zheng 0001
Int. J. Comput. Vis.3
2017 Multi-Timescale Collaborative Tracking
abstract
We present the multi-timescale collaborative tracker for single object tracking. The tracker simultaneously utilizes different types of "forces", namely attraction, repulsion and support, to take advantage of their complementary strengths. We model the three forces via three components that are learned from the sample sets with different timescales. The long-term descriptive component attracts the target sample, while the medium-term discriminative component repulses the target from the background. They are collaborated in the appearance model to benefit each other. The short-term regressive component combines the votes of the auxiliary samples to predict the target's position, forming the context-aware motion model. The appearance model and the motion model collaboratively determine the target state, and the optimal state is estimated by a novel coarse-to-fine search strategy. We have conducted an extensive set of experiments on the standard 50 video benchmark. The results confirm the effectiveness of each component and their collaboration, outperforming current state-of-the-art methods.
Dapeng Chen, Zejian Yuan, Gang Hua 0001, Jingdong Wang 0001, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 Similarity Learning with Spatial Constraints for Person Re-identification
abstract
Pose variation remains one of the major factors that adversely affect the accuracy of person re-identification. Such variation is not arbitrary as body parts (e.g. head, torso, legs) have relative stable spatial distribution. Breaking down the variability of global appearance regarding the spatial distribution potentially benefits the person matching. We therefore learn a novel similarity function, which consists of multiple sub-similarity measurements with each taking in charge of a subregion. In particular, we take advantage of the recently proposed polynomial feature map to describe the matching within each subregion, and inject all the feature maps into a unified framework. The framework not only outputs similarity measurements for different regions, but also makes a better consistency among them. Our framework can collaborate local similarities as well as global similarity to exploit their complementary strength. It is flexible to incorporate multiple visual cues to further elevate the performance. In experiments, we analyze the effectiveness of the major components. The results on four datasets show significant and consistent improvements over the state-of-the-art methods.
Dapeng Chen, Zejian Yuan, Badong Chen, Nanning Zheng 0001
CVPR2
2016 Coordinating Multiple Disparity Proposals for Stereo Computation
abstract
While great progress has been made in stereo computation over the last decades, large textureless regions remain challenging. Segment-based methods can tackle this problem properly, but their performances are sensitive to the segmentation results. In this paper, we alleviate the sensitivity by generating multiple proposals on absolute and relative disparities from multi-segmentations. These proposals supply rich descriptions of surface structures. Especially, the relative disparity between distant pixels can encode the large structure, which is critical to handle the large textureless regions. The proposals are coordinated by point-wise competition and pairwise collaboration within a MRF model. During inference, a dynamic programming is performed in different directions with various step sizes, so the long-range connections are better preserved. In the experiments, we carefully analyzed the effectiveness of the major components. Results on the 2014 Middlebury and KITTI 2015 stereo benchmark show that our method is comparable to state-of-the-art.
Dapeng Chen, Yuanliu Liu, Zejian Yuan
CVPR4
2015 Similarity learning on an explicit polynomial kernel feature map for person re-identification
abstract
In this paper, we address the person re-identification problem, discovering the correct matches for a probe person image from a set of gallery person images. We follow the learning-to-rank methodology and learn a similarity function to maximize the difference between the similarity scores of matched and unmatched images for a same person. We introduce at least three contributions to person re-identification. First, we present an explicit polynomial kernel feature map, which is capable of characterizing the similarity information of all pairs of patches between two images, called soft-patch-matching, instead of greedily keeping only the best matched patch, and thus more robust. Second, we introduce a mixture of linear similarity functions that is able to discover different soft-patch-matching patterns. Last, we introduce a negative semi-definite regularization over a subset of the weights in the similarity function, which is motivated by the connection between explicit polynomial kernel feature map and the Mahalanobis distance, as well as the sparsity constraint over the parameters to avoid over-fitting. Experimental results over three public benchmarks demonstrate the superiority of our approach.
Dapeng Chen, Zejian Yuan, Gang Hua 0001, Nanning Zheng 0001, Jingdong Wang 0001
CVPR2
2015 Saturation-preserving specular reflection separation
abstract
Specular reflection generally decreases the saturation of surface colors, which will be possibly confused with other colors that have the same hue but lower saturation. Traditional methods for specular reflection separation suffer this problem of hue-saturation ambiguity, producing over-saturated specular-free images quite often. We proposed a two-step approach to solve this problem. In the first step, we produce an over-saturated specular-free image by global chromaticity propagation from specular-free pixels to highlighted ones. Then we recover the saturation based on priors of the piecewise constancy of diffuse chromaticity as well as the spatial sparsity and smoothness of specular reflection. We achieve this through increasing the achromatic component of diffuse chromaticity, while the magnitudes of increments are determined by linear programming under the constraints derived from the priors. Experiments on both laboratory and natural images show that our method can separate the specular reflection while preserving the saturation of the underlying surface colors.
Yuanliu Liu, Zejian Yuan, Nanning Zheng 0001, Yang Wu 0001
CVPR2
2015 Illumination Robust Color Naming via Label Propagation
abstract
Color composition is an important property for many computer vision tasks like image retrieval and object classification. In this paper we address the problem of inferring the color composition of the intrinsic reflectance of objects, where the shadows and highlights may change the observed color dramatically. We achieve this through color label propagation without recovering the intrinsic reflectance beforehand. Specifically, the color labels are propagated between regions sharing the same reflectance, and the direction of propagation is promoted to be from regions under full illumination and normal view angles to abnormal regions. We detect shadowed and highlighted regions as well as pairs of regions that have similar reflectance. A joint inference process is adopted to trim the inconsistent identities and connections. For evaluation we collect three datasets of images under noticeable highlights and shadows. Experimental results show that our model can effectively describe the color composition of real-world images.
Yuanliu Liu, Zejian Yuan, Badong Chen, Jianru Xue, Nanning Zheng 0001
ICCV2
2015 Video object segmentation by integrating trajectories from points and regions
Zejian Yuan, Yuehu Liu, Nanning Zheng 0001
Multim. Tools Appl.2
2014 Intrinsic Image Decomposition from Pair-Wise Shading Ordering
Yuanliu Liu, Zejian Yuan, Nanning Zheng 0001
ACCV (5)2
2014 Description-Discrimination Collaborative Tracking
Dapeng Chen, Zejian Yuan, Gang Hua 0001, Yang Wu 0001, Nanning Zheng 0001
ECCV (1)2
2014 Online Nonlinear Granger Causality Detection by Quantized Kernel Least Mean Square
Badong Chen, Zejian Yuan, Nanning Zheng 0001, Andreas Keil, José C. Príncipe
ICONIP (2)3
2014 Robust registration of partially overlapping point sets via genetic algorithm with growth operator
abstract
Recently, genetic algorithm (GA) has been introduced as an effective method to solve the registration problem. It maintains a population of candidate solutions for the problem and evolves by iteratively applying a set of stochastic operators. Accordingly, a key question is how to reduce the population size. In this study, the authors present two techniques for reducing the population size in the GA for registration of partially overlapping point sets. Based on the trimmed iterative closest point algorithm, they introduce a growth operator into the GA. The growth operator, which is also inspired by the biological evolution, can improve the GA efficiency for registration. Furthermore, they present a technique called centre alignment to confirm the value range of all the registration parameters, which can reduce the search space and allow the well‐designed GA to directly solve the registration problem. Experimental results carried out with the m ‐dimensional point sets illustrate its advantages over previous approaches.
Jihua Zhu, Deyu Meng, Zhongyu Li 0002, Shaoyi Du, Zejian Yuan
IET Image Process.5
2013 Salient Object Detection: A Discriminative Regional Feature Integration Approach
abstract
Salient object detection has been attracting a lot of interest, and recently various heuristic computational models have been designed. In this paper, we regard saliency map computation as a regression problem. Our method, which is based on multi-level image segmentation, uses the supervised learning approach to map the regional feature vector to a saliency score, and finally fuses the saliency scores across multiple levels, yielding the saliency map. The contributions lie in two-fold. One is that we show our approach, which integrates the regional contrast, regional property and regional background ness descriptors together to form the master saliency map, is able to produce superior saliency maps to existing algorithms most of which combine saliency maps heuristically computed from different types of features. The other is that we introduce a new regional feature vector, background ness, to characterize the background, which can be regarded as a counterpart of the objectness descriptor [2]. The performance evaluation on several popular benchmark data sets validates that our approach outperforms existing state-of-the-arts.
Huaizu Jiang, Jingdong Wang 0001, Zejian Yuan, Yang Wu 0001, Nanning Zheng 0001, Shipeng Li 0001
CVPR3
2013 Constructing Adaptive Complex Cells for Robust Visual Tracking
abstract
Representation is a fundamental problem in object tracking. Conventional methods track the target by describing its local or global appearance. In this paper we present that, besides the two paradigms, the composition of local region histograms can also provide diverse and important object cues. We use cells to extract local appearance, and construct complex cells to integrate the information from cells. With different spatial arrangements of cells, complex cells can explore various contextual information at multiple scales, which is important to improve the tracking performance. We also develop a novel template-matching algorithm for object tracking, where the template is composed of temporal varying cells and has two layers to capture the target and background appearance respectively. An adaptive weight is associated with each complex cell to cope with occlusion as well as appearance variation. A fusion weight is associated with each complex cell type to preserve the global distinctiveness. Our algorithm is evaluated on 25 challenging sequences, and the results not only confirm the contribution of each component in our tracking system, but also outperform other competing trackers.
Dapeng Chen, Zejian Yuan, Yang Wu 0001, Nanning Zheng 0001
ICCV2
2013 Probabilistic salient object contour detection based on superpixels
abstract
In this paper, we propose a data-driven approach to detect the probabilistic salient object contour, which is formulated as predicting the probability of superpixel boundaries being on the object contour based on the learned regressor. Each superpixel boundary is jointly described by the superpixel saliency, superpixel contrast, and boundary geometry features. Experimental results on the benchmark data set validate the effectiveness of our approach. Furthermore, we demonstrate that the predicted probabilistic salient object contour is useful for improving the multiple segmentations for salient object detection.
Huaizu Jiang, Yang Wu 0001, Zejian Yuan
ICIP3
2013 Kernel minimum error entropy algorithm
Badong Chen, Zejian Yuan, Nanning Zheng 0001, José C. Príncipe
Neurocomputing2
2012 Dense Scene Flow Based on Depth and Multi-channel Bilateral Filter
Dapeng Chen, Zejian Yuan, Nanning Zheng 0001
ACCV (3)3
2012 Detecting occlusion boundaries via saliency network
Dapeng Chen, Zejian Yuan, Nanning Zheng 0001
ICPR2
2012 Learning to describe color composition of visual objects
Yuanliu Liu, Yudong Liang, Zejian Yuan, Nanning Zheng 0001
ICPR3
2012 Video object segmentation by clustering region trajectories
Zejian Yuan, Dapeng Chen, Yuehu Liu, Nanning Zheng 0001
ICPR2
2012 Fine-Grained and Layered Object Recognition
abstract
This paper presents a novel research on promoting the performance and enriching the functionalities of object recognition. Instead of simply fitting various data to a few predefined semantic object categories, we propose to generate proper results for different object instances based on their actual visual appearances. The results can be fine-grained and layered categorization along with absolute or relative localization. We present a generic model based on structured prediction and an efficient online learning algorithm to solve it. Experiments on a new benchmark dataset demonstrate the effectiveness of our model and its superiority against traditional recognition methods.
Yang Wu 0001, Nanning Zheng 0001, Yuanliu Liu, Zejian Yuan
Int. J. Pattern Recognit. Artif. Intell.4
2012 IAIR-CarPed: A psychophysically annotated dataset with fine-grained and layered semantic labels for object recognition
Yang Wu 0001, Yuanliu Liu, Zejian Yuan, Nanning Zheng 0001
Pattern Recognit. Lett.3
2011 Automatic salient object segmentation based on context and shape prior
abstract
We propose a novel automatic salient object segmentation algorithm which integrates both bottom-up salient stimuli and object-level shape prior, i.e., a salient object has a well-defined closed boundary. Our approach is formalized as an iterative energy minimization framework, leading to binary segmentation of the salient object. Such energy minimization is initialized with a saliency map which is computed through context analysis based on multi-scale superpixels. Object-level shape prior is then extracted combining saliency with object boundary information. Both saliency map and shape prior update after each iteration. Experimental results on two public benchmark datasets show that our proposed approach outperforms state-of-the-art methods. 1
Huaizu Jiang, Jingdong Wang 0001, Zejian Yuan, Nanning Zheng 0001
BMVC3
2011 Object detection using discriminative photogrammetric context
abstract
Photogrammetric context captures the relationship between object heights and camera viewpoint, and can be used to reject false detections that appear in wrong locations or scales. In this work, we address the problem of using photogrammetric constraints in object detection when camera poses are unknown. We propose a model to capture both local appearance features and global photogrammetric context, in which the camera pose is treated as a latent variable. We use latent Structural SVM to learn the model parameters. To solve the NP-hard problem in structured prediction, we propose a branch-bound-and-cut algorithm, where cuts of the latent variable are embedded into a branch-and-bound process. The model is experimentally evaluated on INRIA pedestrian dataset. The results show that our model can get significantly better detection performance than models using only appearance features or using photogrammetric context in a graphical model.
Yuanliu Liu, Yang Wu 0001, Zejian Yuan
ICIP3
2011 Learning to Detect a Salient Object
abstract
In this paper, we study the salient object detection problem for images. We formulate this problem as a binary labeling task where we separate the salient object from the background. We propose a set of novel features, including multiscale contrast, center-surround histogram, and color spatial distribution, to describe a salient object locally, regionally, and globally. A conditional random field is learned to effectively combine these features for salient object detection. Further, we extend the proposed approach to detect a salient object from sequential images by introducing the dynamic salient features. We collected a large image database containing tens of thousands of carefully labeled images by multiple users and a video segment database, and conducted a set of experiments over them to demonstrate the effectiveness of the proposed approach.
Zejian Yuan, Jian Sun 0001, Jingdong Wang 0001, Nanning Zheng 0001, Xiaoou Tang, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 A Novel Dual-Probe Adaptive Model for Image Change Detection
abstract
Change detection is the foremost pre-attention process of visual motion analysis. It provides important preprocessing clues for the following complex visual attention selection and pattern recognition process. In this letter, a novel dual-probe adaptive model of the weak image change signal detection is advanced. Then its basic parameter constraints are analyzed and the numerical analysis of its characteristic is discussed. Simulation results show that the related change detector could capture the tiny change signals in synthetic and nature scenes with noisy background.
Nanning Zheng 0001, Zejian Yuan, Xuetao Zhang 0001
IEEE Signal Process. Lett.3
2009 Visual Saliency Based Object Tracking
Zejian Yuan, Nanning Zheng 0001, Xingdong Sheng
ACCV (2)2
2008 Video attention: Learning to detect a salient object sequence
abstract
We study video attention by detecting a salient object sequence from video segment. We formulate salient object sequence detection as energy minimization problem in a conditional random field framework, while static and dynamic salience, spatial and temporal coherence, global topic model are well defined and integrated to identify a salient object sequence. Dynamic programming algorithm is designed to resolve a global optimization, with a rectangle to represent each salient object. We validate our approach on a large number of video segments with the labeled salient object sequence.
Nanning Zheng 0001, Wei Ding 0002, Zejian Yuan
ICPR4
2008 Affine Registration of Point Sets Using ICP and ICA
abstract
This letter proposes a novel algorithm for affine registration of point sets in the way of incorporating an affine transformation into the iterative closest point (ICP) algorithm. At each iterative step of this algorithm, a closed-form solution of the affine transformation is derived. Similar to the ICP algorithm, this new algorithm converges monotonically to a local minimum from any given initial parameters. To get the best affine registration result, good initial parameters are required which are successfully estimated by using independent component analysis (ICA). Experimental results demonstrate the robustness and high accuracy of this algorithm.
Shaoyi Du, Nanning Zheng 0001, Gaofeng Meng, Zejian Yuan
IEEE Signal Process. Lett.4
2006 A Boosting SVM Chain Learning for Visual Information Retrieval
Zejian Yuan, Yanyun Qu, Yuehu Liu
ISNN (1)1
2005 A Cascaded Mixture SVM Classifier for Object Detection
Zejian Yuan, Nanning Zheng 0001, Yuehu Liu
ISNN (1)1
2004 Sequential updating algorithm for extracting the basis of karhunen loeve transformation
Yanyun Qu, Nanning Zheng 0001, Cuihua Li, Zejian Yuan
ICIP4
2003 Level set method for pulmonary vessels extraction
abstract
Pulmonary vessels extraction is of utmost importance in medical study. In this paper, we propose a new scheme to extract pulmonary vessels in CT images based on the level set method. A new region based speed function is designed and integrated with the previous edge based terms to evolve the front. Due to this region based speed term, the front could propagate even in thin vessel branches. The level set evolution is performed on a multiscale space where a new solution passing method is proposed that could induce a fast convergence rate. Meanwhile, an improved Hermes algorithm, we call it fast Hermes, is developed to evolve the level set evolution numerically. Experiments show that the segmentation results using our approach, especially in thin vessel branches extraction, are more promising than that of the usual level set method for segmentation.
Hongmei Zhang 0001, Zhengzhong Bian, Dazong Jiang, Zejian Yuan, Min Ye 0004
ICIP (2)4
2003 A Novel Multiresolution Fuzzy Segmentation Method on MR Image
Hongmei Zhang 0001, Zhengzhong Bian, Zejian Yuan, Min Ye 0004
J. Comput. Sci. Technol.3