Spatiotemporal Dilated Convolution with Uncertain Matching for Video-based Crowd Estimation
Journal
IEEE TRANSACTIONS ON MULTIMEDIA
Journal Volume
24
Pages
261
Date Issued
2021-01-29
Author(s)
Abstract
In this paper, we propose a novel SpatioTemporal convolutional Dense Network
(STDNet) to address the video-based crowd counting problem, which contains the
decomposition of 3D convolution and the 3D spatiotemporal dilated dense
convolution to alleviate the rapid growth of the model size caused by the
Conv3D layer. Moreover, since the dilated convolution extracts the multiscale
features, we combine the dilated convolution with the channel attention block
to enhance the feature representations. Due to the error that occurs from the
difficulty of labeling crowds, especially for videos, imprecise or
standard-inconsistent labels may lead to poor convergence for the model. To
address this issue, we further propose a new patch-wise regression loss (PRL)
to improve the original pixel-wise loss. Experimental results on three
video-based benchmarks, i.e., the UCSD, Mall and WorldExpo'10 datasets, show
that STDNet outperforms both image- and video-based state-of-the-art methods.
The source codes are released at \url{https://github.com/STDNet/STDNet}.
Subjects
Feature extraction; Convolution; Training; Spatiotemporal phenomena; Annotations; Three-dimensional displays; Videos; Crowd counting; density map regression; dilated convolution; patch-wise regression loss; spatiotemporal modeling; MOTION; Computer Science - Computer Vision and Pattern Recognition; Computer Science - Computer Vision and Pattern Recognition; Computer Science - Learning
SDGs
Publisher
IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC
Description
Accepted by IEEE Transactions on Multimedia, 2021
(https://ieeexplore.ieee.org/document/9316927)
(https://ieeexplore.ieee.org/document/9316927)
Type
journal article
