Survey of GaitSet-Based Frameworks for Low-Resolution Gait Recognition
Abstract
Gait recognition remains useful at distances at which facial detail is unavailable, but low-resolution surveillance video removes the spatial cues required by many silhouette models. GaitSet represents a sequence as an unordered set of frames and is relatively tolerant of view variation and temporal misalignment. Its fixed max or mean pooling, however, cannot distinguish reliable silhouettes from frames damaged by downsampling, compression, blur, occlusion, or segmentation error. This focused narrative survey reviews GaitSet and closely related low-resolution gait methods across three intervention points: input enhancement, multi-scale representation learning, and set-level aggregation. Later architectures are separated into a Transformer used after convolutional frame encoding, reliability-aware or multi-scale aggregation that retains a GaitSet-style backbone, and an end-to-end Transformer gait encoder. Source-reported Rank-1 results are compared only within their stated protocols. The literature provides few matched evaluations at both standard 64 × 44 or 64 × 64 resolution and explicit 32 × 32 or compression conditions. Results from CASIA-B, Gait3D, and GREW are also not directly interchangeable because the datasets and evaluation protocols differ substantially. Across the reviewed evidence, reliability-aware aggregation addresses the immediate failure of low-resolution sequences more directly than adding attention alone. A practical design should estimate frame and region quality before set aggregation, then add temporal refinement or a Transformer backbone only when the training data and deployment budget support the added complexity.
Keywords
Download Full Article
PDF format
Recommended Citation
Feng Jiale & Kazem Chamran (2026). Survey of GaitSet-Based Frameworks for Low-Resolution Gait Recognition. Glovento Journal of Integrated Studies (GJIS), 2, Article 87. https://doi.org/10.68246/gjis.v2.87