Literature review /003
- Field
- Piling intelligence
- Journal
- Smart Agricultural Technology, Volume 10, Article 100745
- DOI
- 10.1016/j.atech.2024.100745
- Reviewed
- July 12, 2026
Comparison of strategies for automatic video-based detection of piling behaviour in laying hens
Dan Børge Jensen, Michael Toscano, Esther van der Heide, Matias Grønvig, and Franziska Hakansson
Review contents
Review abstract
Study
Jensen and colleagues compare three secondary neural-network strategies for classifying piling events from commercial laying-hen video. A pre-trained VGG-16 encoder and principal-component analysis feed fully connected, long-short term memory, or convolutional models that classify frames or short sequences.
Findings
Training data came from three Swiss flocks and testing from a fourth held-out flock. The best convolutional model reached major mean accuracy of 90 percent, while the best LSTM reached 86 percent with one-eighth as many principal components and a more intuitive decision threshold.
Interpretation
Temporal information improved piling detection, but the study remains preliminary: all birds were white layers in winter gardens, labels came from one observer, and early stopping monitored the test set. The work informs temporal piling research but does not directly validate broiler-barn detection or nesteye performance.
Research question
Which secondary model best balances event-level piling detection performance and computational practicality when applied to sequences of encoded commercial-flock video?
Evidence profile
Study design
| Subjects | White laying hens from four commercial Swiss flocks |
|---|---|
| Setting | Winter-garden areas recorded by five cameras |
| Imaging | Daytime RGB and nighttime infrared frames converted to grayscale at one frame per second |
| Training split | Four cameras from three flocks |
| Test split | One camera from a fourth held-out flock with 21 piling and 11 non-piling events |
| Encoding | Pre-trained VGG-16 features reduced with principal-component analysis |
| Models | Frame-level FC-ANN and sequence-based LSTM and CNN classifiers |
| Primary metrics | Area under the ROC curve and major mean accuracy at event level |
Density
Duration
Direction
Alert
90%
Best CNN MMA
256 principal components and 16-frame sequences
86%
Best LSTM MMA
32 principal components and 16-frame sequences
4 flocks
Commercial data
Three flocks for training and one held-out flock for testing
Scientific context
Piling is a collective event in which birds form a dense, relatively immobile cluster. It can lead to stress, suffocation, mortality, and production loss. Manual detection is difficult because commercial flocks occupy large spaces and an event may develop between human observations. A single image can also be ambiguous: ordinary crowding, lighting variation, occlusion, or a momentary mass movement may resemble the geometry of a pile without sharing its duration or behavioral organization.
Jensen et al. therefore test whether sequences improve classification. Their design separates generic visual encoding from a task-specific secondary model. This is attractive when annotated agricultural video is limited: a pre-trained network supplies a broad visual representation, while a smaller model learns the particular distinction between piling and non-piling events. The comparison asks whether simple frame-level classification is enough or whether temporal architectures justify their additional complexity.
Methodological reconstruction
Video from four commercial flocks was sampled at one frame per second. Daytime RGB and nighttime infrared frames were converted to grayscale so one model could span lighting modes. VGG-16 encoded each frame, and principal-component analysis compressed the representation. Four cameras from three flocks supplied training data; the fifth camera, representing a fourth flock, supplied the test events. This flock-level split is a meaningful protection against direct frame leakage between training and evaluation.
The FC-ANN receives one encoded frame at a time. The LSTM and CNN receive sequences whose length and retained principal components are varied. Predictions are aggregated over each coherent piling or non-piling event before ROC analysis. The authors evaluate area under the curve and major mean accuracy, defined as the average of sensitivity and specificity. Models are trained with Adam, categorical cross-entropy, and early stopping. Importantly, early stopping tracks loss on the set described as the test set, so that set participates in model selection as well as final reporting.
Principal findings
The strongest FC-ANN reached major mean accuracy of 76 percent with 32 principal components. The best LSTM reached 86 percent with 32 components and a sequence length of 16 frames. The best CNN reached 90 percent with 256 components and the same sequence length. Both temporal approaches outperformed the frame-level model, supporting the proposition that piling contains time-dependent information that is lost when frames are considered independently.
The authors do not select the numerically highest model without qualification. Although CNN performance was slightly higher, the LSTM required far fewer principal components and placed its optimal decision threshold close to 50 percent. The CNN's optimal threshold was substantially lower, suggesting less intuitive calibration. The LSTM therefore offered a stronger practical balance of sensitivity, specificity, compute demand, and threshold behavior for future local deployment.
Critical appraisal
The study's main strength is the comparison of architectures under a shared encoding and event-level evaluation protocol. Holding out an entire flock is more informative than randomly splitting adjacent frames, and event aggregation reflects the operational question better than per-frame accuracy alone. The authors also discuss compute requirements and calibration rather than treating model size as irrelevant. That deployment-aware interpretation is particularly valuable for edge systems.
The evaluation nevertheless combines model development and assessment. Early stopping restores the model with the lowest loss on the test set, and hyperparameter combinations are compared using the same data. Consequently, the held-out flock is not a completely untouched final validation cohort. Reported performance may reflect repeated selection against that flock. A separate validation flock for early stopping and hyperparameter selection, followed by an untouched test flock, would provide stronger evidence.
Limitations and generalizability
The authors explicitly describe the work as preliminary. All included birds were white laying hens in Swiss winter gardens selected partly for favorable image quality. Performance may not transfer to brown birds, interior litter areas, broiler body shapes, different housing systems, dust levels, or camera perspectives. Only one VGG-16 encoder was tested. The event labels were produced by a single observer, creating uncertainty at the transition between piling and non-piling despite high previously reported intra-observer reliability.
The test data contain 21 piling and 11 non-piling events from one flock. Event aggregation reduces frame-level dependence, but the number of independent biological events remains modest. No prospective alerting trial measures latency, false alerts per barn-day, operator response, or prevented harm. The study identifies a promising classifier family; it does not demonstrate a complete detection-and-intervention system or prove that performance remains stable through a production cycle.
Implications for nesteye
For nesteye, the paper supports repeated-frame confirmation and explicit consideration of local compute. A temporal model can reject transient visual similarity and may reduce false alerts compared with a frame-only classifier. The LSTM result also illustrates why the practical model is not always the model with the highest isolated score: memory, inference cost, calibration, detection latency, and maintainability matter inside a barn.
The study does not validate nesteye's reported piling AP, broiler performance, edge hardware, or intervention logic. Layer behavior and broiler behavior should not be treated as interchangeable. Nesteye would require broiler-specific event definitions, multiple independent annotators, barn-disjoint validation, an untouched final test set, and prospective reporting of sensitivity, specificity, detection latency, and false alerts over time. The paper informs experimental design; it is not transferable product evidence.
Conclusion
Jensen et al. show that models capable of interpreting short visual sequences classify piling events more effectively than a model that sees one encoded frame at a time. Their comparison also makes a useful engineering point: a slightly less accurate model may be preferable when it requires fewer inputs, consumes less compute, and yields a better-calibrated threshold.
The evidence remains bounded by a small, favorable, layer-specific dataset and by model selection performed against the reported test flock. The study should therefore be read as support for temporal architecture and disciplined follow-on validation, not as a general piling-detection benchmark. Its most durable contribution is the framing of piling as an event whose duration and evolution are part of the signal.
APA reference
Jensen, D. B., Toscano, M., van der Heide, E., Grønvig, M., & Hakansson, F. (2025). Comparison of strategies for automatic video-based detection of piling behaviour in laying hens. Smart Agricultural Technology, 10, 100745. https://doi.org/10.1016/j.atech.2024.100745
View original paper