Previous model: Machine learning competition¶
The algorithms used by
zamba v1 were based on the winning solution from the
Pri-matrix Factorization machine learning
competition, hosted by DrivenData. Data for the competition was provided by the Chimp&See project and manually labeled by volunteers. The competition had over 300 participants and over 450 submissions throughout the three month challenge. The v1 algorithm was adapted from the winning competition submission, with some aspects changed during development to improve performance.
The core algorithm in
zamba v1 was a stacked ensemble which consisted of a first layer of models that were then combined into a final prediction in a second layer. The first level of the stack consisted of 5
learning models, whose individual predictions were combined in the second level
of the stack to form the final prediction.
In v2, the stacked ensemble algorithm from v1 is replaced with three more powerful single-model options:
european. The new models utilize state-of-the-art image and video classification architectures, and are able to outperform the much more computationally intensive stacked ensemble model.
New geographies and species¶
zamba v2 incorporates data from Western Europe (Germany). The new data is packaged in the pretrained
european model, which can predict 11 common European species not present in
zamba v2 also incorporates new training data from 15 countries in central and west Africa, and adds 12 additional species to the pretrained African models.
Model training is made available
zamba v2, so users can finetune a pretrained model using their own data to improve performance for a specific ecology or set of sites.
zamba v2 also allows users to retrain a model on completely new species labels.