What changed in three years
Until 2023, identifying an animal from a photo meant picking one model per group: a network for birds, another for insects, a third for plants. Each was good at home and lost everywhere else — show a spider to a bird classifier and it hands you back a bird.
Vision-language models trained on the whole tree of life removed that split. SafaRoll runs on BioCLIP 2.5, trained on TreeOfLife, the largest open annotated biological image dataset. A single model covers birds, insects, mammals, reptiles, fish and molluscs, and returns a species name rather than a convenient rank.
The photo does not need to be beautiful
This is the main misconception. A model trained on millions of field images has already seen, during training, exactly the sort of photo you take: badly framed, backlit, slightly blurred. What matters is not sharpness, it is what sits inside the frame.
- The whole subject rather than a close-up: a complete bird shot from afar beats a sharp head.
- A contrasting background: a dark worm on dark soil is the worst possible setup.
- One animal only: two individuals in frame make the model hesitate between them.
- A side or three-quarter view, never top-down for a bird nor head-on for a fish.
What still fails
Juvenile stages first. A caterpillar looks more like other caterpillars than like its own butterfly, and the model reasons on what it sees. Tadpoles, larvae and downy chicks often return a rank broader than species.
Signs of presence next: a footprint, a loose feather, a dropping. SafaRoll identifies the animal, not its traces. And finally domestic breeds — the model returns Felis catus, never \"British Shorthair\".
How many species can really be named?
SafaRoll keeps a collection catalogue of 884 species across ten groups, and a much wider atlas. A species that is recognised but outside the catalogue is still named correctly: it simply joins the atlas rather than the collection. The real ceiling is not how many species the model knows, it is how many you will actually meet.
