
bs
Bacillus subtilis
- Gram
- positive
- Motility
- peritrichous flagella
- Length
- 4–10 µm
NeurIPS 2026 · Track on Evaluations and Datasets
Institute of Artificial Intelligence, University of Central Florida

Every image in the benchmark poses one question: which of the six species does it contain? Pick the species you think are in this field, then check your answer.

Your answer
Six rod-shaped species that differ in cell size, cell wall and motility. Each view is a crop of an image from the dataset.

bs

bt

fj

ka

mx

pf
Each species is grown on its own and confirmed pure, and the cultures are mixed only when the slide is prepared, so the species were never co-cultured and every label is verified at culture level.
Each species grows on its own from a glycerol stock, in nutrient broth at 30 °C and 250 rpm for 72–120 h.
Every culture is checked under the microscope and confirmed pure.
The verified cultures are combined at a controlled volume ratio, only at acquisition time.
The mixture goes straight onto a slide as an unstained wet mount.
Phase-contrast video at 1000× total magnification, with a 100× oil-immersion objective.
Each image carries its culture's species, verified at culture level.
The three decoders share one pipeline: they read frozen features of small tiles and answer for the whole image.
The image is corrected for uneven illumination and cut into a 4 × 4 grid of 224 px tiles.
A frozen DINOv2-S/14 encoder turns each tile into a feature vector.
Each decoder reads every species' presence against fixed species prototypes.
Per-tile scores are averaged over the 16 tiles into one score per species.
Species whose score clears a calibrated threshold are reported present.
Schematic: the example image contains B. subtilis, F. johnsoniae and K. aerogenes; the bar heights are illustrative, not measured.
Each image is labelled with the set of species in its culture. A model must say which of the six are present, including in combinations and with species it never saw in training.
Every combination's images are divided 80/10/10 in acquisition order, so every combination is seen in training.

Fine-tuned backbones and attention-based multiple-instance learning are accurate on the mixtures they were trained on but fail on held-out combinations. The failure comes from how per-image predictions are aggregated, not from the visual representation.
Three lightweight decoders read each species' presence against fixed prototypes over one shared set of frozen tile features, and they hold up under the same shift.
With no further training, the same features reject images that contain an unseen species and group those images into a new class, while the known classes are barely affected.
from datasets import load_dataset
ds = load_dataset("sochastic/PHOEBI", "phoebi6") # random 80/10/10 split
species = ds["train"].features["labels"].feature.names # ['bs', 'bt', 'fj', 'ka', 'mx', 'pf']Splits, protocols and the four-species subset are described on the dataset card.
git clone https://github.com/eternal-f1ame/phoebi && cd phoebi
conda env create -f environment.yml && conda activate phoebi
python tools/fetch_dataset.py # images into data/images/, manifests into data/The code repository has the decoders, every baseline and the scripts behind each result in the paper.
@inproceedings{baranwal2026phoebi,
title = {{PHOEBI}: An Open-World Benchmark for Multi-Label Bacterial Identification in Phase-Contrast Microscopy},
author = {Baranwal, Aaditya and Hasan, Md Jahid and Vyas, Shruti},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},
year = {2026}
}