NeurIPS 2026 · Track on Evaluations and Datasets

PHOEBI: An Open-World Benchmark for Multi-Label Bacterial Identification in Phase-Contrast Microscopy

Aaditya Baranwal Md Jahid Hasan Shruti Vyas

Institute of Artificial Intelligence, University of Central Florida

Phase-contrast micrograph with rod-shaped bacteria of several sizes
A field from the culture that contains all six species, in phase contrast at 1000× total magnification.

Try the task

Every image in the benchmark poses one question: which of the six species does it contain? Pick the species you think are in this field, then check your answer.

A phase-contrast field from one of the 40 cultures

Your answer

Species

Six rod-shaped species that differ in cell size, cell wall and motility. Each view is a crop of an image from the dataset.

Phase-contrast view of Bacillus subtilis

bs

Bacillus subtilis

Gram
positive
Motility
peritrichous flagella
Length
4–10 µm
Phase-contrast view of Bacillus thermoamylovorans

bt

Bacillus thermoamylovorans

Gram
positive
Motility
peritrichous flagella
Length
~4 µm
Phase-contrast view of Flavobacterium johnsoniae

fj

Flavobacterium johnsoniae

Gram
negative
Motility
gliding
Length
5–10 µm
Phase-contrast view of Klebsiella aerogenes

ka

Klebsiella aerogenes

Gram
negative
Motility
peritrichous flagella
Length
1–3 µm, encapsulated
Phase-contrast view of Myxococcus xanthus

mx

Myxococcus xanthus

Gram
negative
Motility
gliding
Length
5–10 µm
Phase-contrast view of Pseudomonas fluorescens

pf

Pseudomonas fluorescens

Gram
negative
Motility
polar flagella
Length
1.5–3 µm

Data collection

Each species is grown on its own and confirmed pure, and the cultures are mixed only when the slide is prepared, so the species were never co-cultured and every label is verified at culture level.

bsfjka
  1. 01

    Grow

    Each species grows on its own from a glycerol stock, in nutrient broth at 30 °C and 250 rpm for 72–120 h.

  2. 02

    Verify

    Every culture is checked under the microscope and confirmed pure.

  3. 03

    Mix

    The verified cultures are combined at a controlled volume ratio, only at acquisition time.

  4. 04

    Mount

    The mixture goes straight onto a slide as an unstained wet mount.

  5. 05

    Image

    Phase-contrast video at 1000× total magnification, with a 100× oil-immersion objective.

  6. 06

    Label

    Each image carries its culture's species, verified at culture level.

Method

The three decoders share one pipeline: they read frozen features of small tiles and answer for the whole image.

FrozenDINOv2-S/14bsbtfjkamxpfbsbtfjkamxpfbsbtfjkamxpf
  1. 01

    Tile

    The image is corrected for uneven illumination and cut into a 4 × 4 grid of 224 px tiles.

  2. 02

    Encode

    A frozen DINOv2-S/14 encoder turns each tile into a feature vector.

  3. 03

    Compare

    Each decoder reads every species' presence against fixed species prototypes.

  4. 04

    Average

    Per-tile scores are averaged over the 16 tiles into one score per species.

  5. 05

    Decide

    Species whose score clears a calibrated threshold are reported present.

Schematic: the example image contains B. subtilis, F. johnsoniae and K. aerogenes; the bar heights are illustrative, not measured.

Benchmark

  • 120,000images
  • 40combinations
  • 6species
  • 1000×magnification

Each image is labelled with the set of species in its culture. A model must say which of the six are present, including in combinations and with species it never saw in training.

Random split
80/10/10 in acquisition order within each combination, for in-distribution results.
Leave-combinations-out
Nine whole combinations are held out while every species still appears in training, so a model must recognise known species in mixtures it has never seen.
Leave-one-species-out
Each species in turn is withheld from training, to test rejecting images that contain an unseen species and discovering it as a new class.

Every combination's images are divided 80/10/10 in acquisition order, so every combination is seen in training.

1 species2 species3 species4 species6 speciesbsbtfjkamxpf

Results

Three panels: a phase-contrast field containing all six species; arrows from each model's F1 on mixtures seen in training to its F1 on unseen mixtures; bar charts for detecting and grouping an unseen species.
Figure 1. PHOEBI at a glance. Left: a phase-contrast field from a culture containing all six species; the benchmark spans 40 combinations of these species. Centre: each arrow runs from a model's F1 on mixtures seen during training (open circle) to its F1 on mixtures it has never seen (filled circle). Standard deep classifiers (red) collapse on the new mixtures; our three lightweight decoders (green, A–C) remain stable. Right: with no further training, the same frozen features detect images that contain a species never seen in training (top) and group those images into a new class (bottom); green bars are the methods we adopt, grey bars the alternatives.

Findings

  1. Per-image classifiers collapse on unseen combinations

    Fine-tuned backbones and attention-based multiple-instance learning are accurate on the mixtures they were trained on but fail on held-out combinations. The failure comes from how per-image predictions are aggregated, not from the visual representation.

  2. Anchor-based decoders stay stable

    Three lightweight decoders read each species' presence against fixed prototypes over one shared set of frozen tile features, and they hold up under the same shift.

  3. New species without retraining

    With no further training, the same features reject images that contain an unseen species and group those images into a new class, while the known classes are barely affected.

Data access

Load with Hugging Face datasets

from datasets import load_dataset

ds = load_dataset("sochastic/PHOEBI", "phoebi6")        # random 80/10/10 split
species = ds["train"].features["labels"].feature.names  # ['bs', 'bt', 'fj', 'ka', 'mx', 'pf']

Splits, protocols and the four-species subset are described on the dataset card.

Or download everything as files

git clone https://github.com/eternal-f1ame/phoebi && cd phoebi
conda env create -f environment.yml && conda activate phoebi
python tools/fetch_dataset.py   # images into data/images/, manifests into data/

The code repository has the decoders, every baseline and the scripts behind each result in the paper.

Citation

@inproceedings{baranwal2026phoebi,
  title     = {{PHOEBI}: An Open-World Benchmark for Multi-Label Bacterial Identification in Phase-Contrast Microscopy},
  author    = {Baranwal, Aaditya and Hasan, Md Jahid and Vyas, Shruti},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},
  year      = {2026}
}