Skip to content

Pair mining

Command: wildmatch mine. Pair mining runs once per dataset and matcher with the pretrained matcher on the training (database) split only. It produces the strong-match index that fine-tuning samples its triplets from. The index is tied to one fixed pretrained checkpoint by design: it defines the candidate pool the adaptation is trained and validated on.

Commands are examples

Locations come from the path profile (see Reproduce). The Slurm script slurm/mine.sbatch targets one specific cluster; adjust its partition and the environment location for yours. Run the commands from the repository root.

Steps

wildmatch mine <step> --dataset <registry key> --backend loma|rdd runs one step; --protocol defaults to legacy. Start with plan, which writes nothing and prints the view, cache, index locations and the exact command of every later step:

wildmatch mine plan --dataset nyala --backend loma
Step What it does
view builds the symlink view from the registry entry's metadata
cache extracts the per-image feature cache (GPU)
check validates the view, the weights and the cache before mining
submit freezes the run and submits the per-split mining arrays and the aggregation

Canonical dataset view

Mining expects a symlink view with train/ and test/ directories (and val/ for CzechLynx) and one folder per individual and collection. The view step builds it from the dataset's registry entry; for WildlifeReID-10k datasets it reads the pre-masked images the paper's runs used and never falls back to unmasked ones.

wildmatch mine view --dataset czechlynx_closed --backend rdd   # CzechLynx, time-closed split
wildmatch mine view --dataset nyala --backend loma             # a WildlifeReID-10k dataset

legacy is the protocol used for the paper: the official training (database) split is used for fine-tuning and the official test (query) split for training-time validation and final reporting. A strict protocol with a held-out validation part exists but is not used in the paper.

Feature cache

All scoring reads a feature cache, never the images. Each image is resized to a long side of 512 pixels (dimensions floored to a multiple of 32 for RDD, 14 for LoMa), run through the detector and descriptor, and stored with up to 512 keypoints. The cache is only valid for one combination of keypoint budget, resolution and preprocessing, which is why the cache directory names encode them.

sbatch slurm/mine.sbatch cache --dataset nyala --backend loma   # one GPU
wildmatch mine check --dataset nyala --backend loma

Images are decoded with PIL to float tensors in [0, 1] and resized with plain bilinear interpolation without antialiasing, the same function fine-tuning and evaluation use, so the features mined here are the features trained on.

Mining the index

wildmatch mine submit --dataset nyala --backend loma --dry-run   # print the Slurm commands
wildmatch mine submit --dataset nyala --backend loma

submit checks the inputs and writes the resolved run, with the code identity and the checksums of the weights and the cache, to a submission file that every job then executes; a job stops if the code or an input changed since submission. Mined files and an existing index are never overwritten without --overwrite.

The submission mines every query collection as one Slurm array task per split and then aggregates the per-query reports into one combined JSON per split, with image paths relative to the view root.

Setting Value Meaning
keypoints per image 512 detector budget, fixed at cache build
long side 512 px resize before extraction
images per collection 20 uniformly sampled; shorter collections contribute all images
positives and negatives kept per anchor 5 each highest-scoring same-identity and different-identity images, spread over distinct collections
anchors kept per collection 10 ranked by their best pair score

Anchors without a positive are dropped. The query's own collection is always excluded from the candidate gallery, so a positive never comes from the same encounter.

Output, under external.mining_outputs: czechlynx-time-closed/<protocol>/<backend>/strong-matches_{train,val,test}_combined.json for CzechLynx, or wildlife-reid-10k/<dataset>/indices/<backend>/strong-matches_{train,test}_combined.json for the other datasets, each with a metadata sidecar recording backend, checkpoint, cache and settings. In the legacy protocol the validation index is an alias of the test index.

Pair sources used in the paper

Each published checkpoint records the index it was trained on. Indices of the two backends are kept in separate directories and never mixed within one run.

Dataset LoMa trained on RDD-LightGlue trained on
CzechLynx (time-closed) LoMa-mined pairs RDD-mined pairs
CzechLynx (time-open, unseen-identity protocol) LoMa-mined pairs LoMa-mined pairs
Hyena, Leopard, Sea star, Turtle LoMa-mined pairs LoMa-mined pairs
Nyala, Whale shark RDD-mined pairs RDD-mined pairs
Salamander LoMa-mined pairs RDD-mined pairs

To train on the other source, mine with that backend and pass matcher_finetune.mined_by=loma|rdd (see Matcher fine-tuning).