mAP Calculator

Score an object detector by PascalVOC and COCO rules.

One Python script that turns a folder of ground-truth boxes and a folder of detections into average precision, with the three interpolation rules the benchmarks use.

Usage

python mAP.py [-d det] [-g gt] [-i iou] [-p points]
              [-n names] [-c confidence] [--coco]

Launch runs the real script from GitHub inside your browser. Nothing is installed and nothing is sent to a server.

About

Average precision is the number every detection paper reports, and small differences in how it is computed change it. This script implements the continuous, 11-point and 101-point rules side by side, so a result can be compared with PascalVOC 2012 or COCO on their own terms.

Background on the metric itself is in Object Detection — A Quick Read.

Try it

Press Launch. The page starts Python, downloads mAP.py from the repository, and runs it on a small demonstration set. The left column says what each stage is doing while it happens. Then move the IoU threshold and watch the score change.

0.50
$ python mAP.py -d detection-results -g ground-truths -i 0.50 -p 0 -n model.names
What is happening
  1. Start Python in this tab
    Pyodide is CPython compiled to WebAssembly. The script runs on your machine, not a server.
  2. Fetch mAP.py from GitHub
    The same file you would get from git clone, unmodified.
  3. preprocessData()
    Reads every ground-truth and detection file and groups the boxes by class and by image.
  4. getMaskBoxes() → vecIoU()
    Measures the overlap of every detection with every ground truth in its image. Pairs above the threshold are claimed best-overlap first, one detection per object. A second detection on the same object counts as false.
  5. calAPClass()
    Sorts detections by confidence and walks down the list, recording precision and recall after each one.
  6. calAveragePrecision()
    Makes precision non-increasing, then measures the area under the curve with the chosen rule.
Output
00.51 00.51 recallprecision Press Launch to draw the curve.
Average precision
—
not run yet
True positives
—
 
False positives
—
 
Script output
(waiting)

Demonstration data: 16 synthetic ground-truth boxes and 22 synthetic detections across 4 frames, authored for this page. The computation is the real script.

#FrameConfidenceResultPrecisionRecall
The ranked detections appear here after a run.

How it works

Every number above can be recomputed by hand from the ranked table under the console.

1. Match. For each image, compute IoU between all ground truths and all detections. Keep pairs above the threshold, take them in order of overlap, and let each ground truth and each detection be used once.

2. Rank. Pool the detections of all images and sort by confidence, highest first.

3. Accumulate. After each detection, precision is true positives over detections so far; recall is true positives over all ground truths.

4. Integrate. Replace each precision with the highest precision at any greater recall, then take the area: exactly (continuous), or sampled at 11 or 101 recall values.

Limitations

  • The first launch downloads the Python runtime, so it takes a few seconds. Later runs on the page are immediate.
  • The demonstration set is synthetic and has one class. It shows the method; it is not a benchmark result.
  • The script calls np.trapz, which recent NumPy renames to np.trapezoid. The launcher adds an alias when needed; the repository file is unchanged.