Pokémon detection, three ways

Computer vision11 ms · per frame for YOLOv11s, at 0.89 F1

Overview

YOLOv11s, RT-DETR and Faster R-CNN trained on the same nine Pokémon classes and compared on speed, classification and box quality.

per frame with YOLOv11s, about 8× faster than Faster R-CNN
11 ms
F1 for YOLOv11s, the best of the three
0.89
mAP 50-95 for Faster R-CNN, the tightest boxes
0.80

The problem

Pick a detector for real-time tracking of drawn, non-photographic characters. A one-stage CNN, a transformer and a two-stage model each trade speed for accuracy in a different way, and the data started out lopsided: 280 images of Pikachu, 22 of Eevee.

The approach

approach.md
  1. 01Balanced the nine classes from Roboflow Universe to 280 training images each, with rotations, flips, blur and brightness changes.
  2. 02Trained YOLOv11s with the first 10 backbone layers frozen, RT-DETR as the transformer, and Faster R-CNN (ResNet-50) fully fine-tuned, all at 640×640 on a T4 GPU.
  3. 03Wrote a loader that converts YOLO-format labels into pixel corners for Faster R-CNN, then compared precision, recall, F1, mAP 50-95 and inference time.

The result

YOLOv11s is the pick for real time: 0.89 F1 at 11 ms per frame, about 8 times faster than Faster R-CNN (88 ms). Faster R-CNN draws the tightest boxes, with a mAP 50-95 of 0.80, but it keeps mistaking background for Pokémon, which drags its F1 down to 0.30.

Stack

  • PyTorch
  • Ultralytics
  • YOLOv11s
  • RT-DETR
  • Faster R-CNN
  • OpenCV
Repository on GitHubgithub.com/ricca200xx/Pokemon-Detection-Model-Comparison
Data ScienceAI EngineeringGenerative AILLM pipelinesMachine LearningStatisticsForecastingOptimizationData ScienceAI EngineeringGenerative AILLM pipelinesMachine LearningStatisticsForecastingOptimizationData ScienceAI EngineeringGenerative AILLM pipelinesMachine LearningStatisticsForecastingOptimizationData ScienceAI EngineeringGenerative AILLM pipelinesMachine LearningStatisticsForecastingOptimization