pokémonpokémonpokémonpokémonYOLOv11s
pokémonpokémonpokémonpokémonRT-DETR
pokémonpokémonpokémonpokémonFaster R-CNN
9 classes, balancedPokémon detection, three ways
Computer vision11 ms · per frame for YOLOv11s, at 0.89 F1
Overview
YOLOv11s, RT-DETR and Faster R-CNN trained on the same nine Pokémon classes and compared on speed, classification and box quality.
- per frame with YOLOv11s, about 8× faster than Faster R-CNN
- 11 ms
- F1 for YOLOv11s, the best of the three
- 0.89
- mAP 50-95 for Faster R-CNN, the tightest boxes
- 0.80
The problem
Pick a detector for real-time tracking of drawn, non-photographic characters. A one-stage CNN, a transformer and a two-stage model each trade speed for accuracy in a different way, and the data started out lopsided: 280 images of Pikachu, 22 of Eevee.
The approach
approach.md
- 01Balanced the nine classes from Roboflow Universe to 280 training images each, with rotations, flips, blur and brightness changes.
- 02Trained YOLOv11s with the first 10 backbone layers frozen, RT-DETR as the transformer, and Faster R-CNN (ResNet-50) fully fine-tuned, all at 640×640 on a T4 GPU.
- 03Wrote a loader that converts YOLO-format labels into pixel corners for Faster R-CNN, then compared precision, recall, F1, mAP 50-95 and inference time.
The result
YOLOv11s is the pick for real time: 0.89 F1 at 11 ms per frame, about 8 times faster than Faster R-CNN (88 ms). Faster R-CNN draws the tightest boxes, with a mAP 50-95 of 0.80, but it keeps mistaking background for Pokémon, which drags its F1 down to 0.30.
Stack
- PyTorch
- Ultralytics
- YOLOv11s
- RT-DETR
- Faster R-CNN
- OpenCV
Next project
Weather, 30 steps ahead