Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

2.5.1 Perception 2D

Authors
Affiliations
Delft University of Technology
Delft University of Technology
Updated: 26 Aug 2026

For determining the type, and possibly the distance of objects, a YOLO26 (‘You Only Look Once-26’) object classification model was used. A dataset was trained by recording videos using the USB Camera included in the MIRTE Master hardware package, of which one in every fifteen frames got extracted to use for the training dataset. A figure of four frames in different environments is shown below.

Figure of four individual frames.

Figure 1:Figure of four individual frames.

Why YOLO26 was used:

The YOLO object-detection family is well established, with YOLO26 Ultralytics, 2026 being the most recent generation. YOLO26 was chosen because the robot has to detect and classify objects from a live camera on the MIRTE Master platform, which has limited processing power as it uses an OrangePi 3B which has limited computational capabilities. Object detection with limited processing power requires a lightweight model with fast inference times while still maintaining accuracy. YOLO26 is suitable for this because it was designed with deployment on low-power devices in mind. It changed the post-processing pipeline to NMS-free inference, simplified bounding-box predictions and supports deployment formats commonly used for low-power devices Sapkota et al., 2026. In this case, the RK3566 NPU of the OrangePi 3B can be used for improving the performance, reducing the total inference time Limaran et al., 2026. For this model, the Nano variant of YOLO26 was chosen to minimize inference times. Another benefit of the newer YOLO26 variant over the older variants is that YOLO26 allows for oriented bounding boxes. Unlike standard bounding boxes, oriented bounding boxes also estimate the rotation of recognized objects Sapkota & Karkee, 2026. This could be useful for complex grasping tasks where orientation of the gripper is crucial. However, this project only uses standard object detection.

Training

Training was done manually using Label Studio. Initially, roughly 280 frames were captured in four different environments. Due to a large amount of blurring within the images, about 100 pictures were deleted. Of the remaining 180 pictures, 60 pictures were imported into Label Studio and labelled per colour: green, red, purple, white, light-grey, dark-grey and black. In total, 1206 bounding boxes were drawn. A test of the perception model performed during training is shown in the figure below.

Figure of tests performed during training.

Figure 2:Figure of tests performed during training.

The curves in the figure below show that the model converged during trading. This means that the model actually learns while training instead of making random or unstable decisions.

Curves of training results per epoch.

Figure 3:Curves of training results per epoch.

The normalised confusion matrix shows how well classes are recognized by the model. As shown in the figure shown below, most classes are classified correctly, as indicated by the strong diagonal values. However, light-grey objects were among the least correctly identified objects, this is a result of exposure and lighting changing enough to where the model will classify them as background. Some of objects were missed as background, a small amount of objects were misclassified, and some objects like dark grey and white where confused with light-grey.

Normalized confusion matrix.

Figure 4:Normalized confusion matrix.

A figure of the F1 confidence curve is shown below. This curve shows how the F1 metric, which combines precision (how many objects where correctly predicted) and recall (how many objects where correctly found), changes with the confidence threshold of the model. The curve shows that a good starting confidence threshold would be 0.3-0.35.

BoxF1 curve

Figure 5:BoxF1 curve

References
  1. Ultralytics. (2026). Ultralytics YOLO26. https://docs.ultralytics.com/models/yolo26/
  2. Sapkota, R., Cheppally, R. H., Sharda, A., & Karkee, M. (2026). YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection. https://arxiv.org/abs/2509.25164
  3. Limaran, A., Wicaksono, A., & Herwanto, P. (2026). Performance enhancement of embedded object detection via neural hardware acceleration. TELKOMNIKA (Telecommunication Computing Electronics and Control), 24, 126. 10.12928/telkomnika.v24i1.27448
  4. Sapkota, R., & Karkee, M. (2026). Ultralytics YOLO Evolution: An Overview of YOLO26, YOLO11, YOLOv8 and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition. https://arxiv.org/abs/2510.09653