Number of Images for multiple different classes

Project Type: Ingredient Recognition System

Hello! I am relatively new to working with datasets and image classification, so I would appreciate some advice. So, we were tasked with creating an Ingredient Recognition System using YOLOv8, where we need to classify around 50 ingredients. Since I am new to dataset preparation, I’m not sure how many images per class are actually needed to achieve good model performance.

We were also given a very short deadline of about one week, so we initially collected a large number of images to make our dataset more diverse. One of my teammates recommended around 500 images per classes, but I am quite unsure about the number of images. Is there any recommended approach for students in this situation? Should we reduce the number of images per class?

Any advice would be greatly appreciated. Thank you so much!

That’s the eternal struggle for all of us! There’s no formula to know how many you need so even experts have to use their best guess. But here are a few tips:

  • This post outlines some basics: How many images do you need to train a model?
  • Start small. Just get 50 detections per class, for just a few of the classes, train a model, see how it does! You can often tell right away classes that need way more examples and ones that learned well from a small dataset. And you’ll have an idea what other classes might need based on those first results [Clarification point - note I said “detections” not images. If you have 1 image with two different apples, a pear, and an orange, you have four detections in a single image.]
  • Adjust based on usage. That is, how will the model be used? One extreme would be if you are going to put a single ingredient on a white surface under bright lights every time, you’ll probably end up good at the 20 range. At the other extreme, if you are trying to find an ingredient that has either been chopped up into a small piece and is sitting in a pile with other ingredients on a marble cutting board or could be only cut in half with the inside exposed but still mixed in a bowl with ten other items - then you might very well need the variety that 500 images brings. (Variety is key - 500 images where 450 are very similar - the same apple on a table from different angles - will not help the model adapt to various situations above.

Good luck!