News

A basic principle for the development of artificial intelligence is the use of a diverse, vast, and structured dataset. In a simple analogy, to teach someone about car models and their characteristics, it is essential for the person to see photos and data of both Volkswagens and Fords, red, black, and white vehicles. With enough experience and familiarity with the styles and trends of the manufacturers, it will be possible to say that an automobile belongs to a certain automaker even without having seen that model before. In machine learning training, the idea is similar to the one mentioned above. However, building these datasets can become a major challenge. Putting into perspective the development of computer vision, an area of knowledge that also applies AI techniques for data extraction from images, we can exemplify it as follows: If the algorithm needs to classify between trucks and cars, perhaps a handful of photos of each of these types of vehicles would be sufficient. You can have several photos of the same car model, of the same truck model, the characteristics are so different that the network will likely classify them relatively quickly and accurately. However, if the goal is to classify compact hatchbacks by color, manufacturer, and model, another level of dataset will be required. Probably hundreds of photos of black Chevrolet Onixes, silver VW Gols, white Ford Kas. A good sample of images of each model, from each automaker with their color variations. So the question remains, how to generate this dataset quickly and at competitive costs? Pix Force has been working collaboratively with its partners. The automotive example was not given in vain; together with Porto Seguro Seguradora (PS), an outstanding piece of work was developed to create a dataset. PS was interested in classifying specific characteristics in some car models, however, to recognize and classify these characteristics, great knowledge in the area was necessary. The traditional path would be for PS to teach Pix Force how to classify the points of interest and provide the database. However, something different was done: Pix Force trained the PS team to use its proprietary annotation software (the image data structuring step) and PS's own experts incorporated this step into their process for a few weeks. This was the fastest and most accurate way to build a structured database. Another example of cooperation in image annotation is RedSoft, a joint venture between Pix Force and iBeef. RedSoft classifies beef carcasses in slaughterhouses regarding their quality, which directly impacts their value. Additionally, the system tracks the meat throughout the processing. Similar to what was done with Porto Seguro, the team of animal scientists and veterinarians from iBeef, under the guidance of Pix Force, classified more than 100,000 carcass images, creating a robust dataset so that the algorithms can be trained and achieve extremely high precision. In the images below, it is possible to notice the evolution of automatic classification systems as the dataset becomes increasingly diverse, vast, and structured. A large part of AI implementation costs are linked to development, and in some cases, the construction of the dataset is significant. Thus, new methods that facilitate the generation of structured data are essential for expanding the use of AI and increasing the competitiveness of the national industry.
Speak with one of our specialists and discover how Pix Force can transform your business

Fabio Caraça
Fábio Caraça is the Chief Growth Officer at Pix Force. He leads Pix Force's transformation into a scalable SaaS operation, combining strategic vision, culture, and high-impact execution.


