Why not all products are found using a barcode scanner - and where our data comes from

In this article, we give you a look behind the scenes at how we process barcodes and where our product data comes from.

How does a barcode work?

A barcode, also known as a stripe code, is a machine-readable identification number. It makes products identifiable for commercial use and retail - ideally unique on a global scale.

The first barcode (Universal Product Code, abbreviated UPC) was introduced in 1973 in the USA. The commonly used UPC-A in the USA consists of 12 digits. It quickly becomes apparent that, aside from the country and the manufacturing company, not much information is contained within those 12 digits.

An EAN13 barcode - the machine-readable part on top, below the encoded 13-digit identification number. No more information is contained directly in the barcode. By VaGla - own work created in Inkscape based on the graphics by Grzexs, CC BY-SA 3.0

In Europe, three years later in 1976, the European Article Number (abbreviated EAN) was introduced. It is 13 digits long and compatible with the UPC system.

Identification scheme since 2015

Since 2015, the identification numbers used worldwide in commerce have been renamed to Global Trade Item Number (GTIN). It is 8, 12, 13, or 14 digits long and always contains a check digit to detect errors during machine reading.

How does one access the product data?

GS1 Germany GmbH based in Cologne is the only official provider of EAN8/EAN13 codes in Germany. Anyone who wants to sell a product with a barcode must purchase a unique barcode from there.
Up to 30 barcodes can be retrieved for free per day from there - however, without all the relevant sizes that are of interest for the storage of food. For example, the product name, the package size, and the nutritional values are missing.

As a test, I used the official barcode search of GS1 Germany called Gepir to look up a randomly selected item: A pack of Haribo from an Edeka in Munich. The barcode 8426617106201 immediately reveals the country that issued the barcode: The two digits on the left stand for Spain (84).

The search in Gepir yields the following result: Company name "HARIBO ESPAÑA S.A.U." as well as an address.

Unfortunately, we can't do much with this yet. Commercial use of the service is costly and does not provide us with the data we are interested in. At a minimum, these would be: name, quantity, nutritional values of the food, and information on any allergens it may contain.

The first pantry app with product data: Crowdsourcing

When we launched the first pantry app in the App Store in 2013, there wasn't even a barcode scanner - and accordingly, no product data was stored.
Each user had to enter their items manually, which was quite laborious. In December 2015, the time had finally come: We created a product database for the pantry, with which users still had to enter food items, but these were then shared among all users! This way, at least common items were usually already in the database, and the recording of stock or shopping became much faster.

However, over time, weaknesses of the approach have become apparent:

Thus, there was only one field per article for the product name - even though the app was also used internationally. People then filled in the name sometimes in Swedish, sometimes in English, sometimes in German.

Although the app had storage locations, it seems that these were not sufficient for organizing items for some users. Thus, cryptic category names like "30 Juice", "31 UHT Milk" quickly emerged, which were only understandable to the respective user.

Some users have meticulously transferred all nutritional values - others have completely skipped this part. The level of detail of the products varied accordingly. Excerpt from the product database of the Pantry App, where product names have been entered by different users in various languages.

The current Pantry WebApp

Before we released our current WebApp for public beta in 2019, we spent a long time searching for the right model regarding the product database. We considered how we could improve the quality of the existing data.

Ultimately, we have actively decided to stop using our own product database and instead switch to the open model of OpenFoodFacts.org. Here too, product data is maintained by volunteer users and stored in a publicly accessible database. However, the project has several advantages, as we will see in the following section.

Products can be accessed directly from OpenFoodFacts and used for one's own purposes. The product data is licensed under the Open Database License - this ensures the data can be used for any purpose - as long as newly added product data is contributed back - thus the crowd-sourcing effect is ensured through the license.

A few facts about OpenFoodFacts

  • OpenFoodFacts was founded in France in May 2012
  • Meanwhile, there are 1,875,095 product data from all over the world listed with barcodes

Compared to our first database, OpenFoodFacts has a very extensive data schema. For instance, name fields are provided for any language.

Additionally, OpenFoodFacts allows for the uploading of photos of the product, the list of ingredients, and the nutritional value table. This enables automated machine quality control of the entered nutritional information, significantly improving the quality of the data.

Everything that can be found on OpenFoodFacts about the bag of Haribo from the example above can be found here.

Pantry App and OpenFoodFacts

Since October 7, 2019, we have been retrieving product data from OpenFoodFacts and, of course, also contributing back for the benefit of all users.

Since then, our users have created 10,184 articles from scratch and edited 20,441 articles, mostly to add missing attributes.

With this, our users have created 0.5% of all articles in the global database - that's an impressive sum, from which users in Germany benefit especially! A heartfelt thank you at this point for the careful entry of data, which will be very conveniently available to all future users.

Although it may happen that individual products are not found when scanning, we believe in the approach: A product database maintained by users is independent of the interests of food manufacturers, can be expanded as needed, and is unrestrictedly usable for all future purposes.

So if you come across an article that is not found, and you have to enter it yourself - think of the many users who will also scan the article, they will thank you for your product entry! 😊