Sunday, September 6, 2026

News

Bielik's Creators Are Teaching Polish AI to Recognize Roadside Shrines and Silesian Dumplings

PolandPatryk Raba
Bielik's Creators Are Teaching Polish AI to Recognize Roadside Shrines and Silesian Dumplings
Fot. Martin Mecnarowski, Wikimedia Commons (CC BY-SA 3.0)

The team behind Bielik, Poland's homegrown language model, is collecting photos of monuments, dishes and roadside shrines from mobile app users to teach its multimodal model to recognize Polish cultural context that global AI systems still get wrong.

Contents
  1. A Photo Instead of Code
  2. The Problem With Global Models
  3. The Risk of a New Stereotype
  4. Technological Sovereignty
  5. What's Next

The team behind Bielik, Poland's homegrown language model, admits that artificial intelligence can conjugate Polish nouns flawlessly and build grammatically correct sentences, yet still has no idea what separates a pork chop (schabowy) from a Wiener schnitzel. As part of the Obywatel Bielik (Citizen Bielik) project, the creators have spent more than a year collecting photos of Polish monuments, dishes and religious symbols from the public, aiming to teach the model to recognize what grammar alone can't capture.

A Photo Instead of Code

The mechanics of the project are simple. A user of the mobile app, available since July 2026, photographs an element of Polish everyday life they know well: a roadside shrine, a regional dish, a historic monument or a distinctive landscape, and adds a short description to the photo. The system first automatically checks whether the material violates the terms of use, and only then does it go through manual moderation by volunteers.

Everyone can easily take part in building Polish artificial intelligence. Just go for a walk, take a photo of a roadside shrine, a regional dish, a monument or a place we know well, and send it along with a short description from your phone - Marcin Dąbrowski, leader of the Obywatel Bielik project

The Problem With Global Models

Marcin Dąbrowski traces the project back to an observation from several years ago: when the team began work on Bielik AI in late 2022, the language models available at the time relied almost exclusively on foreign data, with Polish sources making up only a small fraction of the overall training base. The result is a set of specific, recurring errors that Bielik's creators cite as examples.

Global language models can confuse kluski śląskie (Silesian dumplings) with donuts, or a Wiener schnitzel with a schabowy (breaded pork chop). Sometimes they don't recognize the symbol of Polska Walcząca (Fighting Poland) at all, or describe it incorrectly - Marcin Dąbrowski, leader of the Obywatel Bielik project

Dąbrowski stresses that this isn't just about the team's Polish origin, but about the actual use of Polish sources and texts in training. Bielik is meant to be taught to "think" in Polish, rather than merely translating sentences from English into Polish while retaining a foreign cultural logic.

The Risk of a New Stereotype

The project has also drawn a critical voice, though not one that questions the idea itself. Leszek Chodorowski, founder of the Polska Szkoła AI (Polish School of AI), notes that the sheer number of collected photos means little if the database isn't diverse, and that results can only be judged once it becomes possible to measure whether the model actually makes fewer mistakes in Polish context.

The biggest risk would be replacing the global stereotype of Poland with our own stereotype of Poland - Leszek Chodorowski, founder of the Polska Szkoła AI

Technological Sovereignty

Dr Jan Majewski of the Department of Digital Sociology at the University of Warsaw's Faculty of Sociology comments on the broader context of the project, placing SpeakLeash's initiative within the debate over Poland's and Europe's technological independence from large foreign language model providers.

From the very beginning, AI systems have been skewed toward the English and Chinese languages and their cultures - dr Jan Majewski, Department of Digital Sociology, University of Warsaw

Majewski notes that building such independence can't be limited to hardware, meaning owning computing centers, but must also extend to aligning the AI systems themselves with local values. He points out that SpeakLeash plans to release the project's results as open source, which sets it apart from the closed databases of commercial providers.

What's Next

The project team says it is already in talks with local government representatives interested in joining as honorary patrons, which could speed up the collection of material from smaller towns that are underrepresented online. The database is ultimately meant to reach one million photos, with the data intended not only to improve chatbot responses but also to support use cases in e-commerce, document analysis and work with archival materials. For Polish companies weighing local language models over foreign solutions, this is a practical test of how much work it takes to close the cultural-context gap, as opposed to simply mastering the language.

Share: