Artificial intelligence can identify a famous monument, summarise a Western novel or generate an image of a major city with impressive ease. It often becomes far less reliable when it encounters complex calligraphy, a local reference, an under-documented language or a heritage object missing from large datasets.

This gap is not merely a technical problem. It affects how cultures are described, translated, searched and represented in digital tools.

Models learn from what they are shown

An artificial intelligence model is trained on enormous quantities of text and images. When certain languages, regions or traditions are extensively documented, the model has many examples from which to learn their forms and contexts.

By contrast, a culture that has been poorly digitised or inadequately described appears only in fragments. The model may then produce vague answers, confuse references or fill missing information with stereotypes.

The problem is therefore not an abstract inability to understand a culture. It comes from imbalances in data, annotation practices and research priorities.

Arabic calligraphy as a test case

DuwatBench, a benchmark devoted to recognising Arabic calligraphy, helps measure this difficulty. It contains 1,272 samples covering six calligraphic styles, including Thuluth, Diwani, Naskh and Kufic.

Researchers show that even high-performing multimodal models can make numerous errors when writing becomes an art form. Letters overlap, words change scale and decorative elements complicate reading.

This reveals an important limitation: recognising printed characters does not mean understanding a visual tradition. Calligraphy combines language, composition, history and symbolism.

Visual heritage remains under-documented

Turath-150K was created to address the shortage of properly documented images of Arab heritage. It gathers representations of monuments, objects and cultural sites in order to improve visual search and recognition.

But building a dataset is not simply a matter of collecting files. Images must be described, rights checked, locations identified, historical periods specified and linguistic variants included.

Cultural institutions therefore have a strategic role. Museums, libraries, archives, universities and associations can become producers of high-quality cultural data, provided they have the necessary resources and an ethical framework.

One answer does not fit every culture

CrossCult-KIBench evaluates how well models adapt their answers across several linguistic and cultural contexts. Its findings show that current methods can improve one dimension while weakening another.

A model may learn to avoid some stereotypes but become less precise. It may produce a culturally appropriate answer in one language and fail in another.

This calls for moving beyond a simplistic view of localisation. Translating an interface or adding a few local examples is not enough. Tools must be tested with users, researchers and creators from the contexts concerned.

What are the risks for cultural industries?

An AI system that poorly understands a culture can make mistakes in museum interpretation, translation, content recommendation, education, tourism or image generation.

It can also make works less visible by indexing them incorrectly. A search engine unable to recognise a style, artist or language reduces their discoverability.

These issues complement those discussed in our analysis of AI and the protection of creative work. Pay, consent and rights are essential, but the quality of cultural representation matters too.

Building digital cultural sovereignty

Reducing these biases requires several actions: digitising collections, developing multilingual corpora, funding local research, documenting sources and involving communities in the description of their heritage.

It is also important to prevent cultural data from simply being extracted by large platforms without any benefit for the institutions that produced it.

Digital cultural sovereignty does not mean closing data. It means defining the conditions of access, rights, responsibilities and acceptable uses.

AI as a mirror

The current limits of AI function like a mirror. They reveal which cultures have been extensively digitised, which remain absent and who owns the infrastructure needed to turn heritage into data.

The response cannot therefore be technological alone. It must also be cultural, political and institutional.

Sources