{"title": "PUMA: Polish Unified Multimodal Assessment", "description": "We present PUMA (Polish Unified Multimodal Assessment), a benchmark designed to evaluate the linguistic and cultural competencies of large models in the context of Poland and the Polish language. PUMA consists of 900 hand-crafted questions designed to assess a wide spectrum of multimodal competencies. It is divided into three modalities: images, audio, and documents. Each modality contains three categories targeting distinct competencies or knowledge domains. Therefore, there are nine categories in total, with 100 questions per category. Six of these categories focus on question answering (QA), evaluating the models' knowledge and reasoning abilities in multimodal contexts. We also included three specialized categories testing the models' practical competencies in typical vision and audio tasks: automatic speech recognition (ASR), optical character recognition (OCR), and document structure extraction. The benchmark uses deterministic, rule-based verification process, without external judge.
We evaluated models that support all required modalities, as well as VLMs that support only image and document categories. Results for multimodal models and pairs of models from the same family (e.g. GPT-5.5 + GPT-Audio) are available in the Full benchmark. Results for vision models are available in the Vision only benchmark.", "categories": "
Images | \n History and culture - This category tests the models' knowledge regarding the history, tradition, cultural heritage, and customs of Poland. It also includes art, primarily painting, sculpture, and architecture. Some of the questions concern people or objects of significant historical importance. | \nContemporary life - It covers contemporary life and pop culture. Some of the questions involve modern media such as film, television, and the internet. We have also included questions related to sports, contemporary politics, and show business. Cultural and social issues are also addressed, provided they relate to contemporary topics. | \nGeography and environment - This category verifies the models' knowledge of Poland's geography, including its fauna, flora, and well-known natural landmarks. Questions about man-made structures may also be found in this category. These are primarily related to infrastructure, cities, as well as administrative and socioeconomic issues. | \n
Audio | \n Automatic Speech Recognition (ASR) - This is a specific category designed to test models' ability to transcribe speech, a task of significant practical importance. The recordings used in this category range from about a minute to several minutes in length. For the most part, these are challenging samples containing interference, background noise, complex language, or dialects. | \nSpeech QA - It evaluates advanced speech understanding and knowledge of the Polish cultural context based on audio recordings. This category includes questions about short audio samples lasting from about a minute to several minutes. The scope of tested competencies covers speech comprehension, information extraction, recognition of the speakers' intentions and emotions, as well as the ability to link the content of the recording with broader knowledge about Poland. | \nSound and music QA - This category includes questions about audio recordings where the main content is not speech, but other sounds such as ambient noises, musical instruments, animals, jingles, and the sounds of tools and machinery. All recordings are related to Poland, and the models' task is to recognize the sounds and use broader knowledge to interpret their content. | \n
Documents | \n Optical character recognition (OCR) - This category evaluates the models' ability to convert documents from image formats into plain text. The tasks mostly involve challenging cases, such as handwritten text or low-quality scans. Additionally, some examples contain tables. The metric we use verifies the correct extraction of tables while preserving their structure. | \nDocument QA - This category tests the models' ability to understand and interpret documents in Polish by answering questions about their content. Most of the examples in this category consist of visually rich documents that require the analysis of text, tables, charts, or infographics, among others. | \nStructured extraction - This category evaluates the models' ability to extract relevant information from documents and save it in a structured JSON format. Each example in this category includes a document in the form of an image and defines a JSON schema that specifies the expected format of the model's response. | \n