Skip to content

Every fortress needs its tower

A Polish AI safety model for detecting harmful content.

Baszta-1 was built for the Polish language and real-world AI safety challenges. In the whitepaper, we show how we built it, what worked, what did not, and what we learned along the way.

Model
Baszta-1
Params
124M
Safety categories
5 classes: hate speech, vulgarity, sexual content, crime, self-harm
Training data
26K+ examples
Released
September 2026

Manifest

Poland cannot just use AI. It also needs the ability to build its own models and safeguards.

Artificial intelligence is changing how we work, how we learn and how we decide. As AI advances, organisations also need safeguards that let them use it responsibly while effectively managing the risks it introduces.

As regulations such as the AI Act come into force, AI safety is becoming a foundation of responsible deployment.

We did not want Poland to be only a consumer of foreign AI safety solutions. So we set out to build a model of our own.

Baszta grew out of a need to build AI expertise in Poland.

We focused on our strengths: research and legal expertise, engineering experience, and a deep understanding of the Polish language. Rather than compete on size with the largest models, we specialised.

Our goal was never to produce one more classifier. We wanted to show that Poland can build its own specialised AI solutions - including in AI safety.

We are sharing the process, not just the result.

We are not showing only the successes. We are also showing the experiments that led nowhere, the limits of the model, and the assumptions we had to revise along the way.

We believe AI advances faster when teams share what they learn. When a Polish AI model succeeds, the whole ecosystem benefits. By sharing what we learned, we hope to help other teams build faster, make better decisions, and develop safer systems. Baszta is our contribution to building Polish expertise in AI safety.

Adam GórskiMateusz JąkalakRafał JakubowskiKrystian KoziełPaweł Wcisło

  • Data beat a bigger model

    The largest gain came from fitting the data better, not from another change to the architecture. Synthetic examples reflecting new forms of harmful content helped Baszta handle cases it had not seen before.

  • A safe model should know when it is uncertain

    We found that a model can classify threats correctly while still being too confident when it encounters unfamiliar data. That is why calibration became one of the central parts of Baszta.

  • From lab results to the real world

    We evaluated Baszta both on familiar benchmarks and on new scenarios it had not seen before, to understand where it performs well and where it still needs improvement.

Benchmarks

Baszta-1 against other guard models

We compare Baszta with HerBERT-PL-Guard, Sójka and Qwen3Guard on public safety benchmarks for Polish.

KLEJ CBD

0.669F1

+21.0 pp HerBERT-PL-Guard+23.6 pp Sójka+40.7 pp Qwen3Guard

ModelF1FPR
Baszta-10.66910.9%
HerBERT-PL-Guard0.45926.6%
Sójka0.43313.1%
Qwen3Guard0.2622.8%

Baszta posts the highest F1 in the comparison and the second-lowest FPR.

Gadzi Język — Baszta vs Sójka

0.929micro F1

MetricBaszta-1SójkaDifference
Micro F10.9290.903+2.6 pp
Macro F10.7120.782−7.0 pp

With both models tuned the same way, Baszta’s micro-F1 lead is statistically significant (95% CI: +0.4 to +4.9 pp; p = 0.011), while the macro-F1 difference is not.

Gadzi Język is heavily imbalanced — a naive always-crime baseline reaches 0.910 micro F1.

Gadzi Język — guard model comparison

0.906detection

−8.6 pp HerBERT-PL-Guard−7.9 pp Qwen3Guard+23.1 pp Sójka

ModelDetection rate
HerBERT-PL-Guard0.992
Qwen3Guard0.985
Baszta-10.906
Sójka0.675

Baszta detects far more harmful prompts than Sójka, but trails HerBERT-PL-Guard and Qwen3Guard.

Gadzi Język contains only harmful examples, so we report recall rather than F1.

HateCheck-PL

0.710F1

−7.6 pp HerBERT-PL-Guard−7.4 pp Sójka+7.4 pp Qwen3Guard

ModelF1FPR
HerBERT-PL-Guard0.78659.3%
Sójka0.78474.5%
Baszta-10.71057.8%
Qwen3Guard0.63638.4%

Baszta has lower F1 than HerBERT-PL-Guard and Sójka, but also a lower FPR than both.

BAN-PL

0.667F1

−9.5 pp HerBERT-PL-Guard−8.7 pp Sójka+40.4 pp Qwen3Guard

ModelF1FPR
HerBERT-PL-Guard0.76235.4%
Sójka0.75423.0%
Baszta-10.66785.5%
Qwen3Guard0.2637.8%

On BAN-PL, Baszta trails HerBERT-PL-Guard and Sójka on both F1 and false-positive rate.

PolyGuard-PL

0.562F1

−25.4 pp HerBERT-PL-Guard−20.4 pp Qwen3Guard+8.7 pp Sójka

ModelF1FPR
HerBERT-PL-Guard0.8167.0%
Qwen3Guard0.7668.5%
Baszta-10.56235.7%
Sójka0.47519.7%

Baszta leads Sójka by 8.7 pp F1, but remains behind HerBERT-PL-Guard and Qwen3Guard.

Categories — Baszta vs Sójka

CategoryBaszta-1SójkaDifference
Self-harm0.8120.759+5.3 pp
Crime0.9830.946+3.7 pp
Hate speech0.4900.478+1.2 pp
Sexual content0.4760.727−25.1 pp
Vulgarity0.8001.000−20.0 pp

Baszta scores higher on self-harm, crime and hate speech; Sójka’s clearest lead is on sexual content.

Training-data overlap

BenchmarkOverlapStatus
Gadzi Język0 / 520Unchanged
HateCheck-PL0 / 3,815Unchanged
PolyGuard-PL0 / 1,725Unchanged
KLEJ CBD149 / 1,000Duplicates removed
BAN-PL461 / 24,000Duplicates removed
PL-Guard899 / 900Set excluded
PL-Guard adversarial765 / 900Set excluded

Benchmarks were checked for exact duplicates against the data used to build the model. No exact matches were found in Gadzi Język, HateCheck-PL, or PolyGuard-PL. 149 of 1,000 KLEJ CBD records and 461 of 24,000 BAN-PL records were removed. PL-Guard was excluded because 899 of 900 examples were in the training data. The audit detects exact text matches only; it does not identify paraphrases, near-duplicates, or examples derived from a shared source.

Whitepaper

How Baszta was built - the data, the experiments, and the lessons.

The Baszta whitepaper covers the full process of building a Polish AI safety model - from data preparation and training experiments to out-of-domain evaluation and an analysis of the model’s limitations.

We do not focus only on the best result. We describe the cases where the model failed, the limits of the benchmarks, and the trade-offs we had to make while building the model.

Trust in AI does not come from promises of perfection. It comes from transparency about methods, data, and decisions.

1 / 35
  1. Page 1 of 35
  2. Page 2 of 35
  3. Page 3 of 35
  4. Page 4 of 35
  5. Page 5 of 35
  6. Page 6 of 35
  7. Page 7 of 35
  8. Page 8 of 35
  9. Page 9 of 35
  10. Page 10 of 35
  11. Page 11 of 35
  12. Page 12 of 35
  13. Page 13 of 35
  14. Page 14 of 35
  15. Page 15 of 35
  16. Page 16 of 35
  17. Page 17 of 35
  18. Page 18 of 35
  19. Page 19 of 35
  20. Page 20 of 35
  21. Page 21 of 35
  22. Page 22 of 35
  23. Page 23 of 35
  24. Page 24 of 35
  25. Page 25 of 35
  26. Page 26 of 35
  27. Page 27 of 35
  28. Page 28 of 35
  29. Page 29 of 35
  30. Page 30 of 35
  31. Page 31 of 35
  32. Page 32 of 35
  33. Page 33 of 35
  34. Page 34 of 35
  35. Page 35 of 35
Open the whitepaper full screen
Released
10 September 2026
Authors
Adam Górski, Mateusz Jąkalak, Rafał Jakubowski
Reviewers
Krystian Kozieł, Paweł Wcisło

Playground

Test Baszta-1 on Polish examples of your own.

Model
Baszta-1
CPU
0.5 vCPU
Memory
1 GB RAM
Region
Azure West Europe

Baszta-1Disconnected - the endpoint is not responding

Try one of these, or paste your own Polish text.

The submitted text is processed by the model on servers located in the European Union and deleted immediately after the classification result is returned.

Press Enter to classify. Use Shift + Enter to add a new line. Maximum 4000 characters.

Contact

Partnerships, deployments, research.

If you are interested in using Baszta in your AI system or discussing AI safety, get in touch.