The corpus is the constraint
Pretraining a competent national-language model is a data problem before it is a compute problem. Open web crawls yield Polish in a fraction of the volume they yield English, much of it duplicated or machine-translated, so the work sits in curation: deduplication at scale, quality classification, OCR of archives and legal corpora, and negotiated access to press and publishing back-catalogues. Tokenisers fitted principally on English fragment Polish inflection and inflate tokens per word, taxing training and inference alike; refitting or extending the vocabulary is now standard practice, and forces the second decision — continued pretraining from an English-centric base, or from scratch. Synthetic and translation-augmented data close part of the gap at the cost of translationese and a narrowing distribution.



