research
An open Portuguese language model, trained from scratch, with every data source on the record.
Laborator.io is an independent lab in Brazil. We build the corpus, the tokenizer, the training code and the evaluation ourselves, and we will release weights, code and data documentation openly. This page is the state of the project as measured on the date above, including what is not done yet.
- 2.60Ttokens in the corpus manifest, estimated per source
- 379sources, each with license, origin and hash
- 1.76Ttokens of Portuguese, 67.6% of the corpus
- 13.7%of benchmark items we found inside Portuguese web data
01 corpus
Mostly Portuguese, by design.
| block | tokens | share | scale |
|---|---|---|---|
| Portuguese | 1,760.6B | 67.6% | |
| Mathematics | 294.3B | 11.3% | |
| English | 280.9B | 10.8% | |
| Code | 156.7B | 6.0% | |
| Biomedical | 97.8B | 3.8% | |
| 16 other languages | 13.2B | 0.5% |
The Portuguese block is mostly filtered web: Common Crawl processed by us snapshot by snapshot, plus HPLT, FineWeb-2, ClassiCC-PT, CulturaX and FinePDFs. The rest is public-domain legislative and judicial text (Câmara, Senado, STJ, Diário Oficial), open-access theses and journals (SciELO, RCAAP, CAPES), and books.
What is missing is long-form curated Portuguese. In our last full classification (23/08, 1.82T-token corpus), 97.3% of the Portuguese was web crawl, 2.4% administrative and legal acts, and only 0.29% books, journals, encyclopedias and theses (2.8B tokens). That gap, not volume, is the constraint we are working on, through licensing agreements with libraries, university presses and institutes.
Frozen release. labpt1t was frozen automatically on 25/08/2026 when Portuguese crossed one trillion tokens: 1.85T tokens over 165 sources, 1.58T of them redistributable, 352,469 shards, with a SHA-256 of the file listing. Every number we publish is measured on a frozen release, not on the live corpus.
Token counts in the manifest are per-source estimates. Our own re-measurements have found per-source errors of up to 7x, and a language map that silently missed 44 sources. That is why totals are re-measured on each frozen release and why the method documents every number that was later withdrawn.
02 provenance
No byte without a manifest row.
Each source has one row in data/manifest.csv: license, commercial use, copyleft, redistributable, collection date, document count, token estimate and hash. The preparation step refuses any source that is missing from the manifest or marked verify. This is enforced in code, not in a policy document.
| license | sources | share |
|---|---|---|
| ODC-BY 1.0 | 19 | 28.1% |
| Common Crawl terms of use | 246 | 25.1% |
| CC0 | 15 | 10.6% |
| Public domain (Brazilian law) | 24 | 9.4% |
| NVIDIA data agreement | 3 | 7.8% |
| Inherited from mC4 / OSCAR | 6 | 5.9% |
| CC BY-SA 3.0 and 4.0 | 42 | 6.9% |
| Blue Oak permissive | 2 | 3.0% |
| Not declared upstream | 6 | 2.7% |
| CC BY, Apache 2.0, other | 15 | 0.5% |
- 83.6% of tokens come from redistributable sources. The rest (NVIDIA math data, CulturaX, ClassiCC-PT) can be trained on but not redistributed, and is kept apart in every release.
- Copyleft is tracked, not hidden: 6.9% of tokens are CC BY-SA and carry that flag through the pipeline.
- Excluded by decision: output from commercial LLMs, and newspaper archives whose terms forbid AI training by name.
- Brazilian law is stricter than the EU here. Law 9.610/98 has no text-and-data-mining exception, and the pending AI bill (PL 2338) creates one that excludes commercial use. Our audit is built for that environment.
Where we are not yet at the standard we want
We do not yet apply robots.txt opt-outs retroactively. For web data we rely on Common Crawl honoring robots.txt at crawl time, which is weaker than filtering against today's opt-outs. We also do not yet run an opt-out and PII-removal channel of our own; it is planned for the model card. Both are on our list before any weight release, and this is where we would most like to learn from others.
03 pipeline
Filters that fail loudly.
- Normalization: Unicode NFC, line endings, whitespace.
- PII masking: Brazilian CPF (only when both check digits validate), e-mail and phone, with rules to avoid false positives on dates, versions and code.
- Quality and language: per-document acceptance filters and language identification, with a mapping of every source to a language block that stops the run if any source is unmapped.
- Deduplication: exact and near-duplicate, within each Common Crawl snapshot and not across snapshots, following the FineWeb finding that global dedup yields a worse model. Across snapshots it would have cut the corpus from about 4T to 1.5T tokens.
- Silent-failure guards: most of our errors produced no error message, so the tools now stop with a non-zero exit when an assumption breaks (unmapped source, missing category, incomplete checkpoint) instead of warning and continuing.
04 tokenizer
lab64k, 64,000 pieces.
| tokenizer | vocabulary | tokens per word | vs lab64k |
|---|---|---|---|
| Tucano 2B4 | 32,000 | 1.48 | −14.8% |
| lab64k | 64,000 | 1.73 | base |
| SmolLM3 3B | 128,256 | 1.90 | +9.4% |
| Qwen3 1.7B | 151,669 | 1.97 | +13.6% |
Among open multilingual tokenizers, lab64k spends the fewest tokens on Portuguese with half the vocabulary of SmolLM3. Tucano, trained only on Portuguese, spends fewer still. lab64k is weaker on code and English, which is where we expect to improve it.
1,084 held-out Portuguese documents (Wikipedia and Querido Diário), about 5 million characters, with line breaks normalized. With original line breaks the advantage over SmolLM3 and Qwen3 drops to +2.9% and +6.8%. Measured 15/09/2026.
05 evaluation
Brazilian exams are already in the web crawl.
Before evaluating anything we scan the training data for benchmark items: 13-gram matching in the style of GPT-3, needles taken from the middle of each question so that exam boilerplate does not create false positives, one pass over the whole processed corpus. The full scan (6 to 7 September, 945 shards) found 4,273 of 31,145 items (13.7%).
| benchmark | language | found / tested | share | scale |
|---|---|---|---|---|
| ENEM | pt | 1,314 / 1,432 | 91.8% | |
| OAB exams | pt | 1,380 / 2,037 | 67.7% | |
| BlueX | pt | 423 / 717 | 59.0% | |
| GSM8K | en | 298 / 1,319 | 22.6% | |
| MMLU | en | 746 / 9,828 | 7.6% | |
| ARC-Challenge | en | 49 / 865 | 5.7% | |
| TruthfulQA | en | 5 / 193 | 2.6% | |
| PT hate speech | pt | 3 / 553 | 0.5% | |
| HellaSwag | en | 51 / 9,588 | 0.5% | |
| WinoGrande | en | 4 / 1,267 | 0.3% | |
| GSM8K-PT, ASSIN2, HateBR | pt | 0 / 3,709 | 0.0% |
ENEM, OAB and BlueX are the benchmarks most Portuguese models report, and most of their items circulate on the Portuguese web as study material. Any model trained on a large Portuguese crawl without this scan is likely reporting memorization on them. The English contamination comes mostly from math datasets (FineMath, Nemotron-Math).
Contaminated items are removed from the training data, not only from the evaluation. We hold 15 benchmarks locally (the ones above plus ENEM Challenge and OAB Exams). Multiple-choice items are scored by log-likelihood, raw and normalized by answer length, with binomial error; two more tests are generated by rule, a Portuguese factual-recall test and a Portuguese version of iGSM. So far the suite has only been run on our small models.
06 training
Small scale so far, measured carefully.
- Decoder written from scratch: RMSNorm, RoPE, SwiGLU, grouped-query attention, QK-norm, Muon optimizer with AdamW for embeddings.
- First models: 48M and 155M parameters (July and August), trained on early versions of the corpus.
- Ablations at 400M: seven variants on the same data order. Muon beat the AdamW reference by 0.31 nats; depth-versus-width geometry made no measurable difference; intra-document attention masking won per token but lost 33% of throughput.
- Wind tunnel at 125M: 14 architecture variants with paired comparisons against a seed-noise ruler. Only effects replicated in two independent runs are counted: value residual (−0.106 nats) and attention gating (−0.058 nats).
- First 663M run reached about 9B tokens on two GPUs before a checkpoint transfer failure cost us the weights. The failure, the fix and the verification that now runs before any checkpoint is marked complete are part of the published method.
Next: a dense model in the 3B class, sized by tokens per parameter rather than by headline size, on the frozen corpus. The main open question is the mixture: how much English, math and code a Portuguese-first model needs to keep reasoning ability.
07 working together
What we can share.
- The redistributable Portuguese slices with their per-source audit.
- Contamination lists for Portuguese and English benchmarks, by item.
- The Portuguese evaluation suite and the rule-generated tests.
- The method: every decision with its evidence, and a catalogue of failures that produced no error.
The repository is private until the first weight release. Partners get access to the manifest, the method and the tools on request.