ALIA-40B — Model Documentation
Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d
Model Identification
| Field | Detail |
|---|---|
| Full model name | ALIA-40B |
| Developed by | Barcelona Supercomputing Center (BSC) — Language Technologies Unit |
| Release date | February 2025 |
| Model card | huggingface.co/BSC-LT/ALIA-40b |
| Technical report | arXiv:2502.08489 — Salamandra Technical Report |
| Developer blog | github.com/langtech-bsc/alia |
Architecture and Parameters
| Field | Detail |
|---|---|
| Architecture | Dense (decoder-only transformer) |
| Total parameters | 40.4B (40,433,885,184) |
| Active parameters per token | N/A — dense model (all parameters active) |
| Non-embedding parameters | 38.3B (embedding parameters: 2.1B) |
| Layers | 48 |
| Attention | GQA — 64 Q heads, 8 KV heads |
| Native context length | 32,768 tokens |
| Extended context | Not published by BSC |
| Precision | BF16 |
License and Commercial Use
License: Apache License 2.0
Apache 2.0 is a permissive open-source license that allows free commercial use, distribution, and modification, provided that the original copyright notice and license text are retained and any modifications are documented.
Nextbit's verification: Nextbit has reviewed the license terms applicable to ALIA-40B and confirmed that serving this model via API under a commercial inference service is permitted under Apache 2.0. No separate commercial agreement with the Barcelona Supercomputing Center is required for this use case.
Restrictions relevant to users: None specific to this model beyond standard Apache 2.0 terms. This is a base model (not instruction-tuned); BSC strongly recommends instruction-tuning or alignment fine-tuning before production deployment, as the base model may produce outputs that are inappropriate, biased, or unsafe. Users of Nextbit's API who integrate model outputs into their own products remain responsible for their own compliance with applicable law and any downstream licensing requirements.
Training Data Summary
The following is a summary of publicly available information provided by BSC regarding training data for the ALIA/Salamandra model family, as documented in the Salamandra Technical Report (arXiv:2502.08489).
Pretraining data
The ALIA-40B model was pretrained on approximately 9.37 trillion tokens (with a high-quality final-phase subset of 160 billion tokens and a context-extension phase of 6.3 billion tokens).
Training data sources include:
| Source | Approximate share |
|---|---|
| Colossal OSCAR (web crawl) | 53.05% |
| StarCoder (programming code) | 13.67% |
| FineWeb-Edu (educational web content) | 10.24% |
| HPLT v1/v1.1 (high-quality text) | 4.21% |
| French-PD (public domain books/newspapers) | 3.59% |
| MaCoCu, Legal-ES, EurLex | ~1.4–1.7% each |
| 40+ additional curated sources | remainder |
The additional curated sources include legal, scientific, and domain-specific datasets. Training was conducted in approximately 3.68 epochs over the combined corpus.
Language coverage
The model was trained on data in 35 European languages (including co-official Spanish languages) and 92 programming languages. Key language distribution:
| Language | Share | Notes |
|---|---|---|
| English | 39.31% | — |
| Spanish | 16.12% | Upsampled 2× |
| French | 6.6% | — |
| Russian | 5.56% | — |
| German | 4.79% | — |
| Hungarian | 4.59% | — |
| Code (all languages) | 5.78% | Downsampled 0.5× |
| Catalan | Upsampled 2× | — |
| Basque | Upsampled 2× | — |
| Galician | Upsampled 2× | — |
Knowledge cutoff
Data was collected from April 2023 to April 2024. Some sources include data from as early as 2014. BSC has not published a formal knowledge cutoff date; based on collection periods, the effective knowledge cutoff is approximately April 2024.
Training infrastructure
ALIA-40B was trained on the MareNostrum 5 supercomputer at the Barcelona Supercomputing Center, using 256–512 nodes (1,024–2,048 NVIDIA Hopper GPUs), with the NVIDIA NeMo Framework.
Project funding
The ALIA project received funding from the Government of Catalonia (Aina Project) and the Spanish Ministry for Digital Transformation and Public Service via NextGenerationEU (ILENIA Project, reference: 2022/TL22/00215337).
What is not publicly available
BSC has not published:
- Exact per-source token counts for all 40+ data sources
- Training compute (FLOPs or GPU-hours)
- Data deduplication methodology in full detail
Languages Supported
ALIA-40B natively supports 35 European languages and 92 programming languages. The model is specifically optimized for Iberian languages (Spanish, Catalan, Basque, Galician) and other major European languages (French, German, Russian, Hungarian, Italian, Portuguese, and others).
Full language list: see the Salamandra Technical Report and the model card.
Intended Uses
This model is designed for:
- Multilingual natural language processing tasks, with a focus on European and Iberian languages
- Text generation in Spanish, Catalan, Basque, Galician, and other European languages
- Serving as a base for further instruction-tuning and specialization
- Research and academic applications in European NLP
Important: ALIA-40B is a base model without instruction tuning or safety alignment. It is not intended for direct deployment in consumer-facing applications without further fine-tuning. BSC explicitly warns that the model may generate inappropriate, biased, or misleading content.
Explicitly excluded or cautioned uses (per BSC documentation):
- Direct deployment in production systems without alignment fine-tuning
- Applications requiring strict factual accuracy or safety guarantees without additional safeguards
Systemic Risk Assessment
| Field | Detail |
|---|---|
| Training FLOPs | Not published by BSC |
| Estimated FLOPs (dense model) | ~7 × 10²³ (estimate based on 40.4B parameters × ~9.37T tokens; formula: 6 × N × D) |
| Exceeds 10²⁵ FLOPs threshold? | No — estimated to be approximately 14× below the systemic risk threshold |
| Note on dense architecture | ALIA-40B is a dense model; all 40.4B parameters are active per token. The full parameter count is used for FLOPs estimation. |
| AI Office designation | Not designated as a systemic risk model as of May 2026 |
| Art. 55 obligations apply? | No |
This estimate is based on publicly available information (parameter count from model card; token count from Salamandra Technical Report). BSC has not published official training compute figures. If such figures are published, this section will be updated.
This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. All training data and architectural information is sourced from public documentation published by the Barcelona Supercomputing Center (BSC) Language Technologies Unit.
For questions: [email protected]