ALIA-40B — Model Documentation

Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d


Model Identification

FieldDetail
Full model nameALIA-40B
Developed byBarcelona Supercomputing Center (BSC) — Language Technologies Unit
Release dateFebruary 2025
Model cardhuggingface.co/BSC-LT/ALIA-40b
Technical reportarXiv:2502.08489 — Salamandra Technical Report
Developer bloggithub.com/langtech-bsc/alia

Architecture and Parameters

FieldDetail
ArchitectureDense (decoder-only transformer)
Total parameters40.4B (40,433,885,184)
Active parameters per tokenN/A — dense model (all parameters active)
Non-embedding parameters38.3B (embedding parameters: 2.1B)
Layers48
AttentionGQA — 64 Q heads, 8 KV heads
Native context length32,768 tokens
Extended contextNot published by BSC
PrecisionBF16

License and Commercial Use

License: Apache License 2.0

Apache 2.0 is a permissive open-source license that allows free commercial use, distribution, and modification, provided that the original copyright notice and license text are retained and any modifications are documented.

Nextbit's verification: Nextbit has reviewed the license terms applicable to ALIA-40B and confirmed that serving this model via API under a commercial inference service is permitted under Apache 2.0. No separate commercial agreement with the Barcelona Supercomputing Center is required for this use case.

Restrictions relevant to users: None specific to this model beyond standard Apache 2.0 terms. This is a base model (not instruction-tuned); BSC strongly recommends instruction-tuning or alignment fine-tuning before production deployment, as the base model may produce outputs that are inappropriate, biased, or unsafe. Users of Nextbit's API who integrate model outputs into their own products remain responsible for their own compliance with applicable law and any downstream licensing requirements.


Training Data Summary

The following is a summary of publicly available information provided by BSC regarding training data for the ALIA/Salamandra model family, as documented in the Salamandra Technical Report (arXiv:2502.08489).

Pretraining data

The ALIA-40B model was pretrained on approximately 9.37 trillion tokens (with a high-quality final-phase subset of 160 billion tokens and a context-extension phase of 6.3 billion tokens).

Training data sources include:

SourceApproximate share
Colossal OSCAR (web crawl)53.05%
StarCoder (programming code)13.67%
FineWeb-Edu (educational web content)10.24%
HPLT v1/v1.1 (high-quality text)4.21%
French-PD (public domain books/newspapers)3.59%
MaCoCu, Legal-ES, EurLex~1.4–1.7% each
40+ additional curated sourcesremainder

The additional curated sources include legal, scientific, and domain-specific datasets. Training was conducted in approximately 3.68 epochs over the combined corpus.

Language coverage

The model was trained on data in 35 European languages (including co-official Spanish languages) and 92 programming languages. Key language distribution:

LanguageShareNotes
English39.31%
Spanish16.12%Upsampled 2×
French6.6%
Russian5.56%
German4.79%
Hungarian4.59%
Code (all languages)5.78%Downsampled 0.5×
CatalanUpsampled 2×
BasqueUpsampled 2×
GalicianUpsampled 2×

Knowledge cutoff

Data was collected from April 2023 to April 2024. Some sources include data from as early as 2014. BSC has not published a formal knowledge cutoff date; based on collection periods, the effective knowledge cutoff is approximately April 2024.

Training infrastructure

ALIA-40B was trained on the MareNostrum 5 supercomputer at the Barcelona Supercomputing Center, using 256–512 nodes (1,024–2,048 NVIDIA Hopper GPUs), with the NVIDIA NeMo Framework.

Project funding

The ALIA project received funding from the Government of Catalonia (Aina Project) and the Spanish Ministry for Digital Transformation and Public Service via NextGenerationEU (ILENIA Project, reference: 2022/TL22/00215337).

What is not publicly available

BSC has not published:

  • Exact per-source token counts for all 40+ data sources
  • Training compute (FLOPs or GPU-hours)
  • Data deduplication methodology in full detail

Languages Supported

ALIA-40B natively supports 35 European languages and 92 programming languages. The model is specifically optimized for Iberian languages (Spanish, Catalan, Basque, Galician) and other major European languages (French, German, Russian, Hungarian, Italian, Portuguese, and others).

Full language list: see the Salamandra Technical Report and the model card.


Intended Uses

This model is designed for:

  • Multilingual natural language processing tasks, with a focus on European and Iberian languages
  • Text generation in Spanish, Catalan, Basque, Galician, and other European languages
  • Serving as a base for further instruction-tuning and specialization
  • Research and academic applications in European NLP

Important: ALIA-40B is a base model without instruction tuning or safety alignment. It is not intended for direct deployment in consumer-facing applications without further fine-tuning. BSC explicitly warns that the model may generate inappropriate, biased, or misleading content.

Explicitly excluded or cautioned uses (per BSC documentation):

  • Direct deployment in production systems without alignment fine-tuning
  • Applications requiring strict factual accuracy or safety guarantees without additional safeguards

Systemic Risk Assessment

FieldDetail
Training FLOPsNot published by BSC
Estimated FLOPs (dense model)~7 × 10²³ (estimate based on 40.4B parameters × ~9.37T tokens; formula: 6 × N × D)
Exceeds 10²⁵ FLOPs threshold?No — estimated to be approximately 14× below the systemic risk threshold
Note on dense architectureALIA-40B is a dense model; all 40.4B parameters are active per token. The full parameter count is used for FLOPs estimation.
AI Office designationNot designated as a systemic risk model as of May 2026
Art. 55 obligations apply?No

This estimate is based on publicly available information (parameter count from model card; token count from Salamandra Technical Report). BSC has not published official training compute figures. If such figures are published, this section will be updated.


This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. All training data and architectural information is sourced from public documentation published by the Barcelona Supercomputing Center (BSC) Language Technologies Unit.

For questions: [email protected]

Was this page helpful?