ReMM SLERP L2 13B — Model Documentation

Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d


Model Identification

FieldDetail
Full model nameReMM-SLERP-L2-13B
Developed byUndi95 (merge author); base models in chain: Meta AI (Llama 2 13B), NousResearch (Nous-Hermes-Llama2-13b), jondurbin (airoboros-l2-13b-2.1), The-Face-Of-Goonery (Chronos-Beluga-v2-13b, Huginn-13b-v1.2)
Release date2023
Model cardhuggingface.co/Undi95/ReMM-SLERP-L2-13B
Technical reportNot published by Undi95
Developer bloggithub.com/Undi95/LLM-SLERP-MergeTest

Architecture and Parameters

FieldDetail
ArchitectureDense (decoder-only transformer) — Llama 2 architecture
Total parameters13B
Active parameters per tokenN/A — dense model (all parameters active)
Non-embedding parametersNot published by Undi95
Layers40 (standard Llama 2 13B architecture)
AttentionMulti-head attention (standard Llama 2 13B configuration)
Native context length4,096 tokens (Llama 2 base)
Extended contextNot published by Undi95
PrecisionFP16

License and Commercial Use

License: The model card states CC-BY-NC-4.0 (Creative Commons Attribution Non-Commercial 4.0). However, the foundational base model is Llama 2 13B (Meta Llama 2 Community License), which governs the actual use rights.

License conflict note: The Undi95 model card lists CC-BY-NC-4.0 as the license. However, the base model (Llama 2 13B, via TheBloke/Llama-2-13B-fp16) is subject to the Meta Llama 2 Community License. Since Undi95 cannot grant rights beyond those held in the base model, the more restrictive and foundationally applicable license is the Meta Llama 2 Community License. Additionally, CC-BY-NC-4.0 explicitly prohibits commercial use, while the Llama 2 Community License permits commercial use for organizations below 700 million MAU.

Nextbit's assessment: The effective license for commercial API use is the Meta Llama 2 Community License (the root license of the base model), not CC-BY-NC-4.0. Nextbit is applying the Llama 2 Community License for this model's commercial use determination.

Key terms of the Meta Llama 2 Community License:

  • Commercial use: Permitted for organizations with fewer than 700 million monthly active users (MAU). Organizations whose products or services exceeded 700 million MAU in the preceding calendar month must request a separate license from Meta.
  • Restriction on improving competing models: Llama 2 outputs may not be used to train, improve, or fine-tune competing large language models (other than Llama 2 itself or its derivatives).
  • Attribution: Licensees must retain the Llama 2 copyright notice.
  • Distribution: When sharing Llama materials with third parties, a copy of the Llama 2 Community License must be provided.

Nextbit's verification: Nextbit has reviewed the applicable license terms and confirmed that serving ReMM-SLERP-L2-13B via API for commercial use is permitted under the Meta Llama 2 Community License, subject to the 700 million MAU threshold. Nextbit currently operates well below this threshold.

Restrictions relevant to users: Users of Nextbit's API who access this model are bound by the Meta Llama 2 Community License. Users whose own products exceed 700 million MAU must obtain separate Meta approval. Outputs from this model may not be used to train competing large language models.


Training Data Summary

This model has two distinct layers: the base model pretraining data and the merge-specific data.

(a) Base model pretraining data — Llama 2 13B

ReMM-SLERP-L2-13B is rooted in the Llama 2 13B base model. The pretraining data for Llama 2 is documented by Meta AI. Nextbit does not reproduce that documentation here; for details see the Meta Llama 2 model card and the Llama 2 Technical Report.

Key known facts about Llama 2 13B pretraining:

  • Approximately 2 trillion tokens
  • Publicly available web text, primarily in English
  • Knowledge cutoff: September 2022

(b) Merge data — Undi95's SLERP merge method

ReMM-SLERP-L2-13B is a multi-stage model merge, not a conventional fine-tune on new data. The merge was created in three stages:

Stage 1 — ReML Part 1 (TIES merge): Three models merged via TIES algorithm on top of the Llama 2 13B base:

  • Chronos-Beluga-v2-13b (density 0.42)
  • airoboros-l2-13b-2.1 (density 0.56)
  • Nous-Hermes-Llama2-13b (density 0.30)

Stage 2 — ReML complete (TIES merge): ReML-L2-13B-part1 (density 0.70) merged with Nous-Hermes-Llama2-13b (density 0.30).

Stage 3 — ReMM final (SLERP merge): Huginn-13b-v1.2 and ReML-L2-13B merged at equal weight (0.5/0.5) using Spherical Linear Interpolation (SLERP) — a technique that interpolates smoothly between model weight vectors in spherical space rather than linear space, preserving the directional properties of the weight vectors.

Merge data: Not applicable — all three stages are weight interpolation processes; no new training data was introduced.

Not published by Undi95:

  • Any fine-tuning data applied after the merge

Languages Supported

Primarily English. ReMM-SLERP-L2-13B inherits Llama 2's language distribution, which is predominantly English-language. Multilingual capability is limited. Full language list: see the Llama 2 Technical Report.


Intended Uses

This model is designed for:

  • Roleplaying and creative fiction (inspired by MythoMax's creative writing focus)
  • General instruction following (from Airoboros and Nous-Hermes components)
  • Creative story generation

ReMM-SLERP-L2-13B is described as a recreation of MythoMax with updated component models. It combines instruction-following capabilities from the ReML stack with creative writing capability from Huginn, using SLERP for smoother weight interpolation than standard linear merges.

Use is subject to Meta's Acceptable Use Policy and the Llama 2 Community License.


Systemic Risk Assessment

FieldDetail
Training FLOPsNot published by Meta AI for Llama 2 13B; not applicable for the merge steps
Relevant base for threshold assessmentLlama 2 13B (the pretrained root base model); all merge steps involve weight interpolation only
Estimated FLOPs (Llama 2 13B pretraining)~2.4 × 10²³ (estimate based on 13B parameters × ~2T tokens; formula: 6 × N × D)
Exceeds 10²⁵ FLOPs threshold?No — estimated to be approximately 40× below the systemic risk threshold
Note on dense architectureLlama 2 13B is a dense model; all 13B parameters are active per token.
AI Office designationNot designated as a systemic risk model as of May 2026
Art. 55 obligations apply?No

The FLOPs relevant for systemic risk threshold assessment are those of the root base model (Llama 2 13B pretraining). The multi-stage merge process performed by Undi95 involves weight interpolation only and does not constitute model training in the regulatory sense. This estimate is based on publicly available information from the Llama 2 Technical Report.


This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. The merge was performed by Undi95; the root base model was developed by Meta AI. All information is sourced from public documentation published by Undi95 and Meta AI. This model is served under the Meta Llama 2 Community License.

For questions: [email protected]

Was this page helpful?