UnslopNemo 12B v4.1 — Model Documentation
Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d
Model Identification
| Field | Detail |
|---|---|
| Full model name | UnslopNemo-12B-v4.1 |
| Developed by | TheDrummer (fine-tune author); base model: Mistral AI (Mistral-Nemo-Base-2407) |
| Release date | 2024 |
| Model card | huggingface.co/TheDrummer/UnslopNemo-12B-v4.1 |
| Technical report | Not published by TheDrummer |
| Developer blog | Not published by TheDrummer |
The name "UnslopNemo" is a compound: "Unslop" refers to reducing generic, over-aligned, "sloppy" AI responses; "Nemo" confirms the Mistral NeMo base model. The model targets users who want less sanitized, more authentic output.
Architecture and Parameters
| Field | Detail |
|---|---|
| Architecture | Dense (decoder-only transformer) — Mistral NeMo architecture (MistralForCausalLM) |
| Total parameters | 12B |
| Active parameters per token | N/A — dense model (all parameters active) |
| Non-embedding parameters | Not published by TheDrummer |
| Layers | 40 |
| Attention | GQA — 32 Q heads, 8 KV heads |
| Native context length | 1,024,000 tokens (1M; from config.json; native Mistral NeMo context is 128K) |
| Extended context | Not applicable |
| Precision | BF16 |
License and Commercial Use
License: Apache License 2.0
The base model, Mistral-Nemo-Base-2407, is released under Apache 2.0. The fine-tune author (TheDrummer) has not published a separate license file or model card metadata specifying a different license. As the foundational base model is Apache 2.0, and no superseding license has been published by TheDrummer, Nextbit applies Apache 2.0 as the effective license.
Apache 2.0 is a permissive open-source license that allows free commercial use, distribution, and modification, provided that the original copyright notice and license text are retained and any modifications are documented.
Nextbit's verification: Nextbit has reviewed the base model license (Apache 2.0) and confirmed that serving UnslopNemo-12B-v4.1 via API under a commercial inference service is permitted. No separate commercial agreement with Mistral AI or TheDrummer is required for this use case.
Restrictions relevant to users: None beyond standard Apache 2.0 terms. The model is designed to produce less restricted outputs and does not include safety filtering by default. Users of Nextbit's API who integrate model outputs into their own products remain responsible for their own compliance with applicable law, including any content-related obligations.
Training Data Summary
This model has two distinct layers: the base model pretraining data and the fine-tune data.
(a) Base model pretraining data — Mistral-Nemo-Base-2407
UnslopNemo-12B-v4.1 is built on Mistral-Nemo-Base-2407 (12B, July 2024). The pretraining data for this base model is documented by Mistral AI. Nextbit does not reproduce that documentation here; for details see the Mistral NeMo model card and the Mistral NeMo blog post.
Key known facts about Mistral-Nemo-Base-2407 pretraining:
- Native 128K context window
- Large proportion of multilingual and code data
- Languages: primarily English, with substantial coverage of French, German, Spanish, Italian, Portuguese, Russian, Chinese, and Japanese
- Knowledge cutoff: approximately mid-2024 (July 2024 release)
- Token volume: not published by Mistral AI
(b) Fine-tune data — TheDrummer
Not published by TheDrummer. The model card is minimal and does not describe the fine-tuning datasets, fine-tuning method, or training hyperparameters. Based on the model's purpose (reducing over-aligned "sloppy" outputs and improving authentic narrative quality), the fine-tune likely involved curated creative writing and roleplay datasets with reduced alignment filtering, but no details have been published.
Languages Supported
Primarily English. The base model (Mistral NeMo) has multilingual capability in English, French, German, Spanish, Italian, Portuguese, Russian, Chinese, and Japanese; the fine-tune was likely performed on English-language data. For full base model language list, see the Mistral NeMo model card.
Intended Uses
This model is designed for:
- Creative writing and narrative generation with reduced over-refusal and more natural prose
- Roleplay and interactive fiction with authentic character voices
- General instruction following with less sanitized output style
The "Unslop" design goal is to eliminate the tendency of aligned AI models to produce generic, hedged, or overly cautious ("sloppy") text in creative contexts, resulting in more engaging and authentic storytelling. This model has been used as a base for 27 downstream community merges, indicating strong adoption in creative AI communities.
This model is not specifically designed for factual question answering, mathematical reasoning, or coding tasks. Given the reduced alignment fine-tuning, this model is not appropriate for sensitive consumer-facing applications without additional safety measures.
Systemic Risk Assessment
| Field | Detail |
|---|---|
| Training FLOPs | Not published by Mistral AI for Mistral-Nemo-Base-2407; fine-tune FLOPs not applicable (not published by TheDrummer) |
| Relevant base for threshold assessment | Mistral-Nemo-Base-2407 (the pretrained base model); the fine-tune involves negligible additional compute relative to pretraining |
| Estimated FLOPs (Mistral NeMo 12B pretraining) | ~4.3 × 10²³ (estimate based on 12B parameters × an assumed ~6T training tokens; formula: 6 × N × D; token count is itself an estimate as Mistral AI has not published it) |
| Exceeds 10²⁵ FLOPs threshold? | No — estimated to be approximately 23× below the systemic risk threshold |
| Note on dense architecture | Mistral NeMo 12B is a dense model; all 12B parameters are active per token. |
| AI Office designation | Not designated as a systemic risk model as of May 2026 |
| Art. 55 obligations apply? | No |
The FLOPs relevant for systemic risk threshold assessment are those of the base model (Mistral-Nemo-Base-2407 pretraining). Mistral AI has not published official training token counts or FLOPs for this model. The token estimate is approximate. If official figures are published, this section will be updated.
This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. The fine-tune was performed by TheDrummer; the base model was developed by Mistral AI. All information is sourced from public documentation published by TheDrummer and Mistral AI.
For questions: [email protected]