Rocinante 12B v1.1 — Model Documentation

Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d


Model Identification

FieldDetail
Full model nameRocinante-12B-v1.1
Developed byTheDrummer / BeaverAI (fine-tune author); base model: Mistral AI (Mistral-Nemo-Base-2407)
Release date2024
Model cardhuggingface.co/TheDrummer/Rocinante-12B-v1.1
Technical reportNot published by TheDrummer
Developer blogNot published by TheDrummer

The model name "Rocinante" is a literary reference to the horse of Don Quijote in Cervantes's novel — a deliberately humble but spirited name reflecting the model's character as a versatile creative workhorse.


Architecture and Parameters

FieldDetail
ArchitectureDense (decoder-only transformer) — Mistral NeMo architecture (MistralForCausalLM)
Total parameters12B
Active parameters per tokenN/A — dense model (all parameters active)
Non-embedding parametersNot published by TheDrummer
Layers40
AttentionGQA — 32 Q heads, 8 KV heads
Native context length1,024,000 tokens (1M; from config.json; native Mistral NeMo context is 128K)
Extended contextNot applicable
PrecisionBF16

License and Commercial Use

License: Apache License 2.0

The base model, Mistral-Nemo-Base-2407, is released under Apache 2.0. The fine-tune author (TheDrummer) lists the license as "Other" in the model card metadata, but has not published a separate license file. As the foundational base model is Apache 2.0, and no superseding license has been published by TheDrummer, Nextbit applies Apache 2.0 as the effective license.

Apache 2.0 is a permissive open-source license that allows free commercial use, distribution, and modification, provided that the original copyright notice and license text are retained and any modifications are documented.

Nextbit's verification: Nextbit has reviewed the base model license (Apache 2.0) and confirmed that serving Rocinante-12B-v1.1 via API under a commercial inference service is permitted. No separate commercial agreement with Mistral AI or TheDrummer is required for this use case.

Restrictions relevant to users: None beyond standard Apache 2.0 terms. Users of Nextbit's API who integrate model outputs into their own products remain responsible for their own compliance with applicable law and any downstream licensing requirements. As a fine-tune without content moderation, users should implement their own safety measures for sensitive applications.


Training Data Summary

This model has two distinct layers: the base model pretraining data and the fine-tune data.

(a) Base model pretraining data — Mistral-Nemo-Base-2407

Rocinante-12B-v1.1 is built on Mistral-Nemo-Base-2407 (12B, July 2024). The pretraining data for this base model is documented by Mistral AI. Nextbit does not reproduce that documentation here; for details see the Mistral NeMo model card and the Mistral NeMo blog post.

Key known facts about Mistral-Nemo-Base-2407 pretraining:

  • Native 128K context window
  • Large proportion of multilingual and code data
  • Languages: primarily English, with substantial coverage of French, German, Spanish, Italian, Portuguese, Russian, Chinese, and Japanese
  • Knowledge cutoff: approximately mid-2024 (July 2024 release)
  • Token volume: not published by Mistral AI

(b) Fine-tune data — TheDrummer / BeaverAI

Not published by TheDrummer. The model card does not describe the fine-tuning datasets, fine-tuning method, or training hyperparameters. The model's improved creative vocabulary and narrative quality suggest instruction fine-tuning on creative writing and roleplay datasets, but no details have been published.


Languages Supported

Primarily English. The base model (Mistral NeMo) has multilingual capability in English, French, German, Spanish, Italian, Portuguese, Russian, Chinese, and Japanese; the fine-tune was likely performed on English-language data. For full base model language list, see the Mistral NeMo model card.


Intended Uses

This model is designed for:

  • Creative writing and story generation — the model's primary design purpose
  • Roleplay (RP) — with ChatML prompt format
  • Adventure / instruction stories — with Alpaca prompt format
  • General instruction following — with Mistral NeMo format

The model name references Rocinante, the horse from Don Quijote by Cervantes — reflecting the model's positioning as a versatile, energetic, and creatively agile assistant. Testers describe it as offering "richer and distinct prose," "cranked up creativity," and "engaging and adventure-filled storytelling."

This model is not specifically designed for factual question answering, mathematical reasoning, or coding tasks.


Systemic Risk Assessment

FieldDetail
Training FLOPsNot published by Mistral AI for Mistral-Nemo-Base-2407; fine-tune FLOPs not applicable (not published by TheDrummer)
Relevant base for threshold assessmentMistral-Nemo-Base-2407 (the pretrained base model); the fine-tune involves negligible additional compute relative to pretraining
Estimated FLOPs (Mistral NeMo 12B pretraining)~4.3 × 10²³ (estimate based on 12B parameters × an assumed ~6T training tokens; formula: 6 × N × D; token count is itself an estimate as Mistral AI has not published it)
Exceeds 10²⁵ FLOPs threshold?No — estimated to be approximately 23× below the systemic risk threshold
Note on dense architectureMistral NeMo 12B is a dense model; all 12B parameters are active per token.
AI Office designationNot designated as a systemic risk model as of May 2026
Art. 55 obligations apply?No

The FLOPs relevant for systemic risk threshold assessment are those of the base model (Mistral-Nemo-Base-2407 pretraining). Mistral AI has not published official training token counts or FLOPs for this model. The token estimate is approximate. If official figures are published, this section will be updated.


This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. The fine-tune was performed by TheDrummer / BeaverAI; the base model was developed by Mistral AI. All information is sourced from public documentation published by TheDrummer and Mistral AI.

For questions: [email protected]

Was this page helpful?