Ministral 3B 2512 — Model Documentation

Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d


Model Identification

FieldDetail
Full model nameMinistral 3 3B Instruct 2512
Developed byMistral AI
Release dateDecember 2025 (2512 = year 25, month 12)
Model cardhuggingface.co/mistralai/Ministral-3-3B-Instruct-2512
Technical reportarXiv:2601.08584
Developer blogmistral.ai/news/mistral-3

Architecture and Parameters

FieldDetail
ArchitectureDense (decoder-only transformer) with vision encoder
Total parameters~3.8B (3.4B language model + 0.4B vision encoder)
Active parameters per tokenN/A — dense model (all language model parameters active)
Non-embedding parametersNot published by Mistral AI
LayersNot published by Mistral AI
AttentionNot published by Mistral AI
Native context length256,000 tokens
Extended contextNot applicable
PrecisionFP8 (instruct version); BF16 (base version)
ModalitiesText + Image (multimodal)

License and Commercial Use

License: Apache License 2.0

Apache 2.0 is a permissive open-source license that allows free commercial use, distribution, and modification, provided that the original copyright notice and license text are retained and any modifications are documented.

Nextbit's verification: Nextbit has reviewed the license terms applicable to Ministral 3B 2512 and confirmed that serving this model via API under a commercial inference service is permitted under Apache 2.0. No separate commercial agreement with Mistral AI is required for this use case.

Restrictions relevant to users: None specific to this model beyond standard Apache 2.0 terms. The model card notes that the model must not be used in a manner that infringes, misappropriates, or otherwise violates any third party's rights, including intellectual property rights. Users of Nextbit's API who integrate model outputs into their own products remain responsible for their own compliance with applicable law and any downstream licensing requirements.


Training Data Summary

The following is a summary of publicly available information for the Ministral 3 model family (3B, 8B, and 14B variants). Training data details are shared across the family; individual variant-specific differences have not been published by Mistral AI.

Pretraining data

Mistral AI has not published detailed training data specifications for the Ministral 3 family in the model card or the associated technical report (arXiv:2601.08584). The following is known from public sources:

  • Training method: Cascade Distillation — an iterative pruning and continued training with distillation technique applied to produce the 3B, 8B, and 14B family members
  • Multilingual support: The models support 11 languages: English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, and Arabic, suggesting training data coverage in at least these languages. The blog post references "40+ native languages" for the broader Mistral 3 family

Token volume: Not published by Mistral AI.

Languages: 11 languages explicitly supported; full training data language distribution not published.

Knowledge cutoff

Not published by Mistral AI.

What is not publicly available

Mistral AI has not published:

  • Total pretraining token count
  • Specific datasets or corpora used
  • Language distribution in training data
  • Data filtering or deduplication methodology
  • Knowledge cutoff date
  • Training compute (FLOPs or GPU-hours)

The primary public references are the model card and the Ministral 3 technical report. Nextbit will update this section if Mistral AI publishes additional training data documentation.


Languages Supported

Ministral 3B 2512 explicitly supports 11 languages: English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, and Arabic. The broader Mistral 3 training covers additional languages. Full language coverage: see the Ministral 3 technical report.


Intended Uses

This model is designed for:

  • Edge and on-device deployment (optimized for resource-constrained hardware; fits in 8 GB VRAM in FP8 format)
  • Chat interfaces and local daily-driver AI assistant use cases
  • Image and document description and visual understanding
  • Multilingual content generation and translation
  • Agentic workflows with native function calling and JSON output
  • Fine-tuning and specialization for specific tasks

No explicitly excluded uses are stated in the model card beyond general responsible use principles.


Systemic Risk Assessment

FieldDetail
Training FLOPsNot published by Mistral AI
Estimated FLOPs (dense model)~4 × 10²² (estimate based on 3.4B language model parameters × an assumed ~6T training tokens; formula: 6 × N × D; token count is itself an estimate)
Exceeds 10²⁵ FLOPs threshold?No — estimated to be approximately 250× below the systemic risk threshold
Note on dense architectureMinistral 3B is a dense model; all 3.4B language model parameters are active per token. The full parameter count is used for FLOPs estimation.
AI Office designationNot designated as a systemic risk model as of May 2026
Art. 55 obligations apply?No

This estimate is based on publicly available parameter counts. Neither training token count nor FLOPs have been published by Mistral AI. The token estimate is approximate and will be updated if Mistral AI publishes official figures.


This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. All training data and architectural information is sourced from public documentation published by Mistral AI.

For questions: [email protected]

Was this page helpful?