Phi-4 — Model Documentation
Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d
Model Identification
| Field | Detail |
|---|---|
| Full model name | Phi-4 |
| Developed by | Microsoft Research |
| Release date | December 2024 |
| Model card | huggingface.co/microsoft/phi-4 |
| Technical report | arXiv:2412.08905 |
| Developer blog | Not published by Microsoft Research for this model |
Architecture and Parameters
| Field | Detail |
|---|---|
| Architecture | Dense (decoder-only transformer) |
| Total parameters | 14B |
| Active parameters per token | N/A — dense model (all parameters active) |
| Non-embedding parameters | Not published by Microsoft Research |
| Layers | Not published by Microsoft Research |
| Attention | Not published by Microsoft Research (architecture described as minimal changes from phi-3) |
| Native context length | 16,384 tokens |
| Extended context | Not published by Microsoft Research |
| Precision | Not published by Microsoft Research |
License and Commercial Use
License: MIT License
The MIT License is a permissive open-source license that allows free commercial use, distribution, and modification with minimal restrictions (attribution and license text inclusion required).
Nextbit's verification: Nextbit has reviewed the license terms applicable to Phi-4 and confirmed that serving this model via API under a commercial inference service is permitted under the MIT License. No separate commercial agreement with Microsoft Research is required for this use case.
Restrictions relevant to users: None specific to this model beyond standard MIT License terms. Users of Nextbit's API who integrate model outputs into their own products remain responsible for their own compliance with applicable law and any downstream licensing requirements.
Training Data Summary
The following is a summary of publicly available information provided by Microsoft Research regarding training data for Phi-4, as documented in the technical report (arXiv:2412.08905) and the HuggingFace model card.
Pretraining data
Phi-4 was trained on approximately 9.8 trillion tokens. Microsoft Research's Phi series is notable for its heavy emphasis on synthetic and curated high-quality data. Training data sources include:
- Publicly available documents: Filtered web text and documents, selected for quality
- High-quality educational data and code: Specifically curated for reasoning and coding capability
- Synthetic "textbook-like" data: Synthetically generated content for mathematics, coding, reasoning, and general knowledge — a defining characteristic of the Phi model family
- Acquired academic books and Q&A datasets: Licensed academic and question-answering corpora
- High-quality chat-format supervised data: Used for instruction-following and alignment
Training ran for 21 days on 1,920 H100-80G GPUs (October–November 2024).
Language coverage
Phi-4 is primarily an English-language model. Approximately 92% of training data is in English; multilingual data accounts for approximately 8%. This is intentional: the model is optimized for English-language reasoning tasks.
Knowledge cutoff
June 2024 for publicly available data. Synthetic data generation may extend marginally beyond this date, but June 2024 is the stated cutoff for real-world knowledge.
Post-training
Post-training included:
- Supervised Fine-Tuning (SFT) for instruction following
- Direct Preference Optimization (DPO) for alignment with human preferences
What is not publicly available
Microsoft Research has not published:
- Exact per-source token counts and composition percentages
- Specific datasets or corpora names used
- Detailed data filtering methodology
- Training compute (FLOPs)
- Architecture details (number of layers, attention heads)
The primary public reference is the Phi-4 Technical Report (arXiv:2412.08905).
Languages Supported
Phi-4 is primarily an English-language model (approximately 92% of training data in English). It has limited multilingual capability from the remaining ~8% of multilingual training data. For multilingual use cases, other models in the Nextbit catalog that were trained on broader multilingual corpora (e.g., Qwen3, ALIA-40B) are more appropriate.
Intended Uses
This model is designed for:
- STEM-focused question answering (mathematics, science, engineering)
- Advanced reasoning tasks where data quality matters more than training volume
- Code generation and analysis
- Instruction following in English-language contexts
- Research and academic applications requiring a strong small-parameter reasoning model
The Phi series is designed around the hypothesis that high-quality synthetic training data can produce reasoning capabilities that exceed those of much larger models trained on broader but lower-quality data.
No explicitly excluded uses are stated in the model card beyond standard responsible use guidelines.
Systemic Risk Assessment
| Field | Detail |
|---|---|
| Training FLOPs | Not published by Microsoft Research |
| Estimated FLOPs (dense model) | ~1.7 × 10²³ (estimate based on 14B parameters × 9.8T tokens; formula: 6 × N × D) |
| Exceeds 10²⁵ FLOPs threshold? | No — estimated to be approximately 60× below the systemic risk threshold |
| Note on dense architecture | Phi-4 is a dense model; all 14B parameters are active per token. The full parameter count is used for FLOPs estimation. |
| AI Office designation | Not designated as a systemic risk model as of May 2026 |
| Art. 55 obligations apply? | No |
This estimate is based on publicly available information (parameter count: 14B; token count: 9.8T, both from the technical report and model card). Microsoft Research has not published official FLOPs figures. If such figures are published, this section will be updated.
This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. All training data and architectural information is sourced from public documentation published by Microsoft Research.
For questions: [email protected]