Phi-4 — Model Documentation

Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d


Model Identification

FieldDetail
Full model namePhi-4
Developed byMicrosoft Research
Release dateDecember 2024
Model cardhuggingface.co/microsoft/phi-4
Technical reportarXiv:2412.08905
Developer blogNot published by Microsoft Research for this model

Architecture and Parameters

FieldDetail
ArchitectureDense (decoder-only transformer)
Total parameters14B
Active parameters per tokenN/A — dense model (all parameters active)
Non-embedding parametersNot published by Microsoft Research
LayersNot published by Microsoft Research
AttentionNot published by Microsoft Research (architecture described as minimal changes from phi-3)
Native context length16,384 tokens
Extended contextNot published by Microsoft Research
PrecisionNot published by Microsoft Research

License and Commercial Use

License: MIT License

The MIT License is a permissive open-source license that allows free commercial use, distribution, and modification with minimal restrictions (attribution and license text inclusion required).

Nextbit's verification: Nextbit has reviewed the license terms applicable to Phi-4 and confirmed that serving this model via API under a commercial inference service is permitted under the MIT License. No separate commercial agreement with Microsoft Research is required for this use case.

Restrictions relevant to users: None specific to this model beyond standard MIT License terms. Users of Nextbit's API who integrate model outputs into their own products remain responsible for their own compliance with applicable law and any downstream licensing requirements.


Training Data Summary

The following is a summary of publicly available information provided by Microsoft Research regarding training data for Phi-4, as documented in the technical report (arXiv:2412.08905) and the HuggingFace model card.

Pretraining data

Phi-4 was trained on approximately 9.8 trillion tokens. Microsoft Research's Phi series is notable for its heavy emphasis on synthetic and curated high-quality data. Training data sources include:

  1. Publicly available documents: Filtered web text and documents, selected for quality
  2. High-quality educational data and code: Specifically curated for reasoning and coding capability
  3. Synthetic "textbook-like" data: Synthetically generated content for mathematics, coding, reasoning, and general knowledge — a defining characteristic of the Phi model family
  4. Acquired academic books and Q&A datasets: Licensed academic and question-answering corpora
  5. High-quality chat-format supervised data: Used for instruction-following and alignment

Training ran for 21 days on 1,920 H100-80G GPUs (October–November 2024).

Language coverage

Phi-4 is primarily an English-language model. Approximately 92% of training data is in English; multilingual data accounts for approximately 8%. This is intentional: the model is optimized for English-language reasoning tasks.

Knowledge cutoff

June 2024 for publicly available data. Synthetic data generation may extend marginally beyond this date, but June 2024 is the stated cutoff for real-world knowledge.

Post-training

Post-training included:

  • Supervised Fine-Tuning (SFT) for instruction following
  • Direct Preference Optimization (DPO) for alignment with human preferences

What is not publicly available

Microsoft Research has not published:

  • Exact per-source token counts and composition percentages
  • Specific datasets or corpora names used
  • Detailed data filtering methodology
  • Training compute (FLOPs)
  • Architecture details (number of layers, attention heads)

The primary public reference is the Phi-4 Technical Report (arXiv:2412.08905).


Languages Supported

Phi-4 is primarily an English-language model (approximately 92% of training data in English). It has limited multilingual capability from the remaining ~8% of multilingual training data. For multilingual use cases, other models in the Nextbit catalog that were trained on broader multilingual corpora (e.g., Qwen3, ALIA-40B) are more appropriate.


Intended Uses

This model is designed for:

  • STEM-focused question answering (mathematics, science, engineering)
  • Advanced reasoning tasks where data quality matters more than training volume
  • Code generation and analysis
  • Instruction following in English-language contexts
  • Research and academic applications requiring a strong small-parameter reasoning model

The Phi series is designed around the hypothesis that high-quality synthetic training data can produce reasoning capabilities that exceed those of much larger models trained on broader but lower-quality data.

No explicitly excluded uses are stated in the model card beyond standard responsible use guidelines.


Systemic Risk Assessment

FieldDetail
Training FLOPsNot published by Microsoft Research
Estimated FLOPs (dense model)~1.7 × 10²³ (estimate based on 14B parameters × 9.8T tokens; formula: 6 × N × D)
Exceeds 10²⁵ FLOPs threshold?No — estimated to be approximately 60× below the systemic risk threshold
Note on dense architecturePhi-4 is a dense model; all 14B parameters are active per token. The full parameter count is used for FLOPs estimation.
AI Office designationNot designated as a systemic risk model as of May 2026
Art. 55 obligations apply?No

This estimate is based on publicly available information (parameter count: 14B; token count: 9.8T, both from the technical report and model card). Microsoft Research has not published official FLOPs figures. If such figures are published, this section will be updated.


This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. All training data and architectural information is sourced from public documentation published by Microsoft Research.

For questions: [email protected]

Was this page helpful?