Gemma 4 26B A4B — Model Documentation
Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d
Model Identification
| Field | Detail |
|---|---|
| Full model name | Gemma 4 26B A4B |
| Developed by | Google DeepMind |
| Release date | April 2026 |
| Model card | huggingface.co/google/gemma-4-26B-A4B |
| Technical report | Not published by Google DeepMind as of May 2026 |
| Developer blog | blog.google — Gemma 4 launch |
Architecture and Parameters
| Field | Detail |
|---|---|
| Architecture | Mixture of Experts (MoE) |
| Total parameters | 25.2B |
| Active parameters per token | 3.8B (8 routed experts + 1 shared expert activated, out of 128 total experts) |
| Non-embedding parameters | Not published by Google DeepMind |
| Layers | 30 |
| Attention | Sliding window attention (window size: 1,024 tokens) for local layers; global attention for other layers |
| Native context length | 256,000 tokens |
| Extended context | Not applicable |
| Precision | BF16 |
| Modalities | Text + Image (multimodal) |
| Vision encoder | ~550M parameters |
Note: "A4B" designates approximately 4B active parameters per token, enabling inference performance comparable to a 4B dense model despite the 25.2B total parameter count.
License and Commercial Use
License: Apache License 2.0
Apache 2.0 is a permissive open-source license that allows free commercial use, distribution, and modification, provided that the original copyright notice and license text are retained and any modifications are documented.
Nextbit's verification: Nextbit has reviewed the license terms applicable to Gemma 4 26B A4B and confirmed that serving this model via API under a commercial inference service is permitted under Apache 2.0. No separate commercial agreement with Google DeepMind is required for this use case.
Restrictions relevant to users: None specific to this model beyond standard Apache 2.0 terms. Users of Nextbit's API who integrate model outputs into their own products remain responsible for their own compliance with applicable law and any downstream licensing requirements.
Training Data Summary
The following is a summary of publicly available information provided by Google DeepMind regarding training data for Gemma 4.
Pretraining data
Gemma 4 was pretrained on a diverse multilingual corpus. Training data sources include:
- Web documents: Diverse collection of web text spanning 140+ languages
- Code: Programming language syntax and patterns across multiple languages
- Mathematics: Logical reasoning and symbolic mathematical content
- Images: Wide range of image data for visual understanding tasks
Token volume: Not published by Google DeepMind.
Knowledge cutoff
January 2025. This is the knowledge cutoff published on the model card.
Data safety practices
Google DeepMind applied the following data filtering practices:
- CSAM (Child Sexual Abuse Material) filtering at multiple stages of the pipeline
- Sensitive personal data removal
- Content quality and safety filtering consistent with Google's policies
What is not publicly available
Google DeepMind has not published:
- Total token count for Gemma 4 pretraining
- Exact composition percentages of data sources
- Specific datasets or corpora used
- Data filtering or deduplication methodology in detail
- Training compute (FLOPs or GPU-hours)
The primary public references are the Gemma 4 model card and the Gemma documentation. A formal technical report had not been published as of May 2026.
Languages Supported
Gemma 4 26B A4B supports 140+ languages, as stated in the model card. Full language list: see the Gemma 4 model card and ai.google.dev/gemma/docs.
Intended Uses
This model is designed for:
- Complex reasoning tasks (mathematics, logic, scientific reasoning), including configurable thinking mode via
<|think|>token - Code generation, completion, and correction
- Multimodal tasks: OCR, document parsing, chart comprehension, UI understanding, and general image-text tasks
- Long-context document analysis (up to 256K tokens)
- Multilingual instruction following and translation
- Agentic workflows with native function calling
- General-purpose dialogue
No explicitly excluded uses are stated in the model card beyond Google's standard Gemma Prohibited Use Policy.
Systemic Risk Assessment
| Field | Detail |
|---|---|
| Training FLOPs | Not published by Google DeepMind |
| Estimated FLOPs (active params method) | ~4 × 10²³ (estimate based on 3.8B active parameters × unknown training tokens; using Gemma 2 27B reference of ~13T tokens as a conservative proxy; formula: 6 × N_active × D) |
| Exceeds 10²⁵ FLOPs threshold? | No — estimated to be well below the systemic risk threshold even under conservative assumptions |
| Note on MoE architecture | FLOPs for MoE models are calculated based on active parameters per token, not total parameters. The 3.8B active parameter figure (not the 25.2B total) is the relevant input for FLOPs estimation. |
| AI Office designation | Not designated as a systemic risk model as of May 2026 |
| Art. 55 obligations apply? | No |
This estimate is based on publicly available information. Google DeepMind has not published official training compute figures or total token counts for Gemma 4. The training token estimate is proxied from the Gemma 2 generation for conservatism. If Google DeepMind publishes official figures, this section will be updated.
This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. All training data and architectural information is sourced from public documentation published by Google DeepMind.
For questions: [email protected]