Gemma 4 26B A4B — Model Documentation

Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d


Model Identification

FieldDetail
Full model nameGemma 4 26B A4B
Developed byGoogle DeepMind
Release dateApril 2026
Model cardhuggingface.co/google/gemma-4-26B-A4B
Technical reportNot published by Google DeepMind as of May 2026
Developer blogblog.google — Gemma 4 launch

Architecture and Parameters

FieldDetail
ArchitectureMixture of Experts (MoE)
Total parameters25.2B
Active parameters per token3.8B (8 routed experts + 1 shared expert activated, out of 128 total experts)
Non-embedding parametersNot published by Google DeepMind
Layers30
AttentionSliding window attention (window size: 1,024 tokens) for local layers; global attention for other layers
Native context length256,000 tokens
Extended contextNot applicable
PrecisionBF16
ModalitiesText + Image (multimodal)
Vision encoder~550M parameters

Note: "A4B" designates approximately 4B active parameters per token, enabling inference performance comparable to a 4B dense model despite the 25.2B total parameter count.


License and Commercial Use

License: Apache License 2.0

Apache 2.0 is a permissive open-source license that allows free commercial use, distribution, and modification, provided that the original copyright notice and license text are retained and any modifications are documented.

Nextbit's verification: Nextbit has reviewed the license terms applicable to Gemma 4 26B A4B and confirmed that serving this model via API under a commercial inference service is permitted under Apache 2.0. No separate commercial agreement with Google DeepMind is required for this use case.

Restrictions relevant to users: None specific to this model beyond standard Apache 2.0 terms. Users of Nextbit's API who integrate model outputs into their own products remain responsible for their own compliance with applicable law and any downstream licensing requirements.


Training Data Summary

The following is a summary of publicly available information provided by Google DeepMind regarding training data for Gemma 4.

Pretraining data

Gemma 4 was pretrained on a diverse multilingual corpus. Training data sources include:

  • Web documents: Diverse collection of web text spanning 140+ languages
  • Code: Programming language syntax and patterns across multiple languages
  • Mathematics: Logical reasoning and symbolic mathematical content
  • Images: Wide range of image data for visual understanding tasks

Token volume: Not published by Google DeepMind.

Knowledge cutoff

January 2025. This is the knowledge cutoff published on the model card.

Data safety practices

Google DeepMind applied the following data filtering practices:

  • CSAM (Child Sexual Abuse Material) filtering at multiple stages of the pipeline
  • Sensitive personal data removal
  • Content quality and safety filtering consistent with Google's policies

What is not publicly available

Google DeepMind has not published:

  • Total token count for Gemma 4 pretraining
  • Exact composition percentages of data sources
  • Specific datasets or corpora used
  • Data filtering or deduplication methodology in detail
  • Training compute (FLOPs or GPU-hours)

The primary public references are the Gemma 4 model card and the Gemma documentation. A formal technical report had not been published as of May 2026.


Languages Supported

Gemma 4 26B A4B supports 140+ languages, as stated in the model card. Full language list: see the Gemma 4 model card and ai.google.dev/gemma/docs.


Intended Uses

This model is designed for:

  • Complex reasoning tasks (mathematics, logic, scientific reasoning), including configurable thinking mode via <|think|> token
  • Code generation, completion, and correction
  • Multimodal tasks: OCR, document parsing, chart comprehension, UI understanding, and general image-text tasks
  • Long-context document analysis (up to 256K tokens)
  • Multilingual instruction following and translation
  • Agentic workflows with native function calling
  • General-purpose dialogue

No explicitly excluded uses are stated in the model card beyond Google's standard Gemma Prohibited Use Policy.


Systemic Risk Assessment

FieldDetail
Training FLOPsNot published by Google DeepMind
Estimated FLOPs (active params method)~4 × 10²³ (estimate based on 3.8B active parameters × unknown training tokens; using Gemma 2 27B reference of ~13T tokens as a conservative proxy; formula: 6 × N_active × D)
Exceeds 10²⁵ FLOPs threshold?No — estimated to be well below the systemic risk threshold even under conservative assumptions
Note on MoE architectureFLOPs for MoE models are calculated based on active parameters per token, not total parameters. The 3.8B active parameter figure (not the 25.2B total) is the relevant input for FLOPs estimation.
AI Office designationNot designated as a systemic risk model as of May 2026
Art. 55 obligations apply?No

This estimate is based on publicly available information. Google DeepMind has not published official training compute figures or total token counts for Gemma 4. The training token estimate is proxied from the Gemma 2 generation for conservatism. If Google DeepMind publishes official figures, this section will be updated.


This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. All training data and architectural information is sourced from public documentation published by Google DeepMind.

For questions: [email protected]

Was this page helpful?