Gemma 2 27B IT — Model Documentation

Published by Nextbit256 S.L. | Last updated: May 2026 Compliance reference: AI Act Regulation (EU) 2024/1689, Art. 53.1.c and 53.1.d


Model Identification

FieldDetail
Full model nameGemma 2 27B Instruction Tuned (gemma-2-27b-it)
Developed byGoogle DeepMind
Release dateJune 2024
Model cardhuggingface.co/google/gemma-2-27b-it
Technical reportGemma 2 Technical Report (see ai.google.dev/gemma/docs)
Developer blogai.google.dev/gemma

Architecture and Parameters

FieldDetail
ArchitectureDense (decoder-only transformer) with interleaved local/global attention
Total parameters27B
Active parameters per tokenN/A — dense model (all parameters active)
Non-embedding parametersNot published by Google DeepMind
LayersNot published by Google DeepMind
AttentionInterleaved sliding-window (local) and full (global) attention with Grouped Query Attention (GQA)
Native context lengthNot published by Google DeepMind (refer to model card for max_position_embeddings)
Extended contextNot applicable
PrecisionBF16 (native); supports float32, int8, int4 quantization

License and Commercial Use

License: Gemma Terms of Use (proprietary Google license — NOT Apache 2.0)

The Gemma 2 27B IT model is released under the Gemma Terms of Use, a proprietary license issued by Google LLC. This is a distinct license from Apache 2.0 or MIT and includes specific conditions that differ from standard open-source licenses.

Key terms of the Gemma Terms of Use:

  • Commercial use: Permitted. The Gemma Terms of Use allow commercial use, including serving model outputs via hosted API services, with no user-count limits (unlike some earlier Gemma versions).
  • Downstream use restrictions: Service providers must include the Gemma use restrictions (per Section 3.2 and the Gemma Prohibited Use Policy) as an enforceable provision in any downstream agreement with their own users.
  • Distributing notice: Providers who distribute or deploy Gemma must provide users with a copy of the Gemma Terms of Use.
  • Prohibited Use Policy: Neither providers nor their end users may use the model for restricted uses listed in the Gemma Prohibited Use Policy.
  • Training other models: Permitted under the Gemma Terms of Use, subject to ensuring downstream agreements include the Gemma use restrictions. Model derivatives (including distillation) must propagate the same terms to recipients.

Nextbit's acceptance: Nextbit has accepted the Gemma Terms of Use to serve this model via API.

Important notice for Nextbit API users: Users of Nextbit's API who integrate outputs of this model into their own products or services must independently accept and comply with the Gemma Terms of Use. By using Nextbit's API to access this model, users confirm they have reviewed and accepted the Gemma Terms of Use at ai.google.dev/gemma/terms. Nextbit's Terms of Service incorporate this obligation by reference.


Training Data Summary

The following is a summary of publicly available information provided by Google DeepMind regarding training data for Gemma 2.

Pretraining data

Gemma 2 27B was pretrained on approximately 13 trillion tokens. Training data sources include:

  • Web documents: Diverse web text content, primary language: English
  • Code: Programming language content across multiple languages
  • Mathematics: Mathematical reasoning content

Language coverage

Gemma 2 27B is primarily an English-language model. The training data is predominantly English, with limited representation of other languages. For multilingual use cases, users should evaluate the model's performance carefully in non-English languages.

Knowledge cutoff

Not explicitly published by Google DeepMind. Based on the June 2024 release date, the training data cutoff is estimated to be no later than early 2024, though this has not been officially confirmed. Users should treat this as an approximation.

Data safety practices

Google DeepMind applied the following data filtering practices:

  • CSAM (Child Sexual Abuse Material) filtering
  • Sensitive personal data removal
  • Content quality and safety filtering consistent with Google's policies

Post-training

The instruction-tuned variant (gemma-2-27b-it) was fine-tuned using supervised learning and reinforcement learning from human feedback (RLHF) on top of the gemma-2-27b base model, to improve instruction following and reduce harmful outputs.

Training was conducted on TPUv5p using JAX and Google ML Pathways.

What is not publicly available

Google DeepMind has not published:

  • Exact per-source token counts and composition percentages
  • Specific datasets or corpora used
  • Data filtering or deduplication methodology in detail
  • Training compute (FLOPs or GPU-hours)
  • Number of layers and detailed architecture specifications

The primary public reference is the Gemma documentation and the model card.


Languages Supported

Gemma 2 27B IT is primarily an English-language model. It has limited capability in other languages. Refer to the Gemma model documentation for any updates on multilingual support.


Intended Uses

This model is designed for:

  • General-purpose instruction following and dialogue in English
  • Question answering, summarization, and text generation tasks
  • Code generation and analysis
  • Research and academic applications

No explicitly excluded uses are stated in the model card. Use is governed by the Gemma Prohibited Use Policy, which applies to all Nextbit API users accessing this model.


Systemic Risk Assessment

FieldDetail
Training FLOPsNot published by Google DeepMind
Estimated FLOPs (dense model)~4.7 × 10²⁴ (estimate based on 27B parameters × ~13T training tokens; formula: 6 × N × D)
Exceeds 10²⁵ FLOPs threshold?No — estimated to be approximately 2× below the systemic risk threshold
Note on dense architectureGemma 2 27B is a dense model; all 27B parameters are active per token. The full parameter count is used for FLOPs estimation.
AI Office designationNot designated as a systemic risk model as of May 2026
Art. 55 obligations apply?No

This estimate is based on publicly available information (parameter count: 27B; token count: 13T, from model card). Google DeepMind has not published official FLOPs figures. At the estimated ~4.7 × 10²⁴ FLOPs, this model is below but approaching the 10²⁵ threshold — Nextbit monitors this model for any AI Office review. If official compute figures or AI Office designations are published, this section will be updated.


This documentation is published by Nextbit256 S.L. in accordance with Article 53(1)(c) and 53(1)(d) of Regulation (EU) 2024/1689 (AI Act). Nextbit256 S.L. serves this model via its inference API but did not develop or train it. All training data and architectural information is sourced from public documentation published by Google DeepMind. Nextbit256 S.L. has accepted the Gemma Terms of Use and requires API users accessing this model to do the same.

For questions: [email protected]

Was this page helpful?