外观
Gemma 开放模型使用指南
Google DeepMind 发布了 Gemma 开放语言模型系列,借鉴了 Gemini 的研究与技术。首批包含 20 亿参数和 70 亿参数模型,分别使用 2 万亿和 6 万亿词元训练,提供基础与指令微调检查点。训练上下文长度为 8,192 词元,在多个基准上超过 Llama 2 7B 和 Mistral 7B。
Gemma 基于 Transformer 解码器,采用多项改进:20 亿参数版本使用多查询注意力,70 亿版本使用多头注意力,同时采用 RoPE 位置编码、GeGLU 激活和改进的归一化位置。
根据技术报告,训练数据主要来自网页、数学和代码。与 Gemini 不同,首批 Gemma 没有专门训练多语言或多模态能力。词表约 25.6 万,使用 Gemini 的 SentencePiece 分词器子集,保留空白、拆分数字,并对未知词元使用字节级编码。
指令模型使用合成与人工编写的纯文本问答对进行监督微调,再进行人类反馈强化学习(RLHF)。奖励模型使用带偏好标签的数据训练,策略训练使用高质量提示词。原文说明,这些数据集均为英语。
指令模型还使用特殊控制词元表示对话角色和轮次:

结果
Gemma 7B 在数学、科学和代码任务中表现较强。下图按能力归类,显示学术基准的平均得分。

模型在多个基准上超过 Llama 2 7B 和 Mistral 7B,HumanEval、GSM8K、MATH、AGIEval 等结果较突出,推理、对话、数学和代码能力有所改善。

人工评估表明,Gemma 7B 指令模型在安全性和指令遵循方面超过 Mistral-7B v0.2 Instruct。

研究者也在多项安全基准中与 Mistral 比较。技术报告介绍了去偏差和红队测试,以缓解语言模型的常见风险。负责任开发的说明见模型卡和生成式 AI 工具包。

Gemma 7B 提示词格式
基础模型不要求专门格式,可以用零样本或少样本提示执行任务。指令模型使用以下格式:
<start_of_turn>user
Generate a Python function that multiplies two numbers <end_of_turn>
<start_of_turn>model| 含义 | 控制词元 |
|---|---|
| 用户轮次 | user |
| 模型轮次 | model |
| 对话轮次开始 | <start_of_turn> |
| 对话轮次结束 | <end_of_turn> |
多轮对话也可以使用这些控制词元:
markdown
<start_of_turn>user
What is a good place for travel in the US?<end_of_turn>
<start_of_turn>model
California.<end_of_turn>
<start_of_turn>user
What can I do in California?<end_of_turn>
<start_of_turn>model如何编写提示词
有效使用 Gemma 7B,需要正确运用提示词模板。下面介绍几个任务示例。
零样本提示
可以直接利用模型的零样本能力:
markdown
<start_of_turn>user
Explain why the sky is blue<end_of_turn>
<start_of_turn>model加入总体指令的零样本提示
系统角色或总体指令通常有助于引导模型。Gemma 没有显式的系统角色,但可以把额外指令放入提示词:
markdown
<start_of_turn>user
Answer the following question in a concise and informative manner:
Explain why the sky is blue<end_of_turn>
<start_of_turn>model上述例子加入“请简洁且信息充分地回答下面的问题”,帮助控制结果。
角色扮演
角色扮演既可让回答更个性化,也有助于构建复杂的领域专用语言模型智能体。
提示词:
<start_of_turn>user
You are a helpful 2nd-grade teacher. Help a 2nd grader to answer questions in a short and clear manner.
Explain why the sky is blue<end_of_turn>
<start_of_turn>model模型回答:
Sure, here is an explanation for why the sky is blue:
The sky appears blue because of a phenomenon called Rayleigh Scattering.
Here's the science behind it:
Sunlight is made up of different colors of light, like red, green, blue, and violet.
When sunlight hits the Earth's atmosphere, the different colors of light scatter in different directions.
Blue light scatters more easily than other colors because it has shorter wavelengths.
The scattered blue light travels in all directions, including towards your eyes.
That's why we see the sky as blue.
It's important to note that this explanation is simplified and does not account for all scientific details. However, it gives you a good understanding of the basic science behind why the sky appears blue.原作者没有验证回答的准确性,也没有判断是否适合小学二年级学生。可以把它与下面思维链提示的结果比较。
推理
可通过零样本思维链提示引导模型推理:
提示词:
markdown
<start_of_turn>user
Think and write your step-by-step reasoning before responding.
Explain why the sky is blue.<end_of_turn>
<start_of_turn>model模型回答,已删节:
Step 1: Scattering of Light
The sky appears blue due to the phenomenon of scattering of light. When sunlight hits the Earth's atmosphere, it interacts with the particles of air, primarily nitrogen and oxygen molecules.
Step 2: Rayleigh Scattering
Rayleigh scattering is the specific scattering mechanism that occurs when light interacts with particles much smaller than the wavelength of light. In this process, shorter wavelengths (blue light) are scattered more effectively than longer wavelengths (red light).
Step 3: Blue Scattered Light Travels in All Directions
...
Conclusion
The sky appears blue due to the scattering of light, specifically Rayleigh scattering. Shorter wavelengths of light are scattered more effectively, and this scattered light travels in all directions, with the scattered light that travels opposite to the sun's rays being most visible to our eyes.Gemma 可以自动生成 Markdown。原作者为便于展示略微调整了格式,但内容与模型输出一致;原文没有进一步评估准确性或幻觉。
资源与集成
- Colab 入门笔记本与 Kaggle 示例
- Hugging Face 模型集合
- MaxText
- NVIDIA NeMo
- TensorRT-LLM
- NVIDIA AI Playground 中的 Gemma 7B
根据发布博客,使用条款允许各种规模的组织负责任地开展商业使用和分发。