跳转到正文

Gemma 开放模型使用指南

Google DeepMind 发布了 Gemma 开放语言模型系列,借鉴了 Gemini 的研究与技术。首批包含 20 亿参数和 70 亿参数模型,分别使用 2 万亿和 6 万亿词元训练,提供基础与指令微调检查点。训练上下文长度为 8,192 词元,在多个基准上超过 Llama 2 7B 和 Mistral 7B。

Gemma 基于 Transformer 解码器,采用多项改进:20 亿参数版本使用多查询注意力,70 亿版本使用多头注意力,同时采用 RoPE 位置编码GeGLU 激活和改进的归一化位置

根据技术报告,训练数据主要来自网页、数学和代码。与 Gemini 不同,首批 Gemma 没有专门训练多语言或多模态能力。词表约 25.6 万,使用 Gemini 的 SentencePiece 分词器子集,保留空白、拆分数字,并对未知词元使用字节级编码。

指令模型使用合成与人工编写的纯文本问答对进行监督微调,再进行人类反馈强化学习(RLHF)。奖励模型使用带偏好标签的数据训练,策略训练使用高质量提示词。原文说明,这些数据集均为英语。

指令模型还使用特殊控制词元表示对话角色和轮次:

Gemma 对话控制词元

结果

Gemma 7B 在数学、科学和代码任务中表现较强。下图按能力归类,显示学术基准的平均得分。

Gemma 的各类能力表现

模型在多个基准上超过 Llama 2 7B 和 Mistral 7B,HumanEval、GSM8K、MATH、AGIEval 等结果较突出,推理、对话、数学和代码能力有所改善。

Gemma 学术基准结果

人工评估表明,Gemma 7B 指令模型在安全性和指令遵循方面超过 Mistral-7B v0.2 Instruct。

人工安全与指令遵循评估

研究者也在多项安全基准中与 Mistral 比较。技术报告介绍了去偏差和红队测试,以缓解语言模型的常见风险。负责任开发的说明见模型卡生成式 AI 工具包

安全基准比较

Gemma 7B 提示词格式

基础模型不要求专门格式,可以用零样本或少样本提示执行任务。指令模型使用以下格式:

<start_of_turn>user
Generate a Python function that multiplies two numbers <end_of_turn>
<start_of_turn>model
含义控制词元
用户轮次user
模型轮次model
对话轮次开始<start_of_turn>
对话轮次结束<end_of_turn>

多轮对话也可以使用这些控制词元:

markdown
<start_of_turn>user
What is a good place for travel in the US?<end_of_turn>
<start_of_turn>model
California.<end_of_turn>
<start_of_turn>user
What can I do in California?<end_of_turn>
<start_of_turn>model

如何编写提示词

有效使用 Gemma 7B,需要正确运用提示词模板。下面介绍几个任务示例。

零样本提示

可以直接利用模型的零样本能力:

markdown
<start_of_turn>user
Explain why the sky is blue<end_of_turn>
<start_of_turn>model

加入总体指令的零样本提示

系统角色或总体指令通常有助于引导模型。Gemma 没有显式的系统角色,但可以把额外指令放入提示词:

markdown
<start_of_turn>user
Answer the following question in a concise and informative manner:

Explain why the sky is blue<end_of_turn>
<start_of_turn>model

上述例子加入“请简洁且信息充分地回答下面的问题”,帮助控制结果。

角色扮演

角色扮演既可让回答更个性化,也有助于构建复杂的领域专用语言模型智能体

提示词:

<start_of_turn>user
You are a helpful 2nd-grade teacher. Help a 2nd grader to answer questions in a short and clear manner.

Explain why the sky is blue<end_of_turn>
<start_of_turn>model

模型回答:

Sure, here is an explanation for why the sky is blue:

The sky appears blue because of a phenomenon called Rayleigh Scattering.

Here's the science behind it:

Sunlight is made up of different colors of light, like red, green, blue, and violet.
When sunlight hits the Earth's atmosphere, the different colors of light scatter in different directions.
Blue light scatters more easily than other colors because it has shorter wavelengths.
The scattered blue light travels in all directions, including towards your eyes.
That's why we see the sky as blue.
It's important to note that this explanation is simplified and does not account for all scientific details. However, it gives you a good understanding of the basic science behind why the sky appears blue.

原作者没有验证回答的准确性,也没有判断是否适合小学二年级学生。可以把它与下面思维链提示的结果比较。

推理

可通过零样本思维链提示引导模型推理:

提示词:

markdown
<start_of_turn>user
Think and write your step-by-step reasoning before responding.

Explain why the sky is blue.<end_of_turn>
<start_of_turn>model

模型回答,已删节:

Step 1: Scattering of Light

The sky appears blue due to the phenomenon of scattering of light. When sunlight hits the Earth's atmosphere, it interacts with the particles of air, primarily nitrogen and oxygen molecules.

Step 2: Rayleigh Scattering

Rayleigh scattering is the specific scattering mechanism that occurs when light interacts with particles much smaller than the wavelength of light. In this process, shorter wavelengths (blue light) are scattered more effectively than longer wavelengths (red light).

Step 3: Blue Scattered Light Travels in All Directions

...

Conclusion

The sky appears blue due to the scattering of light, specifically Rayleigh scattering. Shorter wavelengths of light are scattered more effectively, and this scattered light travels in all directions, with the scattered light that travels opposite to the sun's rays being most visible to our eyes.

Gemma 可以自动生成 Markdown。原作者为便于展示略微调整了格式,但内容与模型输出一致;原文没有进一步评估准确性或幻觉。

资源与集成

根据发布博客使用条款允许各种规模的组织负责任地开展商业使用和分发。

参考资料

ChatGPT 中文使用指南 · MIT 许可 · 隐私政策