# ollama-python-Python快速部署Llama 3等大型语言模型最简单方法

> 作者/来源: UCloud 运营管理员
> 发布时间: 2024-04-30T09:52:00.000Z
> 分类: AI专区
> 标签: AI, GPU, Python
> 原文链接: http://117.50.162.249:3000/yun/articles/101

---

# ollama-python-Python快速部署Llama 3等大型语言模型最简单方法

> 来源: https://www.ucloud.cn/yun/131088.html
> 作者: UCloud小助手
> 发布日期: 发布于2024-04-30 17:52

## ollama介绍

![](https://ucloud-blog.cn-bj.ufileos.com/articles/131088/images/131088_000.png)

在本地启动并运行大型语言模型。运行Llama 3、Phi 3、Mistral、Gemma和其他型号。

## Llama 3

Meta Llama 3 是 Meta Inc. 开发的一系列最先进的模型，提供8B和70B参数大小（预训练或指令调整）。

![](https://ucloud-blog.cn-bj.ufileos.com/articles/131088/images/131088_001.png)

Llama 3 指令调整模型针对对话/聊天用例进行了微调和优化，并且在常见基准测试中优于许多可用的开源聊天模型。

![](https://ucloud-blog.cn-bj.ufileos.com/articles/131088/images/131088_002.png)

## ![](https://ucloud-blog.cn-bj.ufileos.com/articles/131088/images/131088_003.png) 安装

```
pip install ollama
```

## 用法

```
import ollamaresponse = ollama.chat(model='llama2', messages=[  {    'role': 'user',    'content': 'Why is the sky blue?',  },])print(response['message']['content'])
```

## 流式响应

可以通过设置stream=True、修改函数调用以返回 Python 生成器来启用响应流，其中每个部分都是流中的一个对象。

```
import ollama

stream = ollama.chat(
    model='llama2',
    messages=[{'role': 'user', 'content': 'Why is the sky blue?'}],
    stream=True,
)

for chunk in stream:
  print(chunk['message']['content'], end='', flush=True)
```

## 应用程序编程接口

Ollama Python 库的 API 是围绕Ollama REST API设计的

## 聊天

```
ollama.chat(model='llama2', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}])
```

## 新增

```
ollama.generate(model='llama2', prompt='Why is the sky blue?')
```

## 列表

```
ollama.list()
```

## 展示

```
ollama.show('llama2')
```

## 创建

```
modelfile='''
FROM llama2
SYSTEM You are mario from super mario bros.
'''

ollama.create(model='example', modelfile=modelfile)
```

## 复制

```
ollama.copy('llama2', 'user/llama2')
```

## 删除

```
ollama.delete('llama2')
Pull
ollama.pull('llama2')
push
ollama.push('user/llama2')
```

## 嵌入

```
ollama.embeddings(model='llama2', prompt='The sky is blue because of rayleigh scattering')
```

## 定制客户端

可以使用以下字段创建自定义客户端：

- host：要连接的 Ollama 主机
- timeout: 请求超时时间

```
from ollama import Client
client = Client(host='http://localhost:11434')
response = client.chat(model='llama2', messages=[
  {
'role': 'user',
'content': 'Why is the sky blue?',
  },
])
```

## 异步客户端

```
import asyncio
from ollama import AsyncClient

async def chat():
  message = {'role': 'user', 'content': 'Why is the sky blue?'}
  response = await AsyncClient().chat(model='llama2', messages=[message])

asyncio.run(chat())
```

设置stream=True修改函数以返回 Python 异步生成器：

```
import asyncio
from ollama import AsyncClient

async def chat():
  message = {'role': 'user', 'content': 'Why is the sky blue?'}
async for part in await AsyncClient().chat(model='llama2', messages=[message], stream=True):
    print(part['message']['content'], end='', flush=True)

asyncio.run(chat())
```

## 错误

如果请求返回错误状态或在流式传输时检测到错误，则会引发错误。

```
model = 'does-not-yet-exist'try:  ollama.chat(model)except ollama.ResponseError as e:  print('Error:', e.error)if e.status_code == 404:    ollama.pull(model)
```

 附高性能NVIDIA RTX 40 系列云服务器购买：  
  
https://www.ucloud.cn/site/active/gpu.html?ytag=seo  
  
https://www.compshare.cn/?ytag=seo