FastChat
在单张昇腾 NPU 上安装 FastChat,并通过 OpenAI 兼容接口调用 Qwen/Qwen2.5-0.5B-Instruct 完成一次对话。
前置条件
硬件
Atlas 900 A2 单卡(Ascend NPU),并按需完成物理机或容器内的设备挂载。
基础软件
在运行本文档示例之前,你的机器上需要已经装好并可用:
可用的 Python 环境
可用的 CANN(参考快速安装昇腾环境)
本文档测试环境使用 Python 3.12、CANN 9.1.0。
本文档配套镜像:
swr.cn-south-1.myhuaweicloud.com/ascendhub/cann:9.1.0-910b-ubuntu22.04-py3.12
加载 CANN 环境
运行示例前加载 CANN 环境变量:
source /usr/local/Ascend/ascend-toolkit/set_env.sh
安装 PyTorch 软件栈
安装与 CANN 匹配的 torch 和 torch_npu(其他环境请参考 Ascend PyTorch 安装文档 选择兼容版本):
python -m pip install "torch==2.9.0" "torch_npu==2.9.0.post2" --extra-index-url https://repo.huaweicloud.com/ascend/repos/pypi
python -c "
import torch
import torch_npu
print('torch:', torch.__version__)
print('torch_npu:', torch_npu.__version__)
"
输出结果如下:
...
torch: 2.9.0+cpu
torch_npu: 2.9.0.post2
安装 FastChat
安装 FastChat:
python -m pip install fschat
python -c "
import fastchat
print('FastChat version:', fastchat.__version__)
"
输出结果如下:
...
FastChat version: xxx
Note
xxx 表示实际安装的 FastChat 版本号。
运行示例:OpenAI 兼容 API
FastChat 用 controller 管理 model worker,并通过 API server 提供 OpenAI 兼容接口。
安装本示例所需的模型 worker、ModelScope 及 API 服务依赖:
python -m pip install "fschat[model_worker]" "transformers==4.57.6" "modelscope==1.37.0" "fastapi==0.141.1" "uvicorn==0.52.0"
controller、model worker 和 API server 都是前台服务,启动后需要保持运行。
启动 controller。 controller 负责注册和调度 model worker:
python -m fastchat.serve.controller
启动 model worker。 worker 在 NPU 上加载模型,并以 Qwen2.5-0.5B-Instruct 为服务名注册到 controller:
FASTCHAT_USE_MODELSCOPE=True python -m fastchat.serve.model_worker \
--model-path Qwen/Qwen2.5-0.5B-Instruct \
--model-names Qwen2.5-0.5B-Instruct \
--revision master \
--device npu
启动 API server。 服务在 http://127.0.0.1:8000/v1 提供 OpenAI 兼容接口:
python -m fastchat.serve.openai_api_server \
--host 127.0.0.1 \
--port 8000
检查模型服务。 model worker 加载完成后,通过 /v1/models 查看已经注册的模型:
curl -fsS http://127.0.0.1:8000/v1/models -o /tmp/fastchat-models.json
python -m json.tool --no-ensure-ascii /tmp/fastchat-models.json
输出结果如下:
{
"object": "list",
"data": [
{
"id": "Qwen2.5-0.5B-Instruct",
"object": "model",
"created": xxx,
"owned_by": "fastchat",
...
}
]
}
Note
xxx 表示动态生成的标识与时间戳,... 表示省略的字段。
发送一次对话请求,调用 OpenAI 兼容的 Chat Completions 接口并打印模型回复:
curl -fsS http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Qwen2.5-0.5B-Instruct","messages":[{"role":"user","content":"你好"}],"max_tokens":64,"temperature":0}' \
-o /tmp/fastchat-chat.json
python -m json.tool --no-ensure-ascii /tmp/fastchat-chat.json
输出结果如下:
{
"id": "xxx",
"object": "chat.completion",
"created": xxx,
"model": "Qwen2.5-0.5B-Instruct",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "xxx"
},
"finish_reason": xxx
}
],
"usage": {
...
}
}
Note
xxx 表示每次请求动态生成的 ID、时间戳、模型回复和 token 统计等内容,... 表示省略的字段。
外部链接
官方仓库:lm-sys/FastChat
官方快速开始:README