llama.cpp
前置条件
硬件
Atlas 800T / 900 A2 训练系列(Ascend 910B)。本文覆盖单卡与多卡。上游已验证的设备列表见 CANN 后端 — Devices。
软件
类别 |
要求 |
|---|---|
CANN |
toolkit + 驱动固件已安装并可 |
编译工具 |
cmake ≥ 3.14、g++(C++17)、make、git |
下载工具 |
curl |
llama.cpp |
本文从 GitHub 源码编译,见下文 |
1. 加载 CANN 环境
source /usr/local/Ascend/ascend-toolkit/set_env.sh
export PATH=/usr/local/sbin:$PATH
2. 检查环境是否就绪
2.1 确认 NPU 在线
npu-smi info
备注
如果 npu-smi 找不到,回到 快速安装昇腾环境 检查驱动与设备挂载。
2.2 确认 CANN 与编译工具
下面确认 CANN 已加载,并且 npu-smi 与 cmake 都在 PATH 里。
test -n "$ASCEND_HOME_PATH"
command -v npu-smi
cmake --version
3. 获取源码并编译
克隆 ggml-org/llama.cpp,开启 CANN 后端后应生成 llama-completion(一次性文本生成)。同一次编译还会生成 llama-cli 和 llama-server,后面的可选步骤会用到。将 <ref> 换成目标分支、tag 或 commit(上游默认分支为 master)。
if [ ! -d llama.cpp/.git ]; then
git clone https://github.com/ggml-org/llama.cpp.git
fi
cd llama.cpp
git checkout <ref>
cmake -B build -DGGML_CANN=on -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j$(nproc)
ls build/bin/llama-completion
4. 准备 GGUF 模型
推理输入是 GGUF 文件;910B 上 CANN 后端支持 FP16、BF16、Q8_0、Q4_0,见 CANN 后端 — DataType。下面从 Hugging Face 下载约 400 MB 的 Q4_0 示例,文件头应为 GGUF。
if [ ! -f qwen2.5-0.5b-instruct-q4_0.gguf ]; then
curl -fL --retry 3 --retry-delay 5 --connect-timeout 30 \
-o qwen2.5-0.5b-instruct-q4_0.gguf \
https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct-GGUF/resolve/main/qwen2.5-0.5b-instruct-q4_0.gguf
fi
head -c 4 qwen2.5-0.5b-instruct-q4_0.gguf
5. 推理
单卡用 -sm none -mg <设备号> 指定全部算子到特定卡上;多卡用 -sm layer 按层拆分。
参数 |
含义 |
|---|---|
|
GGUF 模型路径 |
|
提示词 |
|
最大生成 token 数。 |
|
随机种子。 |
|
offload 到 NPU 的层数。 |
|
关闭自动对话。Instruct 类 GGUF 不加此项会等待键盘输入 |
|
单卡布局:全部算子跑在 0 号卡 |
|
多卡布局:按层拆到可见的多张卡 |
|
限制进程可见的 NPU 编号 |
5.1 单卡推理
cd llama.cpp && ASCEND_RT_VISIBLE_DEVICES=0 ./build/bin/llama-completion \
-m ../qwen2.5-0.5b-instruct-q4_0.gguf \
-p "Building a website can be done in 10 simple steps:" \
-n 64 -no-cnv -ngl 99 -sm none -mg 0 --seed 42 2>&1
完整输出较长,其中应包含:
...
Building a website can be done in 10 simple steps: building a website is a great way to increase your online presence, and this post will help you get started. These steps are all the tools you will need to build your website, and you can learn them one at a time, and you can build a website in the 10 steps as the next step.
...
5.2 多卡推理
有两张及以上 NPU 时,可用 -sm layer 把层拆到多卡。下面示例暴露 0、1 号卡:
cd llama.cpp && ASCEND_RT_VISIBLE_DEVICES=0,1 ./build/bin/llama-completion \
-m ../qwen2.5-0.5b-instruct-q4_0.gguf \
-p "Building a website can be done in 10 simple steps:" \
-n 64 -no-cnv -ngl 99 -sm layer --seed 42 2>&1
完整输出较长,其中应包含:
...
Building a website can be done in 10 simple steps: building a website is a great way to increase your online presence, and this post will help you get started. These steps are all the tools you will need to build your website, and you can learn them one at a time, and you can build a website in the 10 steps as the next step.
You
...
5.3 交互式对话
llama-cli 面向多轮对话,启动后会等待键盘输入。想手动体验模型时可以使用:
cd llama.cpp && ASCEND_RT_VISIBLE_DEVICES=0 ./build/bin/llama-cli \
-m ../qwen2.5-0.5b-instruct-q4_0.gguf \
-ngl 99 -sm none -mg 0
5.4 浏览器对话
llama-server 会启动本地 HTTP 服务,并自带 Web 页面。启动后用浏览器打开 http://127.0.0.1:8080 即可对话。
cd llama.cpp && ASCEND_RT_VISIBLE_DEVICES=0 ./build/bin/llama-server \
-m ../qwen2.5-0.5b-instruct-q4_0.gguf \
-ngl 99 -sm none -mg 0 --host 127.0.0.1 --port 8080
6. 更多文档
上游仓库与总说明:ggml-org/llama.cpp
CANN 后端的设备、精度与环境变量:docs/backend/CANN.md
命令行对话:llama-cli
HTTP 服务与自带 Web 页面:llama-server