快速开始
在两张昇腾 NPU 上验证 Ray 原生 NPU 资源发现、Task/Actor 设备隔离,
并在每个 Ray Worker 中执行一次真实的 torch_npu 运算。本文基于 Ray
上游的 Accelerator Support
文档编写。
前置条件
硬件
Atlas 900 A2 / A3 训练系列产品,至少有两张可用的 Ascend NPU,并已完成
物理机或容器中的设备与驱动配置。CI 使用 linux-aarch64-a2-2 Runner,
由 Runner 自动提供两张配置完好的 NPU。
基础软件
在运行本文档之前,机器上需要已经安装并可用:
Linux aarch64 操作系统;
可用的 Python 环境;
可用的 CANN toolkit 和驱动;
与 CANN 匹配的
torch和torch_npu,且torch.npu.is_available() == True;npu-smi能正常显示 NPU 设备。
CANN 与 PyTorch-NPU 的安装和版本匹配可参考 Ascend PyTorch 安装文档。
本文档示例使用的版本
配套机器:
机器类型:Atlas 900 A2 PODc(Ascend 910B4,64 GB × 2);
操作系统:Ubuntu 22.04。
配套镜像:
swr.cn-south-1.myhuaweicloud.com/ascendhub/cann:9.1.0-910b-ubuntu22.04-py3.12
软件版本:
组件 |
版本 |
|---|---|
Python |
3.12 |
CANN |
9.1.0 |
torch |
2.9.0+cpu |
torch_npu |
2.9.0.post2 |
Ray |
当前最新 release,Linux aarch64 wheel |
NPU |
Ascend 910B4 × 2 |
检查前置是否满足
检查 Python 版本:
python --version
Python 3.12.xxx
检查 Torch、Torch-NPU 和 NPU 设备:
python -c "import torch, torch_npu; print('torch=', torch.__version__); print('torch_npu=', torch_npu.__version__); print('is_available:', torch.npu.is_available()); print('count:', torch.npu.device_count())"
torch= 2.9.0xxx
torch_npu= 2.9.0.post2
is_available: True
count: 2
确认 npu-smi 可以看到两张 NPU:
npu-smi info
如果 npu-smi 不存在或 import torch_npu 失败,请先修复驱动、CANN、
Torch 与 Torch-NPU 的版本匹配问题。
安装 Ray
安装当前最新 Ray release,并验证安装结果:
python -m pip install -q -U "ray[default]"
python -c "import ray; print('ray', ray.__version__)"
ray xxx
验证 Ray 自动发现 NPU
Ray 优先通过 AscendCL 探测设备数量,并以 /dev/davinci* 作为回退,
随后把设备发布为逻辑 NPU 资源。
python - <<'PY'
import ray
from ray._private.accelerators import NPUAcceleratorManager
ray.init(include_dashboard=False, log_to_driver=False)
resources = ray.cluster_resources()
count = int(resources.get("NPU", 0))
assert count == 2, resources
print("Ray NPU resources:", count)
print("Ascend type:", NPUAcceleratorManager.get_current_node_accelerator_type())
ray.shutdown()
PY
Ray NPU resources: 2
Ascend type: xxx
验证 Task 与 Actor 的 NPU 隔离
下面的 Actor 持续占用一张 NPU,Task 应获得另一张。Ray 把分配到的物理
设备 ID 写入 ASCEND_RT_VISIBLE_DEVICES,torch_npu 再通过隔离后的
设备视图执行真实运算。
python - <<'PY'
import os
import ray
ray.init(resources={"NPU": 2}, include_dashboard=False, log_to_driver=False)
def assigned_npu():
import torch
import torch_npu
ids = [str(value) for value in ray.get_runtime_context().get_accelerator_ids()["NPU"]]
assert len(ids) == 1, ids
visible = os.environ["ASCEND_RT_VISIBLE_DEVICES"]
assert visible == ids[0], (visible, ids)
value = torch.ones(4, device="npu:0").sum().cpu().item()
return ids[0], value
@ray.remote(resources={"NPU": 1})
class NPUActor:
def ping(self):
return assigned_npu()
@ray.remote(resources={"NPU": 1})
def npu_task():
return assigned_npu()
actor = NPUActor.remote()
actor_result = ray.get(actor.ping.remote())
task_result = ray.get(npu_task.remote())
ids = sorted([actor_result[0], task_result[0]])
values = sorted([actor_result[1], task_result[1]])
assert ids == ["0", "1"], ids
assert values == [4.0, 4.0], values
print("assigned NPU IDs:", ",".join(ids))
print("tensor sums:", ",".join(str(value) for value in values))
ray.kill(actor)
ray.shutdown()
PY
assigned NPU IDs: 0,1
tensor sums: 4.0,4.0
Ray 的 NPU 是逻辑资源。resources={"NPU": 0.25} 这样的请求允许多个
任务共享一张 NPU,但不会提供显存或算力隔离,应用仍需自行保证共享安全。