Перейти до основного вмісту

Обгортання зовнішньої пам’яті тензора

ПолеЗначення
КатегоріяСпівпроцесинг PCIe
СкладністьСередній
Орієнтовний час читання15 хвилин
МіткиPCIe, C++, tensor, external memory, zero-copy wrapping

Програма створює три повторно використовувані вхідні слоти й надсилає вісім синтетичних кадрів FP32. Слот повертається до черги доступних лише після отримання впорядкованого результату для нього.

Покроковий огляд​

Перевірка контракту моделі​

Створіть модель і прочитайте info().inputs до виділення пам’яті. У цьому посібнику використано архів YOLOv8s з одним входом. Програма перевіряє, що оголошений dtype — FP32, а форма займає рівно size_bytes байтів.

Зовнішнє представлення має відповідати dtype, формі, розміру в байтах та імені відповідного TensorInfo. Не виводьте ці значення з іншої збірки моделі.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
pcie::ConnectionOptions connection;
connection.card_id = card_id;
connection.max_inflight = static_cast<int>(kRingSlots);
pcie::Model model(kModelPath, {}, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1U) {
throw std::runtime_error("tutorial requires a model with one input");
}
validate_fp32_input(info.inputs.front());

Обгортання пам’яті застосунку​

Кожен слот кільця володіє std::shared_ptr<std::vector<float>>. Виклик Tensor::from_external() отримує базовий вказівник, повну кількість елементів базової пам’яті, спільного власника, форму моделі та ім’я маршруту. Оскільки представлення неперервне, хост PCIe може обгорнути його без створення проміжної ділянки пам’яті.

Зберігати лише сирий вказівник недостатньо. Спільний власник обов’язковий, оскільки транспорт може утримувати тензор після повернення з push().

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
std::vector<InputSlot> slots = make_input_ring(info.inputs.front());

Побудова моделі​

Будуйте модель після перевірки контракту та виділення кільця. У прикладі max_inflight дорівнює розміру кільця, тому застосунок і транспорт мають однакову явну межу.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
model.build(kBuildTimeoutMs);

Надсилання й безпечне повторне використання кільця​

Заповніть доступний слот, викличте push() і перенесіть слот до черги активних. Не змінюйте й не використовуйте його пам’ять повторно лише тому, що push() повернув керування. Коли доступних слотів немає, приклад викликає pull() і повертає найстаріший слот до черги доступних тільки після надходження відповідного впорядкованого результату.

Кожному прийнятому надсиланню відповідає одне отримання, включно з фінальним спорожненням. У разі тайм-ауту модель закривається замість повторного використання пам’яті запиту, який усе ще може бути активним.

pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
std::size_t completed = 0;
try {
completed = run_input_ring(model, slots);
} catch (...) {
model.close();
throw;
}
model.close();

Запуск​

Установіть пакет хоста PCIe та завантажте комплект посібників, як описано в розділі Налаштування посібників. Завантажте YOLOv8s до кореня розпакованих додаткових файлів PCIe:

sima-cli modelzoo get yolo_v8s
cp /absolute/path/to/downloaded-yolov8s-archive.tar.gz yolo_v8s_mpk.tar.gz
test -f yolo_v8s_mpk.tar.gz

C++ (prebuilt):

./lib/sima-pcie-host/tutorials/tutorial_028_wrap_external_tensor_memory

C++ (build from source):

./build.sh --target tutorial_028_wrap_external_tensor_memory
./build/tutorials-standalone/tutorial_028_wrap_external_tensor_memory

Типово використовується карта 0 і черга 0. Передавайте --card N лише для іншої карти. Успішний запуск виводить:

input=images
ring_slots=3
completed=8
[OK] 028_wrap_external_tensor_memory

На практиці​

Для прямого обгортання на хості потрібна неперервна пам’ять. Тензор із непослідовними кроками також приймається, якщо його дескриптор коректний, але хост ущільнює його в проміжну ділянку. Кілька окремо виділених входів також пакуються в проміжну пам’ять.

Для моделі з кількома входами уникнути такого виділення можна лише тоді, коли всі тензори є послідовними представленнями однієї спільної упакованої ділянки. Надсилайте їх у порядку з info().inputs і використовуйте ім’я та форму кожного входу. Наприклад, коли обидва входи мають FP32:

const auto& first = info.inputs.at(0);
const auto& second = info.inputs.at(1);
const std::size_t first_count = first.size_bytes / sizeof(float);
const std::size_t second_count = second.size_bytes / sizeof(float);

auto packed =
std::make_shared<std::vector<float>>(first_count + second_count);
pcie::Tensor input0 = pcie::Tensor::from_external(
packed->data(), packed->size(), packed, first.shape, first.name);
pcie::Tensor input1 = pcie::Tensor::from_external(
packed->data(), packed->size(), packed, second.shape, second.name,
static_cast<std::int64_t>(first.size_bytes));

model.push({input0, input1});

Ця оптимізація усуває лише пакувальне копіювання на хості. PCIe однаково копіює упаковане корисне навантаження в транспортну пам’ять карти перед інференсом.

Повний початковий код​

Показати повні програми
pcie_host/tutorials/028_wrap_external_tensor_memory/run_external_tensor_memory.cpp
// Submit application-owned tensor memory without a host staging copy.
//
// Usage:
// tutorial_028_wrap_external_tensor_memory [--card 0]

#include <simaai/neat/pcie/Model.h>

#include <algorithm>
#include <cstdint>
#include <cstdlib>
#include <deque>
#include <filesystem>
#include <iostream>
#include <limits>
#include <memory>
#include <stdexcept>
#include <string>
#include <vector>

namespace pcie = simaai::neat::pcie;

namespace {

constexpr int kBuildTimeoutMs = 180000;
constexpr int kPullTimeoutMs = 30000;
constexpr std::size_t kRingSlots = 3;
constexpr std::size_t kFrameCount = 8;
constexpr char kModelPath[] = "yolo_v8s_mpk.tar.gz";

int parse_card(const int argc, char** argv) {
int card_id = 0;
for (int index = 1; index < argc; ++index) {
const std::string arg = argv[index];
if (arg == "--card" && index + 1 < argc) {
card_id = std::stoi(argv[++index]);
} else if (arg == "-h" || arg == "--help") {
std::cout << "Usage: " << argv[0] << " [--card 0]\n";
std::exit(0);
} else {
throw std::runtime_error("unknown or incomplete argument: " + arg);
}
}
return card_id;
}

std::size_t checked_element_count(const std::vector<std::int64_t>& shape) {
std::size_t count = 1;
for (const std::int64_t dimension : shape) {
if (dimension <= 0 ||
count > std::numeric_limits<std::size_t>::max() / static_cast<std::size_t>(dimension)) {
throw std::runtime_error("model input has an invalid or overflowing shape");
}
count *= static_cast<std::size_t>(dimension);
}
return count;
}

void validate_fp32_input(const pcie::TensorInfo& input) {
if (input.dtype != "FP32" && input.dtype != "FLOAT32") {
throw std::runtime_error("tutorial requires an FP32 model input, got " + input.dtype);
}
const std::size_t element_count = checked_element_count(input.shape);
if (element_count > std::numeric_limits<std::size_t>::max() / sizeof(float) ||
element_count * sizeof(float) != input.size_bytes) {
throw std::runtime_error("model input shape and byte size are inconsistent");
}
}

struct InputSlot {
std::shared_ptr<std::vector<float>> storage;
pcie::Tensor tensor;
};

InputSlot make_input_slot(const pcie::TensorInfo& input) {
auto storage = std::make_shared<std::vector<float>>(checked_element_count(input.shape), 0.0F);
pcie::Tensor tensor = pcie::Tensor::from_external(storage->data(), storage->size(), storage,
input.shape, input.name);
return {.storage = std::move(storage), .tensor = std::move(tensor)};
}

std::vector<InputSlot> make_input_ring(const pcie::TensorInfo& input) {
std::vector<InputSlot> slots;
slots.reserve(kRingSlots);
for (std::size_t index = 0; index < kRingSlots; ++index) {
slots.push_back(make_input_slot(input));
}
return slots;
}

std::size_t run_input_ring(pcie::Model& model, std::vector<InputSlot>& slots) {
std::deque<std::size_t> available;
std::deque<std::size_t> in_flight;
for (std::size_t index = 0; index < slots.size(); ++index) {
available.push_back(index);
}

std::size_t completed = 0;
const auto complete_oldest = [&] {
auto outputs = model.pull(kPullTimeoutMs);
if (!outputs) {
throw std::runtime_error("timed out waiting for an external-memory submission");
}
if (outputs->empty() || in_flight.empty()) {
throw std::runtime_error("received an invalid external-memory completion");
}
available.push_back(in_flight.front());
in_flight.pop_front();
++completed;
};

for (std::size_t frame = 0; frame < kFrameCount; ++frame) {
if (available.empty()) {
complete_oldest();
}

const std::size_t slot_index = available.front();
available.pop_front();
InputSlot& slot = slots[slot_index];

// A slot is writable only while it is not in flight.
std::fill(slot.storage->begin(), slot.storage->end(), static_cast<float>(frame % 10U) / 10.0F);
if (!model.push(slot.tensor)) {
throw std::runtime_error("push rejected frame " + std::to_string(frame));
}
in_flight.push_back(slot_index);
}

while (!in_flight.empty()) {
complete_oldest();
}
return completed;
}

} // namespace

int main(int argc, char** argv) {
try {
const int card_id = parse_card(argc, argv);
if (!std::filesystem::is_regular_file(kModelPath)) {
throw std::runtime_error(std::string("model does not exist: ") + kModelPath);
}

pcie::ConnectionOptions connection;
connection.card_id = card_id;
connection.max_inflight = static_cast<int>(kRingSlots);
pcie::Model model(kModelPath, {}, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1U) {
throw std::runtime_error("tutorial requires a model with one input");
}
validate_fp32_input(info.inputs.front());

std::vector<InputSlot> slots = make_input_ring(info.inputs.front());

model.build(kBuildTimeoutMs);

std::size_t completed = 0;
try {
completed = run_input_ring(model, slots);
} catch (...) {
model.close();
throw;
}
model.close();

std::cout << "input=" << info.inputs.front().name << '\n'
<< "ring_slots=" << slots.size() << '\n'
<< "completed=" << completed << '\n'
<< "[OK] 028_wrap_external_tensor_memory\n";
return 0;
} catch (const std::exception& error) {
std::cerr << "[FAIL] " << error.what() << '\n';
return 1;
}
}

Джерело​