跳至主要内容

以 INT8 張量僅執行 MLA

欄位值
類別PCIe 協同處理
難度中級
預估閱讀時間15 分鐘
標籤PCIe, MLA, INT8, quantization, tensor

一支程式在主機上量化影像、執行僅 MLA 的路徑、反量化輸出頭,並在同一個佇列上與預設路徑進行比對。

操作指南​

檢視僅 MLA 的合約​

在啟用 mla_only 的情況下建構 Model。info() 現在會回報 INT8 的輸入與輸出,每個都帶有 quant.scale 與 quant.zero_point。參考模型只有一個輸入 images,是 INT8 [640, 640, 3] HWC 張量。

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

在主機上量化​

參考模型期望一張像素值在 [0, 1] 範圍內的 RGB 影像。將影像縮放到 640x640,把 BGR 轉成 RGB,除以 255,再以進入端的參數套用量化公式。這個前處理屬於模型,而不是封存檔的合約:其他模型需要自己的處理方式。保留同一組編碼反量化後的值:這正是預設路徑進行同條件比較所需的 FP32 輸入。

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

執行 INT8 路徑​

build() 會啟動只包含 MLA 的卡端管線。run() 接受 INT8 張量,並依 info().outputs 回報的順序與名稱,為每個輸出傳回一個緊密排列的 INT8 張量。此路徑會拒絕其他 dtype;推送 FP32 會失敗,而不會在卡上被量化。

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

反量化並比較​

在不啟用 mla_only 的情況下建構第二個 Model,並將反量化後的 FP32 值送入預設路徑。以每個輸出的參數反量化 INT8 輸出頭,並以該輸出頭的 scale 為單位印出每個輸出頭的最大偏差。兩條路徑以相同的編碼執行相同的 MLA 程式,因此誤差為零。

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

執行​

依照教學設定的說明安裝 PCIe 主機套件並下載教學套件包。

本教學執行 Model Zoo 的 YOLO26n INT8 封存檔,它是為直接 MLA 輸入與輸出而編譯的。將它下載到解壓縮後的 PCIe extras 根目錄:

sima-cli download https://docs.sima.ai/pkg_downloads/SDK2.1.3/models/modalix/yolo26-detection/yolo26n-det-int8-b1.tar.gz

其他封存檔若以 Model SDK 的 tessellate_parameters 設定 enable_mla=True、每個輸入使用 HWC DRAM 版面配置、每個輸出使用 HWC16 編譯,也符合條件。改在 CVU 上進行鑲嵌的封存檔,例如 Model Zoo 的 yolo_v8s 建置,在啟用 mla_only 時會被拒絕:

mla_only does not support stage 'tessellate_quantize_0_MLA_0/...' (tess)

C++ (prebuilt):

./lib/sima-pcie-host/tutorials/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

C++ (build from source):

./build.sh --target tutorial_027_run_mla_only_int8
./build/tutorials-standalone/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

兩個版本都會印出合約,且每個輸出頭的偏差皆為零:

MLA-only contract:
input images INT8 [640, 640, 3] scale=0.00390434 zero_point=-128
output bbox_0 INT8 [80, 80, 4] scale=0.0302856 zero_point=-117
...
Dequantized MLA-only outputs vs the default route (error in scale units):
bbox_0 [80, 80, 4] max_err=0.0000
...
[OK] 027_run_mla_only_int8

預設為卡 0 與佇列 0。只有在使用其他卡時才傳入 --card N。

實務應用​

對於 INT8 封存檔,當應用程式自行掌握量化時,請啟用 mla_only:它已經從感測器或前一個模型產生 INT8、需要原始 INT8 輸出頭進行自己的後處理,或想從卡端延遲中移除量化與反量化階段。每個 scale 與 zero point 都應從 model.info() 讀取,切勿從模型的另一個建置複製。

此路徑是全有或全無。每個輸入都必須符合其對應 info().inputs 項目的 dtype、形狀與位元組大小,且影像前處理或方框解碼無法與 mla_only 併用。當主機持有 FP32 資料且不需要自行進行轉換時,請保留預設路徑。

如需部署診斷,請繼續閱讀 PCIe 模型工作流程。

完整原始碼​

顯示完整原始碼程式
pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
// Quantize on the host, run only the MLA over PCIe, and dequantize the INT8 results.
//
// Usage:
// tutorial_027_run_mla_only_int8 --model yolo26n-det-int8-b1.tar.gz [--card 0]

#include <simaai/neat/pcie/Model.h>

#include <opencv2/imgcodecs.hpp>
#include <opencv2/imgproc.hpp>

#include <algorithm>
#include <cmath>
#include <cstdint>
#include <cstdlib>
#include <filesystem>
#include <iomanip>
#include <iostream>
#include <stdexcept>
#include <string>
#include <utility>
#include <vector>

namespace pcie = simaai::neat::pcie;

namespace {

constexpr int kBuildTimeoutMs = 180000;
constexpr int kRunTimeoutMs = 30000;
constexpr double kMaxErrorScales = 0.05; // one twentieth of a quantization step
constexpr char kImagePath[] = "share/sima-pcie-host/tutorials/assets/street-scene.png";

struct Args {
std::string model;
int card_id = 0;
};

std::string require_value(int argc, char** argv, int& index, const char* option) {
if (index + 1 >= argc) {
throw std::runtime_error(std::string("missing value for ") + option);
}
return argv[++index];
}

Args parse_args(int argc, char** argv) {
Args args;
for (int index = 1; index < argc; ++index) {
const std::string arg = argv[index];
if (arg == "--model") {
args.model = require_value(argc, argv, index, "--model");
} else if (arg == "--card") {
args.card_id = std::stoi(require_value(argc, argv, index, "--card"));
} else if (arg == "-h" || arg == "--help") {
std::cout << "Usage: " << argv[0] << " --model yolo26n-det-int8-b1.tar.gz [--card 0]\n";
std::exit(0);
} else {
throw std::runtime_error("unknown argument: " + arg);
}
}
if (args.model.empty()) {
throw std::runtime_error("--model is required: yolo26n-det-int8-b1.tar.gz from the Model Zoo");
}
return args;
}

std::string shape_string(const std::vector<std::int64_t>& shape) {
std::string text = "[";
for (std::size_t index = 0; index < shape.size(); ++index) {
text += (index == 0 ? "" : ", ") + std::to_string(shape[index]);
}
return text + "]";
}

std::size_t element_count(const std::vector<std::int64_t>& shape) {
std::size_t count = 1;
for (const auto dim : shape) {
count *= static_cast<std::size_t>(dim);
}
return count;
}

// The MLA-only contract quantizes every tensor with one scale and one zero point:
// x = (q - zero_point) * scale
// q = clamp(round(x / scale) + zero_point, -128, 127)
const pcie::QuantParams& require_quant(const pcie::TensorInfo& info) {
if (!info.quant.has_value()) {
throw std::runtime_error("tensor '" + info.name + "' publishes no quantization parameters");
}
return *info.quant;
}

std::int8_t quantize(const float value, const pcie::QuantParams& quant) {
const float code = std::nearbyint(value / quant.scale) + static_cast<float>(quant.zero_point);
return static_cast<std::int8_t>(std::clamp(code, -128.0F, 127.0F));
}

float dequantize(const std::int8_t code, const pcie::QuantParams& quant) {
return static_cast<float>(static_cast<std::int32_t>(code) - quant.zero_point) * quant.scale;
}

const std::int8_t* int8_data(const pcie::Tensor& tensor) {
return reinterpret_cast<const std::int8_t*>(static_cast<const std::uint8_t*>(tensor.data) +
tensor.byte_offset);
}

// Preprocessing of the reference model: one RGB HWC image with pixels in [0, 1]. Resize, BGR to
// RGB, divide by 255, then quantize with the ingress parameters. Another model needs its own
// recipe.
std::vector<std::int8_t> quantize_image(const cv::Mat& bgr, const pcie::TensorInfo& input) {
if (input.shape.size() != 3 || input.shape[2] != 3) {
throw std::runtime_error("the reference model takes a three-channel HWC input, got " +
shape_string(input.shape));
}
const pcie::QuantParams& quant = require_quant(input);
cv::Mat rgb;
cv::resize(bgr, rgb,
cv::Size(static_cast<int>(input.shape[1]), static_cast<int>(input.shape[0])));
cv::cvtColor(rgb, rgb, cv::COLOR_BGR2RGB);
std::vector<std::int8_t> codes(rgb.total() * rgb.channels());
for (std::size_t index = 0; index < codes.size(); ++index) {
codes[index] = quantize(static_cast<float>(rgb.data[index] / 255.0), quant);
}
return codes;
}

} // namespace

int main(int argc, char** argv) {
try {
const Args args = parse_args(argc, argv);
if (!std::filesystem::is_regular_file(args.model)) {
throw std::runtime_error("model does not exist: " + args.model);
}
const cv::Mat image = cv::imread(kImagePath, cv::IMREAD_COLOR);
if (image.empty()) {
throw std::runtime_error(std::string("OpenCV could not decode: ") + kImagePath);
}
pcie::ConnectionOptions connection;
connection.card_id = args.card_id;

// CORE LOGIC
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

std::cout << "[OK] 027_run_mla_only_int8\n";
return 0;
} catch (const std::exception& error) {
std::cerr << "[FAIL] " << error.what() << '\n';
return 1;
}
}

來源​