メインコンテンツまでスキップ

INT8テンソルでMLAのみを実行

項目値
カテゴリPCIe コプロセッシング
難易度中級
推定所要時間15分
ラベルPCIe, MLA, INT8, quantization, tensor

1つのプログラムがホスト上で画像を量子化し、MLAのみのルートを実行し、ヘッドを逆量子化して、同じキュー上のデフォルトルートと照合します。

ウォークスルー​

MLAのみのコントラクトを確認する​

mla_onlyを有効にしてModelを構築します。info()はINT8の入力と出力を報告し、それぞれにquant.scaleとquant.zero_pointが含まれます。リファレンスモデルの入力はimagesの1つで、INT8の[640, 640, 3] HWCテンソルです。

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

ホスト上で量子化する​

リファレンスモデルは、ピクセルが[0, 1]の範囲にある1枚のRGB画像を期待します。画像を640x640にリサイズし、BGRをRGBに変換し、255で割り、入口のパラメータで量子化の式を適用します。この前処理はアーカイブのコントラクトではなくモデルに属するものです。別のモデルには独自のレシピが必要です。同じコードを逆量子化した値を保持しておきます。これは、デフォルトルートが同条件の比較に必要とする正確なFP32入力です。

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

INT8ルートを実行する​

build()はMLAのみを含むカードのパイプラインを起動します。run()はINT8テンソルを受け付け、info().outputsが報告した順序と名前で、出力ごとに1つの密なINT8テンソルを返します。このルートは他のdtypeを拒否します。FP32のプッシュはカード上で量子化されるのではなく、失敗します。

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

逆量子化して比較する​

mla_onlyなしで2つ目のModelを構築し、逆量子化したFP32の値をデフォルトルートに送信します。各出力のパラメータでINT8ヘッドを逆量子化し、ヘッドごとの最大偏差をそのヘッドのスケール単位で出力します。両方のルートは同じコードで同じMLAプログラムを実行するため、誤差はゼロです。

pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

実行​

PCIeホストパッケージをインストールし、チュートリアルの設定で説明されているように、チュートリアルバンドルをダウンロードします。

このチュートリアルは、MLAの入出力を直接扱うようにコンパイルされたModel ZooのYOLO26n INT8アーカイブを実行します。抽出したPCIeエクストラのルートにダウンロードします。

sima-cli download https://docs.sima.ai/pkg_downloads/SDK2.1.3/models/modalix/yolo26-detection/yolo26n-det-int8-b1.tar.gz

その他のアーカイブは、Model SDKのtessellate_parametersでenable_mla=Trueを指定し、すべての入力にHWCのDRAMレイアウト、すべての出力にHWC16を指定してコンパイルされていれば条件を満たします。Model Zooのyolo_v8sビルドのように代わりにCVUでテッセレーションを行うアーカイブは、mla_onlyを有効にすると拒否されます。

mla_only does not support stage 'tessellate_quantize_0_MLA_0/...' (tess)

C++ (prebuilt):

./lib/sima-pcie-host/tutorials/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

C++ (build from source):

./build.sh --target tutorial_027_run_mla_only_int8
./build/tutorials-standalone/tutorial_027_run_mla_only_int8 \
--model yolo26n-det-int8-b1.tar.gz

両方のバージョンがコントラクトと、すべてのヘッドでゼロの偏差を出力します。

MLA-only contract:
input images INT8 [640, 640, 3] scale=0.00390434 zero_point=-128
output bbox_0 INT8 [80, 80, 4] scale=0.0302856 zero_point=-117
...
Dequantized MLA-only outputs vs the default route (error in scale units):
bbox_0 [80, 80, 4] max_err=0.0000
...
[OK] 027_run_mla_only_int8

デフォルトはカード0とキュー0です。別のカードを使用する場合にのみ--card Nを渡します。

実践​

INT8アーカイブでは、アプリケーションが量子化を所有する場合にmla_onlyを有効にします。センサーや前段のモデルからすでにINT8を生成している場合、独自の後処理のために生のINT8ヘッドが必要な場合、またはカード側のレイテンシから量子化と逆量子化のステージを取り除きたい場合です。すべてのスケールとゼロポイントはmodel.info()から読み取り、モデルの別のビルドからコピーしないでください。

このルートは全か無かです。すべての入力は対応するinfo().inputsエントリのdtype、形状、バイトサイズに一致する必要があり、画像の前処理やボックスデコードをmla_onlyと組み合わせることはできません。ホストがFP32データを保持し、変換を自分で行う必要がない場合は、デフォルトルートを使用してください。

デプロイメントの診断については、PCIeモデルワークフローに進んでください。

完全なソース​

完全なソースプログラムを表示
pcie_host/tutorials/027_run_mla_only_int8/run_mla_only_int8.cpp
// Quantize on the host, run only the MLA over PCIe, and dequantize the INT8 results.
//
// Usage:
// tutorial_027_run_mla_only_int8 --model yolo26n-det-int8-b1.tar.gz [--card 0]

#include <simaai/neat/pcie/Model.h>

#include <opencv2/imgcodecs.hpp>
#include <opencv2/imgproc.hpp>

#include <algorithm>
#include <cmath>
#include <cstdint>
#include <cstdlib>
#include <filesystem>
#include <iomanip>
#include <iostream>
#include <stdexcept>
#include <string>
#include <utility>
#include <vector>

namespace pcie = simaai::neat::pcie;

namespace {

constexpr int kBuildTimeoutMs = 180000;
constexpr int kRunTimeoutMs = 30000;
constexpr double kMaxErrorScales = 0.05; // one twentieth of a quantization step
constexpr char kImagePath[] = "share/sima-pcie-host/tutorials/assets/street-scene.png";

struct Args {
std::string model;
int card_id = 0;
};

std::string require_value(int argc, char** argv, int& index, const char* option) {
if (index + 1 >= argc) {
throw std::runtime_error(std::string("missing value for ") + option);
}
return argv[++index];
}

Args parse_args(int argc, char** argv) {
Args args;
for (int index = 1; index < argc; ++index) {
const std::string arg = argv[index];
if (arg == "--model") {
args.model = require_value(argc, argv, index, "--model");
} else if (arg == "--card") {
args.card_id = std::stoi(require_value(argc, argv, index, "--card"));
} else if (arg == "-h" || arg == "--help") {
std::cout << "Usage: " << argv[0] << " --model yolo26n-det-int8-b1.tar.gz [--card 0]\n";
std::exit(0);
} else {
throw std::runtime_error("unknown argument: " + arg);
}
}
if (args.model.empty()) {
throw std::runtime_error("--model is required: yolo26n-det-int8-b1.tar.gz from the Model Zoo");
}
return args;
}

std::string shape_string(const std::vector<std::int64_t>& shape) {
std::string text = "[";
for (std::size_t index = 0; index < shape.size(); ++index) {
text += (index == 0 ? "" : ", ") + std::to_string(shape[index]);
}
return text + "]";
}

std::size_t element_count(const std::vector<std::int64_t>& shape) {
std::size_t count = 1;
for (const auto dim : shape) {
count *= static_cast<std::size_t>(dim);
}
return count;
}

// The MLA-only contract quantizes every tensor with one scale and one zero point:
// x = (q - zero_point) * scale
// q = clamp(round(x / scale) + zero_point, -128, 127)
const pcie::QuantParams& require_quant(const pcie::TensorInfo& info) {
if (!info.quant.has_value()) {
throw std::runtime_error("tensor '" + info.name + "' publishes no quantization parameters");
}
return *info.quant;
}

std::int8_t quantize(const float value, const pcie::QuantParams& quant) {
const float code = std::nearbyint(value / quant.scale) + static_cast<float>(quant.zero_point);
return static_cast<std::int8_t>(std::clamp(code, -128.0F, 127.0F));
}

float dequantize(const std::int8_t code, const pcie::QuantParams& quant) {
return static_cast<float>(static_cast<std::int32_t>(code) - quant.zero_point) * quant.scale;
}

const std::int8_t* int8_data(const pcie::Tensor& tensor) {
return reinterpret_cast<const std::int8_t*>(static_cast<const std::uint8_t*>(tensor.data) +
tensor.byte_offset);
}

// Preprocessing of the reference model: one RGB HWC image with pixels in [0, 1]. Resize, BGR to
// RGB, divide by 255, then quantize with the ingress parameters. Another model needs its own
// recipe.
std::vector<std::int8_t> quantize_image(const cv::Mat& bgr, const pcie::TensorInfo& input) {
if (input.shape.size() != 3 || input.shape[2] != 3) {
throw std::runtime_error("the reference model takes a three-channel HWC input, got " +
shape_string(input.shape));
}
const pcie::QuantParams& quant = require_quant(input);
cv::Mat rgb;
cv::resize(bgr, rgb,
cv::Size(static_cast<int>(input.shape[1]), static_cast<int>(input.shape[0])));
cv::cvtColor(rgb, rgb, cv::COLOR_BGR2RGB);
std::vector<std::int8_t> codes(rgb.total() * rgb.channels());
for (std::size_t index = 0; index < codes.size(); ++index) {
codes[index] = quantize(static_cast<float>(rgb.data[index] / 255.0), quant);
}
return codes;
}

} // namespace

int main(int argc, char** argv) {
try {
const Args args = parse_args(argc, argv);
if (!std::filesystem::is_regular_file(args.model)) {
throw std::runtime_error("model does not exist: " + args.model);
}
const cv::Mat image = cv::imread(kImagePath, cv::IMREAD_COLOR);
if (image.empty()) {
throw std::runtime_error(std::string("OpenCV could not decode: ") + kImagePath);
}
pcie::ConnectionOptions connection;
connection.card_id = args.card_id;

// CORE LOGIC
pcie::ModelOptions options;
options.mla_only = true;
pcie::Model model(args.model, options, connection);
const pcie::ModelInfo info = model.info();
if (info.inputs.size() != 1) {
throw std::runtime_error("the reference model has exactly one input");
}
std::cout << "MLA-only contract:\n";
for (const auto& input : info.inputs) {
const pcie::QuantParams& quant = require_quant(input);
std::cout << " input " << input.name << " " << input.dtype << " "
<< shape_string(input.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}
for (const auto& output : info.outputs) {
const pcie::QuantParams& quant = require_quant(output);
std::cout << " output " << output.name << " " << output.dtype << " "
<< shape_string(output.shape) << " scale=" << quant.scale
<< " zero_point=" << quant.zero_point << '\n';
}

const pcie::TensorInfo& ingress = info.inputs.front();
std::vector<std::int8_t> codes = quantize_image(image, ingress);
std::vector<float> fp32_input(codes.size());
for (std::size_t index = 0; index < codes.size(); ++index) {
fp32_input[index] = dequantize(codes[index], *ingress.quant);
}
pcie::TensorList int8_inputs;
int8_inputs.push_back(pcie::Tensor::from_vector(std::move(codes), ingress.shape, ingress.name));

model.build(kBuildTimeoutMs);
const pcie::TensorList outputs = model.run(int8_inputs, kRunTimeoutMs);
model.close();

pcie::Model reference(args.model, {}, connection);
reference.build(kBuildTimeoutMs);
pcie::TensorList fp32_tensors;
fp32_tensors.push_back(
pcie::Tensor::from_vector(std::move(fp32_input), ingress.shape, ingress.name));
const pcie::TensorList reference_outputs = reference.run(fp32_tensors, kRunTimeoutMs);
reference.close();

std::cout << "Dequantized MLA-only outputs vs the default route (error in scale units):\n";
for (std::size_t index = 0; index < outputs.size(); ++index) {
const auto& spec = info.outputs[index];
const pcie::QuantParams& quant = *spec.quant;
const std::int8_t* codes = int8_data(outputs[index]);
const auto& card = reference_outputs[index];
if (card.route.name != spec.name || card.dtype != pcie::TensorDType::Float32) {
throw std::runtime_error("default route output '" + spec.name + "' is not FP32");
}
const auto* card_values = reinterpret_cast<const float*>(
static_cast<const std::uint8_t*>(card.data) + card.byte_offset);
double max_error = 0.0;
for (std::size_t element = 0; element < element_count(spec.shape); ++element) {
const double host = dequantize(codes[element], quant);
max_error = std::max(max_error, std::fabs(host - card_values[element]) / quant.scale);
}
std::cout << " " << spec.name << " " << shape_string(spec.shape) << " max_err=" << std::fixed
<< std::setprecision(4) << max_error << '\n';
if (max_error > kMaxErrorScales) {
throw std::runtime_error("output '" + spec.name + "' deviates from the default route");
}
}

std::cout << "[OK] 027_run_mla_only_int8\n";
return 0;
} catch (const std::exception& error) {
std::cerr << "[FAIL] " << error.what() << '\n';
return 1;
}
}

ソース​