BoxDecode Decode Types
nodes::SimaBoxDecode converts raw detection-head tensors into detection results. It runs after model inference, applies the decode math for the selected model family, filters low-confidence boxes, runs NMS, and emits a tensor payload that starts with decoded boxes. Detection models can parse that payload as boxes; pose and segmentation models can also parse the keypoints or masks that follow the boxes.
For normal model-pack usage, prefer the Model-aware constructor. The model archive supplies the tensor order, layout, quantization, class count, resize metadata, and score-domain hints needed by the decoder. Your application usually only chooses the decode family and filtering thresholds.
Quick start
using namespace simaai::neat;
Model model("/path/to/yolov8_model.tar.gz");
auto boxdecode = nodes::SimaBoxDecode(
model,
BoxDecodeType::YoloV8,
/* detection_threshold */ 0.25,
/* nms_iou_threshold */ 0.45,
/* top_k */ 100);
For standalone stage usage:
simaai::neat::stages::BoxDecodeOptions opt(simaai::neat::BoxDecodeType::YoloV8);
opt.detection_threshold = 0.25;
opt.nms_iou_threshold = 0.45;
opt.top_k = 100;
Arguments
| Argument | Meaning |
|---|---|
decode_type | Model family/head format, such as BoxDecodeType::YoloV8 or BoxDecodeType::YoloX. Required. |
detection_threshold | Minimum score required to keep a detection. Use a model-appropriate value such as 0.25. |
nms_iou_threshold | IoU threshold used by non-maximum suppression. |
top_k | Maximum number of detections to keep. 0 uses the backend/model default. |
original_width, original_height | Source image size for coordinate mapping when using the raw-geometry constructor. |
model_width, model_height | Model input size override. With the Model constructor this changes spatial decode knobs, not the packaged tensor contract. |
resize_mode_override | Use only when no upstream Preproc stage writes resize metadata and you need to specify stretch/letterbox/crop behavior explicitly. |
decode_type_option | Advanced sub-layout selector. Leave as Auto for model-pack usage unless you know the exported head layout. |
Inputs and outputs
Input: raw detection tensors from the model. The expected tensor shapes depend on the model family. With an MPK/model archive, Neat reads those details from the packaged contract.
Output: one BoxDecode tensor containing decoded detections. Detection models use the standard BBOX payload. Pose and segmentation models keep the same leading boxes and append their task-specific payload:
| Model task | C++ helper | Python helper | Decoded tensors |
|---|---|---|---|
| Detection | decode_bbox(...) | pyneat.decode_bbox(...) | [N, 6] float32 boxes: x1, y1, x2, y2, score, class_id |
| Pose | decode_pose(...) | pyneat.decode_pose(...) | boxes [N, 6] and keypoints [N, 17, 3] float32: x, y, visibility |
| Segmentation | decode_segmentation(...) | pyneat.decode_segmentation(...) | boxes [N, 6] float32 and masks [N, 160, 160] uint8 |
| Segmentation + pose | decode_segmentation_pose(...) | pyneat.decode_segmentation_pose(...) | boxes [N, 6] float32, masks [N, 160, 160] uint8, keypoints [N, 17, 3] float32 |
| SuperPoint | decode_superpoint(...) | pyneat.decode_superpoint(...) | keypoints [N,2], scores [N], descriptors [N,D] |
Detection-display graphs can feed the result to SimaRender. Application code that only needs boxes can continue to use decode_bbox(...) on BoxDecode outputs.
SuperPoint
SuperPoint remains part of the BoxDecode product surface, but emits feature points rather than pretending that they are boxes. The minimal A65-default configuration is:
BoxDecodeOptions options{BoxDecodeType::SuperPoint};
options.superpoint.descriptor_output_dtype = TensorDType::Float32;
auto decoder = nodes::SimaBoxDecode(model, options);
Python uses the same defaults:
options = pyneat.BoxDecodeOptions(pyneat.BoxDecodeType.SuperPoint)
options.superpoint.descriptor_output_dtype = pyneat.TensorDType.Float32
decoder = pyneat.nodes.sima_box_decode(model, options=options)
A65V1 is the default profile. Select another profile explicitly when the
model requires different numerical behavior; Neat does not infer behavior from
tensor shapes or values:
| Profile | When to select it | Production status |
|---|---|---|
LightGlueV1 | LightGlue-compatible detector, NMS, coordinate, and descriptor behavior | Supported |
MagicLeapDemoV1 | The pinned Magic Leap demo behavior | Supported |
A65V1 | Compatibility with the former A65 SuperPoint decoder | Supported; default |
PaperBicubicV1 | Reserved numeric ID for a future fully specified bicubic policy | Rejected until production-defined |
Numerical behavior and output encoding are independent. For example, select A65 numerical behavior with the default V1 output:
BoxDecodeOptions options{BoxDecodeType::SuperPoint};
options.superpoint.profile = SuperPointProfile::A65V1;
options.superpoint.output_format = SuperPointOutputFormat::FeaturePointsV1;
The legacy byte layout is opt-in and has additional constraints:
options.superpoint.profile = SuperPointProfile::A65V1;
options.superpoint.output_format = SuperPointOutputFormat::LegacyA65InterleavedV0;
options.superpoint.descriptor_output_dtype = TensorDType::Int8;
SuperPointProfile::Auto first uses authoritative MPK superpoint.profile metadata. If neither
the API, Model::Options.superpoint.profile, nor MPK supplies a profile, it resolves to A65V1.
Neat never guesses a profile from tensor shapes, values, filenames, or downstream nodes.
When their public sentinel values are left unchanged, detection_threshold=0.0, top_k=0,
nms_radius=-1, and border_margin=-1 resolve from the selected profile. A65V1 resolves to a
threshold of 0.1, Top-K 600, NMS radius 4, and border margin 0. LightGlueV1 and
MagicLeapDemoV1 use thresholds 0.0005 and 0.015, respectively; both use Top-K 600, NMS
radius 4, and border margin 4.
nms_iou_threshold does not apply to SuperPoint; use the pixel-radius
superpoint.nms_radius. The default output is the versioned FEATURE_POINTS_V1 structure-of-arrays
payload. LegacyA65InterleavedV0 is an explicit migration format and requires 256-dimensional INT8
descriptors. Use decode_superpoint rather than decode_bbox or BoxDecodeResults.
Versioned MPK superpoint schema v1 records are fail-closed. They must name the profile, distinct
detector and descriptor tensor IDs, a sha256: fingerprint with 64 hexadecimal digits, and the
supported input representations raw-logits-65 and coarse-pre-l2. Schema 0 remains accepted only
as a migration/manual record; omitted schema-0 representation fields are canonicalized to those two
raw-input representations and recorded as defaults in diagnostics. Unknown schema versions or
representation tokens fail contract compilation.
If an API profile override conflicts with a fingerprint stamped for a different MPK profile,
re-stamp the MPK for the selected profile; Neat does not discard or reinterpret that provenance.
BBOX wire payload
Detection decode emits one tensor tagged BBOX per input frame. That tensor is
a rank-1 UInt8 byte buffer:
| Field | Value |
|---|---|
semantic.detection.format | "BBOX" |
dtype | UInt8 |
shape | [N_bytes], where N_bytes is the packed buffer capacity from the model archive |
The tensor shape is a byte count, not a detection count. The payload uses little-endian layout:
offset size content
------ ---- -------
0 4 uint32 N = valid detections in this frame
4 24 RawBox[0]
28 24 RawBox[1]
. . ...
. . RawBox[N-1]
trailing bytes are padding and must be ignored
Each RawBox record is 24 bytes:
| Offset | Size | Type | Field | Meaning |
|---|---|---|---|---|
| 0 | 4 | int32 | x | Top-left x in source pixels. |
| 4 | 4 | int32 | y | Top-left y in source pixels. |
| 8 | 4 | int32 | w | Width in source pixels. |
| 12 | 4 | int32 | h | Height in source pixels. |
| 16 | 4 | float32 | score | Post-NMS confidence in [0.0, 1.0]. |
| 20 | 4 | int32 | class_id | Model-defined class id. |
The matching Python struct format for one record is "<iiiifi".
Coordinates are in original-image pixels when upstream preprocessing metadata is
present. They are not normalized to [0, 1] and are not expressed in the
model's internal letterboxed input space.
Combined segmentation + pose payload
yolox-seg-pose emits a single buffer carrying three regions. Every region is
strided by the same slot count, top_k:
| Region | Offset | Stride | Contents |
|---|---|---|---|
| header | 0 | 4 | int32 detection count |
| boxes | 4 | 24 | BoundingBoxOut records, as above |
| masks | 4 + 24*top_k | mask_w * mask_h | uint8, one plane per slot |
| poses | 4 + (24 + mask_w*mask_h)*top_k | 204 | 17 x {uint32 x, uint32 y, float32 visibility} |
Use decode_segmentation_pose(...) to get boxes, masks and keypoints. Row i in each tensor describes the same detection. Use decode_bbox(...) or BoxDecodeResults(...) if you only need boxes.
There are 17 keypoint slots per detection. The backend zeros unused slots and keypoints for classes excluded by pose_classes. The helper copies those values unchanged.
Core derives num_classes from the class-head depth minus one, because channel 0 holds objectness. For example, 30 channels means 29 classes. An explicit num_classes must match.
Keypoint classes
Set the class IDs that have keypoints:
options = pyneat.BoxDecodeOptions(pyneat.BoxDecodeType.YoloXSegPose)
options.yolox_seg_pose.pose_classes = [0, 5]
ModelOptions.yolox_seg_pose accepts the same setting.
- Classes outside the list get zero keypoint coordinates and visibility.
- An empty
BoxDecodeOptionslist inherits the model's settings. With no list configured, every class gets keypoints. - IDs must be unique and in
[0, num_classes). This option is only supported forYoloXSegPose.
When model.run returns raw heads
Some model routes return raw feature-map heads from model.run(...) instead of
a decoded BBOX tensor. That is not a failed run. It means the model executed,
but the route did not include BoxDecode at the point where you read output.
Use this rule:
detections=...or aBBOXtensor: parse the packed BBOX payload or use the decode helpers.raw_output_heads=...: add a BoxDecode stage, inspect the model route, or consume the raw tensors with model-specific postprocessing.
Do not parse raw heads as boxes. The raw tensor layout depends on the exported model family and model archive contract.
Override contract
The model archive can provide defaults for decode type, thresholds, top_k, and
source geometry. Runtime arguments override those defaults only when you pass a
non-empty or positive value.
| Runtime argument | Value passed | Behavior |
|---|---|---|
decode_type | empty / Unspecified | Preserve model archive or route-planner inference where supported. |
decode_type | concrete type | Override the decode family for this run. |
original_width / original_height | 0 | Preserve packaged geometry or upstream preprocess metadata. |
original_width / original_height | positive integer | Override source dimensions for coordinate mapping. |
detection_threshold / score_threshold | 0.0 | Preserve packaged threshold. |
detection_threshold / score_threshold | > 0.0 | Override the score gate. |
nms_iou_threshold | 0.0 | Preserve packaged NMS IoU. |
nms_iou_threshold | > 0.0 | Override NMS IoU. |
top_k | 0 | Preserve packaged top-K. |
top_k | > 0 | Override the maximum kept detections. |
num_classes compares the value configured by the caller with the class count
derived from the MPK tensor contract:
| Model family | Configured num_classes | MPK-derived num_classes | Behavior |
|---|---|---|---|
| Any supported model | 0 | positive, inferable value | Use the MPK-derived class count. |
| Any supported model | positive integer | same value | Use the configured class count. |
| Model with an ambiguous class split | positive integer | unavailable | Use the configured class count. This is required when a split single-class head cannot be inferred reliably. |
| YOLOv5 or YOLO26 | positive integer | different value | Fail before pipeline construction and report both values. These raw-head layouts derive their class count from tensor depth. |
| SSD or another pre-YOLO26 non-pose YOLO family | positive integer | different value | Apply the existing family-specific explicit-override behavior. Pose decoders and SuperPoint retain their family-specific rules. |
detection_threshold is the name used by the BoxDecode node/stage
constructors. ModelOptions.score_threshold is the model-route option that
feeds the same control.
Decode type mapping
| API enum | Backend token | Typical model family |
|---|---|---|
BoxDecodeType::Yolo | yolo | Generic YOLO-style heads |
BoxDecodeType::YoloV5 | yolov5 | YOLOv5 detection |
BoxDecodeType::YoloV5Seg | yolov5-seg | YOLOv5 segmentation |
BoxDecodeType::YoloV7 | yolov7 | YOLOv7 detection |
BoxDecodeType::YoloV7Seg | yolov7-seg | YOLOv7 segmentation |
BoxDecodeType::YoloV8 | yolov8 | YOLOv8 detection |
BoxDecodeType::YoloV8Seg | yolov8-seg | YOLOv8 segmentation |
BoxDecodeType::YoloV8Pose | yolov8-pose | YOLOv8 pose |
BoxDecodeType::YoloV9 | yolov9 | YOLOv9 detection |
BoxDecodeType::YoloV9Seg | yolov9-seg | YOLOv9 segmentation |
BoxDecodeType::YoloV10 | yolov10 | YOLOv10 detection |
BoxDecodeType::YoloV10Seg | yolov10-seg | YOLOv10 segmentation |
BoxDecodeType::YoloV26 | yolo26 | YOLO26 detection |
BoxDecodeType::YoloV26Pose | yolo26-pose | YOLO26 pose |
BoxDecodeType::YoloV26Seg | yolo26-seg | YOLO26 segmentation |
BoxDecodeType::YoloV6 | yolov6 | YOLOv6 detection |
BoxDecodeType::YoloX | yolox | YOLOX detection |
BoxDecodeType::YoloXSegPose | yolox-seg-pose | YOLOX packed export carrying box, mask and keypoint heads together |
BoxDecodeType::Ssd | ssd | Exact prepared SSD300, SSD-Mobile-300, SSD-Mobile-320, or SSDlite-Mobile-320 contract, selected from ordered head geometry |
BoxDecodeType::SuperPoint | superpoint | SuperPoint detector and descriptor postprocessing |
BoxDecodeType::Detr | detr | DETR-style transformer detection |
BoxDecodeType::EffDet | effdet | EfficientDet detection |
BoxDecodeType::RcnnStage1 | rcnn-stage1 | R-CNN proposal stage |
BoxDecodeType::Centernet | centernet | CenterNet detection |
BoxDecodeType::Unspecified is an unset sentinel and fails before runtime. SSD recipe identity is
an internal Core contract (ssd300-v1, ssd-mobile-300-v1, ssd-mobile-320-v1, or
ssdlite-mobile-320-v1), not another public decode type or
backend token. Core resolves it before lowering, while the installed object decoder continues to
receive its supported ssd family token and selects the corresponding fixed implementation from
the already validated head geometry.
Choosing the right type
- If you are using a SiMa-provided or SiMa-compiled model pack, choose the
BoxDecodeTypethat matches the model family and leavedecode_type_optionasAuto. - If your detections are missing or all scores are unexpectedly low, first verify that the decode family matches the exported model head. YOLOX, YOLOv6, and YOLO26 use raw/logit-style heads and should not be treated like probability-only YOLO heads.
- If boxes are shifted or scaled incorrectly, check the image resize policy. Use
resize_mode_overrideonly when your graph does not have an upstreamPreprocstage writing resize metadata. - If you are authoring a custom model pack, ensure the archive describes the detection heads accurately: tensor order, logical shape, physical storage, dtype/quantization, score domain, class count, and any sliced outputs. Application code should not need to compensate for these details.
Shape and layout guidance
Different detection models expose different head layouts. Some use one tensor per feature-map level; others split boxes, objectness, classes, keypoints, or masks into separate tensors. Some model outputs are dense HWC tensors; others are packed or sliced by the compiler/runtime.
For model-pack flows this is handled by the packaged contract. For manually wired tensors, the key rule is: match the exported head format exactly. Do not choose a decode type based only on rank or channel count.
Advanced tensor-contract rules:
-
YOLO-family decode types other than
YoloV5detection (Yolo,YoloV7,YoloV8,YoloV9,YoloV10, and segmentation/pose variants) expect either decoupled heads or packed heads that match the model family. -
Packed YOLO heads must keep class count and head depth consistent across feature levels.
-
YoloV5detection accepts exactly three undecoded packed heads ordered P3/P4/P5. Their grids must have stride-8/16/32 geometry and each logical depth must be3 * (num_classes + 5). BoxDecode applies sigmoid, the grid and stride transform, and the standard YOLOv5 anchors ({10,13},{16,30},{33,23};{30,61},{62,45},{59,119};{116,90},{156,198},{373,326}). Custom AutoAnchor tables and decoded six-tensor box/class exports must use another contract. -
YoloV26uses grouped raw l/t/r/b bbox heads plus class-score heads. -
Ssdis not a generic SSD decoder. It resolves exactly four prepared profiles from the complete ordered loc/conf H/W/C signature at compile time. Any other head set or order is rejected with an error that prints the observed and supported signatures:- SSD300 (
dboxes300_coco): 300×300 input, feature maps{38,19,10,5,3,1}, priors-per-cell{4,6,6,6,4,4}, confidence channel orderclass*A + anchor, class scores via softmax over the class dimension (background at index 0 included). - SSD-Mobile-300-v1 (
ssd_anchor_generator): 300×300 input, feature maps{19,10,5,3,2,1}, priors-per-cell{3,6,6,6,6,6}, confidence channel orderanchor*C + class, class scores via per-class sigmoid (background ignored). - SSD-Mobile-320-v1 (
ssd_anchor_generator): 320×320 input, feature maps{20,10,5,3,2,1}, priors-per-cell{3,6,6,6,6,6}, confidence channel orderanchor*C + class, class scores via per-class sigmoid (background ignored). - SSDlite-Mobile-320-v1 (TorchVision
DefaultBoxGenerator): 320×320 input, feature maps{20,10,5,3,2,1}, six priors per cell at every level, localization orderanchor*4 + {dx,dy,dw,dh}, confidence orderanchor*C + class, and class scores via softmax over all 91 classes including background.
All recipes use grouped per-level localization heads (depth =
4 * priors-per-cell) paired with class-confidence heads (depth =num_classes * priors-per-cell), FasterRcnnBoxCoder variance scaling (scale_xy 0.1,scale_wh 0.2), and a stretch (anisotropic) preprocessing resize. The score activation is fixed by the recipe (matching the on-device decoder), and the grouped-by-role layout is selected automatically — leavedecode_type_optionasAuto. A non-grouped layout token is rejected.The model frame is part of the profile, not just the head geometry. SSD300-v1 and SSD-Mobile-300-v1 require 300×300; both 320-v1 profiles require 320×320. A resolved preprocess resize target or model-dimension override of any other size is rejected at build time, because the prior tables and the stretch back-projection are only valid at that frame.
Raw/standalone
SimaBoxDecodeconstruction never invents a resize mode. Keep the upstreamPreprocmetadata requirement, or use the explicit raw overload to assert externally performedResizeMode::Stretch; Letterbox and Crop are rejected.num_classescontract. The encoded class count is always derived from the confidence-head depth (conf_depth / priors-per-cell, background at index 0 included). SSD300-v1 permits a contiguous prefix selection such as the prepared 81-to-8 route; the other three profiles require the exact encoded count. An invalid selection is rejected at build time. Leave it unset to use the profile default. - SSD300 (
-
Detrinfers class channels from the maximum head depth and requires a valid class dimension. -
EffDet,RcnnStage1, andCenternetuse their model-family contracts; do not route them through a YOLO decode type. -
*-segdecode types produce box-leading output plus task-specific mask data.
If a custom model pack does not match either complete ordered signature, prepare a new explicitly supported profile rather than weakening the matcher.
Python note
When configuring model options from Python, use the typed enum rather than a string when available:
opt = pyneat.ModelOptions()
opt.decode_type = pyneat.BoxDecodeType.YoloV8
Parse outputs with the helper that matches the model task:
outputs = model.run([image])
boxes = pyneat.decode_bbox(outputs)[0].to_numpy()
pose = pyneat.decode_pose(outputs)[0]
pose_boxes = pose.boxes.to_numpy()
keypoints = pose.keypoints.to_numpy()
seg = pyneat.decode_segmentation(outputs)[0]
seg_boxes = seg.boxes.to_numpy()
masks = seg.masks.to_numpy()
Upgrading
Replace the preview API's top-level pose_classes with yolox_seg_pose.pose_classes.
The new options change C++ object layouts. Core uses ABI 5, libsima_neat.so.5. Rebuild C++ applications, plugins and Python bindings with matching headers and libraries. Do not link ABI 4 binaries to ABI 5 through a compatibility symlink.