Dataset layout#
data/generated/<dataset>/
|-- dataset_manifest.json
`-- scene_001_<label>/
|-- scene_info.json
|-- provenance.json
|-- source_buildings.json
|-- scene.xml
|-- mesh/
|-- terrain/
|-- measurement_surfaces/
|-- vox_slices/
|-- vox_depth/
|-- 2D_Building_Height_Map.npy
`-- cases/
`-- case_0001_<antenna>_<n>tx_<frequency>ghz/
|-- case_info.json
|-- arrays/
| |-- path_gain/path_gain.npz
| |-- rss/rss.npz
| `-- sinr/sinr.npz
`-- images/
|-- path_gain_db.png, path_gain_db_masked.png
|-- rss_db.png, rss_db_masked.png
|-- sinr.png, sinr_masked.png
`-- association.png
Manifest#
dataset_manifest.json holds the validated request, the expanded plan with every case’s
seed and transmitter parameters, the software version, the resume contract, per-scene
status, and a generation_runs list with the initial, generated and failed case counts of
each run.
Scene directory#
Each scene directory is a complete compiled scene (Scene layout) plus the derived
assets that its case space needs: one measurement surface per requested cell size and
receiver height, voxel slices and depth at the requested pitch. With include_terrain=True
the terrain/ arrays are written for every scene and are reloaded for resume without a new
provider query.
Case directory#
case_info.json is mandatory and records the resolved simulation request, the solver
details and warnings, the data provenance of buildings, terrain and surfaces, and the
artifact list. The arrays are compressed NPZ files containing the linear float32 metric
array with shape (n_tx, rows, columns), rows south to north. PNGs are north-up, one pixel
per cell, with transparent no-data and no interpolation; the _masked variants hide building
pixels. association.png is rendered from the discrete serving-transmitter labels.
Metric |
Array unit |
Image scale |
|---|---|---|
|
linear ratio |
dB |
|
W |
dBm |
|
linear ratio |
dB |
association |
integer transmitter index |
categorical |
Reading a dataset#
import json
import numpy as np
from pathlib import Path
case = Path("data/generated/my-radio-dataset/scene_001_amsterdam/cases/case_0001_sector_2tx_3.5ghz")
info = json.loads((case / "case_info.json").read_text())
rss_w = np.load(case / "arrays/rss/rss.npz")["rss"] # (n_tx, rows, columns)
best_dbm = 10 * np.log10(rss_w.max(axis=0)) + 30 # strongest transmitter per cell
scene = case.parents[1]
local_elevation = np.load(scene / "terrain/local_elevation_m.npy")
surface_classes = np.load(scene / "terrain/surface_classes.npy")
Elevation arrays are grid nodes and surface classes are cells, so their shapes differ by one in each dimension. Use the coordinates stored with each product when aligning them with the north-up radio images.