Planning and generation#

The request#

A dataset request is a validated document with a name, a seed, an output selection and a list of scenes. Each scene carries its geometry options and a case space. The Python interface, the batch page and POST /api/batches accept the same schema.

from openworld_radio_twin.dataset import generate_dataset_sync, plan_dataset

request = {
    "dataset_name": "my-radio-dataset",
    "random_seed": 2026,
    "scenes": [
        {
            "label": "Amsterdam",
            "latitude": 52.3762,
            "longitude": 4.8993,
            "radius_m": 250,
            "include_buildings": True,
            "building_source": "auto",
            "include_terrain": True,
            "material_profile": "itu",
            "terrain_resolution_m": 5.0,
            "receiver_height_m": 1.5,
            "voxel_pitch_m": 3.0,
            "cases": {
                "strategy": "random",         # "fixed" | "random"
                "count": 8,
                "engine": "sionna",           # "preview" | "sionna"
                "engine_config": {"device": "cuda"},
                "transmitter_count": {"minimum": 1, "maximum": 3},
                "frequency_ghz": {"minimum": 3.5, "maximum": 3.5},
                "power_dbm": {"minimum": 24, "maximum": 36},
                "altitude_m": {"minimum": 15, "maximum": 35},
                "offset_radius_m": {"minimum": 0, "maximum": 250},
                "azimuth_deg": {"minimum": 0, "maximum": 359},
                "downtilt_deg": {"minimum": 0, "maximum": 12},
                "antenna_patterns": ["sector"],
                "resolutions_m": [5],
                "max_depths": [3],
                "samples_per_tx": [1_000_000],
                "association_metrics": ["sinr"],
            },
        }
    ],
}

plan = plan_dataset(request)
print(plan.total_cases, plan.total_receiver_cells, plan.total_initial_ray_samples)

result = generate_dataset_sync(request, output_root="data/generated")
print(result.output_directory, result.manifest_path)

plan_dataset expands the request without provider access or solver execution, so the exact case names, solver seeds, receiver-cell count and initial ray budget can be inspected first. generate_dataset_sync is for scripts; await generate_dataset(...) serves notebooks and services. Both own their HTTP client and worker thread, and a request may also be given as a BatchDatasetRequest model or a path to a JSON file. Copy examples/generate_dataset.py for a self-contained, environment-configurable script.

Scene fields#

Field

Default

Range

label

required

1 to 80 characters; slugified into the directory name

scene_mode

"geospatial"

"empty" disables every environment query

latitude, longitude

required

WGS84

radius_m

750

100 to 5,000

receiver_height_m

1.5

0.1 to 100

voxel_pitch_m

3.0

0.5 to 25

include_buildings

true

building_source

"auto"

any registered ID (Building sources)

include_terrain

false

DEM-following ground when true, flat plane when false

material_profile

"itu"

"itu" or "uniform"

terrain_resolution_m

5.0

1 to 100; the terrain grid may not exceed 250,000 cells

Case space#

Field

Default

Range

strategy

"random"

"fixed" uses the minimum of every range and the first entry of every list

count

8

1 to 10,000 per scene

engine

"preview"

"preview" or "sionna"

engine_config

{}

Sionna keys (Sionna RT)

transmitter_count

1 to 3

1 to 8

frequency_ghz

3.5

0.1 to 100; 1 to 10 for Sionna

power_dbm

24 to 36

-20 to 80

altitude_m

15 to 35

0 to 1,000, above ground

offset_radius_m

0 to 250

0 to the scene radius

azimuth_deg

0 to 359

0 to 360 exclusive

downtilt_deg

0 to 12

-30 to 90

antenna_patterns

["sector"]

"sector", "isotropic"

resolutions_m

[15]

1 to 200; grid limits depend on the solver

max_depths

[3]

0 to 8

samples_per_tx

[1000000]

10,000 to 100,000,000

association_metrics

["sinr"]

"path_gain", "rss", "sinr"

The 10,000-case limit per scene is an operational validation limit rather than a Sionna constraint. A dataset accepts up to 20 scenes.

Seeded expansion#

Random plans use NumPy’s PCG64. Each case derives an independent stream from SeedSequence([dataset_seed, scene_index, case_index]) with one-based indices; the fully expanded transmitter coordinates, physical parameters and solver seed are written to the manifest. The first transmitter anchors the shared scene origin, and additional transmitters are sampled uniformly by area within the offset annulus. Rebuilding a plan with the same request and software version produces the same cases, and cases 1 to 100 stay unchanged when a scene’s count later grows to 200.

Empty scenes#

scene_mode="empty" is a true free-space scene. It performs no building, terrain or surface query and writes a Sionna scene with no ground or building shapes; the coordinates still define the local frame and transmitter positions.

{
    "label": "Empty free space",
    "scene_mode": "empty",
    "latitude": 52.3762,
    "longitude": 4.8993,
    "radius_m": 250,
    "cases": {"engine": "sionna", "count": 1},
}

For a flat ground plane without buildings, keep scene_mode="geospatial" and set include_buildings=False and include_terrain=False.

Selecting case artifacts#

Case artifacts are chosen once per dataset. The default keeps every array and PNG; empty lists produce a metadata-only case while case_info.json stays for provenance and resume:

request["output"] = {
    "array_metrics": ["rss"],        # subset of path_gain, rss, sinr
    "image_metrics": [],
    "association_image": False,
}

The example script exposes full, arrays_only, rss_only and metadata_only profiles through OWRT_EXAMPLE_OUTPUT_PROFILE. Output selection is part of the resume contract, so a changed selection needs a new dataset name. Shared scene assets, including terrain arrays, are retained regardless of this policy.

Execution model#

Cases run sequentially within a dataset so GPU memory stays bounded and scene geometry can be reused safely. Sionna availability is checked before any provider access. A non-resume run never overwrites an existing dataset; a dataset directory is locked while it is being written. Progress callbacks receive the completed count, total count and a message. OWRT_DATASET_ROOT sets the default output root (data/generated).