Resume and extend#

Dataset generation supports interruption recovery and deterministic extension. Generate 100 cases, raise the same scene’s cases.count to 200 later, and resume the existing dataset:

result = generate_dataset_sync(request, output_root="data/generated", resume=True)

The batch page exposes the same behaviour through Resume / extend, and the HTTP API through POST /api/batches?resume=true.

How resume decides what to run#

Resume scans case artifacts rather than trusting counters in the manifest. A case is complete when case_info.json and every artifact selected by the output policy exist; failed, interrupted, corrupt or incomplete cases are generated again. Because each case has an independent seed derived from the dataset seed and its scene and case indices, existing cases keep their parameters when the target count grows.

Compatibility checks#

To protect consistency, resume only accepts a request whose following parts are unchanged:

  • the dataset seed,

  • every scene definition, including the resolved building source and its version,

  • solver settings and parameter ranges,

  • the output selection,

  • the software version and, for Sionna cases, the radio-map contract version.

Per-scene case counts may stay the same or increase. Persisted building and terrain inputs rebuild the runtime scene, so a new provider release is never mixed into an old dataset.

Version 0.2.1 introduced radio-map contract version 2, which changed SINR and SINR-based association. Resuming a Sionna dataset with an older or absent contract version is rejected: keep the old dataset and use a new output directory, or continue it with its original generator. Preview-only datasets from 0.2.0 can resume under 0.2.1.

Each run appends a record to generation_runs in the manifest with its initial, generated and failed case counts, so the history of a dataset stays visible.