Scene compile performance#

Loading buildings and compiling a scene were the slowest steps of load_scene, the explorer and dataset generation. This page records what the compile path cost before the perf/scene-compile-speed changes, what was changed, how the outputs were checked, and the measurements after the change. All numbers come from one AMD EPYC 7282 node (64 cores) with Python 3.11, NumPy 2.4, Shapely 2.1 and pyarrow 25; the network figures also depend on the remote services on the day of the run.

Summary#

Step

Before

After

Compile Bergen, 3,574 footprint buildings, r = 750 m (offline)

23.4 s

4.3 s

Compile Amsterdam, 2,282 3DBAG LoD2.2 shells, r = 500 m (offline)

36.3 s

11.9 s

Compile Los Angeles, 1,120 footprint buildings, r = 450 m (offline)

5.8 s

1.8 s

1 m measurement surface, Bergen (2.25 million cells)

181.9 s

4.7 s

Voxels at 3 m, Amsterdam shells

19.2 s

5.0 s

Provider queries, Bergen (Overture buildings, Mapzen terrain, Overture surfaces)

28.8 s

10.5 s

Provider queries, Los Angeles (Overture, USGS 3DEP, Overture surfaces)

23.7 s

4.5 s

Provider queries, Amsterdam (3DBAG API, Mapzen, Overture surfaces)

222 s

39 s

Cold load_scene compile, Bergen (queries + compile)

50.3 s

16.0 s

Cold load_scene compile, Amsterdam (queries + compile)

253 s

55 s

Explorer read-back of the cached Amsterdam scene (load and scene_response)

12.7 s

4.7 s

Offline figures are medians of three runs of scripts/benchmarks/benchmark_scene_compile.py (two runs for the baseline); network figures are medians of the runs listed under Network measurements. “Before” is commit b5d483d.

The compiled scenes are byte-for-byte the same. Every PLY mesh, the height map, the terrain arrays, the voxel and measurement-surface products, and the parsed content of every JSON and XML file were compared between the old and the new code on the three cached scenes above: 7,173 files for Bergen, 4,589 for Amsterdam and 2,264 for Los Angeles, with no difference.

Where the time went#

Profiling build_scene_assets on the cached provider records showed that the compile step was bound by Python loops, not by geometry:

  • _regular_grid_mesh built the terrain mesh one cell at a time with a NumPy cross product per triangle: 12.7 s of the 35 s profile for Bergen. The same function builds every measurement surface, which is why a 1 m receiver grid took three minutes.

  • _terrain_footing_elevation created a Shapely polygon for every terrain triangle under every building footprint and tested them one by one: 9.8 s for Bergen.

  • _write_height_map allocated a full-scene image per building and rasterized into it, then indexed the whole height map to apply the roof height: 3.7 s for Bergen.

  • For native shells, _shell_intervals intersected vertical rays with each triangle in a Python loop with a per-triangle meshgrid: 14 s for Amsterdam. dataclasses.asdict deep-copied every shell vertex for provenance.json, and the indented JSON encoder (pure Python) then spent 6.7 s on it; the compiler read that file back and rewrote it with the same encoder to append the compilation record.

  • write_binary_ply packed vertices and faces with struct.pack per element, and the scene XML was pretty-printed through minidom.

The provider queries were sequential: buildings, then terrain, then the four Overture surface themes one after another. Each Overture read downloaded the release’s STAC index and opened a new S3 client, and read every column of the matching row groups. The 3DBAG API was paged ten items at a time along next links, which is 229 sequential requests for the Amsterdam scene. On the read-back side, load_scene_assets opened both PLY files of every building to read their headers, parsed the building records although the placement records were already present, and scene_response re-derived the terrain footing of every native shell to build the display meshes.

What changed#

Compile step (simulation/scene_builder.py, simulation/native_geometry.py):

  • Terrain and measurement-surface meshes are built with array indexing; triangle areas use numpy.linalg.norm over the last axis, which rounds like the previous per-triangle norm.

  • Building footing tests all candidate terrain triangles with one vectorized Shapely intersects call; footprint rings are converted with one transformer call per ring.

  • Footprints are rasterized inside their pixel bounding box and shifted by whole pixels, so the scan conversion is unchanged; height map, receiver slice and voxel products index only that window.

  • Ray-shell intersections are evaluated for all (triangle, grid point) pairs of a building at once, in chunks of at most two million pairs; columns are paired in one pass, and the few columns with crossings closer than the tolerance keep the sequential rule. Shell components use label propagation instead of a Python union-find.

  • provenance.json and source_buildings.json are written with the native JSON encoder, one record per line, so the files stay readable and parse to the same objects. Placement records are built without deep copies. The scene XML is indented with ElementTree.

  • PLY files are written from NumPy buffers with the same little-endian layout.

  • Native shells are split into roof and wall faces with a boolean mask; footprint roofs are triangulated once and oriented with the same signed-area rule as before.

Acquisition (compiler.py, providers/):

  • Buildings and the environment grid are queried at the same time; the first failure cancels the other and is reported as before.

  • providers/overture_reader.py keeps one S3 client per process and the STAC index of each release in memory and on disk (OWRT_OVERTURE_INDEX_ROOT), retries the index download briefly when the server throttles, and projects each read to the columns the providers decode. The row filter is the same bounding-box predicate. The index is the only way objects are selected: when it cannot be fetched, the Overture query fails, which sends buildings to the OpenStreetMap fallback and reports the surface query as failed, instead of the upstream reader’s scan of the whole theme directory.

  • The Overture building and building-part objects, and the four surface themes, are read concurrently; features are assembled in the original order.

  • USGS 3DEP export and catalog lookup run together.

  • The 3DBAG provider requests 100 items per page, the service maximum. When the service reports numberMatched, the remaining pages up to the building limit are fetched by offset, four at a time, and decoded in offset order; the last page’s next link is followed as before. Decoding a CityJSON face orients all of its triangles in one array operation instead of one Shapely and NumPy call per triangle; on 4,825 faces of real 3DBAG pages the triangles are identical.

Service responsiveness (compiler.py, scenes.py, main.py, providers/three_d_bag.py):

  • The 3DBAG provider decoded every page on the server’s event loop, so a large view request (a 2.5 km view radius over Amsterdam returns about 7,200 shells) stalled every other request for minutes, which the explorer reported as failed fetches. Decoding, scene building, reading a compiled scene back and building its response now run in worker threads; during such a request the service keeps answering within a few seconds. The provider also keeps at most four page requests in flight instead of fetching all pages before decoding.

Read-back (scene_builder.load_scene_assets, scenes.py, main.py, models.py):

  • Building mesh triangle counts come from the placement records the compiler wrote; the PLY headers are read only for legacy caches without records, and the building records are parsed only on that legacy path.

  • The explorer and the simulation response build native display meshes from the placed buildings of the scene instead of placing them again.

  • A cache hit in compile_cached_scene no longer loads the scene and rewrites scene_info.json when no derived asset was requested.

  • BuildingGeometry validates triangle indices with NumPy.

Verification#

  • tests/test_scene_builder_vectorized.py compares each vectorized function with the sequential algorithm it replaced on random inputs: grid mesh, cell selection, PLY bytes, terrain footing, footprint windows, ray-shell intervals including the near-duplicate rule, and triangle orientation.

  • tests/test_jsonio.py, tests/test_compiler_acquisition.py and tests/test_overture_reader.py cover the row writer, concurrent acquisition with failure ordering, index caching, retries and column projection.

  • The full suite passes (pytest: 253 passed, 8 skipped), including the 3DBAG pagination, roundtrip, native-geometry, roof-coverage and measurement-surface tests.

  • Old and new outputs of compile_scene with a replaying provider were compared on the three cached scenes as described above.

Reproducing the measurements#

conda activate owrt
# offline compile, derived assets and read-back, from the records of cached scenes
python scripts/benchmarks/benchmark_scene_compile.py .cache/scenes/scene_* --repeats 3
# provider queries only (network)
python scripts/benchmarks/benchmark_acquisition.py --site bergen --repeats 3
# cold load_scene into an empty cache (network + compile)
python scripts/benchmarks/benchmark_end_to_end.py --site amsterdam --repeats 2

To measure the previous code, check out b5d483d into a second working tree and run the same scripts with PYTHONPATH=<worktree>/src.

Offline results by scene (median seconds):

Scene

compile

load assets

1 m surface

voxels 3 m

read-back

Los Angeles, before

5.84

0.27

61.15

0.67

0.36

Los Angeles, after

1.75

0.12

1.62

0.46

0.22

Bergen, before

23.42

1.22

181.89

4.33

1.40

Bergen, after

4.32

0.42

4.71

1.26

0.90

Amsterdam 3DBAG, before

36.33

3.09

76.49

19.23

12.68

Amsterdam 3DBAG, after

11.94

1.49

1.72

5.03

4.71

“read-back” is load_compiled_scene plus scene_response, the two calls behind the explorer’s scene endpoint.

Network measurements#

Runs of 18 September 2026, one after another with a minute between them, from the same node. “Provider queries” is _acquire_inputs with a fresh HTTP client and provider each time; “cold compile” is compile_cached_scene into an empty cache.

Measurement

Before (runs)

After (runs)

Provider queries, Bergen

28.6, 28.9

10.5, 11.4, 9.1

Provider queries, Los Angeles

24.6, 22.8

6.0, 4.2, 4.5

Provider queries, Amsterdam (3DBAG)

222.0

39.4, 38.0

Cold compile, Bergen

50.3

16.8, 15.3

Cold compile, Amsterdam (3DBAG)

253.3

55.0

The Bergen and Los Angeles scenes take the global route: Overture buildings, Overture surface classes, and Mapzen or USGS terrain. Amsterdam takes the 3DBAG route: 2,282 buildings arrive in 46 pages of at most 100 items instead of 229 pages of 10, four pages in flight at a time, and decoding the LoD2.2 faces is now the larger part of that time.

The STAC index server (stac.overturemaps.org) rate-limits bursts of requests. The previous code fetched the index six times per scene, and a run of concurrent benchmarks on this node was throttled for about an hour, during which every Overture read fell back to scanning the whole theme directory on S3 and took several times longer. The new reader fetches the index once per release and keeps it on disk, retries briefly on a 429, and otherwise falls back the same way.

What still costs time#

  • Amsterdam’s compile spends about three seconds in localize_buildings (the union of ground faces per shell and the footing test) and three seconds in the height map (column_intervals), and about two seconds encoding the 43 MB provenance.json and the 53 MB source_buildings.json of native geometry. Reading those files back is what remains of the explorer’s cache hit.

  • Every building still becomes two PLY files, as the scene layout requires; writing 7,000 small files costs about one second on this file system.

  • The provider queries are bound by the remote services: one S3 dataset open costs about 1.5 s, and the 3DBAG API serves one page in roughly 0.7 s. Decoding the Amsterdam LoD2.2 shells still takes about 30 s of CPU, mostly Shapely polygon construction and triangulation of some 200,000 faces.

  • sionna.rt.load_scene on the compiled Bergen scene takes 8.2 s with merge_shapes=True (2.7 s without); that is Mitsuba’s parsing and is unchanged.

The acquisition figures reported in the paper’s scene inventory were measured with the previous code and remain valid for that version; the table above is the reference for this branch.