geoml.data

The containers, the variables they hold, and the geometry they are cut against. Everything is addressed by tree path — container.values("assay/Zn/prediction") — rather than by attribute chain.

Point containers

The point-based containers: _SpatialData (what every container is), _PointBased, PointData, GaussianData, DirectionalData and Section3D, with the batching contract the models read.

class geoml.data.containers.PointData(data, coordinates)[source]

Bases: _PointBased

Data represented as points in arbitrary locations.

The container samples come in, from a table with a column per coordinate, and the one most predictions go out to when the locations are not on a lattice. Variables and metadata are added to it by name.

__init__(data, coordinates)[source]
Parameters:
  • data (_pd.DataFrame)

  • coordinates (str or list)

as_data_frame(metadata=True, include='**', simulations=False, columns='flat')[source]

Conversion of a spatial object to a data frame.

Metadata first (bare names, the way HOLEID is read back), then the coordinates, then every filled column of every variable, named by its path – assay_Zn_prediction. include chooses what comes (“**/prediction”, “assay/**”), simulations how many realizations, and columns=”multi” keeps the path as one MultiIndex level per segment instead of flattening – for staying in pandas; written to CSV it makes several header rows, which other software reads as data.

static default_coordinate_labels(n_dim)[source]
Return type:

list[str]

classmethod from_array(coordinates, coordinate_labels=None)[source]
subset_region(min_val, max_val, include_min=None, include_max=None)[source]
as_pyvista(simulations=False, include='**')[source]

Converts this object to a pyvista one, carrying its variables.

Parameters:

simulations – Which simulations to include: False for none (the default, since each one is a full-length array in the exported object), True for all of them, an int for the first n, or a sequence of indices.

classmethod from_geoh5(workspace, name=None)[source]

Reads a Points object from a geoh5 workspace.

The columns come in as measured variables: float data as continuous variables and referenced (coded) data as categorical ones, named as the file names them, with geoh5’s “Unknown” code reading as not measured. What the file has no room to say — which column was a prediction, whose tree it belonged to — is not guessed at: a geoML container round-trips whole through to_zarr, and this reader is for data somebody else made.

name says which Points object to read, and may be left out when the workspace holds exactly one; geoml.data.geoh5.contents lists what there is to name. Needs the geoh5py package: pip install geoml[geoh5].

Parameters:
  • workspace (_types.PathLike | _GeoH5Workspace) – Path of the workspace to read, or an open geoml.data.geoh5.Workspace.

  • name (str | None) – The Points object to read, when the file holds more than one.

Returns:

data (PointData) – The vertices as locations, the columns as variables.

Raises:

ValueError – If the workspace holds no such Points object — the message lists what it does hold.

Return type:

PointData

to_geoh5(workspace, name='Points', include='**', simulations=False, replace=True, folder=None)[source]

Writes this object into a geoh5 workspace, as a Points object.

A workspace is what Geoscience ANALYST — a free viewer for the format — opens as one project, so a model’s pieces belong in one file: writing into an existing path adds the object beside what is already there, and several exports in a row go fastest through an open geoml.data.geoh5.Workspace, which holds the file open across them.

The columns are the ones every export carries, named as the pyvista export names them, with categorical columns as geoh5’s own referenced data; the mapping from each name back to the path that produced it rides in the object’s metadata under geoml_paths, since no rendered name parses back. Needs the geoh5py package: pip install geoml[geoh5].

Parameters:
  • workspace (_types.PathLike | _GeoH5Workspace) – Path of the workspace to write into, or an open geoml.data.geoh5.Workspace.

  • name (str) – The name the object gets in the workspace.

  • include (str) – A path pattern selecting the columns to carry, as in as_pyvista.

  • simulations (bool | int | Sequence[int]) – Which simulations to include: False for none (the default, since each one is a full-length array in the file), True for all of them, an int for the first n, or a sequence of indices.

  • replace (bool) – Whether an existing Points object of this name, in this folder, makes way — what a re-run export script means. False keeps both, and reading that name back then requires saying which.

  • folder (str | None) – Where the object sits in ANALYST’s project tree, as a path — “Data/Assays” — each segment a group, created when it does not exist and reused when it does. None is the root.

decluster(on=None, cell=None, name='declustering')[source]

Computes cell-declustering weights and keeps them, for everything downstream to share.

Samples are rarely laid down evenly — drilling follows the ore — and every statistic that gives them equal votes describes the sampling rather than the field. This stores one weight per location as the metadata column “declustering”, which is where the declustered consumers look first: the variogram figure, and the warping initializers a model runs at construction. One call here turns declustering on everywhere, with one consistent set of weights — computed independently, each consumer would sweep its own cell size on its own values and quietly disagree.

The weights are a snapshot of these locations, exactly as a fold column is: a subset carries its parent’s values, which are then not the subset’s own weights — recompute after subsetting when it matters.

Parameters:
  • on (str | None) – The continuous variable (or component) whose values drive the automatic cell-size choice — the sweep keeps the cell whose declustered mean departs furthest from the naive one, which needs values to mean anything. Left out, the only continuous variable is used; holding several, this refuses to guess. Not needed when cell is given.

  • cell (float | None) – The cell side, fixing it instead of sweeping.

  • name (str) – The metadata column written. An existing column of this name is replaced.

Returns:

  • weights (array) – One per location, summing to the number of locations, as math.geometry.declustering_weights returns them.

  • cell (float) – The cell side used, whether given or chosen.

Return type:

tuple[ndarray, float]

See also

math.geometry.declustering_weights

the arithmetic, and the cell-sweep rule.

spatial_k_fold(test_data, k=5, groups=None, seed=None, name='fold')[source]

Builds cross-validation folds that mimic a prediction task.

A random fold is answered by its neighbours and flatters every score; folds pushed as far from the training data as possible overshoot the other way, testing an extrapolation nobody asked for. What decides how hard a location is to predict is how far its nearest training point sits, so the folds chosen here are the ones whose held-out-to-training distances are distributed like the distances from test_data – the object the model is actually meant to predict – to this data. This is the nearest-neighbour distance matching idea of Linnenbrink et al. (2024), built on discrete groups: continuous per-sample weightings can match the distributions perfectly while the folds are spatially wrong.

The data is first gathered into small groups that are never split across folds – the samples of one drill hole stand or fall together – and the candidate partitions come from cutting a Ward dendrogram of the group centroids at every count from k clusters up to one cluster per group, each cut’s clusters dealt to the emptiest fold largest-first. The cut whose Wasserstein distance to the target distribution is smallest wins, and the result is written to a metadata column – "fold" unless name says otherwise, which is also what models.cross_validate reads by default.

Parameters:
  • test_data – The spatial object the model is meant to predict – a grid, a block model, or any container with coordinates.

  • k (int) – Number of folds.

  • groups (str, optional) – Name of a metadata column whose labels must never be split across folds (a drill hole id). Without one, the data is pre-clustered into many small spatial groups.

  • seed (int) – Passed to sklearn.cluster.KMeans for a reproducible pre-clustering when groups is not given. The rest of the search is deterministic.

  • name (str) – The metadata column to write the folds to. An existing column with this name is replaced, so two calls with two names give two labellings to compare.

Returns:

  • w (float) – The Wasserstein distance between the two distributions below, in coordinate units. Zero is a perfect match.

  • target_distances (array) – Distance from each of test_data’s locations to its nearest data point – the prediction task.

  • fold_distances (array) – Distance from each data point to its nearest training point when its fold is held out – the task the cross-validation poses.

class geoml.data.containers.GaussianData(data, coordinates_mean, coordinates_variance)[source]

Bases: PointData

Points whose locations are uncertain, with a variance per coordinate.

Each location is a Gaussian about its given coordinates – a sample from an uncertain survey, a composite smeared along its length. A model whose input is a GaussianInput carries the variance into the kernel; any other reads the coordinates alone.

classmethod from_array(coordinates, coordinates_variance, coordinate_labels=None)[source]
as_data_frame(metadata=True, **kwargs)[source]

Conversion of a spatial object to a data frame.

Metadata first (bare names, the way HOLEID is read back), then the coordinates, then every filled column of every variable, named by its path – assay_Zn_prediction. include chooses what comes (“**/prediction”, “assay/**”), simulations how many realizations, and columns=”multi” keeps the path as one MultiIndex level per segment instead of flattening – for staying in pandas; written to CSV it makes several header rows, which other software reads as data.

get_batched_variance(index=None)[source]

Variance of the input locations, mirroring get_batched_coordinates.

Zero unless the object was built with an explicit variance — see GaussianData. Only the requested batch is built: deriving the zeros from the coordinates would cost O(n_data) on every batch, which dominates prediction on large objects.

class geoml.data.containers.DirectionalData(data, coordinates, directions)[source]

Bases: PointData

Points carrying a direction each: where a field’s gradient is known.

Structural measurements – a bedding’s dip vector, a contact’s normal – as a unit vector per location, in columns beside the coordinates. What a GradientConstrainedInput takes to make a field follow them. It is training data, not something a model predicts into.

Parameters:
  • data – A table with the coordinates and the direction’s components.

  • coordinates (Any) – The names of the coordinate columns.

  • directions – The names of the direction’s columns, one per coordinate.

__init__(data, coordinates, directions)[source]
Parameters:
  • data (_pd.DataFrame)

  • coordinates (str or list)

as_data_frame(**kwargs)[source]

Conversion of a spatial object to a data frame.

classmethod from_azimuth(data, coordinates, azimuth)[source]
classmethod from_planes(data, coordinates, azimuth, dip)[source]
classmethod from_azimuth_and_dip(data, coordinates, azimuth, dip)[source]
classmethod from_normals(data, coordinates, azimuth, dip)[source]
geoml.data.containers.batch_index(n_data, batch_size)[source]
geoml.data.containers.export_planes(coordinates, dip, azimuth, filename, size=1)[source]
class geoml.data.containers.Section3D(center, azimuth, dip, width, height, n_x, n_y, coordinate_labels=('X', 'Y', 'Z'))[source]

Bases: PointData

A planar section through a 3-D model, as a lattice of points.

A rectangle of width by height about center, turned to azimuth and dip, with n_x by n_y nodes on it. What a model predicts into to draw a cross-section that does not follow the coordinate axes.

Parameters:
  • center – The section’s centre.

  • azimuth – Its orientation, in degrees.

  • dip – Its orientation, in degrees.

  • width – Its size along and across strike.

  • height – Its size along and across strike.

  • n_x – The number of nodes along each side.

  • n_y – The number of nodes along each side.

  • coordinate_labels – The coordinates’ names.

as_pyvista(simulations=False, include='**')[source]

Converts this object to a pyvista one, carrying its variables.

Parameters:

simulations – Which simulations to include: False for none (the default, since each one is a full-length array in the exported object), True for all of them, an int for the first n, or a sequence of indices.

Grids

The regular grids: _GriddedData, Grid1D/2D/3D, GridND and RotatedGrid3D, plus aggregate (one implementation for every kind of variable) and the from_data box-fitting shared with the block classes.

class geoml.data.grids.Grid1D(start, n, step=None, end=None, labels=None)[source]

Bases: _GriddedData

Equally spaced points in 1D.

Nodes along one axis at a fixed spacing, from start, n of them.

The coordinates are never stored; they are generated a batch at a time, so a grid of millions of nodes costs little to hold. What a model predicts into for maps and sections, and what from_data builds around another object’s extent.

step_size

Distance between grid nodes.

Type:

list

grid

The grid coordinates.

Type:

list

grid_size

The number of points in grid.

Type:

list

__init__(start, n, step=None, end=None, labels=None)[source]

Initializer for Grid1D.

Parameters:
  • start (float) – Starting point for grid.

  • n (int) – Number of grid nodes.

  • step (float | None) – Spacing between grid nodes. One number: this grid has one axis.

  • end (float | None) – Last grid point.

  • labels (str | None) – The label for the coordinate.

Either step or end must be given. If both are given, end is ignored.

classmethod from_bounding_box(box, step, margin=0.1, rounding_decimals=0)[source]
class geoml.data.grids.Grid2D(start, n, step=None, end=None, labels=None)[source]

Bases: _GriddedData

Equally spaced points in 2D.

Nodes on a regular lattice in the plane, n along each axis at the spacing step, from start.

The coordinates are never stored; they are generated a batch at a time, so a grid of millions of nodes costs little to hold. What a model predicts into for maps and sections, and what from_data builds around another object’s extent.

step_size

Distance between grid nodes.

Type:

list

grid

The grid coordinates.

Type:

list

grid_size

The number of points in grid.

Type:

list

__init__(start, n, step=None, end=None, labels=None)[source]

Initializer for Grid2D.

Parameters:
  • start (length 2 array, list, or tuple) – Starting point for grid.

  • n (length 2 array, list, or tuple of ints) – Number of grid nodes.

  • step (length 2 array, list, or tuple) – Spacing between grid nodes.

  • end (length 2 array, list, or tuple) – Last grid point.

  • labels (list) – The labels for the coordinates.

Either step or end must be given. If both are given, end is ignored.

classmethod from_bounding_box(box, step, margin=0.1, rounding_decimals=0)[source]
class geoml.data.grids.Grid3D(start, n, step=None, end=None, labels=None)[source]

Bases: _GriddedData

Equally spaced points in 3D.

Nodes on a regular lattice in space, n along each axis at the spacing step, from start.

The coordinates are never stored; they are generated a batch at a time, so a grid of millions of nodes costs little to hold. What a model predicts into for maps and sections, and what from_data builds around another object’s extent.

step_size

Distance between grid nodes.

Type:

list

grid

The grid coordinates.

Type:

list

grid_size

The number of points in grid.

Type:

list

__init__(start, n, step=None, end=None, labels=None)[source]

Initializer for Grid3D.

Parameters:
  • start (length 2 array, list, or tuple) – Starting point for grid.

  • n (length 2 array, list, or tuple of ints) – Number of grid nodes.

  • step (length 2 array, list, or tuple) – Spacing between grid nodes.

  • end (length 2 array, list, or tuple) – Last grid point.

  • labels (list) – The labels for the coordinates.

Either step or end must be given. If both are given, end is ignored.

as_pyvista(simulations=False, include='**')[source]

Converts this object to a pyvista one, carrying its variables.

Parameters:

simulations – Which simulations to include: False for none (the default, since each one is a full-length array in the exported object), True for all of them, an int for the first n, or a sequence of indices.

rotation_matrix()[source]
assign_from_surface(surface, name, labels=('above', 'below'), uncovered=None)[source]

As _SpatialData.assign_from_surface, reading the sheet once for each column of cells rather than once for each cell.

A grid repeats the same (x, y) at every level — _generate varies the first axis fastest, so the pair cycles with period n_x * n_y — and a sheet depends on nothing else, which makes interpolating it n_z times over the same arithmetic n_z times. The columns are generated by the same method that generates the rows, so the two cannot fall out of step. A RotatedGrid3D does not lie on an axis-aligned lattice and takes the general path.

classmethod from_bounding_box(box, step, margin=0.1, rounding_decimals=0)[source]
class geoml.data.grids.GridND(start, n, step=None, end=None, labels=None)[source]

Bases: _GriddedData

Implicit grid in N dimensions.

geoml.data.grids.rotate(data, origin, azimuth=0.0, dip=0.0, rake=0.0, reverse=False)[source]
class geoml.data.grids.RotatedGrid3D(start, n, step, azimuth=0.0, dip=0.0, rake=0.0, labels=None)[source]

Bases: Grid3D

A regular lattice in space, turned about its first node.

Grid3D with an azimuth, a dip and a rake. The nodes are counted in the lattice’s own frame and turned into the world’s as they are generated, so a lattice can follow a deposit’s strike and dip without holding its coordinates.

Parameters:
  • start – The first node, which the lattice turns about.

  • n – The number of nodes along each of the lattice’s own axes.

  • step – The spacing along each of them.

  • azimuth – The rotation, in degrees.

  • dip – The rotation, in degrees.

  • rake – The rotation, in degrees.

  • labels – The coordinates’ names.

rotate(other)[source]
rotation_matrix()[source]
index_data(data)[source]
as_pyvista(simulations=False, include='**')[source]

Converts this object to a pyvista one, carrying its variables.

Parameters:

simulations – Which simulations to include: False for none (the default, since each one is a full-length array in the exported object), True for all of them, an int for the first n, or a sequence of indices.

classmethod from_bounding_box(box, step, margin=0.1, rounding_decimals=0)[source]
classmethod from_data(data, step, margin=0.1, decimals=0)[source]

A rotated grid fitted to another object’s spread.

The rotation is fitted to the data’s own points (a drillhole’s desurveyed cloud serves where there are no point coordinates), and the angles are rounded to decimals before anything is built from them – a grid at 47.3182 degrees is nobody’s intention – so the box is measured in the rounded frame and the data stays covered. The world origin is rounded to the same decimals.

Parameters:
  • data (_SpatialData) – Any spatial object with 3-dimensional coordinates, drillholes included.

  • step (float | ArrayLike) – The step size, one number or one per direction.

  • margin (float or array) – A fraction of the unrotated box’s extent; see Grid3D.from_data.

  • decimals (int) – Decimals for the origin and for the azimuth, dip and rake, in degrees.

classmethod from_points(points, step, margin=0.1, rounding_decimals=0, labels=None)[source]

Blocks

Blocks3D is the regular block model and RotatedBlocks3D the same one turned; BlockSet3D is the variable-size one, where every block’s origin and size are whole numbers of a base cell so that splitting keeps it tiling exactly. BlockSet3D.as_blocks3d hands a refined model back as a regular one at its coarsest level, a gathered block averaged from its parts by volume. Design record: Variable block sizes in geoML — analysis, measurements, plan.

Block models: the _blockdata fan-out shared by Blocks1D/2D/3D, the variable-size BlockSet3D on its integer lattice, RotatedBlockSet3D, and the sub-block geometry behind the mesh assignments and crossed_by.

class geoml.data.blocks.Blocks1D(start, n, step=None, end=None, labels=None, discretization=None, **kwargs)[source]

Bases: Grid1D

Blocks along one axis, each averaged over its sub-blocks.

A block model: Grid1D’s lattice, each node the centre of a block of the lattice’s spacing.

A prediction is the mean over each block, taken from its discretization sub-blocks, a regular lattice inside it; the variance within the block is recorded as dispersion. What a resource is estimated on when every block is the same size; BlockSet3D is the one whose blocks are not.

assign_from_solid(solid, name, labels=('outside', 'inside'), fraction=None)

As _SpatialData.assign_from_solid, measuring the partial blocks on request.

fraction behaves as it does in assign_from_surface: the flag in name follows the block centre, while the column named here holds the share of each block’s sub-blocks falling inside the body — the share of its volume, for the regular sub-blocks a discretization defines.

Parameters:
  • solid (Surface3D) – The closed body to test against.

  • name (str) – Name of the metadata column holding the whole-block flag.

  • labels (tuple) – What to call the two sides, in the order (outside, inside).

  • fraction (str, optional) – Name of a second metadata column, to hold the share of each block inside the body. Costs prod(discretization) queries per block.

assign_from_surface(surface, name, labels=('above', 'below'), fraction=None, uncovered=None)

As Grid3D.assign_from_surface, measuring the partial blocks on request.

The flag in name follows the block centre, as a whole-block code does everywhere else. Name a fraction column as well and the share of each block lying below the sheet is measured over the sub-blocks discretization already defines — what a tonnage near surface needs, where counting a half-buried block whole is the error.

Where the sheet reaches part of a block but not all of it, the sub-blocks past its edge count as not below. Where it does not reach the block at all — the centre included, which is what leaves the flag empty — the fraction is uncovered instead of a measurement.

Parameters:
  • surface (Surface3D) – The sheet to compare against.

  • name (str) – Name of the metadata column holding the whole-block flag.

  • labels (tuple) – What to call the two sides, in the order (above, below).

  • fraction (str, optional) – Name of a second metadata column, to hold the share of each block below the sheet. Costs prod(discretization) queries per block.

  • uncovered (float or "raise") – What the fraction column records for a block the sheet does not reach: numpy.nan for None, the default, so it cannot pass for a block genuinely above ground, or 0.0 to count it as nothing. Pass “raise” to refuse a surface that does not cover every block.

discretized_coordinates(index)
get_batched_coordinates(index)
inducing_grid(index)
property rows_per_location

Rows the model evaluates for each location of this object.

One, except where a location fans out into several — a block with discretization. Prediction divides the batch size by this, so that prediction_batch_size counts the rows actually handed to the model rather than meaning something different for every container.

class geoml.data.blocks.Blocks2D(start, n, step=None, end=None, labels=None, discretization=None, **kwargs)[source]

Bases: Grid2D

Blocks in the plane, each averaged over its sub-blocks.

A block model: Grid2D’s lattice, each node the centre of a block of the lattice’s spacing.

A prediction is the mean over each block, taken from its discretization sub-blocks, a regular lattice inside it; the variance within the block is recorded as dispersion. What a resource is estimated on when every block is the same size; BlockSet3D is the one whose blocks are not.

assign_from_solid(solid, name, labels=('outside', 'inside'), fraction=None)

As _SpatialData.assign_from_solid, measuring the partial blocks on request.

fraction behaves as it does in assign_from_surface: the flag in name follows the block centre, while the column named here holds the share of each block’s sub-blocks falling inside the body — the share of its volume, for the regular sub-blocks a discretization defines.

Parameters:
  • solid (Surface3D) – The closed body to test against.

  • name (str) – Name of the metadata column holding the whole-block flag.

  • labels (tuple) – What to call the two sides, in the order (outside, inside).

  • fraction (str, optional) – Name of a second metadata column, to hold the share of each block inside the body. Costs prod(discretization) queries per block.

assign_from_surface(surface, name, labels=('above', 'below'), fraction=None, uncovered=None)

As Grid3D.assign_from_surface, measuring the partial blocks on request.

The flag in name follows the block centre, as a whole-block code does everywhere else. Name a fraction column as well and the share of each block lying below the sheet is measured over the sub-blocks discretization already defines — what a tonnage near surface needs, where counting a half-buried block whole is the error.

Where the sheet reaches part of a block but not all of it, the sub-blocks past its edge count as not below. Where it does not reach the block at all — the centre included, which is what leaves the flag empty — the fraction is uncovered instead of a measurement.

Parameters:
  • surface (Surface3D) – The sheet to compare against.

  • name (str) – Name of the metadata column holding the whole-block flag.

  • labels (tuple) – What to call the two sides, in the order (above, below).

  • fraction (str, optional) – Name of a second metadata column, to hold the share of each block below the sheet. Costs prod(discretization) queries per block.

  • uncovered (float or "raise") – What the fraction column records for a block the sheet does not reach: numpy.nan for None, the default, so it cannot pass for a block genuinely above ground, or 0.0 to count it as nothing. Pass “raise” to refuse a surface that does not cover every block.

discretized_coordinates(index)
get_batched_coordinates(index)
inducing_grid(index)
property rows_per_location

Rows the model evaluates for each location of this object.

One, except where a location fans out into several — a block with discretization. Prediction divides the batch size by this, so that prediction_batch_size counts the rows actually handed to the model rather than meaning something different for every container.

class geoml.data.blocks.Blocks3D(start, n, step=None, end=None, labels=None, discretization=None, **kwargs)[source]

Bases: Grid3D

Blocks in space, each averaged over its sub-blocks.

A block model: Grid3D’s lattice, each node the centre of a block of the lattice’s spacing.

A prediction is the mean over each block, taken from its discretization sub-blocks, a regular lattice inside it; the variance within the block is recorded as dispersion. What a resource is estimated on when every block is the same size; BlockSet3D is the one whose blocks are not.

as_pyvista(simulations=False, include='**')[source]

Converts this object to a pyvista one, carrying its variables.

Parameters:

simulations – Which simulations to include: False for none (the default, since each one is a full-length array in the exported object), True for all of them, an int for the first n, or a sequence of indices.

to_geoh5(workspace, name='Blocks', include='**', simulations=False, replace=True, folder=None)[source]

Writes this block model into a geoh5 workspace, as a BlockModel.

The geoh5 BlockModel is the regular-grid sibling of the Octree a BlockSet3D writes; a uniform model maps onto it exactly. Everything else is as PointData.to_geoh5: the workspace accumulates, the columns ride with their path table, replace makes a re-run export mean what it says. Needs the geoh5py package: pip install geoml[geoh5].

Parameters:
  • workspace (_types.PathLike | _GeoH5Workspace) – Path of the workspace to write into, or an open geoml.data.geoh5.Workspace.

  • name (str) – The name the object gets in the workspace.

  • include (str) – A path pattern selecting the columns to carry.

  • simulations (bool | int | Sequence[int]) – Which simulations to include: False for none (the default), True for all, an int for the first n, or a sequence of indices.

  • replace (bool) – Whether an existing BlockModel of this name, in this folder, makes way.

  • folder (str | None) – Where the object sits in ANALYST’s project tree, as a path — each segment a group, created when it does not exist and reused when it does. None is the root.

classmethod from_geoh5(workspace, name=None)[source]

Reads a BlockModel from a geoh5 workspace.

Only a uniform spacing has a geoML container: a true tartan grid — uneven cells along an axis — is refused with the axis named. An unrotated model comes back as a Blocks3D; a rotated one as a RotatedBlockSet3D with max_levels=0 — the blocks, their rotation and block-support prediction all work, nothing can be refined, and as_blocks3d makes it a RotatedBlocks3D. Float cell data becomes continuous variables and referenced data categorical ones, as in PointData.from_geoh5. Needs the geoh5py package: pip install geoml[geoh5].

Parameters:
  • workspace (_types.PathLike | _GeoH5Workspace) – Path of the workspace to read, or an open geoml.data.geoh5.Workspace.

  • name (str | None) – The BlockModel to read, when the file holds more than one.

Returns:

blocks (Blocks3D or RotatedBlockSet3D) – Whichever the file’s rotation calls for.

Raises:

ValueError – If the workspace holds no such BlockModel, or its spacing is uneven.

Return type:

Blocks3D | RotatedBlockSet3D

assign_from_solid(solid, name, labels=('outside', 'inside'), fraction=None)

As _SpatialData.assign_from_solid, measuring the partial blocks on request.

fraction behaves as it does in assign_from_surface: the flag in name follows the block centre, while the column named here holds the share of each block’s sub-blocks falling inside the body — the share of its volume, for the regular sub-blocks a discretization defines.

Parameters:
  • solid (Surface3D) – The closed body to test against.

  • name (str) – Name of the metadata column holding the whole-block flag.

  • labels (tuple) – What to call the two sides, in the order (outside, inside).

  • fraction (str, optional) – Name of a second metadata column, to hold the share of each block inside the body. Costs prod(discretization) queries per block.

assign_from_surface(surface, name, labels=('above', 'below'), fraction=None, uncovered=None)

As Grid3D.assign_from_surface, measuring the partial blocks on request.

The flag in name follows the block centre, as a whole-block code does everywhere else. Name a fraction column as well and the share of each block lying below the sheet is measured over the sub-blocks discretization already defines — what a tonnage near surface needs, where counting a half-buried block whole is the error.

Where the sheet reaches part of a block but not all of it, the sub-blocks past its edge count as not below. Where it does not reach the block at all — the centre included, which is what leaves the flag empty — the fraction is uncovered instead of a measurement.

Parameters:
  • surface (Surface3D) – The sheet to compare against.

  • name (str) – Name of the metadata column holding the whole-block flag.

  • labels (tuple) – What to call the two sides, in the order (above, below).

  • fraction (str, optional) – Name of a second metadata column, to hold the share of each block below the sheet. Costs prod(discretization) queries per block.

  • uncovered (float or "raise") – What the fraction column records for a block the sheet does not reach: numpy.nan for None, the default, so it cannot pass for a block genuinely above ground, or 0.0 to count it as nothing. Pass “raise” to refuse a surface that does not cover every block.

discretized_coordinates(index)
get_batched_coordinates(index)
inducing_grid(index)
property rows_per_location

Rows the model evaluates for each location of this object.

One, except where a location fans out into several — a block with discretization. Prediction divides the batch size by this, so that prediction_batch_size counts the rows actually handed to the model rather than meaning something different for every container.

class geoml.data.blocks.RotatedBlocks3D(start, n, step=None, end=None, labels=None, discretization=None, **kwargs)[source]

Bases: RotatedGrid3D

A regular block model turned about its first block.

Blocks3D with an azimuth, a dip and a rake, as RotatedGrid3D is Grid3D with them. The blocks are counted in the model’s own frame and turned about the first block’s centre where coordinates leave – the centres, the sub-blocks a prediction averages over, the exported cells – and turned back where they come in (index_data, and so aggregate). BlockSet3D.as_blocks3d returns one for a RotatedBlockSet3D, with the same angles and the same pivot.

Parameters:
  • start (array-like) – Centre of the first block, which the model turns about.

  • n (array-like) – Number of blocks along each of the model’s own axes.

  • step (array-like) – Size of a block along each of them.

  • labels (list, optional) – Coordinate names.

  • discretization (array-like, optional) – Sub-blocks per axis, averaged over to predict a block. One by default, which predicts each block at its centre.

  • azimuth (float) – The rotation in degrees, as RotatedGrid3D takes it. Given by keyword.

  • dip (float) – The rotation in degrees, as RotatedGrid3D takes it. Given by keyword.

  • rake (float) – The rotation in degrees, as RotatedGrid3D takes it. Given by keyword.

as_pyvista(simulations=False, include='**')[source]

Converts this object to a pyvista one, carrying its variables.

The cells of a Blocks3D of the same shape, turned into place as one piece, so that each cell’s centre is its block’s.

Parameters:

simulations – Which simulations to include: False for none (the default, since each one is a full-length array in the exported object), True for all of them, an int for the first n, or a sequence of indices.

assign_from_solid(solid, name, labels=('outside', 'inside'), fraction=None)

As _SpatialData.assign_from_solid, measuring the partial blocks on request.

fraction behaves as it does in assign_from_surface: the flag in name follows the block centre, while the column named here holds the share of each block’s sub-blocks falling inside the body — the share of its volume, for the regular sub-blocks a discretization defines.

Parameters:
  • solid (Surface3D) – The closed body to test against.

  • name (str) – Name of the metadata column holding the whole-block flag.

  • labels (tuple) – What to call the two sides, in the order (outside, inside).

  • fraction (str, optional) – Name of a second metadata column, to hold the share of each block inside the body. Costs prod(discretization) queries per block.

assign_from_surface(surface, name, labels=('above', 'below'), fraction=None, uncovered=None)

As Grid3D.assign_from_surface, measuring the partial blocks on request.

The flag in name follows the block centre, as a whole-block code does everywhere else. Name a fraction column as well and the share of each block lying below the sheet is measured over the sub-blocks discretization already defines — what a tonnage near surface needs, where counting a half-buried block whole is the error.

Where the sheet reaches part of a block but not all of it, the sub-blocks past its edge count as not below. Where it does not reach the block at all — the centre included, which is what leaves the flag empty — the fraction is uncovered instead of a measurement.

Parameters:
  • surface (Surface3D) – The sheet to compare against.

  • name (str) – Name of the metadata column holding the whole-block flag.

  • labels (tuple) – What to call the two sides, in the order (above, below).

  • fraction (str, optional) – Name of a second metadata column, to hold the share of each block below the sheet. Costs prod(discretization) queries per block.

  • uncovered (float or "raise") – What the fraction column records for a block the sheet does not reach: numpy.nan for None, the default, so it cannot pass for a block genuinely above ground, or 0.0 to count it as nothing. Pass “raise” to refuse a surface that does not cover every block.

discretized_coordinates(index)
get_batched_coordinates(index)
inducing_grid(index)
property rows_per_location

Rows the model evaluates for each location of this object.

One, except where a location fans out into several — a block with discretization. Prediction divides the batch size by this, so that prediction_batch_size counts the rows actually handed to the model rather than meaning something different for every container.

class geoml.data.blocks.BlockSet3D(start, n, step, discretization=(2, 2, 2), max_levels=3, labels=('X', 'Y', 'Z'))[source]

Bases: PointData

Blocks of several sizes, on one integer lattice.

A block model where the interesting ground can be carried finely and the rest coarsely. On a real deposit the ground worth resolving is a small part of the volume, and a uniform model at the resolution that part needs spends almost all of its cells saying nothing: refining 5 m only where it is wanted takes a 29-million-cell model to under 700 000.

Every block’s position and size are whole numbers of a base cell, the finest the model may go, which is step / discretization ** max_levels. Working in those integers rather than in metres is what makes it exact: blocks meet without a tolerance, a block is a whole number of its own children, and regrouping conserves mass to the last digit. It also means the model can say which of its answers rest on coarse blocks, which is the one thing a mixed-support model has to be able to prove – see docs/variable-block-models.md.

It is built full: the blocks tile their box exactly, and every operation keeps them that way. Ground to leave out is filtered, not removed, so that grouping is always safe – a half-populated group would average over blocks that are not there and quietly weigh the answer wrong.

discretization does two jobs, and they are the same job. It is how finely a block is sampled to average it, and it is how a block splits: each sub-block becomes a child. So the refinement ratio is the discretization, per axis and not necessarily two – [2, 2, 1] refines in plan and leaves the bench height alone. Being the same at every level is what lets a block of any size fan out into the same number of rows, so the model sees one shape whatever it is looking at and nothing downstream has to know that levels exist. It costs a coarse block some accuracy in its own average, always by overstating how variable it is, which errs towards splitting it – and splitting is what removes the error.

Parameters:
  • start (array-like) – Centre of the first (coarsest) block.

  • n (array-like) – Number of blocks along each axis, at the coarsest level.

  • step (array-like) – Size of a coarsest-level block.

  • discretization (array-like) – Sub-blocks per axis, used at every level, and so also the ratio a block splits by. An axis given 1 is never refined.

  • max_levels (int) – How many times a block may be split. Fixes the base cell, and so the lattice everything else is counted in.

  • labels (list) – Coordinate names.

block_size

(n_data, 3), each block’s size in the coordinates’ own units. Not called step_size: that name means one size for the whole object, and anything reading it would take the product of this array for a volume.

Type:

array

block_volume

(n_data,), what each block is worth in a tonnage.

Type:

array

level

(n_data,), 0 for a coarsest block up to max_levels for a base one.

Type:

array

property block_size
property block_volume
property level
property rows_per_location

Rows the model evaluates for each location of this object.

One, except where a location fans out into several — a block with discretization. Prediction divides the batch size by this, so that prediction_batch_size counts the rows actually handed to the model rather than meaning something different for every container.

is_full()[source]

Whether the blocks tile their box exactly – no gap, no overlap.

Volume alone cannot answer: a gap and an overlap of the same size cancel. So the base cells are counted, in an array the shape of the lattice, which costs a byte per base cell (29 MB for a 30-million-cell model). Construction is full and both split and group preserve it, so this is for checking something that came from elsewhere rather than for routine use.

Return type:

bool

subset_region(min_val, max_val, include_min=None, include_max=None)[source]
split(mask, carry=True, labels=('X', 'Y', 'Z'))[source]

A new block set with each marked block cut into its own sub-blocks.

A block becomes prod(discretization) children, one per sub-block and in the same order, so the values a coarse prediction already holds for those sub-blocks describe the blocks this makes.

A block that was not split keeps what was predicted for it: it is the same block on the same support, so its value is still the right answer and arriving at it again would be work for nothing. The children start missing, and unpredicted() says which they are, so predict(…, where=…) visits only them. A parent’s value is never handed down – that would manufacture children agreeing exactly, which is the one thing refining is meant to find out rather than assume.

Parameters:
  • mask (array-like) – One boolean per block, or the indices of the blocks to split.

  • carry (bool) – Whether to bring the variables and metadata across. False gives bare geometry, for building a mesh to predict onto from scratch.

Return type:

BlockSet3D

group(mask, carry=True, labels=('X', 'Y', 'Z'))[source]

The inverse of split: whole families of children, back into the parent they came from.

A block is grouped with its siblings, so the mask must name every child of a parent or none of them. A partial family would average over children that are not there and mis-weight the parent – the mass-conservation error the lattice exists to prevent – so it is refused rather than approximated. That check is what makes conversion between supports two-directional: group undoes split exactly, and the blocks tile their box afterwards as they did before.

A block that was not grouped keeps what it holds, exactly as in split and for the same reason: it is the same block on the same support, so its value is still the right answer for it. The parents are the ones on new ground, and they come back missing – unpredicted() names them and predict(…, where=…) fills them.

A parent’s value is never averaged from its children. Coarsening is a change of support and almost nothing survives it: a parent’s spread is not its children’s, its within-block dispersion is larger by exactly what the grouping absorbed, and a category has no mean. The one thing that would come across exactly is a realization, and a variable is more than its realizations. Metadata does come across, from the first child – it describes the ground rather than the model, and where the children disagree about it there is no right answer to be had.

Parameters:
  • mask (array-like) – One boolean per block, or the indices of the blocks to group.

  • carry (bool) – Whether to bring across what the blocks that were not grouped hold. False gives bare geometry.

  • labels (list) – Coordinate names.

Return type:

BlockSet3D

as_blocks3d()[source]

This model at its coarsest level, as a regular block model.

One block per coarsest-level block, in Blocks3D’s own order and with this model’s discretization, so a block that was never split is the same block on the same support and keeps every column exactly. Where the model was refined, each coarse block is gathered from the blocks inside it, weighted by volume:

  • a column that is a mean over a block’s sub-blocks – prediction, the latent moments, noise_variance, the shares below each cut-off, a category’s probability and entropy – takes the mean of its parts’, which is that same mean over the whole block;

  • every realization is averaged index by index, the parts’ realization i making the block’s realization i;

  • what is read off the realizations is read again: the quantiles, the probabilities, the dispersion – the parts’ mean dispersion plus the spread between them – and the predicted category, the winner of the averaged probabilities;

  • anything else is missing, none of it following from the parts: divided, the measurements, a binary variable’s entropy.

A coarse block with any part that holds nothing – a block where= kept from being predicted, say – holds nothing either, so a partial answer never passes for a whole one. Metadata is gathered as aggregate gathers it, by volume: a number averages, a coded column keeps the label holding most of the block (none on a tie), and a boolean one holds where it held throughout.

Returns:

Blocks3D or RotatedBlocks3D – A RotatedBlocks3D for a RotatedBlockSet3D, with the same angles and turning about the same point.

Return type:

Blocks3D | RotatedBlocks3D

See also

group

the coarsening that leaves a parent missing instead, since the change of support is otherwise the model’s to re-predict.

Notes

A derived variable is averaged like any other, which suits a quantity that adds up over a volume. One that does not needs derive again on the result, where its function sees the coarse blocks.

to_geoh5(workspace, name='Blocks', include='**', simulations=False, replace=True, folder=None)[source]

Writes this block model into a geoh5 workspace, as an Octree.

The lattice maps one to one when discretization is (2, 2, 2) — the default — since a geoh5 octree subdivides strictly two by two by two; any other discretization is refused rather than resampled. A workspace is what Geoscience ANALYST — a free viewer for the format — opens as one project, so a model’s pieces belong in one file, as in PointData.to_geoh5; several exports in a row go fastest through an open geoml.data.geoh5.Workspace. Needs the geoh5py package: pip install geoml[geoh5].

Parameters:
  • workspace (_types.PathLike | _GeoH5Workspace) – Path of the workspace to write into, or an open geoml.data.geoh5.Workspace.

  • name (str) – The name the object gets in the workspace.

  • include (str) – A path pattern selecting the columns to carry, as in as_pyvista.

  • simulations (bool | int | Sequence[int]) – Which simulations to include: False for none (the default), True for all, an int for the first n, or a sequence of indices.

  • replace (bool) – Whether an existing Octree of this name, in this folder, makes way — what a re-run export script means. False keeps both, and reading that name back then requires saying which.

  • folder (str | None) – Where the object sits in ANALYST’s project tree, as a path — each segment a group, created when it does not exist and reused when it does. None is the root.

Raises:

ValueError – If the discretization is not (2, 2, 2), or the model carries a dip or rake — geoh5 rotates about the vertical axis alone.

classmethod from_geoh5(workspace, name=None)[source]

Reads an Octree from a geoh5 workspace, as a block model.

The octree cells become blocks on a (2, 2, 2) lattice, in the file’s own row order, with float cell data as continuous variables and referenced data as categorical ones. A rotated octree comes back as a RotatedBlockSet3D; an octree written with its vertical axis running downward is normalized on the way in. name says which Octree to read when the workspace holds more than one; geoml.data.geoh5.contents lists what there is to name. Needs the geoh5py package: pip install geoml[geoh5].

A foreign file is not ours to trust, and the always-full design does not bend for it: cells must sit aligned to their own power-of-two size and free of overlaps, and whatever the cells do not cover inside their box is filled with unvalued blocks, a metadata column named imported marking the difference — a vendor octree usually carries only the domain of interest.

Parameters:
  • workspace (_types.PathLike | _GeoH5Workspace) – Path of the workspace to read, or an open geoml.data.geoh5.Workspace.

  • name (str | None) – The Octree object to read, when the file holds more than one.

Returns:

blocks (BlockSet3D or RotatedBlockSet3D) – Whichever the file’s rotation calls for.

Raises:

ValueError – If the workspace holds no such Octree, or its cells are not an octree this lattice can hold (misaligned, overlapping, or reflected under a rotation).

Return type:

BlockSet3D

block_shares(split_on=None)[source]

How often each decision cuts a block in two.

A continuous variable contributes one column per cut-off it declares, a categorical one per category, and both mean the same thing: the share of realizations in which this block’s sub-blocks fall on both sides of something that matters. A grade is judged against the grades someone declared; an indicator against zero, that being where one category stops winning and its rival starts.

Returns a dict of name -> (n_data,) array, empty where nothing declared a decision to make.

Return type:

dict[str, ndarray]

needs_splitting(split_on=None, tolerance=None)[source]

Which blocks a decision’s surface runs through, and so are worth cutting.

A block is marked where the prediction at its sub-blocks falls on both sides of a cut-off, or of a category’s boundary: the surface a contour of the prediction draws passes through it, and cutting the block lets that surface bend there.

Note what this does not mark: a block the model is merely unsure about. The prediction is the mean over the realizations, and where the data do not reach, each realization crosses the cut-off somewhere of its own while their mean crosses it nowhere – cutting there would buy blocks and no surface.

The test is over every decision at once and any one is enough, which is the cautious way round on purpose: a block worth splitting for one variable is worth splitting whatever the others say. Name split_on to narrow it – letting every element of a polymetallic deposit vote marks most of the model and gives back the saving.

Blocks already at the finest level are never marked, there being nothing to cut them into.

Parameters:
  • split_on (str or list, optional) – Which variables get a say. All of them by default.

  • tolerance (float, optional) – Deprecated, and without effect: a block is divided or it is not. It was the share of realizations that had to find a block divided, and will be removed.

Return type:

ndarray

unbalanced(gap=1)[source]

Blocks with a neighbour more than gap levels finer than they are.

A block whose own sub-blocks agree is never marked by needs_splitting, and rightly so – cutting it would not change the answer it gives. But the field can still turn sharply inside it, and nothing in the block itself says so. That a neighbour was cut twice while this block was not cut at all is the evidence, and it lives outside the block.

It matters for what is drawn rather than for what is decided. A contour reads a block through its eight corners, so a coarse block beside much finer ones is a crude straight guess across a long span laid right where the surface runs. Levelling the jump measured three times closer to the true surface for 35% more blocks, where refining a whole level deeper without it bought almost nothing for 2.6 times as many – deeper refinement widens the jumps as fast as it narrows the blocks. models.refine therefore cuts these as it goes.

Blocks already at the finest level are never marked, there being nothing to cut them into.

Parameters:

gap (int) – How many levels of difference to tolerate. One is the usual 2:1 balance: a block may meet blocks one level finer, not two.

Return type:

ndarray

classmethod from_data(data, step, margin=0.1, decimals=0, discretization=(2, 2, 2), max_levels=3)[source]

A block model covering another object’s bounding box.

As Grid3D.from_data, counting blocks rather than nodes: the margined box’s lower corner is floored to decimals, and enough blocks follow to cover the far side, so the corner is round and the margin never shrinks.

Parameters:
  • data – Any spatial object, drillholes included.

  • step – The coarse block size, one number or one per direction.

  • margin (float or array) – A fraction of the data’s extent; see Grid3D.from_data.

  • decimals (int) – Decimals to floor the box corner to.

  • discretization – As the constructor takes them.

  • max_levels – As the constructor takes them.

index_data(data)[source]

Which block each of data’s locations falls in.

One row index per location, -1 for anything outside the box. Note this is not what a grid’s index_data returns – a cell index per axis – because blocks of several sizes have no per-axis index to return. Which block is the answer here.

The lattice makes it cheap: a location’s base cell is arithmetic, and the block covering that cell is the one whose origin is the cell’s ancestor at that block’s own level, so one searchsorted per level finds it and every location is settled within max_levels + 1 of them.

Return type:

ndarray

aggregate(data, variables=None, metadata=True)[source]

Carries another object’s measurements onto the blocks holding them.

As Grid3D.aggregate – one method, the operation following each variable’s kind – over blocks of several sizes.

assign_from_surface(surface, name, labels=('above', 'below'), fraction=None, uncovered=None)[source]

As Blocks3D.assign_from_surface, over blocks of several sizes.

The flag in name follows the block centre; naming a fraction column measures the share of each block below the sheet over the sub-blocks discretization defines, scaled to each block’s own size. crossed_by asks the same question and answers with the blocks worth cutting.

assign_from_solid(solid, name, labels=('outside', 'inside'), fraction=None)[source]

As Blocks3D.assign_from_solid, over blocks of several sizes.

fraction holds the share of each block’s volume inside the body, and crossed_by turns the same test into the blocks worth cutting.

crossed_by(mesh)[source]

Which blocks a mesh passes through, and so which are worth cutting.

A block is crossed when its sub-blocks fall on both sides of the mesh – the question needs_splitting asks of a cut-off, asked of geometry instead. A topography, a vein wall, a lease boundary: a block the surface runs through holds two answers whatever is predicted into it, and no amount of prediction will separate them.

One entirely above or entirely below is left alone however close it lies, which is what keeps this from refining a whole domain. Blocks already at the finest level are never marked, as elsewhere.

Hand it to split, or give the mesh to models.refine, which unions this with the other two criteria and repeats until nothing is left to cut:

blocks = blocks.split(blocks.crossed_by(topography))

A sheet that covers only part of the model counts a sub-block past its edge as not below, the way fraction does, so a block straddling the sheet’s own boundary reads as crossed. That is usually wanted – the edge is a real feature of the ground being described – but it is why a sheet should reach across the model when it is not.

Parameters:

mesh (Surface3D or Solid3D) – The sheet or closed body to test against. A Mesh3D that is neither has no sides and is refused.

Returns:

array – One boolean per block.

Return type:

ndarray

unpredicted(variable=None)[source]

Which blocks have nothing in them yet.

What split leaves behind, and what a cancelled prediction did not reach: hand it to predict(…, where=…) and only those blocks are visited. Read off the missing values, as on every container, once the set holds a variable; before anything has been predicted into it, the blocks the last split made. Naming a variable reads that one’s missing values alone.

Return type:

ndarray

get_batched_coordinates(index=None)[source]
as_data_frame(metadata=True, **kwargs)[source]

Conversion of a spatial object to a data frame.

Metadata first (bare names, the way HOLEID is read back), then the coordinates, then every filled column of every variable, named by its path – assay_Zn_prediction. include chooses what comes (“**/prediction”, “assay/**”), simulations how many realizations, and columns=”multi” keeps the path as one MultiIndex level per segment instead of flattening – for staying in pandas; written to CSV it makes several header rows, which other software reads as data.

Return type:

DataFrame

as_pyvista(simulations=False, include='**')[source]

Converts this object to a pyvista one, carrying its variables.

One hexahedron per block, written out corner by corner, rather than the ImageData a regular block model exports: implicit geometry can only say one spacing, and the point here is that there is more than one. The cells are welded, so blocks that touch share the corners they meet at – which is what lets anything be contoured across them.

Parameters:

simulations – Which simulations to include: False for none (the default, since each one is a full-length array in the exported object), True for all of them, an int for the first n, or a sequence of indices.

get_contour(path, value, supersample=0, simplify=None, close=False)[source]

Isosurface through blocks of more than one size.

The field is painted onto slabs of a regular grid and contoured by flying edges (_painted_contour) – the model itself never carries every block at the finest size, but a transient slab of the field can, and contouring it is what an unstructured mesh of welded hexahedra used to be built for, at most of the call’s cost. The answer is the one a model carried at the finest size throughout would have given.

Values live on the cells and an isosurface needs them on the corners, so what is painted are the corners – the blocks meeting at each one decide where the surface passes, one vote per block whatever its size.

The blocks the surface runs through are cut to the finest size the lattice allows before any of that, in the mesh handed to VTK and not in the model: a coarse block cannot see the corners its finer neighbours place in the middle of the face they share, so the two sides draw different curves there and the surface tears along every such interface it crosses. Cutting its neighbourhood to one size puts the interfaces out of the way. Nothing is predicted – a child reads its parent’s value and the shape its corners carry – and the surface comes back the one a model carried at the finest size throughout would have given. See _cut_to_contour.

Blocks holding no value – ground a where= left out, or a split not yet predicted – vote at no corner. One whose corners valued blocks all share is read across from them; beyond that the ground has no level to draw, and the surface stops at it, or with close closes against it.

Parameters:
  • path (str) – The column to contour, named the way the tree names it: “grade/prediction” is that column, “grade” alone defaults to the variable’s prediction, and “Elements/Zn” reaches a component (only the components of a composition hold a grade). A single bare segment that is no variable of its own is searched for anywhere in the tree, so “Zn” still finds “Elements/Zn” as long as only one variable holds a Zn.

  • value (float) – The value to draw the surface at.

  • supersample (int) – How many levels past the model’s own finest block to cut the mesh to. Costs prod(discretization) times the cells per level, around the surface only. On a smooth field one level buys a rounder and genuinely closer surface – what VTK reads between block corners is trilinear, creasing at every face, and averaging onto corners again at each finer level composes into a reconstruction worth roughly predicting a model several times the size. On a near-binary field – an indicator contoured close to zero – the surface is corner-locked either way, and the extra lattice buys almost nothing: measured on a real 908k-block model, one level cost 3.5-5.8x the time and moved the volumes by 0.03-0.17%. The default is 0 for that reason; pass 1 when contouring a smooth grade at a cut-off well inside its range.

  • simplify (float, optional) – A geometric error budget, in coordinate units: the surface is simplified until pushing further would move it more than this (see Mesh3D.simplify). Pairs naturally with supersample, which buys accuracy in triangles this then spends back where the surface is flat. None returns the full triangulation.

  • close (bool or str) – Whether to close the surface where it runs out of the model, so that what comes back is a body rather than a sheet with a hole in the side – “above” (or True) keeps the region where the values exceed value, a grade shell, and “below” the region under it, as on a grid. Done with a shell of ghost cells carrying the boundary blocks’ values on past the box, each its partner’s own size, and the box then cutting the body off, so its faces, edges and corners on the box are the box’s own.

Returns:

surf (Solid3D, Surface3D or Mesh3D) – Whichever the geometry calls for, as get_contour on a grid.

class geoml.data.blocks.RotatedBlockSet3D(start, n, step, azimuth=0.0, dip=0.0, rake=0.0, discretization=(2, 2, 2), max_levels=3, labels=('X', 'Y', 'Z'))[source]

Bases: BlockSet3D

A variable-size block model rotated about its starting block.

The lattice is BlockSet3D’s, untouched: splitting, grouping, the refinement criteria and the integer arithmetic all happen in the unrotated frame, which is what keeps them exact. The rotation is applied where coordinates leave – the block centres, the sub-block fan-out a prediction reads, the exported hexahedra – and removed where coordinates come in (index_data, and so aggregate). Every mesh test and assignment reads sub-block positions through get_batched_coordinates, so geometry against surfaces and solids works in world coordinates with nothing overridden.

rotation_matrix()[source]
index_data(data)[source]

Which block each of data’s locations falls in.

One row index per location, -1 for anything outside the box. Note this is not what a grid’s index_data returns – a cell index per axis – because blocks of several sizes have no per-axis index to return. Which block is the answer here.

The lattice makes it cheap: a location’s base cell is arithmetic, and the block covering that cell is the one whose origin is the cell’s ancestor at that block’s own level, so one searchsorted per level finds it and every location is settled within max_levels + 1 of them.

get_batched_coordinates(index=None)[source]
classmethod from_data(data, step, margin=0.1, decimals=0, discretization=(2, 2, 2), max_levels=3)[source]

A rotated block model fitted to another object’s spread.

As RotatedGrid3D.from_data – the rotation fitted to the points and rounded to decimals (degrees) before the box is measured – counting blocks rather than nodes.

Meshes

Triangulated meshes: Mesh3D the primitive, Surface3D and Solid3D as siblings, DTM3D the terrain, mesh3d picking by geometry, the booleans, the DXF round trip, and the adapters every container’s assignments read (_sheet_interpolator, _closed_body, _side_codes). The arithmetic itself is geoml.math.geometry; what lives here is what touches a container or holds an error message.

class geoml.data.meshes.Mesh3D(points, triangles, normals)[source]

Bases: _PointBased

A triangulated surface: vertices, the triangles indexing them, normals.

The primitive Surface3D and Solid3D are built on, and the only one of the three that promises nothing about its shape — which is what a mesh must be allowed to be while it is still being repaired. What it does do is measure itself as it is built, so that everything downstream can ask rather than work it out again: area, and whether it is closed and consistent. Those cost a few milliseconds on a mesh of tens of thousands of triangles.

mesh3d(points, triangles, normals) builds whichever of the three the geometry calls for, and is what the readers use.

area

The surface area, whether or not the mesh closes.

Type:

float

closed

Whether every edge is shared by two triangles, so that the mesh bounds a volume. Vacuously true of an empty mesh.

Type:

bool

consistent

Whether the triangles agree about which way is out. A closed mesh that is not consistent bounds nothing that can be tested.

Type:

bool

provenance

What the mesh was made from, where whatever made it says so: a contour records the column, the level, the side it closes on and its budgets, and a MeshSet adds its limits and the realization. Empty otherwise. It travels with to_zarr and to_geoh5.

Type:

dict

provenance: dict
split()[source]

The mesh’s connected pieces, each as an object of its own.

A boolean operation readily answers with a body in several pieces — an ore shell cut in two by a fault — and each piece is a body in its own right, while together they are still one legitimate mesh. This is how to take them apart; each piece comes back as whichever class its own geometry calls for.

Returns:

pieces (list) – One mesh per connected piece, longest-standing order. A mesh already in one piece returns [self].

heal(hole_size=None)[source]

A repaired copy of this mesh.

Four things are put right, in the order that works: coincident vertices are welded, so that seams stop reading as boundaries; the faces that bound nothing are dropped; holes smaller than hole_size are covered over; and the triangles are made to agree about which way is out, then turned to face outward. That last step is not optional — filling a hole leaves the new triangles wound however they came, which would leave the mesh closed and still untestable.

The dropping has to come first, and this method did not do it until 0.6.7. A zero-thickness flap or a zero-area sliver makes an edge run twice the same way round, so it reads as a winding failure — and it is the one winding failure reorienting cannot mend, the surface being non-manifold there for VTK to walk. Measured on a box carrying one flap: clean, compute_normals(consistent_normals= True, auto_orient_normals=True) and triangulate left all three of its reversed edges exactly as they were, and every error message in this module sends the user here to have them fixed.

What comes back is whichever class the repaired geometry calls for, which may be the same one, and may be an empty Mesh3D if nothing survived. Healing is not guaranteed: a mesh with a hole larger than hole_size, or one self-intersecting, can come back no better.

Parameters:

hole_size (float, optional) – The largest hole to cover, in the mesh’s own units. None to weld and reorient only, leaving every boundary where it is.

Returns:

mesh (Mesh3D, Surface3D or Solid3D)

simplify(max_error)[source]

The same shape on as few triangles as the error budget allows.

Built for what get_contour returns: a contoured surface carries a triangle for every block corner it crosses, most of them slivers saying nothing the budget would miss. The argument is geometric – how far, in the mesh’s own units, the simplified surface may sit from the original – so the same call means the same thing on a coarse shell and a fine one, which a fraction of triangles does not.

The caller’s kind is kept: a body stays a body, a terrain a terrain. If a cut breaks the kind’s own promise – a solid opened, a terrain folded over – the mesh is cut more gently until the promise holds, since a gentler cut only sits closer to the original; if no cut survives, the mesh comes back as it came, with a warning saying so. Simplification never trades the shape away.

Parameters:

max_error (float) – The largest distance the simplified surface may sit from the original, in the mesh’s own units, either way. Enforced by measurement: the simplified faces are probed against the original surface and the original’s vertices against the simplified one, and the decimation tightened until the promise holds.

Returns:

mesh (the same class as this one.)

smooth(iterations=20, pass_band=0.1)[source]

A smoothed copy, by Taubin’s non-shrinking filter.

Cosmetic, and priced honestly: applied to a block-model contour this was measured to take away a sixth of the faceting while moving the surface 50% further from the true level set – the creases go, and accuracy goes with them, which is why no contour smooths itself. For a surface that is both rounder and closer to the truth, contour with supersample instead; smooth when the look of the mesh is what matters.

The caller’s kind is kept, as in simplify.

Parameters:
  • iterations (int) – Passes of the filter; more is smoother.

  • pass_band (float) – The filter’s pass band, in (0, 2): lower smooths more.

Returns:

mesh (the same class as this one.)

classmethod from_dxf(filename)[source]

Reads a triangulated surface from a DXF file.

Three ways of writing a triangulation are understood. The MESH entity that export_dxf writes already holds a vertex list and the faces that index into it, and is taken as it stands. POLYFACE meshes and loose 3DFACE entities instead repeat the coordinates of every corner they share, and are welded back into shared vertices, matched to six decimal places. Faces with more than three corners are split into a fan of triangles.

Every mesh in the file is read and the results are concatenated, so a file holding several bodies comes back as one surface in several disconnected pieces. Each MESH entity keeps its own vertices, while the welded entities share one vertex list, so pieces that meet there are joined. Entities nested inside blocks are not searched.

Only the geometry is read: see export_dxf on what a DXF file has no room for.

Parameters:

filename (str) – Path of the file to read.

Returns:

mesh (Surface3D, Solid3D or Mesh3D) – Whichever the geometry read calls for, with normals computed from the triangles (see geometry.vertex_normals), since a DXF file carries none.

export_dxf(filename, offset=None)[source]

Writes this surface to a DXF file, as a single MESH entity.

A MESH holds the vertex list and the triangles that index into it, so the surface comes back from from_dxf exactly as it went out – nothing is welded and there is no ceiling on the number of vertices, unlike the POLYFACE mesh DXF is more often written as.

Only the geometry travels. A DXF file has nowhere to put the variables and metadata a surface carries: to_zarr keeps a container whole, and as_pyvista carries the values onto a mesh object.

Parameters:
  • filename (str) – Path of the file to write.

  • offset (array-like) – Added to the coordinates on the way out, as in export_micromine, for writing into a local grid. It is not recorded in the file, so reading it back gives the shifted coordinates.

classmethod from_geoh5(workspace, name=None)[source]

Reads a triangulated surface from a geoh5 workspace.

A workspace holds any number of named objects; name says which Surface to read, and may be left out when the file holds exactly one. geoml.data.geoh5.contents lists what there is to name.

Only the geometry is read, as in from_dxf, with the normals computed from the triangles. Needs the geoh5py package: pip install geoml[geoh5].

Parameters:
  • workspace (_types.PathLike | _GeoH5Workspace) – Path of the workspace to read, or an open geoml.data.geoh5.Workspace.

  • name (str | None) – The Surface object to read, when the file holds more than one.

Returns:

mesh (Surface3D, Solid3D or Mesh3D) – Whichever the geometry read calls for.

Raises:

ValueError – If the workspace holds no such Surface — the message lists what it does hold.

Return type:

Mesh3D

to_geoh5(workspace, name='Surface', replace=True, folder=None)[source]

Writes this surface into a geoh5 workspace, as a Surface object.

A workspace is what Geoscience ANALYST — a free viewer for the format — opens as one project, so a model’s pieces belong in one file: writing into an existing path adds the object beside what is there, and several exports in a row go fastest through an open geoml.data.geoh5.Workspace, which holds the file open across them.

Only the geometry travels, as in export_dxf: to_zarr is what keeps a container whole. Needs the geoh5py package: pip install geoml[geoh5].

Parameters:
  • workspace (_types.PathLike | _GeoH5Workspace) – Path of the workspace to write into, or an open geoml.data.geoh5.Workspace.

  • name (str) – The name the object gets in the workspace.

  • replace (bool) – Whether an existing Surface of this name, in this folder, makes way — what a re-run export script means. False keeps both, and reading that name back then requires saying which.

  • folder (str | None) – Where the object sits in ANALYST’s project tree, as a path — “Surfaces/Ore” — each segment a group, created when it does not exist and reused when it does. None is the root.

export_micromine(points_filename='points', triangles_filename='triangles', offset=[0, 0, 0], **kwargs)[source]
as_pyvista(simulations=False, include='**')[source]

Converts this object to a pyvista one, carrying its variables.

Parameters:

simulations – Which simulations to include: False for none (the default, since each one is a full-length array in the exported object), True for all of them, an int for the first n, or a sequence of indices.

class geoml.data.meshes.Surface3D(points, triangles, normals)[source]

Bases: Mesh3D

A mesh that does not close: a sheet, with an edge to it.

A topography, a seam roof, a weathering front, a fault plane — anything that has two sides rather than an inside. The promise is checked where it is made, so assign_from_surface need only be given one of these.

intersection(other)[source]

The part of this sheet lying inside a body, or under a terrain.

Against a body, the sheet is cut where it crosses the body’s surface, so what comes back follows the body’s shape rather than the triangles’ — the piece of a fault plane inside an ore envelope, say. Against a single-valued sheet — a topography — the cut is against the ground below it, keeping what lies under. A sheet lying wholly outside comes back empty.

Parameters:

other (Solid3D or Surface3D) – The body to cut against, or the terrain whose underneath to keep.

Returns:

surface (Surface3D)

difference(other)[source]

The part of this sheet lying outside a body, or over a terrain.

The complement of intersection: together the two hold the whole sheet. A sheet lying wholly inside comes back empty.

Parameters:

other (Solid3D or Surface3D) – The body to cut away, or the terrain whose underneath to cut away.

Returns:

surface (Surface3D)

clip_meshes(meshes)[source]

Everything below this sheet, each mesh cut to its own kind.

The batch form of cutting against a terrain: one ground body is extruded under the sheet and serves every cut, where cutting one by one would rebuild it per mesh. A body comes back a closed body (the boolean engines see to it), a sheet comes back a sheet — a shell that runs out of its model is open, and stays open here; contour it with close= first if a body is what is wanted.

Parameters:

meshes (sequence of Mesh3D) – Bodies and sheets to cut below this one.

Returns:

list – One cut mesh per input, in the same order.

class geoml.data.meshes.Solid3D(points, triangles, normals)[source]

Bases: Mesh3D

A mesh that closes: a body, with an inside.

An ore envelope, a stope, a dyke, a contoured shell. Both promises are checked where they are made — the mesh must close, and its triangles must agree which way is out — so assign_from_solid need only be given one of these, and volume always means something.

A body wound inwards is turned round on the way in rather than refused: nothing about it is ambiguous, only reversed. The triangles are what get reversed; the normals are left as they were given.

volume

The volume enclosed, always positive. Zero for an empty body, which is what an intersection of two bodies that do not meet comes to.

Type:

float

union(other)[source]

A body covering everything either of these two covers.

Two bodies that do not meet make a union in two pieces, which is one legitimate body; split() takes it apart.

intersection(other)[source]

A body covering what both of these two cover, empty where they do not meet at all.

difference(other)[source]

A body covering what this one covers and the other does not.

Where other lies wholly inside this one the answer is this body with a cavity in it, which is written as both surfaces, the inner one turned inwards — so volume comes to the difference of the two, and a location in the cavity tests as outside.

class geoml.data.meshes.DTM3D(points, triangles, normals)[source]

Bases: Surface3D

A terrain: a sheet standing at one height over each (x, y).

A digital terrain model, and the shape most of the surfaces in a project have — a topography, a seam roof, a weathering front. The promise is that it never folds back over itself, checked where the object is made, which is what lets a body be divided into what lies under it and what lies over it, and what makes “the elevation here” a question with one answer.

Not what mesh3d returns: an ordinary sheet is a Surface3D unless a terrain is asked for, this being a promise to make rather than a fact to detect. Triangles standing exactly vertical are allowed, a cliff being single valued everywhere but along the line of its face.

geoml.data.meshes.mesh3d(points, triangles, normals)[source]

A mesh of whichever class its geometry calls for.

A Solid3D where the triangles close and agree which way is out, a Surface3D where they do not close, and a plain Mesh3D where they close but disagree — the one case that is neither a sheet nor a body, and what Mesh3D.heal exists for.

Parameters:
  • points (array) – An (n, 3) array of vertex coordinates.

  • triangles (array) – An (m, 3) array of vertex indices.

  • normals (array) – An (n, 3) array of vertex normals.

Returns:

mesh (Surface3D, Solid3D or Mesh3D)

Mesh sets

Every contour of one column at once: a block model contoured at each of a variable’s cut-offs, or once per category, for the prediction and for each realization, as one read-only mapping – shells[0.5], shells["BIF"], shells.simulations[4][0.5]. The set cuts its meshes to limits, checks that shells nest and categories do not overlap, and measures volumes, bands, tonnage and how the realizations’ volumes spread. Design record: Mesh sets.

Mesh sets: every contour of one column of a block model or a grid, keyed by cut-off – or by name, for a categorical variable – the prediction’s and one set per realization, with the reports only a whole set can make.

class geoml.data.meshsets.MeshSet(data, path, cutoffs=None, close='above', limits=None, exclude=None, simulations=True, supersample=0, simplify=None, rule='largest', repair=False, workers=None, store=None)[source]

Bases: Mapping

Every contour of one column, at every cut-off, as one set.

Contours a block model or a grid at each of a variable’s cut-offs, or a categorical variable once per category, and holds the bodies as a read-only mapping: shells[0.5] is the body where the prediction clears 0.5, shells[“BIF”] the body a category holds. The same is done for every realization the variable carries, each realization being a set of its own – shells.simulations[4][0.5] – whose meshes wait in a Zarr store until asked for.

Every mesh is a closed Solid3D, cut to the limits the set was given: a sheet keeps what lies below it, a body what lies inside it, and the exclusions take theirs away. A terrain is extruded into the ground under it once for the whole set.

In principle the shells nest – the body above a higher cut-off lies inside the body above a lower one – and categories do not overlap. check measures how far that fails, exactly, and repair enforces it. The volumes, bands and differences are worked out on the triangles by Manifold, in one frame for the whole set.

Parameters:
  • data (geoml.data.blocks.BlockSet3D | geoml.data.grids.Grid3D | None) – The block model or grid carrying the column: a BlockSet3D, or a regular three-dimensional grid.

  • path (str) – What to contour, named the way the tree names it: a variable or component, whose prediction is contoured (“Comp/Fe”), a column (“Comp/Fe/latent_variance”), or a categorical variable, for one body per category.

  • cutoffs (_types.Cutoffs | None) – The levels to contour a continuous column at. The variable’s own cut-offs when left out; a column that is not a prediction has none of its own and needs them here.

  • close (str) – “above” keeps the ground where the values clear each cut-off, a grade shell; “below” the ground under it. A categorical set is always the ground each category holds.

  • limits (dict[str, geoml.data.meshes.Mesh3D]) – Sheets and bodies every mesh is cut to, by name: a sheet keeps what lies below it, a body what lies inside it.

  • exclude (dict[str, Mesh3D] | None) – Sheets and bodies taken away from every mesh, by name: what lies below a sheet, or a body’s inside.

  • simulations (bool | int | Sequence[int]) – Which realizations to contour as well: True for every one the variable carries, False for none, an int for the first n, or a sequence of realization numbers. Each costs about what the prediction’s own contours do, spread over workers.

  • supersample (int) – How many levels past a block model’s finest block each contour is cut to, as in BlockSet3D.get_contour.

  • simplify (float | None) – A geometric error budget every mesh is simplified to after it is cut, in coordinate units, as in Mesh3D.simplify. Everything the set reports is measured on the meshes it holds.

  • rule (str | None) – How a realization of a categorical variable decides which category holds a location: “largest”, the category with the largest draw, which is the rule of likelihood.CategoricalGaussianIndicator; or “priority”, the latest category in the variable’s order whose draw is positive, which is the rule of likelihood.HierarchicalGaussianIndicator.

  • repair (bool) – Whether to make every set consistent as it is made – each shell cut to the one outside it, each category’s body cut away from the ones before it – rather than only reporting what check finds.

  • workers (int | None) – How many processes contour the realizations; 1 contours them in this one. Defaults to the number of CPUs, eight at most, and no more than memory holds: each worker costs about what the first realization did, which the set contours in this process and measures before the pool starts, or the prediction’s worst contour where that was more, and a warning says when that leaves fewer. A worker that dies before it is done raises a RuntimeError.

  • store (_types.PathLike | None) – The path of a Zarr store to keep the set in, which MeshSet.open reads back. A temporary store, removed with the set, holds the realizations when left out.

data

The container contoured; None for a set read back without one.

Type:

geoml.data.blocks.BlockSet3D | geoml.data.grids.Grid3D | None

path

The column contoured, or the categorical variable.

Type:

str

kind

“cutoff”, keyed by cut-off, or “category”, keyed by name.

Type:

str

close

The side each body keeps, “above” or “below”.

Type:

str

limits, excluded

The sheets and bodies the meshes were cut to and cut away from, as given.

Type:

dict

realization

Which realization this set is; None for the prediction’s.

Type:

int or None

provenance

What the set was made from and how, which each mesh also carries.

Type:

dict

repairs

The volume taken from each mesh to make the set consistent, on a set that was.

Type:

pandas.Series or None

See also

BlockSet3D.get_contour

one contour, which a set makes many of.

Solid3D

what every mesh of a set is.

Examples

shells = geoml.data.MeshSet(blocks, "Comp/FeO_total",
                            cutoffs=[40, 50, 60],
                            limits={"topography": topography})
shells[50.0].volume
shells.bands[(40.0, 50.0)]
shells.simulations[4][50.0]
shells.volume_dispersion()
data: BlockSet3D | Grid3D | None
path: str
kind: str
close: str
supersample: int
rule: str | None
limits: dict[str, Mesh3D]
excluded: dict[str, Mesh3D]
realization: int | None
provenance: dict
repairs: Series | None
complete: bool = True
property simulations: _Realizations | None

The realizations contoured, each a MeshSet of its own.

A sequence in the order of the realization numbers: when every realization was contoured, position and number agree, and numbers says which they are when only some were. Each set is read from the store when asked for, and not kept here.

property unit: str | float | None

What the contoured variable is measured in, where it says so.

property failures: list[dict]

The realization shells that could not be made, and why.

property bands: _Bands

The bodies between consecutive cut-offs, keyed by (low, high).

Above, the band from 0.3 to 0.5 is the shell at 0.3 less the shell at 0.5, and the last band is the top shell itself, (0.5, inf); below, the bands run up from (-inf, lowest). One band per cut-off, in the same order. Each is worked out when first asked for.

table(density=None, grade=None)[source]

What each mesh holds, one row per cut-off or per category.

volume is the mesh as held, raw_volume the contour before any limit cut it, and one removed: <name> column per limit and exclusion says what each took, in the order they cut. pieces and largest count the bodies a mesh is in and the largest one’s share. For a cut-off set, band and band_volume give the ground from each cut-off to the next on the side kept, and crossing how much of each shell lies outside the one around it; a categorical set gives each body’s overlap with the others instead. blocks_volume adds up the blocks whose own value is on the kept side of the cut-off, or whose predicted category it is, for comparison with the raw contour.

With a density or a grade, each band or body is also measured against the blocks: tonnage, the weighted mean of the grade and the metal it comes to, and – where the grade carries realizations – metal_p10, metal_p50 and metal_p90, every realization’s grade filling the same bands. A block a surface passes through counts the share of its sub-blocks inside.

Parameters:
  • density (float | str | None) – A number, the name of a metadata column, or the path of a variable whose prediction is the density; realizations of a density are paired one to one with the grade’s.

  • grade (str | None) – The variable or column to measure in each band or body, by path.

Returns:

pandas.DataFrame

Return type:

DataFrame

check()[source]

Where the set fails its own promise, measured exactly.

For a cut-off set, one row per pair of consecutive shells: the volume of the inner shell lying outside the outer one, which nested shells would hold none of. For a categorical set, one row per pair of categories – the volume both bodies claim – and one for the gap, the ground within the model and its limits that no body holds. Part of a gap is always the model’s own edges: a body closed against the box has its caps on the faces but rounds the edges where two faces meet by about half a boundary block, so bodies that fill the box between them still leave those strips.

Returns:

pandas.DataFrame – kind (“crossing”, “overlap” or “gap”), first, second, volume, and share: of the inner shell, of the smaller body, or of the ground.

Return type:

DataFrame

volume_dispersion()[source]

The realizations’ mesh volumes against the prediction’s.

One row per cut-off or category. prediction is the volume of the prediction’s mesh; mean, sd, p10, p50 and p90 are of the realizations’. rank is the share of realizations whose mesh is smaller than the prediction’s, ties counted half: far from one half, the prediction’s mesh is not a typical realization’s, which is what a smooth field does at a cut-off in either tail. gained and lost are the mean volume a realization’s mesh adds outside the prediction’s and leaves out of it, as shares of the prediction’s – large while the volumes agree means the same volume in different places. pieces and pieces_p50 compare how fragmented they are.

Returns:

pandas.DataFrame

Raises:

ValueError – If no realization was contoured.

Return type:

DataFrame

realization_volumes()[source]

Every realization’s mesh volumes, as measured when it was made.

Returns:

pandas.DataFrame – One row per realization, one column per cut-off or category.

Return type:

DataFrame

realization_table(density=None, grade=None, simulations=True)[source]

What each realization’s own meshes hold, measured against the blocks.

table fills the prediction’s bands with every realization’s grade, which is what mining the prediction’s shells would recover; this is the other spread, each realization’s own bands filled with that realization’s grade and density – how much material there is. A block a band’s surface passes through counts the share of its sub-blocks inside, so each band of each realization costs about what one band of table(grade=) does.

Parameters:
  • density (float | str | None) – A number, the name of a metadata column, or the path of a variable whose prediction is the density; a density with realizations is read realization by realization.

  • grade (str | None) – The variable or column to measure in each band or body, by path; one with realizations is read realization by realization, one without is the same in all of them.

  • simulations (bool | int | Sequence[int]) – Which realizations: True for all, an int for the first n, or a sequence of realization numbers.

Returns:

pandas.DataFrame – One row per realization and band or body: volume, and with a density tonnage, with a grade mean and metal.

Return type:

DataFrame

connectivity()[source]

How many pieces each mesh is in, and the largest one’s share.

Read along the cut-offs, the largest piece’s share is a connectivity curve: where it drops, the ground above the cut-off breaks into pods. With realizations, largest_p10 to largest_p90 and pieces_p50 say the same of theirs.

Returns:

pandas.DataFrame

Return type:

DataFrame

spacing()[source]

How far apart consecutive shells sit.

For each pair, the distance from every vertex of the inner shell to the surface of the outer one: tight means the grade climbs fast, a sharp contact; wide, a gradational one.

Returns:

pandas.DataFrame – One row per pair: inner, outer, and the distances’ min, p10, p50, p90 and max.

Return type:

DataFrame

compare(other)[source]

How another set’s meshes differ from these, key by key.

gained is the volume of the other’s mesh lying outside this one’s, lost the volume of this one’s the other leaves out – exact, both ways – and moved_p50, moved_p90 and moved_max how far the other’s vertices sit from this one’s surface. Two models of the same ground, or one model before and after a drilling campaign.

Parameters:

other (MeshSet) – A set with keys in common with this one.

Returns:

pandas.DataFrame – One row per key the two have in common.

Return type:

DataFrame

section(axis, value)[source]

Where each mesh crosses a plane across one axis.

Parameters:
  • axis (int | str) – The coordinate held fixed, by index or by label (“X”).

  • value (float) – Where along it the plane sits.

Returns:

dict – For every key, a list of polylines, each an (n, 3) array of points in order; a closed line repeats its first point last.

Return type:

dict[Any, list[ndarray]]

repair(priority=None)[source]

The set made consistent, and what that took.

Each shell is cut to the one outside it, so the shells nest; or each category’s body is cut away from the ones before it in priority, so no ground is claimed twice. What each mesh lost is in repairs on the set returned. A gap is not filled – that would be inventing ground.

Parameters:

priority (Sequence[str] | None) – For a categorical set, the order in which categories keep their ground, first first. The set’s own order when left out.

Returns:

MeshSet – The prediction’s meshes, repaired; realizations are not carried.

Return type:

MeshSet

clip(mesh, name=None)[source]

Every mesh cut to one more limit: what lies below a sheet, or inside a body.

Parameters:
  • mesh (Mesh3D) – The sheet or body.

  • name (str | None) – What to call it among the limits.

Returns:

MeshSet – The prediction’s meshes, cut; realizations are not carried – build the set with the limit to have them cut too.

Return type:

MeshSet

exclude(mesh, name=None)[source]

Every mesh with one more piece taken away: what lies below a sheet, or inside a body.

Parameters:
  • mesh (Mesh3D) – The sheet or body.

  • name (str | None) – What to call it among the exclusions.

Returns:

MeshSet – The prediction’s meshes, cut; realizations are not carried.

Return type:

MeshSet

simplify(max_error)[source]

Every mesh on as few triangles as max_error allows, still nested.

Each mesh is simplified on its own (Mesh3D.simplify), which can move two shells closer than twice the budget across each other, so a cut-off set is nested again afterwards, each shell cut to the one outside it; repairs says what that took.

Parameters:

max_error (float) – The largest distance a simplified surface may sit from its original, in coordinate units.

Returns:

MeshSet – The prediction’s meshes, simplified; realizations are not carried.

Return type:

MeshSet

drop_pieces(min_volume)[source]

Every mesh without the pieces smaller than min_volume.

For pods below a mining unit, say. A cut-off set is nested again afterwards: dropping a piece of an outer shell takes whatever of the inner shells it held.

Parameters:

min_volume (float) – The smallest body to keep, in cubic coordinate units.

Returns:

MeshSet – The prediction’s meshes, without the small pieces; realizations are not carried.

Return type:

MeshSet

classmethod probability(data, path, cutoff, levels=(0.1, 0.5, 0.9), side='above', limits=None, exclude=None, supersample=0, simplify=None)[source]

The bodies where a variable clears a cut-off with given probability.

For each level, the ground where at least that share of the realizations clears cutoff on side: at 0.9 the ground the model is sure of, at 0.1 the ground it cannot rule out. Worked out from the realizations a band of blocks at a time, and never written to the container. The set is keyed by the levels, and nests like any cut-off set.

Parameters:
  • data (BlockSet3D | Grid3D) – The block model or grid.

  • path (str) – The variable or component, by path.

  • cutoff (float) – The cut-off the probability is of.

  • levels (Sequence[float]) – The probabilities to contour at.

  • side (str) – “above”, the probability of exceeding the cut-off, or “below”, of falling under it.

  • limits (dict[str, Mesh3D] | None) – As for MeshSet.

  • exclude (dict[str, Mesh3D] | None) – As for MeshSet.

  • supersample (int) – As for MeshSet.

  • simplify (float | None) – As for MeshSet.

Returns:

MeshSet

Return type:

MeshSet

assign(container, name, fraction=None)[source]

Writes which band or body each location falls in, as metadata.

A coded column named name: for a cut-off set the band, counted by how many shells hold the location, labelled < 0.3, 0.3–0.5, ≥ 0.5; for a categorical set the first body in the set’s order holding it, missing where none does. Grade-shell domaining for drillholes, say.

Parameters:
  • container (PointData) – Anything with coordinates: points, drillhole composites, a grid, blocks.

  • name (str) – The metadata column to write.

  • fraction (str | None) – For blocks, a prefix: one more column per band or body, named “<fraction> <label>”, holding the share of each block’s sub-blocks inside it.

crossed_by(blocks)[source]

Which blocks any mesh of the set passes through.

BlockSet3D.crossed_by asked of every mesh at once, to refine a model against the whole set: blocks.split(shells.crossed_by(blocks)).

Returns:

array – One boolean per block.

Return type:

ndarray

to_geoh5(workspace, folder=None, replace=True, simulations=False)[source]

Writes every mesh into a geoh5 workspace, one Surface each.

Named after the variable and the cut-off, with its unit where the variable declares one, or after the category, each carrying its provenance. The limits and exclusions go in a limits folder beside them, and each realization asked for in simulations/<n>. Empty meshes are left out. Needs geoh5py: pip install geoml[geoh5].

Parameters:
  • workspace (_types.PathLike | _GeoH5Workspace) – The path of the workspace, or an open geoml.data.geoh5.Workspace.

  • folder (str | None) – Where the set goes in ANALYST’s project tree, “Shells/Fe”.

  • replace (bool) – Whether objects of the same names in the same folders make way.

  • simulations (bool | int | Sequence[int]) – Which realizations to write as well: True for all, an int for the first n, or a sequence of realization numbers.

export_dxf(filename, offset=None)[source]

Writes every mesh into one DXF file, a layer each.

Each mesh is a MESH entity on a layer named after the variable and the cut-off, or after the category, coloured in turn. Only the geometry travels, as in Mesh3D.export_dxf.

Parameters:
  • filename (str | PathLike) – The file to write.

  • offset (ArrayLike | None) – Added to every coordinate on the way out, and not recorded.

as_pyvista()[source]

Every mesh as one pyvista MultiBlock, named as to_geoh5 names them.

Return type:

MultiBlock

plot(plotter=None, opacity=0.35, colors=None, show_limits=False, **kwargs)[source]

Adds every mesh to a pyvista scene, translucent, a colour each.

Parameters:
  • plotter (Plotter | None) – The scene to add to; a new pyvista.Plotter when left out.

  • opacity (float) – How opaque each mesh is: nested shells need to be seen through.

  • colors (Sequence[str] | None) – One colour per mesh, in order. The package palette by default.

  • show_limits (bool) – Whether to draw the limits and exclusions too, as wireframes.

  • **kwargs – Passed to add_mesh.

Returns:

pyvista.Plotter

Return type:

Plotter

to_zarr(path)[source]

Writes the set into a Zarr store, which MeshSet.open reads back.

Every mesh goes in a group of its own – prediction/0.5, simulations/4/0.5, limits/topography – each a container that Solid3D.open also reads on its own, with what the set measured beside them. The realizations are copied from the set’s store.

A store already at path is replaced, except for the root attributes other programs wrote there: every one whose key does not start with geoml is kept. Writing a set to its own store rewrites its description in place.

Parameters:

path (str | PathLike) – Where to write the store.

Returns:

The path written.

Raises:

ValueError – If path lies inside or around the store this set reads its meshes from, or is that store and this set is one realization of another.

Return type:

str | PathLike

classmethod open(path, data=None)[source]

A set written by to_zarr, its meshes read when asked for.

Parameters:
  • path (str | PathLike) – The store.

  • data (BlockSet3D | Grid3D | None) – The block model the set was contoured from, for the reports that measure against the blocks; they are left out without it.

Returns:

MeshSet

Return type:

MeshSet

Variables

What a container holds at each location: measurements, and everything a model writes back.

A variable may declare what it is measured in — unit="g/t", unit="%", or a number to divide by. On a variable a model reads directly the unit is a label: it names an axis and rides an export, and no number changes. On a part of a CompositionalVariable it is also the divisor, because parts in different units cannot be added up and a composition is defined by its sum; there the parts are stored, predicted and simulated in the units they were assayed in, becoming fractions of the whole only where the model reads and writes them. UNITS is the table of names the package can divide by.

The variable family: _Variable and the concrete kinds a container holds (continuous, vector, compositional, categorical, binary), each declaring its own columns for the tree machinery in base to fold over. Constructed by a container’s add_*_variable methods, never directly.

class geoml.data.variables.ContinuousVariable(name, coordinates, measurements=None, unit=None)[source]

Bases: _Variable

Representation of a continuous random variable.

measurements

The raw measurements.

Type:

_Attribute

latent_mean

The mean of the latent Gaussian representation.

Type:

_Attribute

latent_variance

The variance of the latent Gaussian representation.

Type:

_Attribute

dispersion

How much the locations inside each block differ among themselves – the variance over a block’s sub-blocks, averaged over the realizations, in the variable’s own units rather than the latent ones. A different question from latent_variance, which is how sure the model is of the block: a well-known block can still be heterogeneous, and that is what decides whether cutting it finer would tell anyone anything. Filled only where the container discretizes; elsewhere a location has no interior and this stays missing rather than zero.

Type:

_Attribute

noise_variance

How far a fresh measurement here would fall from the value above – the likelihood noise carried into the variable’s own units, averaged over the realizations. The third of three variances and the third question: latent_variance is how sure the model is of the value, dispersion is how much the ground varies inside a block, and this is how much a sample of it would scatter. A prediction reports the ground, with the noise integrated out, so this is what has to be added back to compare against an assay. Missing where the prediction was made with include_noise=False, there being no integration to read it from.

Type:

_Attribute

simulations

Draws from the variable’s posterior distribution, in a single (n_data, n_sim) array. Use simulation() to get one of them as an _Attribute.

Type:

ArrayStore

quantiles

The variables quantiles, indexed by the corresponding percentile.

Type:

dict

probabilities

Cumulative distribution probabilities, indexed by the corresponding quantile.

Type:

dict

responsibilities

Under a Mixture likelihood, how likely each measurement is to have come from each of its noise components, indexed by the component’s position. Empty otherwise; written by set_responsibilities.

Type:

dict

unit

What the values are measured in – “%”, “ppm”, “g/t”, or a number. On a variable a model reads directly this is a label: the values go to the likelihood as they stand, and the unit travels so that a figure can say what an axis is in and an export can record it. On a part of a CompositionalVariable it is also the divisor that turns the value into a fraction of the whole, since parts in different units cannot be added up. None is undeclared.

Type:

str, float or None

measurements: _Attribute
unit: str | float | None
latent_mean: _Attribute
latent_variance: _Attribute
prediction: _Attribute
dispersion: _Attribute
noise_variance: _Attribute
quantiles: dict[float, _Attribute]
probabilities: dict[float, _Attribute]
cutoffs: list[float] | None
proportions: dict[float, _Attribute]
divided: dict[float, _Attribute]
responsibilities: dict[int, _Attribute]
set_unit(unit)[source]

What the values are measured in.

A label here: the values reach the likelihood as they stand, and the unit travels with the variable so that a figure can say what an axis is in and an export can record it. Anything is accepted, UNITS holding only the ones that can also be divided by – which is what a part of a composition needs, and what _Component.set_unit insists on.

Return type:

ContinuousVariable

divisor()[source]

What to divide this variable’s values by to make them fractions.

Return type:

float

set_cutoffs(cutoffs)[source]

The grades this variable is judged against.

They travel with the variable, so a model trained on data that declares them hands them to every block model predicted from it, and refine knows what the blocks have to be resolved against without being told a second time. The proportions and divided columns of a cut-off no longer declared are dropped.

Return type:

ContinuousVariable

prediction_input()[source]
get_measurements()[source]
get_simulations()[source]
get_predictions()[source]
reset_quantiles(probabilities=None)[source]

Resets the variable’s quantiles.

Parameters:

probabilities (ArrayLike | None) – Probabilities between 0 and 1, exclusive, at which to take the quantiles.

reset_probabilities(quantiles=None)[source]

Resets the variable’s probabilities.

Parameters:

quantiles (ArrayLike | None) – Values in the variable’s own units, at which to take the cumulative probabilities.

classmethod from_variable(coordinates, variable)[source]
update(idx, **kwargs)[source]
allocate_simulations(n_sim)[source]
compute_metrics(alpha=0.05)[source]

Scores this variable’s prediction against its own measurements.

Parameters:

alpha – Significance level for the interval-based scores.

Returns:

dict – One entry per score, named.

Notes

The spread-based scores here – goodness, coverage, CRPS, the interval score – are of the ground: a container’s simulations have the likelihood’s noise integrated out, so they describe a quantity no sample observes, while the measurements they are compared against carry it. Those scores therefore read pessimistic on held-out data, by the share of the variance the model calls noise, and increasingly so for a model with more capacity, which calls less of it noise. Measured on Jura, a nominal 90% band read 0.59 here against 0.94 through the measurement distribution.

For calibration on data the model has not seen, use geoml.models.cross_validate(), whose scores come from geoml.models.VGPNetwork.predict_measurements(), or the accuracy figure, which asks the model for the same thing. The location-wise scores (rmse, mae, bias) are unaffected: integrating the noise out changes the spread, not the value.

See also

geoml.models.VGPNetwork.predict_measurements

the distribution an assay is drawn from.

geoml.models.cross_validate

out-of-fold scores, of measurements.

class geoml.data.variables.DerivedVariable(name, coordinates, parents=None, unit=None)[source]

Bases: ContinuousVariable

A variable computed from others, realization by realization.

The middle ground between metadata (a constant the models never see) and a modelled variable (measured, likelihooded, written by a model): it carries a full set of simulations and everything built on them – quantiles, cut-offs, contours, grade-tonnage – but every bit of its uncertainty is inherited from the variables it was derived from. Built by derive on the container, never fed to a model. Applying the function to each realization and summarizing afterwards is what keeps a nonlinear function honest: f(E[grades]) is not E[f(grades)], and the second is the answer.

The recipe – the function itself – lives in the script that ran derive, not here: functions do not survive a Zarr store honestly. A reloaded container has the values, fully usable; re-deriving is running the script again. parents records which paths it came from.

training_input(idx=None)[source]
get_measurements()[source]
update(idx, **kwargs)[source]
class geoml.data.variables.LatentVariable(name, coordinates, labels)[source]

Bases: _Variable

What a node inside a model’s tree says, at every location.

Written by geoml.models.VGPNetwork.predict_node(): one part per output of the node, each holding the node’s mean and variance there and its realizations. Everything is on the latent scale – a node has no likelihood, so nothing is back-transformed, and there are no measurements, units or cut-offs.

The moments are the node’s own, carried up the tree from the inputs. The realizations carry only the variance a GP node’s inducing points explain, so their spread can fall short of latent_variance; above a nonlinear node (Exponentiation, Multiply, GaussianMixture) the moments are an approximation and the realizations are the reference.

A model does not train on it: it holds a prediction, not measurements.

components

One part per output, keyed by label, each with latent_mean, latent_variance and simulations.

Type:

dict

components: dict[str, _LatentPart]
classmethod from_variable(coordinates, variable)[source]
property n_sim

Number of simulations available, or 0 if none were drawn.

training_input(idx=None)[source]
get_measurements()[source]
get_simulations()[source]
get_predictions()[source]
allocate_simulations(n_sim)[source]
update(idx, **kwargs)[source]

Writes one batch: latent_mean and latent_variance as (rows, size), simulations as (rows, size, n_sim).

unpredicted()[source]

One boolean per location: True where any part is missing its mean.

class geoml.data.variables.VectorVariable(name, coordinates, labels, measurements=None, units=None)[source]

Bases: _Variable

components: dict[str, ContinuousVariable]
uncertainty: _Attribute
responsibilities: dict[int, _Attribute]
get_measurements()[source]
prediction_input()[source]

The components’ cut-offs, as one row each.

They are declared per component – two grades are judged against two different numbers – but the model sees the variable whole, so they travel as a matrix with a row per component. A component declaring fewer than the widest is padded with infinity, which nothing is ever above, so its spare columns come back empty and update drops them.

get_simulations()[source]
get_predictions()[source]
classmethod from_variable(coordinates, variable)[source]
classmethod from_data_frame(name, coordinates, df, columns=None, units=None, *args, **kwargs)[source]
update(idx, elementwise=False, **kwargs)[source]
allocate_simulations(n_sim)[source]
reset_quantiles(probabilities=None)[source]
reset_probabilities(quantiles=None)[source]
compute_metrics(alpha=0.05)[source]
class geoml.data.variables.CompositionalVariable(name, coordinates, labels, measurements=None, units=None)[source]

Bases: VectorVariable

A composition, each part in its own unit.

The parts are stored, reported and simulated in the units they were measured in – percent, ppm, g/t – and turned into fractions of the whole only where the model reads them, since parts in different units cannot be added up. add_compositional_variable is the door that prepares them; the two crossings are get_measurements here and _Component.update on the way back.

divisors()[source]

What each part is divided by to become a fraction, in order.

from_model_units(values)[source]

Model-space values (fractions) in the parts’ own units.

values has the component axis second, as everything the model hands back does.

get_measurements()[source]
classmethod from_variable(coordinates, variable)[source]
classmethod from_data_frame(name, coordinates, df, columns=None, units=None, *args, **kwargs)[source]
update(idx, **kwargs)[source]
compute_metrics(alpha=0.05)[source]
class geoml.data.variables.RockTypeVariable(name, coordinates, labels=None, measurements_a=None, measurements_b=None)[source]

Bases: _Variable

components: dict[str, _Category]
predicted: _Attribute
entropy: _Attribute
uncertainty: _Attribute
measurements_a: _Attribute
measurements_b: _Attribute
boundary: _Attribute
get_measurements()[source]
allocate_simulations(n_sim)[source]
classmethod from_variable(coordinates, variable)[source]
classmethod from_data_frame(name, coordinates, df, col_a=None, col_b=None, *args, **kwargs)[source]
update(idx, **kwargs)[source]
training_input(idx=None)[source]
compute_metrics(decluster=False)[source]

Scores this variable’s prediction against its own measurements.

Every score is of one category against the rest, so the table has a column per category. The locations scored are the ones the confusion matrix counts: measured, predicted, and off the contacts, where a location holds two measurements and no one truth.

  • Balanced accuracy, Jaccard, Matthews, Cohen’s kappa, precision, recall and F1 score read the predicted category. Recall is the share of the locations measured as the category that the model calls it; precision is the share of the locations the model calls it that were measured as it.

  • Quantity and allocation disagreement split the category’s errors into the part a wrong proportion explains and the part a wrong place explains (Pontius and Millones, 2011). Summed over the categories and halved, they add up to one minus the accuracy.

  • The Brier score and the log score read the predicted probability, as a forecast of whether a location is the category. Both are proper: neither hedging nor overconfidence improves them. Lower is better.

Parameters:

decluster (bool) – Weight each location by the container’s “declustering” metadata column, which the container’s decluster method writes, so that densely drilled ground does not dominate.

Returns:

pandas.DataFrame – One row per score, one column per category.

Raises:

ValueError – If no location holds both a measurement and a prediction, or if decluster is asked for and the container keeps no weights.

Return type:

DataFrame

Notes

A score with no value is NaN: the precision of a category the model never calls, the recall of one never measured. At the locations a model was trained on, every score flatters it; the out-of-fold container geoml.models.cross_validate() fills is the honest input.

References

Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1-3.

Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37-46.

Pontius, R. G., & Millones, M. (2011). Death to Kappa: birth of quantity disagreement and allocation disagreement for accuracy assessment. International Journal of Remote Sensing, 32(15), 4407-4429.

class geoml.data.variables.CategoricalVariable(name, coordinates, labels=None, measurements=None)[source]

Bases: RockTypeVariable

classmethod from_data_frame(name, coordinates, df, measurements_col=None, *args, **kwargs)[source]
class geoml.data.variables.OrderedRockType(name, coordinates, labels=None, measurements_a=None, measurements_b=None)[source]

Bases: RockTypeVariable

get_measurements()[source]
class geoml.data.variables.BinaryVariable(name, coordinates, labels=None, measurements=None)[source]

Bases: _Variable

indicator: _Attribute
measurements: _Attribute
weights: _Attribute
predicted: _Attribute
probability: _Attribute
entropy: _Attribute
uncertainty: _Attribute
latent_mean: _Attribute
latent_variance: _Attribute
get_measurements()[source]
classmethod from_variable(coordinates, variable)[source]
update(idx, **kwargs)[source]
allocate_simulations(n_sim)[source]
classmethod from_data_frame(name, coordinates, df, col, positive_class)[source]
class geoml.data.variables.AnomalyVariable(name, coordinates, label, measurements=None)[source]

Bases: BinaryVariable

classmethod from_data_frame(name, coordinates, df, col, positive_class)[source]
classmethod from_variable(coordinates, variable)[source]

Paths

How a variable’s columns are named and addressed. Design record: Naming, reaching and exporting a variable’s parts — analysis and plan.

The container tree: the errors, the bounding box, the path grammar (VariablePath, render), the traversal (_TreeNode) and the leaf it carries (_Attribute). Everything a container or a variable is built on, and nothing that is one.

exception geoml.data.base.NoDataError[source]

Bases: Exception

Exception raised when a data object is empty.

exception geoml.data.base.NotGriddedDataError[source]

Bases: Exception

Exception raised when expecting a gridded data object.

exception geoml.data.base.NotClosedError[source]

Bases: ValueError

A mesh that does not bound a volume was asked to.

exception geoml.data.base.InconsistentMeshError[source]

Bases: ValueError

A closed mesh whose triangles disagree about which way is out.

exception geoml.data.base.NotSingleValuedError[source]

Bases: ValueError

A sheet that folds over was asked which of its heights to use.

exception geoml.data.base.MeshTypeError[source]

Bases: ValueError

Two meshes were combined in a way that means nothing.

exception geoml.data.base.DimensionMismatchError[source]

Bases: Exception

Exception raised when the dimensionality of objects does not match.

class geoml.data.base.BoundingBox(min_values, max_values)[source]

Bases: object

An n-dimensional box.

__init__(min_values, max_values)[source]

An n-dimensional box.

Parameters:
  • min_values (array) – The box’s minimum values in each direction.

  • max_values (array) – The box’s maximum values in each direction.

property n_dim
property diagonal
property min
property max
property center
as_array()[source]
as_data_frame()[source]
overlaps_with(other)[source]

Checks if box overlaps with another box.

Parameters:

other (BoundingBox) – The other box.

Returns:

check (bool) – The checking result.

classmethod from_array(array)[source]

Builds bounding box from an array with minimum and maximum coordinates.

Parameters:

array (array) – Array with the minimum and maximum coordinates.

Returns:

box (BoundingBox) – A box object.

contains_points(array)[source]

Checks if points are contained within the box.

Parameters:

array (array) – A set of coordinates.

Returns:

check (bool) – True if all points are contained within the box.

class geoml.data.base.VariablePath(parts=())[source]

Bases: object

Where a piece of data sits inside a container.

A container holds variables, a variable holds components or attributes, and an attribute holds one array per location. This names a place in that tree the way a file system names a file – assay/Zn/noise_variance – so that one string can serve the lookup, the persistence key and the exported column name.

Built from a string, from parts, or from another path; / composes, as it does for pathlib:

>>> VariablePath("assay") / "Zn" / "prediction"
VariablePath('assay/Zn/prediction')
parts
property name

The last segment, or ‘’ for the root.

property parent
geoml.data.base.render(path, style='path')[source]

An addressable path as a name in some flat namespace.

Purely mechanical – the segments joined, nothing else. That is the point: the four spellings this replaced each had rules of their own about which role was abbreviated and which was dropped, and no two agreed.

path is what the store and every internal caller use, flat is for data-frame columns, CSV and mining software (identifier-safe), and pretty is for pyvista and ParaView, where the name is read by a person.

geoml.data.base.render_all(paths, style='flat')[source]

{path: name} for a whole namespace at once, every name distinct.

A path cannot collide – / is refused inside a segment – so a collision is made by the join, and only flat makes one readily: _ is in nearly every role name, so a variable noise with a component variance lands on the same column as a leaf called noise_variance.

The rule has to be deterministic or an export changes shape between runs, so a colliding group is sorted by path and the ones after the first take a suffix. Only the group is touched: suffixing all of them would penalize the innocent column to spare the pathological one. Adding a variable can therefore rename a column inside a colliding group, which is why this warns rather than quietly putting it right.