pelutils.array package

Utilities and various algorithms for working with numpy arrays.

The flagship is unique(), a drop-in for numpy.unique that runs in linear time (rather than sorting), making it dramatically faster on large arrays.

Quick start

import numpy as np
from pelutils.array import unique

x = np.random.randint(0, 100, size=10_000_000)
values = unique(x)
values, index, inverse, counts = unique(
    x, return_index=True, return_inverse=True, return_counts=True,
)

SparseGridBlobDetection is used for detecting blobs in a sparse, N-dimensional grid. It also works for dense grids (effetively boolean numpy arrays), which can trivially be converted to sparse grids using np.where and np.column_stack, as shown in the example.

A blob is defined as an area of touching grid entries which are all True. An obvious example is finding areas that are left in an image after thresholding on pixel values.

greyscale_image = ...  # Numpy array of shape height x width
thresholded = greyscale_image > 100
# Find the indices in the thresholded image which are True and stack them into an n x 2 array
coords_above_threshold = np.column_stack(np.where(thresholded))

blob_detector = SparseGridBlobDetection(coords_above_threshold)
blobs = blob_detector.find_all_blobs()

# The pixel coordinates in the first blob are fetched with
coords_above_threshold[blobs[0]]

Both unique() and SparseGridBlobDetection are implemented in C, making them blazingly fast compared to what you usually expect from Python.

class pelutils.array.SparseGridBlobDetection(grid_coords: IntArray, wrap_axes: dict[int, int] | None = None)[source]

Detect blobs (continuous regions) in a sparse grid, implemented in C for high performance.

The grids can have an arbitrary number of dimensions, but a common usecase might be in image analysis. Imaging thresholding an image based on pixel values. This class offers blazingly fast detection of all blobs in the resulting boolean image.

While images are a prime use case, this class works for any dimensionality of arrays. It provides two methods: find_all_blobs and find_single_blob. The first, as demonstrated above, detects, for each blobs, all pixels belonging to that blob. The second finds only a single blob and optionally accepts a starting pixel.

Adjacency is defined by a plus-shaped kernel (generalised to the dimensions of the array), not a solid square-shaped kernel. To get the behaviour of a square-shaped kernel, run a dilation with a square kernel over the array first.

Note that this class is stateful, meaning that for each call to either of the afforementioned methods, trying to detected an already detected blob will raise a RuntimeError.

Parameters:
  • grid_coords (IntArray) –

    Coordinates in the grid that are part of any blob. For an d-dimensional grid, it has shape n x d where n is the number of nodes in the grid belonging to any blob. For a grid representated by a boolean numpy array, the coordinates can be gotten with np.column_stack(np.where(grid)).

    Note that even though grid_coords commonly would represent axis coordinates in a numpy array, negative coordinates are fully supported.

  • wrap_axes (dict[int, int] | None, optional) –

    A dictionary mapping axis numbers to their corresponding axis lengths in the grid. Values will be wrapped on that axis. Negative axis numbers are supported.

    As an example, if wrap_axes = {2: 10}, the third column in grid_coords will be wrapped to the range (0, 9), and the coordinates 0 and 9 will be considered next to each other. This is effectively the same as doing grid_coords[:, 2] %= 10, and pretending the ends touch each other.

property done: bool

True if all coordinates have been explored and there are no more blobs to discover.

find_all_blobs() → list[IntArray][source]

Detect all blobs in grid_coords.

For each detected blob, an array is returned containing the rows in grid_coords which are part of that blob.

find_single_blob(init_index: int | None = None) → IntArray[source]

Detect one blob starting at grid_coords[init_index], or grid_coords[self.next_index] if init_index is not provided.

Return the index of each row in grid_coords which is part of the blob.

property next_index: int

Return the index of the first unvisited coordinate.

This starts at 0 and advances as blobs are detected, whether by find_single_blob() or find_all_blobs(). It reaches len(grid_coords) when all coordinates have been visited, and there are no more blobs to be found. This is equivalent to done() being True.

When a blob is detected, if the current value is part of that blob, it will increase to the index of the first coordinate which has not been visited.

property visited_coords: BoolArray

Return 1D boolean array of length len(grid_coords) indicating which grid coordinates have been visited.

pelutils.array.array_bytes(x: ArrayLike) → int[source]

Calculate the size of a numpy array or torch tensor in bytes.

pelutils.array.array_ptr(arr: ArrayLike) → c_void_p[source]

Return a pointer to a numpy array or torch tensor which can be used to interact with it in low-level languages like C/C++/Rust.

This function is mostly useful when not using Python’s C api and instead interfacing with .so files directly with ctypes.

pelutils.array.unique(array: ArrayLike, *, return_index: bool = False, return_inverse: bool = False, return_counts: bool = False, axis: int = 0)[source]

Return unique elements in a given numpy array.

This function works very similar to np.unique, but it runs in linear time, making it significantly faster for large arrays. Unlike with np.unique, the returned unique values are unsorted. Because of this, it can also be used for detecting uniqueness along axes when ordering along the respective axes matters.