BRAIDGROUP
RESEARCH & DEV
64. ML Library Docs

Tensor Library

The Braid tensor library provides N-dimensional array operations with automatic memory management, GPU backends, and autograd integration. Every tensor is a first-class VM object with shape, strides, dtype, and optional gradient tracking.

Data Types

Tensors support the following dtypes defined by the TensorDType enum:

TENSOR_FLOAT64  = 0   // double precision
TENSOR_FLOAT32  = 1   // single precision
TENSOR_FLOAT16  = 2   // half precision
TENSOR_BFLOAT16 = 3   // brain float
TENSOR_INT8     = 4   // 8-bit integer
TENSOR_INT16    = 5   // 16-bit integer
TENSOR_INT32    = 6   // 32-bit integer
TENSOR_INT64    = 7   // 64-bit integer
TENSOR_TERNARY  = 8   // {-1, 0, +1} packed as 2-bit values

Tensor Properties

Each ObjTensor exposes:

  • ndim — number of dimensions
  • shape — array of dimension sizes
  • strides — byte/element strides for each dimension
  • size — total number of elements
  • dtype — element data type
  • data — raw pointer to element storage
  • grad — gradient tensor (populated when requires_grad is set)
  • grad_node — autograd graph node

Tensor Creation

tensor_create(ndim, shape, dtype)

Allocates a new tensor with the given shape and dtype. Uses the memory pool (or falls back to calloc).

// C: 2x3 float32 tensor
int64_t shape[] = {2, 3};
ObjTensor* t = tensor_create(2, shape, TENSOR_FLOAT32);

// Braid:
let t = std.tensor.create([2, 3], 0);

zeros(shape)

Creates a tensor filled with zeros.

// Braid std/ml:
let z = zeros([3, 4])          // float64 zeros

// C equivalent:
int64_t shape[] = {3, 4};
ObjTensor* z = tensor_create(2, shape, TENSOR_FLOAT64);
tensor_zero(z);

ones(shape)

Creates a tensor filled with ones.

let o = ones([128, 256])

randn(shape, mean?, stddev?)

Creates a tensor with values drawn from a normal distribution using the Box-Muller transform.

let r = randn([1024])          // N(0, 1)
let r2 = randn([64, 64], 0.0, 0.02)  // N(0, 0.02)

arange(start, end, step?)

Creates a 1-D tensor with evenly spaced values.

let a = arange(0, 10)         // [0, 1, 2, ..., 9]
let b = arange(0, 1, 0.1)     // [0, 0.1, 0.2, ..., 0.9]

from_buffer(ndim, shape, data, dtype)

Creates a tensor by copying data from an existing buffer.

// C:
double buf[] = {1, 2, 3, 4, 5, 6};
int64_t shape[] = {2, 3};
ObjTensor* t = tensor_from_buffer(2, shape, buf, TENSOR_FLOAT64);

clone(t)

Creates a deep copy of a tensor with its own data storage.

ObjTensor* copy = tensor_clone(original);

Manipulation

reshape(t, ndim, shape)

Returns a view with a new shape (total elements must match). Creates a shallow view sharing the underlying data.

// Reshape 2x3 to 6x1
int64_t shape[] = {6};
ObjTensor* v = tensor_reshape(t, 1, shape);

transpose(t, axis1, axis2)

Returns a view with two axes swapped. Swaps both shape and strides.

// Transpose 2x3 to 3x2
ObjTensor* v = tensor_transpose(t, 0, 1);

slice(t, starts, sizes)

Returns a view over a sub-region of the tensor.

// Slice first 2 rows, all columns
int64_t starts[] = {0, 0};
int64_t sizes[] = {2, 3};
ObjTensor* v = tensor_slice(t, starts, sizes);

flatten(t)

Reshapes the tensor to 1-D. Equivalent to reshape(t, 1, [t->size]).

ObjTensor* flat = tensor_flatten(t);

unsqueeze(t, axis)

Inserts a dimension of size 1 at the given axis.

// shape [3, 4] -> [1, 3, 4]
ObjTensor* v = tensor_unsqueeze(t, 0);

squeeze(t)

Removes all dimensions of size 1.

// shape [1, 3, 1, 4] -> [3, 4]
ObjTensor* v = tensor_squeeze(t);

contiguous(t)

Returns a contiguous copy if the tensor is non-contiguous (e.g., after transpose). Otherwise returns a clone.

ObjTensor* c = tensor_contiguous(view);

broadcast(a, b)

Returns a view of a expanded to match the broadcast-compatible shape with b. Uses zero-strides for dimensions that need repeating.

// a: [3, 1], b: [1, 4] -> broadcast to [3, 4]
ObjTensor* v = tensor_broadcast(a, b);

Element Access

tensor_get(t, indices), tensor_set(t, indices, val)

Access elements by N-dimensional index tuple. Handles all dtypes via double conversion.

int64_t idx[] = {1, 2};
double val = tensor_get(t, idx);
tensor_set(t, idx, 42.0);

tensor_get_flat(t, flat_idx), tensor_set_flat(t, flat_idx, val)

Access elements by linear (row-major) index. Automatically computes N-dimensional offset using strides.

double val = tensor_get_flat(t, 5);
tensor_set_flat(t, 5, 3.14);

Save / Load

Tensors are serialized with a binary format:

  • Magic bytes: "TENS" (4 bytes)
  • Version: uint32_t (currently 1)
  • ndim: uint32_t
  • shape: int64_t[ndim]
  • dtype: uint8_t
  • data: contiguous raw bytes
// Save
tensor_save(t, "tensor.bin");

// Load
ObjTensor* loaded = tensor_load("tensor.bin");
if (!loaded) { /* error */ }

Under the hood, tensor_save forces the tensor to contiguous layout before writing. tensor_load validates magic and version, then reconstructs the tensor from disk.