BRAIDGROUP
RESEARCH & DEV
66. ML Library Docs

Neural Network Operations

The tensor NN library provides SIMD-accelerated activation functions, normalizations, attention mechanisms, loss functions, ternary quantized operations, and pooling — all with typed dispatch for float64 and float32.

Activation Functions

tensor_relu(t)

Rectified Linear Unit. AVX2-accelerated for float64 (4-wide) and float32 (8-wide) using comparison and masked move.

ObjTensor* r = tensor_relu(t);
// out[i] = in[i] > 0 ? in[i] : 0

tensor_sigmoid(t)

Computes 0.5 * (1 + tanh(x / 2)) for numerical stability.

ObjTensor* s = tensor_sigmoid(t);

tensor_tanh(t)

Element-wise hyperbolic tangent.

ObjTensor* h = tensor_tanh(t);

tensor_gelu(t)

Gaussian Error Linear Unit using the tanh approximation:

// gelu(x) = 0.5 * x * (1 + tanh(sqrt(2/pi) * (x + 0.044715 * x^3)))
ObjTensor* g = tensor_gelu(t);

tensor_silu(t) / Swish

SiLU (Sigmoid Linear Unit): x * sigmoid(x).

ObjTensor* s = tensor_silu(t);

tensor_log(t)

Element-wise natural logarithm.

ObjTensor* l = tensor_log(t);

Normalization

tensor_softmax(t, axis)

Numerically stable softmax. Subtracts the row max before exponentiation and normalizes by the sum. Clones the input and operates in-place on the clone.

ObjTensor* sm = tensor_softmax(logits, -1);  // over last axis

tensor_log_softmax(t, axis)

Computes log(softmax(x)) by calling softmax then log.

ObjTensor* lsm = tensor_log_softmax(logits, -1);

tensor_layer_norm(t, gamma, beta, eps)

Layer normalization over the last dimension. Computes mean and variance, normalizes, then applies affine transform with gamma (scale) and beta (shift).

ObjTensor* ln = tensor_layer_norm(x, weight, bias, 1e-5);

tensor_rms_norm(t, gamma, eps)

Root Mean Square normalization. Uses AVX2 FMA for the sum-of-squares computation. Does not center (no mean subtraction).

ObjTensor* rms = tensor_rms_norm(x, gamma, 1e-6);

tensor_batch_norm(t, gamma, beta, running_mean, running_var, eps, training)

Batch normalization over the channel dimension (axis 1). In training mode computes running statistics with momentum 0.9.

ObjTensor* bn = tensor_batch_norm(x, gamma, beta, rm, rv, 1e-5, 1);

Matrix Operations

tensor_matmul(a, b)

General matrix multiply supporting 1-D (dot product), 2-D, and batched N-D (>=3). AVX2-accelerated for float64 with FMA. Batch dimensions must match (or broadcast) between operands.

// 2-D: [M, K] x [K, N] -> [M, N]
ObjTensor* c = tensor_matmul(a, b);

// Batched: [B, M, K] x [B, K, N] -> [B, M, N]
ObjTensor* batched = tensor_matmul(q, k);

Attention

tensor_scaled_dot_product_attention(Q, K, V, mask, dropout_p, is_causal)

Computes softmax(Q * K^T / sqrt(d)) * V with optional causal masking and dropout. Internally transposes K, computes scores, applies scale, applies causal mask (upper triangle set to -1e9), adds optional mask tensor, applies softmax, dropout, then final matmul with V.

ObjTensor* out = tensor_scaled_dot_product_attention(
    Q, K, V, NULL, 0.0, 1);  // causal, no dropout

Dropout

tensor_dropout(t, p, training)

During training, zeros out elements with probability p and scales remaining elements by 1/(1-p). Uses xorshift128+ PRNG for random number generation.

ObjTensor* dropped = tensor_dropout(x, 0.1, 1);  // 10% dropout

Embedding

tensor_embedding_lookup(weight, indices)

Looks up rows from a 2-D weight table using integer indices. Supports INT32 and INT64 index dtypes with bounds clamping.

ObjTensor* emb = tensor_embedding_lookup(embed_weight, token_ids);

Loss Functions

tensor_cross_entropy_loss(pred, target, label_smoothing)

Computes cross-entropy loss with optional label smoothing. Handles both integer class indices and one-hot targets. Uses log-softmax internally.

ObjTensor* loss = tensor_cross_entropy_loss(pred, target, 0.1);

tensor_mse_loss(pred, target)

Mean squared error: sum((pred - target)^2) / N.

ObjTensor* loss = tensor_mse_loss(pred, target);

tensor_bce_with_logits_loss(pred, target)

Binary cross-entropy with logits. Uses the numerically stable formulation.

ObjTensor* loss = tensor_bce_with_logits_loss(pred, target);

Pooling

tensor_avg_pool2d(t, kernel_h, kernel_w, stride_h, stride_w, pad_h, pad_w)

ObjTensor* pooled = tensor_avg_pool2d(x, 3, 3, 2, 2, 1, 1);

tensor_max_pool2d(t, kernel_h, kernel_w, stride_h, stride_w, pad_h, pad_w)

ObjTensor* pooled = tensor_max_pool2d(x, 2, 2, 2, 2, 0, 0);

Ternary Quantized Operations

tensor_embedding_lookup with ternary weights

When weights use TENSOR_TERNARY dtype, each weight value is packed as a 2-bit entry {-1, 0, +1}, reducing memory by 16x vs float32.

Ternary matmul and embedding

Ternary operations decode packed values and compute dot products using integer arithmetic, avoiding floating-point multiplies in the critical path.

Autograd Integration

Every NN op creates the corresponding AgNode for automatic differentiation when the input tensor has requires_grad set. The grad_node field on ObjTensor tracks the computation graph for backward pass.

Complete Example

// Forward pass through a transformer block
ObjTensor* x = /* [B, T, d_model] */;
ObjTensor* w_q = /* [d_model, d_model] */;
ObjTensor* gamma = /* [d_model] */;

// Layer norm -> attention -> FF
ObjTensor* ln = tensor_rms_norm(x, gamma, 1e-6);
ObjTensor* q = tensor_matmul(ln, w_q);
// ... project K, V, compute attention
ObjTensor* attn_out = tensor_scaled_dot_product_attention(
    q, k, v, NULL, 0.0, 1);
ObjTensor* residual = tensor_add(x, attn_out);
ObjTensor* output = tensor_relu(residual);