BRAIDGROUP
RESEARCH & DEV
30. Documentation

Optimization & Performance

Compiler Optimizations

Braid's compiler applies several optimization passes during compilation to improve generated code quality.

Constant Folding

The optimizer evaluates constant expressions at compile time. For example, an AddI64 operation where both operands are constants is replaced with a single constant. This is implemented in opt.c and runs before code generation.

// Before optimization:
let result = 100 + 200
// IR: %0 = ConstantI64 100
//     %1 = ConstantI64 200
//     %2 = AddI64 %0, %1

// After constant folding:
let result = 300
// IR: %0 = ConstantI64 300

Dead Code Elimination

Unreachable code and unused variable assignments are removed from the IR. Variables that are assigned but never read, and code paths that can never be reached, are eliminated.

// Before DCE:
fn compute() -> int {
    let unused = 42
    let result = 10
    return result
    let dead = 99    // unreachable
}

// After DCE:
fn compute() -> int {
    let result = 10
    return result
}

Supercompilation

The supercompiler pass (run_supercompiler in supercompiler.c) performs compile-time function evaluation. Pure functions with constant arguments are evaluated at compile time, and their results are inlined as constants. This is Phase 5 of the C compilation pipeline.

// Supercompiler evaluates this at compile time:
fn square(n: int) -> int {
    return n * n
}

fn main() {
    let x = square(5) + square(10)
    // Supercompiler inlines: (5*5) + (10*10)
    // After fold: 25 + 100
    // After fold: 125
    print(x)
}

Runtime Optimizations

Tensor Operation Vectorization (AVX2)

The Braid runtime uses AVX2 SIMD instructions for tensor operations. The CMake default flags include -mavx2 -mfma -march=native to enable compiler auto-vectorization. Critical tensor ops (matmul, element-wise add, activation functions) are implemented with explicit SIMD intrinsics in cpu_streaming.c and ternary.c.

// Runtime uses AVX2 for tensor matmul:
// - 256-bit vectors process 4 float64 or 8 float32 elements per instruction
// - FMA (fused multiply-add) for dot products
// - Cache-blocked tiling for large matrices

CPU Streaming Engine

The CPU streaming module (cpu_streaming.c) processes tensor data in streaming fashion, overlapping computation with memory loads. This minimizes cache misses and keeps the CPU pipeline full during ML training loops.

// Streaming engine processes data in chunks:
// while stream.has_data() {
//     chunk = stream.next_batch(1024)  // prefetched
//     result = compute(chunk)           // process while next chunk loads
// }

Memory Pool Allocation

Tensor memory pools (cpu_memory.c, test_mempool_c.c) provide O(1) allocation and bulk deallocation. During training, intermediate tensors are allocated from the pool and the pool is reset at each step, eliminating individual free() calls.

// Pool allocation pattern:
fn train_epoch(data: tensor) {
    let pool = create_mempool(1024 * 1024 * 100)  // 100MB pool
    let step = 0
    while step < 1000 {
        let batch = data[step * 32 : (step + 1) * 32]
        let hidden = layer1.forward(batch, pool)  // allocates from pool
        let output = layer2.forward(hidden, pool)
        pool.reset()  // O(1) bulk free, no individual frees
        step = step + 1
    }
    pool.destroy()
}

Benchmark Results

Performance metrics from the CPU training benchmark suite (bench-cpu-training):

OperationThroughputNotes
Ternary matmul (512x512)~250 GFLOPSAVX2 ternary packing
Float32 matmul (1024x1024)~45 GFLOPSCache-blocked, AVX2+FMA
Memory pool alloc/free~50 ns / opvs ~300 ns for malloc/free
13M param forward pass~2 msBatch size 8, seq len 8
13M param training step~15 msForward + backward + update

Tips for Writing Fast Braid Code

  • Use pool allocation in tight loops: call pool.reset() instead of letting ARC free tensors individually
  • Keep tensors on the same device: minimize CPU-GPU transfers by using with device("cuda") blocks
  • Batch operations: process data in vectorized batches rather than element-wise loops
  • Prefer compile-time evaluation: constant expressions and pure functions with constant args are supercompiled
  • Minimize closure creation: closures in hot loops prevent inlining and add ARC overhead
  • Use native functions for performance-critical operations; native fn calls have zero marshalling overhead
  • Avoid nil checks in hot paths: use the type system to guarantee non-nil values
// Optimized training loop:
fn train_fast(model, data: tensor, pool) {
    with device("cuda") {
        let step = 0
        while step < 1000 {
            let batch = data[step * 32 : (step + 1) * 32]
            let logits = model.forward(batch, pool)
            let loss = cross_entropy(logits, labels)
            model.backward(loss)
            model.update(0.001)
            pool.reset()    // O(1), avoids ARC overhead
            step = step + 1
        }
    }
}