BRAIDGROUP
RESEARCH & DEV
21. Documentation

Memory Model

Automatic Reference Counting

Braid uses Automatic Reference Counting (ARC) as its primary memory management strategy. Every heap-allocated object carries a ref_count field in its Object header that tracks the number of strong references pointing to it. When a reference is copied, the count increments; when a reference goes out of scope or is overwritten, the count decrements. When the count reaches zero, the object is deallocated immediately.

let x = [1, 2, 3]    // ref_count = 1
let y = x            // ref_count = 2 (shared reference)
y = nil              // ref_count = 1
x = nil              // ref_count = 0 -> deallocated

Cycle Detection

ARC alone cannot handle reference cycles (e.g., two objects referencing each other). Braid's runtime includes a cycle collector that traverses object graphs using the next pointer chain and the marked flag on each Object. Periodically, the collector scans for unreachable cycles and reclaims them.

struct Node {
    value: int
    next: Node     // forward reference
}

let a = Node { value: 1, next: nil }
let b = Node { value: 2, next: a }
a.next = b         // cycle: a -> b -> a
// Cycle collector detects and reclaims both when unreachable

Weak References and Strong Cycles

To break intentional cycles, Braid supports weak references via weak struct fields. A weak reference does not increment the object's ref count. The cycle collector ignores weak edges when tracing, allowing cycles to be freed. Accessing a weak reference returns nil if the target has been deallocated.

struct Parent {
    name: string
    children: Array<Child>
}

struct Child {
    name: string
    parent: weak Parent    // weak back-reference
}

let p = Parent { name: "Alice", children: [] }
let c = Child { name: "Bob", parent: p }
p.children = [c]
// No cycle in ARC counts: Parent owns Child, Child has weak ref to Parent

Stack vs Heap Allocation

Primitive types (int, float, bool) and small structs are allocated on the stack when they are local variables. Large or dynamically-sized data (strings, arrays, tensors, closures) are heap-allocated via malloc through the memory pool system. The compiler decides allocation strategy based on type size and escape analysis.

fn compute() -&gt; int {
    let x = 42              // stack-allocated int
    let name = "hello"      // heap-allocated string (ARC-managed)
    let data = [1, 2, 3]    // heap-allocated array
    return x
}

Tensor Memory Pools

Tensor objects use a MemPool allocator to reduce malloc overhead. Each ObjTensor has an is_pooled flag and a pool pointer. When a tensor is created, its data buffer is carved from a pre-allocated pool. Pools support fast O(1) allocation and bulk deallocation, critical for ML training loops.

fn train_step(batch: tensor) -&gt; float {
    let x = batch[0:32]      // pooled tensor view (no alloc)
    let logits = model(x)    // intermediate tensors use pool
    let loss = cross_entropy(logits, labels)
    return loss
    // pool reset at end of step reuses all memory
}

GPU Memory Management

Tensors can reside on GPU devices. The ObjTensor struct contains a GpuBackend enum, a gpu_handle pointer, and a gpu_backend field. Transfers between CPU and GPU use .to() methods. GPU memory is reference-counted separately and freed when the tensor is no longer referenced.

fn train_on_gpu() {
    let data = [[1, 2], [3, 4]]
    let gpu_data = data.to("cuda")    // transfer to GPU
    with device("cuda") {
        let result = gpu_data * gpu_data
    }
    let cpu_result = result.to("cpu") // transfer back
}

Memory-Safe Patterns

Braid prevents use-after-free and double-free by construction: ARC guarantees that an object lives as long as any reference exists. The compiler enforces that local variables are initialized before use and that pointers to stack data do not escape their scope.

fn safe_patterns() {
    let a = [10, 20, 30]      // heap, ref_count = 1
    {
        let b = a              // ref_count = 2
        print(b[0])            // safe: a is alive
    }                          // b dies, ref_count = 1
    print(a[0])                // safe: a is still alive
}

Performance Notes

ARC eliminates the need for a tracing garbage collector, avoiding stop-the-world pauses. The cycle collector runs incrementally in the background. Memory pools for tensors reduce allocation overhead by ~90% in training workloads. GPU memory is managed with explicit transfer boundaries, minimizing PCI-e traffic.