Memory Model
Automatic Reference Counting
Braid uses Automatic Reference Counting (ARC) as its primary memory management strategy. Every heap-allocated object carries a ref_count field in its Object header that tracks the number of strong references pointing to it. When a reference is copied, the count increments; when a reference goes out of scope or is overwritten, the count decrements. When the count reaches zero, the object is deallocated immediately.
let x = [1, 2, 3] // ref_count = 1
let y = x // ref_count = 2 (shared reference)
y = nil // ref_count = 1
x = nil // ref_count = 0 -> deallocatedCycle Detection
ARC alone cannot handle reference cycles (e.g., two objects referencing each other). Braid's runtime includes a cycle collector that traverses object graphs using the next pointer chain and the marked flag on each Object. Periodically, the collector scans for unreachable cycles and reclaims them.
struct Node {
value: int
next: Node // forward reference
}
let a = Node { value: 1, next: nil }
let b = Node { value: 2, next: a }
a.next = b // cycle: a -> b -> a
// Cycle collector detects and reclaims both when unreachableWeak References and Strong Cycles
To break intentional cycles, Braid supports weak references via weak struct fields. A weak reference does not increment the object's ref count. The cycle collector ignores weak edges when tracing, allowing cycles to be freed. Accessing a weak reference returns nil if the target has been deallocated.
struct Parent {
name: string
children: Array<Child>
}
struct Child {
name: string
parent: weak Parent // weak back-reference
}
let p = Parent { name: "Alice", children: [] }
let c = Child { name: "Bob", parent: p }
p.children = [c]
// No cycle in ARC counts: Parent owns Child, Child has weak ref to ParentStack vs Heap Allocation
Primitive types (int, float, bool) and small structs are allocated on the stack when they are local variables. Large or dynamically-sized data (strings, arrays, tensors, closures) are heap-allocated via malloc through the memory pool system. The compiler decides allocation strategy based on type size and escape analysis.
fn compute() -> int {
let x = 42 // stack-allocated int
let name = "hello" // heap-allocated string (ARC-managed)
let data = [1, 2, 3] // heap-allocated array
return x
}Tensor Memory Pools
Tensor objects use a MemPool allocator to reduce malloc overhead. Each ObjTensor has an is_pooled flag and a pool pointer. When a tensor is created, its data buffer is carved from a pre-allocated pool. Pools support fast O(1) allocation and bulk deallocation, critical for ML training loops.
fn train_step(batch: tensor) -> float {
let x = batch[0:32] // pooled tensor view (no alloc)
let logits = model(x) // intermediate tensors use pool
let loss = cross_entropy(logits, labels)
return loss
// pool reset at end of step reuses all memory
}GPU Memory Management
Tensors can reside on GPU devices. The ObjTensor struct contains a GpuBackend enum, a gpu_handle pointer, and a gpu_backend field. Transfers between CPU and GPU use .to() methods. GPU memory is reference-counted separately and freed when the tensor is no longer referenced.
fn train_on_gpu() {
let data = [[1, 2], [3, 4]]
let gpu_data = data.to("cuda") // transfer to GPU
with device("cuda") {
let result = gpu_data * gpu_data
}
let cpu_result = result.to("cpu") // transfer back
}Memory-Safe Patterns
Braid prevents use-after-free and double-free by construction: ARC guarantees that an object lives as long as any reference exists. The compiler enforces that local variables are initialized before use and that pointers to stack data do not escape their scope.
fn safe_patterns() {
let a = [10, 20, 30] // heap, ref_count = 1
{
let b = a // ref_count = 2
print(b[0]) // safe: a is alive
} // b dies, ref_count = 1
print(a[0]) // safe: a is still alive
}Performance Notes
ARC eliminates the need for a tracing garbage collector, avoiding stop-the-world pauses. The cycle collector runs incrementally in the background. Memory pools for tensors reduce allocation overhead by ~90% in training workloads. GPU memory is managed with explicit transfer boundaries, minimizing PCI-e traffic.