Synchronization & Memory Management

⏳ Vulkan Synchronization & Memory Management

In legacy graphics APIs, the graphics driver secretly handled synchronization and memory allocation on background threads. In Vulkan, the driver does zero background management.

If you write to a buffer from the CPU while the GPU is reading it, or if you present an image before the fragment shader finishes writing to it, your screen will flicker, tear, or crash the GPU driver.

Vulkan gives you explicit synchronization primitives and direct access to physical GPU memory heaps.


🚦 The Three Synchronization Primitives

Primitive Synchronization Scope Waiter Common Use Case
VkFence Coarse-grained (Execution done) CPU waits on GPU Waiting for Frame N rendering to complete before recording Frame N+1
VkSemaphore Coarse-grained (Queue operations) GPU waits on GPU Waiting for swapchain image acquisition before rendering; waiting for render before present
VkPipelineBarrier Fine-grained (Stages & Memory) GPU waits on GPU Image layout transitions and cache flushes between render passes

🔄 Frames in Flight & The CPU-GPU Render Loop

To prevent the CPU from idling while the GPU renders, games use Double Buffering (2 Frames in Flight):

Frame 0: [ CPU Records Commands ] ──(Submit)──> [ GPU Renders Frame 0 ]
Frame 1:                         [ CPU Records Commands ] ──(Submit)──> [ GPU Renders Frame 1 ]

The Standard Per-Frame Loop:

  1. vkWaitForFences(device, 1, &in_flight_fence, true, UINT64_MAX): CPU pauses until the previous frame using this command buffer has finished.
  2. vkResetFences(device, 1, &in_flight_fence): Reset the fence back to unsignaled state.
  3. vkAcquireNextImageKHR(...): Ask swapchain for an available image index. Signals image_available_semaphore.
  4. Record Command Buffer: Bind pipelines, push constants, and draw calls.
  5. vkQueueSubmit(...): Submit command buffer.
    • Waits on: image_available_semaphore.
    • Signals: render_finished_semaphore and in_flight_fence.
  6. vkQueuePresentKHR(...): Send image to display.
    • Waits on: render_finished_semaphore.

🚧 Pipeline Barriers & Image Layout Transitions

GPU memory layouts for textures are hardware-dependent and tiled for cache locality. An image cannot be drawn to and displayed with the same internal layout.

You must transition layouts using an Image Memory Barrier (vkCmdPipelineBarrier or VkDependencyInfo in Vulkan 1.3):

// Transitioning Swapchain Image from UNDEFINED to COLOR_ATTACHMENT_OPTIMAL
barrier := vk.ImageMemoryBarrier{
    sType               = .IMAGE_MEMORY_BARRIER,
    oldLayout           = .UNDEFINED,
    newLayout           = .COLOR_ATTACHMENT_OPTIMAL,
    srcQueueFamilyIndex = vk.QUEUE_FAMILY_IGNORED,
    dstQueueFamilyIndex = vk.QUEUE_FAMILY_IGNORED,
    image               = swapchain_images[image_index],
    subresourceRange    = {aspectMask = {.COLOR}, levelCount = 1, layerCount = 1},
    srcAccessMask       = {}, // No memory write to flush
    dstAccessMask       = {.COLOR_ATTACHMENT_WRITE}, // Subsequent writes must wait
}

vk.CmdPipelineBarrier(
    cmd,
    {.TOP_OF_PIPE},                    // Source stage
    {.COLOR_ATTACHMENT_OUTPUT},        // Destination stage
    {}, 0, nil, 0, nil, 1, &barrier,
)

💾 GPU Memory Allocation & Staging Buffers

GPUs have distinct physical memory heaps:

  1. DEVICE_LOCAL: Fast on-board VRAM (e.g. 16GB GDDR6 on an RTX 4080). The GPU reads this at hundreds of GB/s, but the CPU cannot directly write to it.
  2. HOST_VISIBLE: System RAM or PCIe BAR (ReBAR). The CPU can map this memory (vkMapMemory) and copy data with memcpy.

The Staging Buffer Pattern (Uploading Meshes & Textures)

Because device-local VRAM is much faster for vertex and index buffers, game engines never render directly from host-visible memory:

[ CPU System RAM ]
       │  (memcpy)
       ▼
[ Staging Buffer (HOST_VISIBLE | HOST_COHERENT) ]
       │
       │  (vkCmdCopyBuffer on GPU Transfer Queue over PCIe)
       ▼
[ Vertex Buffer (DEVICE_LOCAL VRAM) ] ──> Fast GPU Rendering!

Once the transfer finishes, the staging buffer can be freed or reused.


⚠️ The 4,096 Allocation Limit (Why You Need VMA)

In Vulkan, calling vkAllocateMemory() is an OS kernel-level operation. Most graphics drivers enforce a strict limit of 4,096 maximum allocations.

If you call vkAllocateMemory() for every individual mesh, texture, and uniform buffer, your engine will crash as soon as you load a realistic level.

The Solution: Sub-Allocation

Professional engines allocate massive blocks of memory (e.g., 64MB to 256MB chunks) and sub-allocate small slices (with required byte offsets) for individual meshes and textures.