Skip to main content
←Projects
Side Projects
PA

palloc

pomagrenate/palloc

Overview

An ultra-fast, lightweight, and thread-safe general-purpose memory allocator. Drop-in malloc replacement built for high throughput and low latency.

Technologies & Concepts

CMemory AllocatorLow Latency

Project Analysis

🔴 Problem

mimalloc limitation: Microsoft's mimalloc is an excellent general-purpose allocator, but it's not optimized for vector/embedding workloads where memory allocations are contiguous and huge.

Vector allocation patterns: Vector operations in machine learning and embedding workloads require large, contiguous memory blocks that are allocated and freed in batch patterns, which general-purpose allocators don't handle optimally.

Performance bottleneck: The overhead of individual malloc/free operations for large contiguous blocks creates significant performance penalties in high-throughput vector processing scenarios.

⚪ Baseline

Starting point: mimalloc (Microsoft Research) - a high-performance general-purpose memory allocator

Baseline characteristics: Excellent for general workloads, but not optimized for:

  • Large contiguous memory blocks
  • Batch allocation/deallocation patterns
  • Vector and embedding workloads
  • SIMD-aligned memory access

Expected performance: Standard malloc/free overhead for large allocations

🔵 Change

Forked mimalloc: Created palloc as a specialized fork of mimalloc optimized for vector/embedding workloads

Key architectural changes:

  • Arena-based allocation with O(1) reset vs O(N) individual frees
  • 64-byte alignment guarantees for SIMD operations (AVX-512)
  • Contiguous memory allocation for better TLB utilization
  • Reduced fragmentation in batch allocation scenarios
  • Optimized for large, contiguous memory blocks

Implementation approach: Drop-in malloc replacement via LD_PRELOAD/DYLD_INSERT_LIBRARIES/Windows DLL redirect

🟣 Measurement

Test environment: Linux with POSIX override, Dell Latitude E5440, Intel Core i5 (2 cores), 8GB RAM

Benchmark suite: Adversarial benchmarks for vector/embedding workloads including:

  • Vector batch churn (batch sizes: 32, 128, 512, 4096)
  • SIMD latency tests (hot/cold cache, aligned/unaligned)
  • TLB pressure tests (working sets up to 8GB)
  • Comparison: palloc vs system allocator

Metrics collected: Throughput, latency, memory usage, cache performance, TLB efficiency

🟢 Result

Exceptional batch performance: 18.4x average speedup, 60.53x peak speedup in batch churn scenarios

SIMD improvements: 3.57x speedup (hot aligned), 4.10x speedup (cold unaligned) for memory operations

Latency reductions: p50 latency reduced from 37-165ns (system) to 6-37ns (palloc)

TLB benefits: 1.33x improvement in TLB pressure tests with large working sets

Exceeded expectations: Significantly outperformed expected 2-4x batch improvements, achieving 18.4x average

Performance Visualizations

Speedup Comparison
Speedup comparison chart
Throughput Comparison
Throughput comparison chart
Latency Distribution
Latency distribution chart
Memory Usage Analysis
Memory usage chart

🟡 Lesson

Architecture matters for specific workloads: General-purpose allocators like mimalloc are excellent, but specialized allocators can provide order-of-magnitude improvements for specific workload patterns.

Arena-based allocation power: The O(1) reset operation vs O(N) individual frees creates dramatic performance differences at scale, especially for batch allocation patterns common in ML workloads.

Memory alignment benefits: Guaranteed 64-byte alignment provided measurable SIMD benefits (3.57x-4.10x), validating the design decision to prioritize alignment for vector operations.

Benchmark design complexity: I learned that benchmarking memory allocators is non-trivial. Workload patterns, warm-up iterations, and measurement methodology all significantly impact results. The adversarial benchmark suite was crucial for testing real-world vector/embedding workloads.

Performance vs expectations: I initially expected 2-4x improvements for batch operations, but achieved 18.4x average with 60.53x peak. This taught me that well-designed specialized allocators can outperform general-purpose ones by orders of magnitude for specific workloads.

Explore This Project